Claude Opus 5.5 vs. GPT-6 Sol/Luna: Same-Day Price Cuts, Neither Official Chart Reaches the Other's New Model
On September 22, 2026, Anthropic and OpenAI both cut prices on the same day. Benchmarks take weeks to run before a launch, so neither company could have tested against a rival model announced that same morning — this piece lays out what each company's pricing and comparison charts actually show, where they stop, and what's actually worth comparing as a reader.
Tuesday, September 22, 2026. Anthropic released Claude Opus 5.5, and OpenAI released GPT-6 Sol and GPT-6 Luna — the same outlet, TechCrunch, covered the two announcements about ninety minutes apart (9:30 AM and 11:00 AM Pacific). Same day, two companies, both moving their price lists down a notch.
That’s not really a surprise — benchmark evals have to be locked in weeks before a launch, so neither company could have gotten its hands on a rival model that was only announced that same morning. Lay each company’s own comparison chart next to the other and that shows up directly: Anthropic’s chart has no GPT-6 Sol, no GPT-6 Luna. OpenAI’s chart has no Claude Opus 5.5. Two documents published the same day, each one benchmarked against the rival’s previous-generation model. Which points to something more practical for readers: neither official comparison chart can actually tell you which of the two new models is stronger.
What actually got cheaper
Opus 5.5’s pricing is straightforward: $4 and $20 per million tokens for input and output, 20% below Opus 5. Cache reads drop to $0.20 per million, 60% below Opus 5. Anthropic says that at default settings, typical-workload costs fall 40% overall, and output generation runs more than 30% faster. On the subscription side, something else changed: the five-hour usage limit on Pro, Max, Team, and seat-based Enterprise plans went up, plus a new rate-limit reset you can save and deploy whenever you choose. That’s a quota change, not a price change — a different decision from the API price cuts, and not one that belongs on the same scale.
Anthropic frames Opus 5.5 as “the first model in our new Claude 5.5 family”, matching the tier above it, Fable 5.1, on most work while costing 40% less than Opus 5 to run. The announcement also carries a line of context: last week, Anthropic CEO Dario Amodei argued that AI progress should be paced so that safety practices stay ahead of model capabilities — Opus 5.5 is the company’s first release since that call. Anthropic backs the release with concrete anecdotes from early testers: one engineer used it to complete a 680,000-line code migration in under a day, work that would have taken an engineering team weeks; in another test, asked to cut page-load times across a web app, Opus 5.5 succeeded 39 out of 40 times, while Opus 5 made smaller gains that also altered the app’s behavior.
OpenAI positions GPT-6 Sol and Luna below its flagship, GPT-6 Astra, as the value tier — Astra launched September 3 as a new model, with no price cut attached to that release. Sol’s pricing goes from $4 to $2 on input and $20 to $10 on output, an exact 50% across the board. Luna’s input also halves, from $0.20 to $0.10, but its output drops from $1.20 to $0.50 — a 58% cut, more than the press release’s blanket “50% cheaper” line. That reduction is measured against GPT-5.6’s promotional pricing, not list pricing. Sol and Luna are now live in ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users; Free and Go users get Luna only in the desktop app, and neither model is in the regular Chat product yet.
| Claude Opus 5.5 | GPT-6 Sol | GPT-6 Luna | |
|---|---|---|---|
| Released | 2026-09-22 | 2026-09-22 | 2026-09-22 |
| Input (per million tokens) | $4 (-20%) | $2 (-50%) | $0.10 (-50%) |
| Output (per million tokens) | $20 (-20%) | $10 (-50%) | $0.50 (-58%) |
| Cache reads | $0.20 (-60%) | Only a 90% discount disclosed, no absolute price | Same |
| Speed / positioning | -40% typical cost, +30% output speed | ~Half the error rate of predecessor on internal factuality eval | ~1/100th the cost of predecessor at higher effort |
Who each company benchmarks against
The pricing table above is uncontroversial — both companies published those numbers directly. The performance comparisons are constrained differently: both are locked in by eval timelines, so both stop at the previous generation.
Anthropic’s comparison chart lists five models: Opus 5.5, Fable 5.1, Opus 5, GPT-6 Astra, and GPT-5.6 Sol — not one cell mentions GPT-6 Sol or GPT-6 Luna. The clearest illustration is on CursorBench: Anthropic writes that Opus 5.5 “beats GPT-5.6 Sol by 11 points for about a third of the cost”, but that “11 points” figure uses default (medium) effort scores. The main table’s max/xhigh scores put the gap at 57.8% versus 41.7% — a 16.1-point difference. Same benchmark, different effort level, different number. The methodology note underneath says the GPT-6 Astra and GPT-5.6 Sol figures are as reported by OpenAI — not independently reproduced.
OpenAI’s chart runs the same way in reverse: not one cell mentions Claude Opus 5.5. On AutomationBench, GPT-6 Sol at xhigh effort scores 33.2% at $0.27 per task, against “Claude Fable 5.1 w/ Opus 5 Fallback” at 31.4% — more than 8.9x Sol’s cost, with the fallback’s own cost left out of that number; OpenAI’s own text notes the fallback triggered on roughly 40% of tasks. OpenAI’s methodology note reads almost identically to Anthropic’s: competitor scores were taken from publicly available reports, not run in-house.
Both charts are doing the same thing, symmetrically: each company measures its newest model against whatever the rival had already shipped by eval time. That’s a function of how release cycles work, not either company dodging the other — but the result is the same: neither of these two “same-day” documents constitutes a real head-to-head test of what actually launched that day.
Alignment: each company grading its own homework
Both releases lead with alignment claims. Anthropic calls Opus 5.5 “the strongest-performing model we’ve tested to date”, based on its own “automated behavioral audit” — an internal evaluation suite that tests model behavior across scenarios. The announcement separately notes that Opus 5.5 was also evaluated before release by outside organizations, Frontier Design and METR — a distinct, pre-release external check, not the same exercise as the internal audit claim above it.
Elsewhere in the announcement, Anthropic gives a more specific number: on its primary evaluation suite, an audit spanning nearly 2,000 scenarios, Opus 5.5 scored better than any recent Claude model on nearly every measure of misaligned behavior, and is also its strongest model on most measures of honesty.
OpenAI carries Astra’s alignment training forward into Sol and Luna, saying both show improvement over their GPT-5.6 counterparts, including lower rates of misleading claims about coding work. That, too, is a number OpenAI is reporting about its own internal evaluation — not something an independent third party verified.
The same script, one month running
Zoom out to the calendar and this is actually the third round this month. On September 1, Anthropic cut cache-read pricing 75% on its existing Fable 5.1 and Mythos 5.1 models. On September 3, OpenAI launched its new flagship, GPT-6 Astra — no price cut attached, just a new model flexing capability. It took until September 22 for both companies to actually move together on the value tier: Anthropic’s across-the-board cut to Opus 5.5, OpenAI’s halving of Sol and Luna. Within one month, both companies ran the same playbook — flex the flagship first, then push the entry price down.
What’s actually worth comparing
Read only the headlines and this “price war” sounds like two companies going head-to-head, racing each other on speed and cost. Look at what the documents actually contain, though, and what’s really moving isn’t the outcome of any single matchup — it’s the AI inference cost curve itself, bending down. In one month, both companies each flexed a flagship and each cut a price once — and each price cut made the same tier of intelligence cheaper than before.
As for which company to actually use, neither official comparison chart can answer that, because each one stops at the rival’s previous generation. What can answer it is the workload in front of you — how many tokens it burns, whether it can tolerate retries, how much of the bill is cache hits. The price lists will keep moving; the next round could land next month. What’s worth remembering isn’t any single chart either company posted on launch day — it’s the direction itself: AI inference keeps getting cheaper.