Anthropic and OpenAI released new models ninety minutes apart
Ronni Holmvig Strøm · 2026-09-23
Anthropic shipped Claude Opus 5.5 on Tuesday. Ninety minutes later, OpenAI shipped GPT-6 Sol and Luna. Neither company led with a capability claim. Anthropic's first sentence says Opus 5.5 performs at the level of Claude Fable 5.1 and [costs 40% less to
Anthropic shipped Claude Opus 5.5 on Tuesday. Ninety minutes later, OpenAI shipped GPT-6 Sol and Luna. Neither company led with a capability claim. Anthropic's first sentence says Opus 5.5 performs at the level of Claude Fable 5.1 and costs 40% less to run than Opus 5. The only bolded line in OpenAI's opening, per ZDNET's read of the release, is a 50% API price cut. Two labs, one afternoon, the same pitch.
Further down Anthropic's post, the company grades its own leaderboard. "At these levels of capability we've found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest."
That is a lab posting 66.4% on Terminal-Bench 4.0, against 57.9% for GPT-6 Astra, and telling you in the same breath that the eight-point gap does not mean what a reader would assume.
What replaces the margin is drawn right there in the charts. Anthropic plots accuracy against cost per attempt, cost on a log axis. Opus 5.5 at default effort beats Opus 5 at maximum effort for about a fifth of the price. It matches Astra on Terminal-Bench at roughly 40% of the cost, and on FrontierCode its default-effort score of 54.6% edges Astra's best result for about a fifth of the cost per task.
OpenAI published the same shape of argument. On DeepSWE 1.1, GPT-6 Luna at medium effort scored comparably to Opus 5 and Fable 5 while costing 93% and 96% less per task. At higher effort, Luna matches GPT-5.6 Sol at about a hundredth of its cost. GPT-5.6 shipped in July.
The unit of account changed this week. The comparison artifact both labs now publish is a curve, and the axis that moved is cost per completed task.
It is a better instrument than a benchmark score, for a reason that has nothing to do with capability. A leaderboard margin is contested and contaminated, and no deployer can audit one. A cost-per-task figure shows up on your own invoice at the end of the month. The field spent three years arguing about whether evals measured anything real. The number that quietly replaced them is the one finance was already tracking.
The practical consequence is a widening of what is worth attempting. An early tester audited and fixed a 200,000-line codebase in under three hours, against more than twenty hours and 2.5 times the tokens for Opus 5. In an internal test, Opus 5.5 translated HAProxy from C into Rust in 9.5 hours where Fable 5.1 took twelve, at 51% lower cost, both passing nearly all of HAProxy's own regression tests. These are unglamorous jobs. They are also exactly the work that sat undone at last quarter's prices, because nobody could justify the line item.
Opus 5.5 is the first release since Anthropic called for pacing the frontier. It was evaluated before release by METR and Frontier Design, and it scored higher than any model Anthropic has tested on its automated behavioral audit, with better resistance to prompt injection and less tendency toward hard-to-reverse actions. A roughly held ceiling, a halved price, and the best alignment audit on record, all in one release.
Both labs also shipped better prose. Anthropic notes that clearer output made Opus 5.5's work easier to follow and check, "which is a safety benefit as well as a practical one." Legibility is an oversight property. A model whose reasoning you can verify in a quarter of the time is a model you can supervise at four times the scale, and supervision capacity is the binding constraint on how much of this work any organization can responsibly hand over.
One question the twelve-page OpenAI release leaves open, and nobody has yet tested: whether a 50% API cut reaches the people on subscription plans at all. The API buyers can already read what Tuesday bought them off a cost-per-task curve. The subscription tier is the next place that curve needs to be drawn.
Sources: Anthropic, 2026, TechCrunch, 2026, ZDNET, 2026.