Anthropic ships Claude Opus 5.5: flagship coding at 40% lower cost — and a slower frontier
Anthropic's first release since Dario Amodei's "pace the frontier" essay is a cheaper, faster Opus with benchmark leads on coding evals. The figures are all vendor-reported — and some come with asterisks worth reading.

On September 22, Anthropic released Claude Opus 5.5 — the first model of the Claude 5.5 family, and the company's first release since chief executive Dario Amodei's August essay arguing that the breakneck era of frontier progress is ending, and that labs should pace deployments for reliability. The model that followed is, fittingly, less a leap than a deliberate repositioning: flagship-level coding performance, measurably faster output, and pricing cut deep enough that Anthropic claims a typical workload costs about 40 percent less than on Opus 5.
It arrived the same day OpenAI unveiled GPT-6 Sol and Luna — a dual launch that tells you the industry's 2026 playbook more honestly than any benchmark chart. The race is no longer only about who tops the leaderboard. It is about who delivers the point at the lowest price per task.
The price of the frontier, repriced#
The headline number is the price list. Opus 5.5 costs $4 per million input tokens and $20 per million output tokens on the Anthropic API, against $5 and $25 for Opus 5 — a straight 20 percent cut. Cache reads, the line item that quietly dominates heavy agentic workloads, reportedly fell from $0.50 per million to $0.20 per million, a 60 percent drop. Anthropic says typical workloads land about 40 percent cheaper than on Opus 5, and claims roughly Fable-5.1-level performance on most tasks at about 60 percent lower cost than Fable 5.1.
Read the fine print before budgeting. Independent coverage notes that the 40 percent figure compares Opus 5.5 at its default effort setting — medium — with Opus 5 running at high effort, while the headline benchmarks below were generated at maximum effort. Those are different races measured on different tracks. Anthropic also flags four breaking API changes that teams will need to handle when migrating, which matters for anyone planning the switch on launch day.
The list price is also the latest chapter in the token-price war the industry has been quietly fighting all year. Cache reads at $0.20 per million put Opus 5.5's effective cost for long-context agentic loops well under anything Anthropic has shipped — and invite the obvious comparison with the cache-tiering moves from OpenAI and Google earlier this year. Every lab is now discounting the same line items.
The numbers — and the asterisks#
On coding, Anthropic's own benchmarks put Opus 5.5 clearly ahead of the field. These are vendor-run figures — treat them as the company's best case until third parties reproduce them.
| Benchmark | Opus 5.5 | Comparison (vendor figures) |
|---|---|---|
| Terminal-Bench 4.0 | 66.4% | Fable 5.1: 55.8% · GPT-6 Astra: 57.9% |
| FrontierCode | 54.4% | Best reported score, per Anthropic |
| GDPval-AA v2.1 | 1,846 | Fable 5.1: 1,735 · Opus 5: 1,708 |
| OSWorld 2.0 | 81.8% | New reported high |
| CursorBench | +11 pts over GPT-5.6 Sol | At roughly a third of the cost |
The exceptions matter. GPT-6 Astra still leads on AutomationBench and Terminal-Bench-Science 0.1, where it scores 64.6 against Opus 5.5's 58.7 — a reminder that "best coder" depends on which coding you mean. And Anthropic itself has cautioned that once frontier models saturate these evals, the margins get less reliable: models are increasingly good at recognizing when they are being tested, which means benchmark-aware behavior can flatter results without improving real-world performance.
Faster, clearer, and a 680,000-line anecdote#
Speed is the part of the story that needs no benchmark. Anthropic reports output running more than 30 percent faster than Opus 5, and says the model writes more clearly — a direct answer to one of the most consistent complaints about its predecessor. The company pairs this with customer anecdotes: one customer migrated a 680,000-line codebase in under a day, and another ran the model 2.4 times longer on a harder task at the same price as before.
Anecdotes are marketing, but they are chosen to make a real point about where the value lands. For teams already using Opus-class models, the pitch is not "smarter" — it is faster, cheaper, and less verbose on the same hardware budget. That is the efficiency playbook the whole industry is converging on this year.
Safety, with receipts requested#
Anthropic leaned harder on external evaluation this time, with pre-release testing by Frontier Design and METR — a welcome move in an era when most capability claims are graded by the vendor. On Gray Swan's prompt-injection suite, the company reports Opus 5.5 had the lowest attack success rate of any Claude model, tied with Fable 5.1. Containment-boundary circumvention attempts fell roughly 85 percent versus Opus 5 or Mythos 5.1 — though that figure comes from Anthropic's own internal evaluation, not a third party.
The model ships with Fable-5.1-level safeguards, additional jailbreak defenses, and expanded vetting for researchers. Anthropic is also routing around its own frontier edges: cyber queries the model handles better than Opus are redirected to Opus 4.8, and higher-risk biology work goes through the company's Life Sciences Verification Program. The structure is honest about a tension the industry rarely names — the most capable model is not always the one that should answer.
Availability#
Opus 5.5 is available on the Anthropic API and through AWS, Google Cloud, and Azure. Anthropic says Sonnet 5.5 and Haiku 5.5 will follow in the coming weeks — and that is where the volume story gets interesting. The flagship gets the headlines, but the mid-tier and lightweight models are where API spend actually lives.
Why this matters#
Step back from the scoreboard and the pricing page, and Opus 5.5 is a thesis about where the frontier goes next. Amodei's August essay argued the era of doubling capabilities is ending; Anthropic's first release since is a model that does not try to double anything. It codes as well as the best, runs faster, and costs substantially less per task — then gets undercut the same day by OpenAI's Sol and Luna at half their predecessors' prices.
This is what the slower frontier looks like: not stagnation, but a shift in the axis of competition. When every lab can field a model within a few points of the best, the differentiator becomes economics — tokens per dollar, latency per query, containment per deployment. Opus 5.5 is Anthropic's bid to win that game on the coding workloads its customers actually run.
There is a quieter signal here too. This is Anthropic's first release under the doctrine Amodei laid out in August — that the frontier's pace should be deliberately managed, not maximized. Fittingly, Opus 5.5 is the first flagship-era launch that asks to be judged on reliability, cost, and containment rather than raw capability. Whether the market rewards that restraint is the experiment now running.
What to watch#
- Independent reproduction. Every figure above is vendor-graded. The first serious third-party runs — Artificial Analysis, GDPval's own leaderboard updates — are this story's real next chapter.
- Effort-level honesty. The 40 percent cost claim rests on a medium-vs-high effort comparison. Watch whether it holds when both models run at the same effort — and whether Anthropic keeps publishing apples-to-apples numbers.
- The sibling models. Sonnet 5.5 and Haiku 5.5 land in the coming weeks. The mid-tier pricing is where the efficiency story either pays off for builders or quietly narrows.
- OpenAI's answer. Astra still holds the AutomationBench crown. The next benchmark round — and the next price cut — is a matter of weeks, not months.
The frontier is getting slower and cheaper at the same time. Opus 5.5 is the first model built for that world — and the first test of whether the new economics hold.