Claude Opus 5.5: Benchmarks, Pricing, What Changed
Anthropic released Claude Opus 5.5 on Sept 22, priced 20% below Opus 5 with Fable-class benchmarks and roughly 40% lower cost on agentic coding tasks.
Anthropic refreshed its flagship again barely two months after the last one. On Monday, September 22, 2026, the company released Claude Opus 5.5, a mid-cycle upgrade to the Opus line that it says matches the quality of its top-end Claude Fable 5.1 on the hardest agentic work while running materially cheaper than the Opus 5 it replaces. The pitch is the same one Anthropic has leaned on all year — more work per dollar — but this time it is backed by an outright price cut rather than a flat sticker.
The release continues the compressed release cadence that has defined 2026’s frontier race. Opus 5.5 arrives less than two months after Claude Opus 5 shipped in late July, and days after competing updates from OpenAI and xAI. Anthropic’s decision to ship a “point-five” rather than wait for a full generation signals how fast the field is moving: the company would rather push incremental efficiency gains into production immediately than hold them for a marquee launch.
What Anthropic shipped
Claude Opus 5.5 is a reasoning model with extended thinking on by default, the test-time-compute approach standard across the current generation. On paper the headline specs are unchanged from Opus 5 — a 1-million-token context window and a 128,000-token maximum output — which keeps the model’s working envelope identical for agents that hold large codebases or long documents in memory. What changed is underneath: Anthropic says Opus 5.5 reaches its scores while spending fewer tokens to get there, the efficiency lever that turns a modest benchmark bump into a real cost reduction.
The company is positioning the model squarely at its core audience of developers and agent builders, emphasizing agentic coding, knowledge work, and computer use — the long-horizon, tool-driven tasks where a model has to plan, act, and recover over many steps rather than answer a single prompt. That is the workload we survey in our look at the state of AI coding assistants, and it is the segment Anthropic has fought hardest to hold as challengers undercut it on price.
The pricing story
The most deliberate part of the announcement is the number on the invoice. Claude Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens — roughly 20% below the $5 / $25 Opus 5 charged. Prompt-cache reads fall further, dropping about 60% to $0.20 per million tokens, a change aimed directly at agent workloads that re-read the same context hundreds of times across a task.
Anthropic frames the combined effect as roughly 40% lower execution cost than Opus 5 on real agentic runs, because the model both charges less per token and burns fewer tokens to finish a job. That two-part savings — a lower rate multiplied by fewer tokens — is the claim that matters to buyers running the model at scale, where a 40% cut in the cost of an agent loop compounds across millions of calls.
The move reads as a direct response to a market that has spent the year watching per-token prices fall. Chinese open-weight labs and aggressive challengers have pushed capable models to a few dollars per million tokens, and after holding the Opus price flat in July, Anthropic has now blinked on the sticker itself — betting it can defend a premium tier only by making each generation cheaper to run, not just more capable.
The benchmarks
On the figures Anthropic published in its September 22 release table, Opus 5.5 improves on Opus 5 across the board and, by the company’s account, edges ahead of both Fable 5.1 and OpenAI’s flagship on the tasks it cares most about.
- On SWE-bench Pro, the harder variant of the most-watched software-engineering benchmark, Anthropic reported 89.9%, a large jump from the 79.2% it credited to Opus 5 at launch.
- On Terminal-Bench 4.0, which measures how well a model drives a real terminal over long horizons, it reported 66.4%.
- On CursorBench, an editor-integrated coding evaluation, it reported 57.8%, and on FrontierCode, 54.4%.
- On GDPval-AA, a knowledge-work benchmark scored in Elo, it reported 1,846.
Anthropic says Opus 5.5 outscores Opus 5 on all nine benchmarks in the release table and leads Fable 5.1 and GPT-6 Astra on agentic coding, knowledge work, and computer use. As always, these are vendor-reported numbers from the company’s own testing; independent replications will follow, and the field’s evaluation practices reward a skeptical read, as we discuss in what an LLM eval is. The pattern to note is not any single score but the shape of the claim: near-top-end quality delivered by a cheaper model, rather than a new capability ceiling.
Where it sits in the lineup
Opus 5.5’s arrival slots it beneath Fable 5.1, Anthropic’s premium flagship, while claiming Fable-class results on the workloads most developers actually run. That is an unusual position — a mid-tier model that the vendor says matches its own top tier on coding and computer use — and it reflects a deliberate segmentation. Anthropic appears to be reserving Fable for the frontier headline and using the Opus line as the volume workhorse, pricing it to win the day-to-day agent traffic that competitors like xAI’s Grok have been chasing with “Opus-class at a fraction of the price” messaging.
For existing Opus 5 users, the migration is the easy kind: same context window, same output ceiling, lower price, higher scores. The friction, if any, will be in behavior — a new model that plans and recovers differently inside an agent loop can change how tuned prompts and tool chains perform, and teams running Opus in production will want to re-validate their harnesses before assuming the cheaper model is a drop-in.
What it means
The clear winner is the buyer running agents at scale. A 20% rate cut, a 60% drop in cache-read pricing, and fewer tokens per task combine into a cost structure that makes long-running agent workloads meaningfully cheaper to operate, and Anthropic is betting that lower unit economics keep developers on Opus rather than defecting to cheaper challengers. For a company whose revenue is increasingly tied to token volume, trading price for retained usage is a rational bet — and one that pressures its own margins if the efficiency gains don’t fully offset the lower rate.
The pointed target is the rest of the frontier field. By shipping Fable-class quality in a cheaper, faster package barely two months into Opus 5’s life, Anthropic is trying to collapse the space competitors use to undercut it — the gap between “almost as good” and “much cheaper” that rivals have exploited all year. If Opus 5.5’s independent benchmarks hold up, the standard sales pitch of an Opus-class-for-less challenger gets harder to make, because Opus itself just got cheaper.
Two things are worth watching from here. First, whether outside evaluations confirm the 40% cost claim in production, where token-efficiency gains often shrink once real prompts, tool errors, and retries enter the picture. And second, how OpenAI and xAI answer a mid-cycle price cut from the model they were trying to undercut — because in a market where a flagship can be refreshed every eight weeks, the release cadence itself has become a competitive weapon, and Anthropic just used it.
Tagged
Keep reading
Chisato · · 6 min read Claude Opus 5: Benchmarks, Pricing, and 1M Context
Anthropic launched Claude Opus 5 on July 24 with a 1M-token context, a new xhigh effort mode, and unchanged $5/$25 pricing. Benchmarks, specs, and what changed.
Chisato · · 3 min read Is There a Claude Sonnet 5? Anthropic's 2026 Lineup
Looking for Claude Sonnet 5? Here's the honest answer — plus a clear map of Anthropic's 2026 models: Haiku 4.5, Sonnet 4.6, Opus 4.8, and the new Fable 5.
Chisato · · 6 min read Anthropic Claude Formalizes Fermat's Last Theorem in Lean
Anthropic says Claude agents produced the first complete, machine-checked proof of Fermat's Last Theorem in Lean — 13M lines of code in about 11 days.