On September 22, 2026, Anthropic released Claude Opus 5.5, the newest version of its Opus model. In short: it is smarter at coding and long, multi-step tasks, it answers faster, and it costs less than the Opus it replaces.
What's new, in plain English
Anthropic makes several Claude models. Opus is its main "do the hard work" model, and Fable is its top-of-the-line model, priced well above Opus. The surprise in this release is that the new Opus now beats Fable on most of the coding and computer-task tests Anthropic published, while costing less than half as much.
Three things stand out:
- It gets more done with fewer steps. Anthropic and its early customers say Opus 5.5 finishes the same work in fewer back-and-forth turns, so long tasks like writing and fixing code, analysing documents or working through a spreadsheet wrap up sooner.
- It is faster. Anthropic says replies arrive about 30% faster than with Opus 5.
- It is cheaper. The price per word is 20% lower, and because it needs fewer steps, Anthropic says it costs about 40% less than Opus 5 in practice.
It also keeps Opus 5's very long memory: it can read and reason over roughly 550,000 words in one conversation, according to Anthropic's documentation, enough for several long reports or a large codebase at once.
AnthropicHow it compares
Anthropic tested Opus 5.5 against its own Fable 5.1 and Opus 5, and against OpenAI's GPT-6 Astra. Each bar below is a different test; higher is better.
Opus 5.5 against Fable 5.1, Opus 5 and GPT-6 Astra
Launch-day scores published by Anthropic on 22 September 2026, Opus 5.5 at max effort. Higher is better. Self-reported by the vendor, not independently audited.
Opus 5.5 leads the coding and computer-use rows. GPT-6 Astra stays ahead on Terminal-Bench-Science and AutomationBench.
What the tests measure, and what the results say:
- Coding tests (Terminal-Bench, CursorBench, FrontierCode) ask the model to complete real programming tasks. Opus 5.5 posts the highest published score on all three (OpenAI has not published a CursorBench score for GPT-6 Astra). On Terminal-Bench it scores 66.4%, against 55.8% for Fable 5.1 and 57.9% for GPT-6 Astra.
- Computer use (OSWorld) asks the model to operate a computer the way a person would, clicking and typing through apps. Opus 5.5 scores 81.8%, just ahead of Fable 5.1.
- Hard questions (Humanity's Last Exam) are expert-level questions across many subjects. Opus 5.5 scores 67.7%, the highest of the four.
- Where others lead: GPT-6 Astra stays ahead on a science-task test (64.6% against 58.7%) and, narrowly, on a business automation test (41.4% against 40.0%).
One caution: these numbers come from Anthropic itself, and independent testing often lands a little lower. The overall picture, a clear step up over Opus 5 and a close race at the top, is still useful.

What it costs
Every Opus price went down. The chart shows Anthropic's list prices per million tokens (a token is a short chunk of text, a little over half a word on average for this model).
Every Opus rate falls, cache reads the most
USD per million tokens, standard API pricing. Input, output and cache writes fall 20%; cache reads fall 60%.
Cache reads drop from 0.1x to 0.05x of the input rate, the same lever Anthropic pulled for Fable 5.1.
The biggest cut, 60%, is on "cache reads": the cost of re-reading material the model has already seen earlier in a conversation, such as a long document you uploaded. That matters most in long conversations and long-running tasks, where the same material is read again and again.
The new Opus beats Anthropic's pricier flagship on most coding tests, at less than half the price.
On the Opus 5.5 release
What early users said
Anthropic shared results from more than twenty companies that tried Opus 5.5 before launch. Most of them describe getting the same or better work done with less effort:
- Quantium said a coding task that previously took 38 prompts over four days was done in 11 prompts over three hours.
- Deloitte said that even on its lowest effort setting, Opus 5.5 caught 72% of known bugs in a code review, compared with 56% for Opus 5 on a high setting.
- Box said answers were 40% shorter without losing accuracy.
These examples were chosen by Anthropic, so treat them as highlights rather than a controlled comparison.

Safety
Anthropic says Opus 5.5 is the best-behaved model it has tested on its safety checks. It reports that the model tried to get around the limits placed on it "around 85% less often" than Opus 5. Like Anthropic's other recent models, it has built-in safety checks that may decline some requests in sensitive areas such as cybersecurity or biology. On Metir, if that happens, you get a clear message explaining why and what to try instead, rather than a blank reply.
What this means for you
For everyday use, Opus 5.5 is simply a better Opus: quicker, cheaper and stronger on hard work. The notable part is that the gap between Anthropic's regular and premium models has almost closed for most tasks, which makes Opus 5.5 the sensible default for serious work.
It is also a reminder that no single AI model wins everything. Opus 5.5 leads on coding and computer tasks, while GPT-6 Astra still leads on some science and automation tests. Having several models in one place, and switching between them freely, is the easiest way to always use the best one for the job.
How to use Claude Opus 5.5 on Metir
Sources:
- Introducing Claude Opus 5.5 - Anthropic
- Claude Opus 5.5 model overview - Claude Docs
- What's new in Claude Opus 5.5 - Claude Docs
- Pricing - Claude Docs
- Introducing Claude Fable 5.1 and Claude Mythos 5.1 - Anthropic
Image credits
Header image and both in-body images: artwork from Anthropic's announcement Introducing Claude Opus 5.5. Images: Anthropic.
