On July 24, 2026, Anthropic released Claude Opus 5, the fourth new Claude 5-generation flagship to ship in under two months. It follows Mythos 5 (limited to approved organizations), Fable 5, and Sonnet 5, and it now stands as the default model on Claude Max and the strongest option available on Claude Pro.
The headline is less about a single new capability than about where Opus 5 lands on the price/performance map. It posts benchmark gains that roughly double its predecessor's scores on agentic coding, while pricing stays exactly where it was: $5 per million input tokens and $25 per million output tokens, unchanged from Anthropic's Claude Opus 4.8.
AnthropicWhat's new in Claude Opus 5
Opus 5 ships with a 1 million token context window as both the default and the maximum, up from the 200K/1M split structure of earlier Opus releases, and a maximum output of 128,000 tokens. Extended thinking is on by default, and Anthropic added a new "xhigh" effort tier that sits above "high" and below an uncapped "max" level, aimed at long-horizon agentic and coding work that can run for 30 minutes or more. According to BitsMinds' coverage of the launch, xhigh effort is reported to outperform max on several evaluations while using roughly 25% fewer output tokens, and thinking becomes mandatory at both the xhigh and max settings.
Pricing carries over three tiers: standard API access at $5/$25 per million tokens, a "fast mode" running at roughly 2.5x speed for double the base price ($10/$50), and Batch API access at $2.50/$12.50. That fast-mode rate happens to land exactly on Fable 5's standard pricing, giving buyers a rough sense of what a step up in either speed or raw capability costs relative to Opus 5's baseline.
Benchmark results: a large jump over Opus 4.8
The clearest generational gain shows up on Frontier-Bench v0.1, Anthropic's agentic coding evaluation. Opus 5 scores 43.3%, more than double Opus 4.8's 18.7%, and ahead of both Claude Fable 5 (33.7%) and OpenAI's GPT-5.6 Sol, which different trackers place in the mid-30s (34.4% to 37.5% depending on the source).
Opus 5 more than doubles Opus 4.8 on agentic coding
Frontier-Bench v0.1 scores at launch, July 2026. Data from Digital Applied and Vellum. Higher is better.
Opus 5 (green) jumps to 43.3%, ahead of both Fable 5 and GPT-5.6 Sol, and more than double Opus 4.8’s 18.7% from earlier in 2026.
On ARC-AGI-3, a benchmark built around novel problem-solving that resists memorization, Opus 5 reaches 30.2% against GPT-5.6 Sol's 7.8%, roughly a fourfold gap, and against Opus 4.8's 1.5% according to Digital Applied's benchmark breakdown. On SWE-bench Pro, a harder successor to the widely cited SWE-bench Verified suite, Opus 5 posts 79.2%, about 10 points above Opus 4.8. And on GDPval-AA v2, which scores economically valuable knowledge work using an Elo-style rating, Opus 5 reaches 1,861 versus 1,736 for GPT-5.6 Sol, with Fable 5 at 1,747 and Opus 4.8 at 1,593.

Anthropic also reports Opus 5 leading on Zapier's AutomationBench and on OSWorld 2.0, a computer-use benchmark, where it scores 70.6% against Fable 5's 66.1%, Opus 4.8's 55.7%, and GPT-5.6 Sol's 62.6%. Not every comparison favors Opus 5: on OpenAI's DeepSWE v1.1 coding benchmark, GPT-5.6 Sol scores higher, a reminder that benchmark leadership is distributed across evaluations rather than concentrated in one model.
The price/performance frontier is compressing
The more interesting story sits in how Opus 5 is priced relative to what it scores. On CursorBench 3.2, a coding-agent benchmark, Opus 5 lands within 0.5 percentage points of Fable 5's peak score at maximum effort, according to both Vellum's analysis and Anthropic's own materials, while costing roughly half as much per task. Fable 5 is priced at $10/$50 per million tokens; Opus 5 held at $5/$25.
Opus 5 moves the price/performance frontier, not just the score
Blended cost per million tokens (weighted 3:1 input:output) versus Frontier-Bench v0.1 score. Toward the top-left is more capability for less money. Pricing and scores from provider pages and independent benchmark trackers, July 2026.
Opus 5 (green) sits at the same price as Opus 4.8 but far higher on the score axis, landing close to Fable 5’s performance at half its blended cost.
Opus 5 does not claim to beat Fable 5 outright. It claims to get close enough, for less money, that the gap stops mattering for most work.
Framing drawn from Anthropic's Opus 5 launch positioning
That is the practical takeaway for buyers: "near-parity at half the price" is not the same claim as "better than the frontier model." It means a large share of day-to-day agentic and coding work, the kind that previously required paying frontier-tier rates, can now run on a cheaper model without a meaningful quality drop. For teams running high request volumes, that difference compounds quickly across a monthly bill. It doesn't eliminate the case for Fable 5 or GPT-5.6 Sol on tasks where the last few points of capability matter, but it narrows the set of tasks where paying the premium is clearly justified.
An accelerating release cadence
Opus 5 is notable simply for how quickly it arrived. Mythos 5, Fable 5, Sonnet 5, and now Opus 5 have all shipped within roughly eight weeks, according to Digital Applied's account of the release sequence. That cadence mirrors a broader industry pattern in 2026: rather than a single yearly flagship release, labs are now shipping a family of differently sized and priced models in quick succession, letting a single generation cover multiple points on the cost/capability curve rather than one model trying to serve every use case.
Anthropic also published an automated behavioral audit score of 2.30 for Opus 5, described as the lowest, and by that measure best, of recent Claude models, alongside a note that certain misuse-related classifiers trigger roughly 85% less often than they did for Fable 5. Whether that reflects a durable safety trend or a one-model data point is not yet possible to say from a single release, and Anthropic's own reporting should be read as self-published rather than independently audited.
No single model wins every task
The pattern across every comparison here, Opus 5 ahead on Frontier-Bench and ARC-AGI-3, GPT-5.6 Sol ahead on DeepSWE and HealthBench Professional, Fable 5 marginally ahead on raw CursorBench ceiling, reinforces something that has held for most of 2026: no model wins across the board. Model choice increasingly depends on the specific task, and the right answer for a coding agent is not always the right answer for a knowledge-work summarization job or a long-document analysis.
That's part of why platforms that give access to multiple frontier models side by side, rather than locking a workflow to a single provider, have become more useful as the pace of releases has picked up. Metir gives users access to Opus 5 alongside GPT, Gemini, and Grok models in one place, so a task can be routed to whichever model actually performs best on it without juggling separate subscriptions.
Sources:
- Anthropic news
- Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing - MarkTechPost
- Claude Opus 5 vs GPT-5.6 - CodersEra
- Claude Opus 5 vs GPT-5.6 Sol comparison - LLM Stats
- Claude Opus 5 Benchmarks Explained - Vellum
- Claude Opus 5: Frontier Intelligence at Half the Price - Digital Applied
- Claude Opus 5 Is Officially Here: 1M Context, a New 'xhigh' Mode - BitsMinds
- Claude Opus 5 Launch: Benchmarks, Price, Fast Mode - ExplainX
Image credits
Header image: Anthropic co-founder and CEO Dario Amodei at 10 Downing Street, London, May 2023, by the UK Prime Minister's Office (Simon Walker / No 10 Downing Street) via Wikimedia Commons, licensed under CC BY 2.0. In-body photo of Dario Amodei speaking at TechCrunch Disrupt 2023 by TechCrunch via Wikimedia Commons, licensed under CC BY 2.0.
