metir
metir
Download on App StoreGet it on Google PlayF1 FantasyLoginSign Up
Back to Blog
Claude Opus 5
Anthropic AI
AI Benchmarks 2026
Claude vs GPT-5.6
AI Model Pricing

Claude Opus 5: Anthropic's New Flagship AI Model Benchmarks (2026)

Anthropic's Claude Opus 5 launched July 24, 2026 with unchanged Opus pricing, a 1M-token context window, and benchmark scores that close in on Claude Fable 5 at half the cost.

Metir AI TeamJuly 26, 20266 min read
Claude Opus 5: Anthropic's New Flagship AI Model Benchmarks (2026)

On July 24, 2026, Anthropic released Claude Opus 5, the fourth new Claude 5-generation flagship to ship in under two months. It follows Mythos 5 (limited to approved organizations), Fable 5, and Sonnet 5, and it now stands as the default model on Claude Max and the strongest option available on Claude Pro.

The headline is less about a single new capability than about where Opus 5 lands on the price/performance map. It posts benchmark gains that roughly double its predecessor's scores on agentic coding, while pricing stays exactly where it was: $5 per million input tokens and $25 per million output tokens, unchanged from Anthropic's Claude Opus 4.8.

Anthropic logoAnthropic
OpenAI logoOpenAI
Claude Opus 5 launches into a market where OpenAI's GPT-5.6 tier remains the closest frontier-model comparison.

What's new in Claude Opus 5

Opus 5 ships with a 1 million token context window as both the default and the maximum, up from the 200K/1M split structure of earlier Opus releases, and a maximum output of 128,000 tokens. Extended thinking is on by default, and Anthropic added a new "xhigh" effort tier that sits above "high" and below an uncapped "max" level, aimed at long-horizon agentic and coding work that can run for 30 minutes or more. According to BitsMinds' coverage of the launch, xhigh effort is reported to outperform max on several evaluations while using roughly 25% fewer output tokens, and thinking becomes mandatory at both the xhigh and max settings.

Pricing carries over three tiers: standard API access at $5/$25 per million tokens, a "fast mode" running at roughly 2.5x speed for double the base price ($10/$50), and Batch API access at $2.50/$12.50. That fast-mode rate happens to land exactly on Fable 5's standard pricing, giving buyers a rough sense of what a step up in either speed or raw capability costs relative to Opus 5's baseline.

43.3%Frontier-Bench v0.1 scorevs. 18.7% for Opus 4.8
30.2%ARC-AGI-3 novel reasoning~4x GPT-5.6 Sol's 7.8%
79.2%SWE-bench Pro~10 points above Opus 4.8
$5 / $25per-million-token pricingunchanged since Opus 4.8

Benchmark results: a large jump over Opus 4.8

The clearest generational gain shows up on Frontier-Bench v0.1, Anthropic's agentic coding evaluation. Opus 5 scores 43.3%, more than double Opus 4.8's 18.7%, and ahead of both Claude Fable 5 (33.7%) and OpenAI's GPT-5.6 Sol, which different trackers place in the mid-30s (34.4% to 37.5% depending on the source).

Opus 5 more than doubles Opus 4.8 on agentic coding

Frontier-Bench v0.1 scores at launch, July 2026. Data from Digital Applied and Vellum. Higher is better.

Opus 5 (green) jumps to 43.3%, ahead of both Fable 5 and GPT-5.6 Sol, and more than double Opus 4.8’s 18.7% from earlier in 2026.

On ARC-AGI-3, a benchmark built around novel problem-solving that resists memorization, Opus 5 reaches 30.2% against GPT-5.6 Sol's 7.8%, roughly a fourfold gap, and against Opus 4.8's 1.5% according to Digital Applied's benchmark breakdown. On SWE-bench Pro, a harder successor to the widely cited SWE-bench Verified suite, Opus 5 posts 79.2%, about 10 points above Opus 4.8. And on GDPval-AA v2, which scores economically valuable knowledge work using an Elo-style rating, Opus 5 reaches 1,861 versus 1,736 for GPT-5.6 Sol, with Fable 5 at 1,747 and Opus 4.8 at 1,593.

Dario Amodei, co-founder and CEO of Anthropic, speaking on stage at TechCrunch Disrupt 2023
Anthropic co-founder and CEO Dario Amodei at a 2023 conference appearance. Anthropic has now shipped four Claude 5-generation flagships, Mythos 5, Fable 5, Sonnet 5, and Opus 5, in under two months. Photo by TechCrunch via Wikimedia Commons, CC BY 2.0.

Anthropic also reports Opus 5 leading on Zapier's AutomationBench and on OSWorld 2.0, a computer-use benchmark, where it scores 70.6% against Fable 5's 66.1%, Opus 4.8's 55.7%, and GPT-5.6 Sol's 62.6%. Not every comparison favors Opus 5: on OpenAI's DeepSWE v1.1 coding benchmark, GPT-5.6 Sol scores higher, a reminder that benchmark leadership is distributed across evaluations rather than concentrated in one model.

The price/performance frontier is compressing

The more interesting story sits in how Opus 5 is priced relative to what it scores. On CursorBench 3.2, a coding-agent benchmark, Opus 5 lands within 0.5 percentage points of Fable 5's peak score at maximum effort, according to both Vellum's analysis and Anthropic's own materials, while costing roughly half as much per task. Fable 5 is priced at $10/$50 per million tokens; Opus 5 held at $5/$25.

Opus 5 moves the price/performance frontier, not just the score

Blended cost per million tokens (weighted 3:1 input:output) versus Frontier-Bench v0.1 score. Toward the top-left is more capability for less money. Pricing and scores from provider pages and independent benchmark trackers, July 2026.

Opus 5 (green) sits at the same price as Opus 4.8 but far higher on the score axis, landing close to Fable 5’s performance at half its blended cost.

“

Opus 5 does not claim to beat Fable 5 outright. It claims to get close enough, for less money, that the gap stops mattering for most work.

Framing drawn from Anthropic's Opus 5 launch positioning

That is the practical takeaway for buyers: "near-parity at half the price" is not the same claim as "better than the frontier model." It means a large share of day-to-day agentic and coding work, the kind that previously required paying frontier-tier rates, can now run on a cheaper model without a meaningful quality drop. For teams running high request volumes, that difference compounds quickly across a monthly bill. It doesn't eliminate the case for Fable 5 or GPT-5.6 Sol on tasks where the last few points of capability matter, but it narrows the set of tasks where paying the premium is clearly justified.

An accelerating release cadence

Opus 5 is notable simply for how quickly it arrived. Mythos 5, Fable 5, Sonnet 5, and now Opus 5 have all shipped within roughly eight weeks, according to Digital Applied's account of the release sequence. That cadence mirrors a broader industry pattern in 2026: rather than a single yearly flagship release, labs are now shipping a family of differently sized and priced models in quick succession, letting a single generation cover multiple points on the cost/capability curve rather than one model trying to serve every use case.

Anthropic also published an automated behavioral audit score of 2.30 for Opus 5, described as the lowest, and by that measure best, of recent Claude models, alongside a note that certain misuse-related classifiers trigger roughly 85% less often than they did for Fable 5. Whether that reflects a durable safety trend or a one-model data point is not yet possible to say from a single release, and Anthropic's own reporting should be read as self-published rather than independently audited.

No single model wins every task

The pattern across every comparison here, Opus 5 ahead on Frontier-Bench and ARC-AGI-3, GPT-5.6 Sol ahead on DeepSWE and HealthBench Professional, Fable 5 marginally ahead on raw CursorBench ceiling, reinforces something that has held for most of 2026: no model wins across the board. Model choice increasingly depends on the specific task, and the right answer for a coding agent is not always the right answer for a knowledge-work summarization job or a long-document analysis.

That's part of why platforms that give access to multiple frontier models side by side, rather than locking a workflow to a single provider, have become more useful as the pace of releases has picked up. Metir gives users access to Opus 5 alongside GPT, Gemini, and Grok models in one place, so a task can be routed to whichever model actually performs best on it without juggling separate subscriptions.

Sources:

  • Anthropic news
  • Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing - MarkTechPost
  • Claude Opus 5 vs GPT-5.6 - CodersEra
  • Claude Opus 5 vs GPT-5.6 Sol comparison - LLM Stats
  • Claude Opus 5 Benchmarks Explained - Vellum
  • Claude Opus 5: Frontier Intelligence at Half the Price - Digital Applied
  • Claude Opus 5 Is Officially Here: 1M Context, a New 'xhigh' Mode - BitsMinds
  • Claude Opus 5 Launch: Benchmarks, Price, Fast Mode - ExplainX

Image credits

Header image: Anthropic co-founder and CEO Dario Amodei at 10 Downing Street, London, May 2023, by the UK Prime Minister's Office (Simon Walker / No 10 Downing Street) via Wikimedia Commons, licensed under CC BY 2.0. In-body photo of Dario Amodei speaking at TechCrunch Disrupt 2023 by TechCrunch via Wikimedia Commons, licensed under CC BY 2.0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Blog
  • Enterprise

Company

  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational