metir
metir
Download on App StoreGet it on Google PlayF1 FantasyLoginSign Up
Back to Blog
AI Models
Benchmarks
Intelligence Index
AI Strategy
Model Routing

The AI Frontier Compressed in July 2026: Six Labs Above 50

Four frontier model launches in eight days pushed six AI labs above 50 on the Artificial Analysis Intelligence Index, up from two in early June, while price per task fell sharply. Here is what the compression actually means for buyers.

Metir AI TeamJuly 21, 20269 min read
The AI Frontier Compressed in July 2026: Six Labs Above 50

For most of the last two years, the frontier of AI capability was a two-lab story. In early June 2026, only Anthropic and OpenAI had a model scoring above 50 on the Artificial Analysis Intelligence Index, the composite benchmark that has become one of the most watched scoreboards in the industry. By mid-July, six labs did. Four separate frontier launches landed in an eight-day window, and the gap between the best model available and the sixth-best model shrank to nine points on a hundred-point-style scale.

That is the real story of July 2026: not any single release, but how quickly the field closed around the top of the leaderboard, and how much cheaper it got to buy intelligence that used to be exclusive to one or two providers.

The compression, by the numbers

6 labsnow field a model above 50up from 2 in early June
4 launchesin just 8 days
60top Intelligence Index scoreClaude Fable 5
~1/3 the costKimi K3 & GPT-5.6 Sol vs. Fable 5, per task
Anthropic logoAnthropic
OpenAI logoOpenAI
xAI logoxAI
Meta logoMeta
Moonshot AI logoMoonshot AI
Z.ai logoZ.ai
The six labs now fielding a model above 50 on the Artificial Analysis Intelligence Index.

Anthropic's Claude Fable 5 has held the top spot on the Intelligence Index since June 9, at a score of 60. What changed over the following weeks is everything underneath it. OpenAI shipped GPT-5.6 as three simultaneous tiers, Sol, Terra and Luna, scoring 59, 55 and 51 respectively. Moonshot AI's Kimi K3 landed at 57. SpaceXAI's Grok 4.5 scored 54. Meta's Muse Spark 1.1 and Z AI's GLM-5.2 both sit at 51.

Six labs, one narrow band: every model now above 50

Artificial Analysis Intelligence Index v4.1 (10 evaluations), mid-July 2026. Data from Artificial Analysis. Higher is better.

AnthropicOpenAIMoonshot AISpaceXAIMetaZ AI

Nine points separate #1 (Claude Fable 5, 60) from the six models tied at #6 (51). In early June only two labs cleared 50 at all.

Laid out on a chart, the picture is less a ladder than a cluster. Eight models from six labs now occupy a nine-point band at the top of the index, using a benchmark, version 4.1, that Artificial Analysis has deliberately weighted toward longer-horizon agentic tasks rather than static question answering. That detail matters: the compression is happening on a test that is getting harder to game, not an easier one.

Four launches, eight days

The timing is what makes this a genuine inflection rather than a slow drift. SpaceXAI's Grok 4.5 launched July 8 at a score of 54. OpenAI's three-tier GPT-5.6 family followed July 9 to 10, adding Sol, Terra and Luna to the board in a single announcement alongside Meta's Muse Spark 1.1. Moonshot AI's Kimi K3 closed the window on July 16 at 57, a score Artificial Analysis characterized as comparable to Claude Opus 4.8 and GPT-5.5, though still behind Fable 5 and Sol.

“

Nine points now separate the best model in the world from the sixth-best. In early June, only two labs cleared 50 at all.

Reading of Artificial Analysis Intelligence Index v4.1 data, July 2026

None of these four launches individually reset the top of the leaderboard. Fable 5 still leads. What they did collectively was eliminate the idea that frontier-class intelligence was scarce. A buyer evaluating "the best available model" in early June had a two-lab shortlist. By late July, that shortlist is six labs deep, and the models near the bottom of that list are separated from the top by a margin that is easy to close with prompting, tool access or simply picking a different task.

Why the gap at the top stopped mattering

A useful way to read a compressing frontier is to separate two questions that used to have the same answer: which model is smartest, and which model should you actually use. When the field was two labs wide and the gap between them was large, those questions collapsed into one, buyers picked the top scorer and the decision was largely settled.

That logic breaks down once six labs cluster within a single-digit range on a benchmark that already carries real measurement noise. A three-point difference on the Intelligence Index is not obviously visible in day-to-day output quality, but it is very visible on the invoice. Kimi K3, three points behind Fable 5, and GPT-5.6 Sol, one point behind it, each deliver their results at roughly a third of Fable 5's reported cost per representative agentic task, according to Artificial Analysis's own task-based cost accounting. Grok 4.5, Muse Spark 1.1 and GPT-5.6 Luna go further, each landing at or below the per-task price that GLM-5.2 had set only a week earlier.

The real story is price

Similar intelligence, very different prices

Cost per task on Artificial Analysis's representative agentic evaluation, mid-July 2026. Data from Artificial Analysis. Only models with a confirmed published cost-per-task figure are shown.

AnthropicOpenAIMoonshot AISpaceXAIMetaZ AI

Kimi K3 and GPT-5.6 Sol land within three and one points of the top score at roughly a third of Claude Fable 5's cost per task. Grok 4.5, Muse Spark 1.1 and GPT-5.6 Luna all match or beat the price GLM-5.2 set a week earlier.

Plot intelligence against cost per task and the leaderboard reorganizes itself again. The top-left corner, high capability at low cost, is no longer empty or occupied by a single outlier. It is getting crowded, and it is getting crowded fast enough that a price set one week becomes the baseline the next model has to beat.

This is the pattern economists would recognize from any technology that reaches a capability plateau before it reaches a cost floor: competition stops being primarily about "who is smartest" and starts being about "who delivers a given quality bar most cheaply." The Intelligence Index compression is the visible symptom. The price data is the actual mechanism.

Close-up of wiring and status lights inside a high-performance computing data center
Detail of wiring inside the high-performance computing data center at the U.S. National Renewable Energy Laboratory's Energy Systems Integration Facility. The falling cost-per-task figures behind this compression trace back to the compute infrastructure that trains and serves these models. Photo by Dennis Schroeder / NREL, U.S. Department of Energy.

What buyers should take from this

The practical implication is that committing to a single provider now carries a real opportunity cost, in both directions. Standardizing on the top scorer means paying a premium that a three-point quality gap rarely justifies for most workloads. Standardizing on the cheapest model in the cluster means leaving capability on the table for the handful of tasks where the top few points genuinely matter, complex reasoning chains, ambiguous multi-step agent work, anything close to the edge of what any model can currently do.

The alternative that a compressed, price-divergent frontier rewards is routing: sending each task to whichever model currently clears the bar for that job at the lowest cost, and revisiting that assignment as the leaderboard keeps moving, which it now does on a roughly weekly cadence. That is a harder discipline to maintain by hand than it sounds, since it means tracking pricing and benchmark movement across six labs rather than one. It is also the specific problem model-agnostic platforms like Metir AI are built to solve, giving a single workspace access to Claude, GPT-5.6, Grok, Kimi and other frontier models so the routing decision does not have to be a manual subscription-by-subscription exercise.

The honest limits of a single number

A neutral read requires caveats. The Intelligence Index is a composite of ten evaluations with specific weightings, and any composite benchmark compresses genuinely different capability profiles into one number; a model that trails on the index can still be the better choice for a task its underlying evaluations do not emphasize. Reported cost-per-task figures are also representative averages on Artificial Analysis's chosen agentic evaluation, not a guarantee of what any specific real-world workload will cost. And a benchmark cluster this tight is inherently sensitive to measurement noise, a rerun with a slightly different task mix could reorder several of these models without any of them actually changing.

The takeaway

Six labs above 50, four launches in eight days, and per-task prices falling faster than intelligence scores are rising: July 2026 is the month the AI frontier stopped being a two-horse race and became a genuine market. The top score still belongs to Claude Fable 5, and that is worth noting. But the more consequential number is nine, the point spread that now separates the best model in the world from six credible alternatives, several of them a third of the price.


Work across every frontier model in one place

When six labs cluster near the top of the Intelligence Index and prices diverge this much, the winning move is rarely committing to one provider. With Metir AI you get unified access to Claude, GPT-5.6, Grok, Kimi and other leading models in a single workspace, so you can route each task to whichever model clears the bar for the best price this week, not whichever one you subscribed to last quarter. Try Metir AI free and let the right model handle every job.

Sources:

  • Four frontier launches in eight days: six labs now field a model above 50 on the Artificial Analysis Intelligence Index | Artificial Analysis
  • Kimi K3 achieves 3rd in the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5 | Artificial Analysis
  • Artificial Analysis Intelligence Index v4.1 | Artificial Analysis
  • GPT-5.6 has landed | Artificial Analysis

Image credits

Header image: server room of BalticServers, rack-mounted servers under blue LED status lighting, via Wikimedia Commons, licensed under CC BY-SA 3.0. In-body photograph: detail of wiring inside the high-performance computing data center at NREL's Energy Systems Integration Facility, taken March 14, 2013, by Dennis Schroeder / NREL, U.S. Department of Energy, via Wikimedia Commons, public domain.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Blog
  • Enterprise

Company

  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational