On August 12, 2026, SpaceXAI released Grok 4.6, an incremental update to the Grok 4.5 flagship it shipped a month earlier. In xAI's own framing, Grok 4.6 "builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work." The independent benchmarking firm Artificial Analysis summarized it more bluntly: the model "returns SpaceXAI to the intelligence frontier and leads on cost efficiency."
Both statements are worth unpacking, because Grok 4.6 is less a story about a new intelligence ceiling and more a story about what happens to a market when several labs reach roughly the same ceiling at once. The interesting numbers here are not the ones at the top of the leaderboard. They are the ones next to the price and the token counts.
Where Grok 4.6 lands
On the Artificial Analysis Intelligence Index, a composite of hard evaluations spanning reasoning, coding, and agentic tasks, Grok 4.6 scores 61. That ties it with OpenAI's GPT-5.6 Sol and places it just behind Anthropic's two leaders, Claude Fable 5 at 62 and Claude Opus 5 at 63. It is a five-point gain over Grok 4.5, which scored 56 on the same index.
Grok 4.6 rejoins a tightly packed frontier
Artificial Analysis Intelligence Index at the Grok 4.6 launch. Grok 4.6 ties GPT-5.6 Sol and sits about two points off the top. Data from Artificial Analysis. Higher is better.
Grok 4.6 gains +5 points over Grok 4.5 (56 to 61). The four leading models now sit within roughly two points of each other.
The headline is not that Grok 4.6 is the smartest model, because it is not. The headline is how little separates the top four. A two-point spread across four models from three different labs is a statement about the shape of the frontier in late 2026: raw general intelligence, as these composite benchmarks measure it, is converging. When the leaders cluster this tightly, the score stops being the thing that distinguishes them, and attention shifts to everything else, cost, speed, token efficiency, and reliability on long tasks.
xAI
AnthropicThat said, a composite score hides real per-benchmark variation, and it is worth being precise about it. On the DeepSWE v1.1 software-engineering test, Grok 4.6 scored 65.9%, a large jump from Grok 4.5's 54% but still behind GPT-5.6 Sol's stronger configurations. On CursorBench v3.2 it scored 69.9%, a hair behind Claude Fable 5 and ahead of GPT-5.6 Sol. On the GDPval-AA v2 measure of real-world agentic knowledge work it posted an Elo of 1753, behind only Claude Opus 5. In other words, Grok 4.6 is not uniformly first or last anywhere. It trades places with its peers depending on the task, which is exactly what a converged frontier looks like up close.
The actual differentiator: cost and token efficiency
If intelligence is roughly tied, price is where Grok 4.6 draws a sharp line. It is priced at $2 per million input tokens and $6 per million output tokens, with cache hits at $0.50. For comparison, Artificial Analysis lists Claude Opus 5 at $5 / $25 and GPT-5.6 Sol at $5 / $30. On output, the tokens that dominate cost for generative and agentic work, Grok 4.6 is roughly four to five times cheaper than the models it ties or narrowly trails.
Frontier-tier intelligence at a fraction of the price
Published output token price, dollars per million tokens. Grok 4.6 matches the frontier on intelligence while pricing output four to five times below its nearest peers. Input prices follow the same pattern ($2 for Grok 4.6 against $5 for the others).
Cache hits fall to $0.50 per million tokens. Lower unit price compounds on long agentic runs that consume many tokens.
Sticker price is only half of it, and arguably the less important half for agents. The cost of an agentic task is the price per token multiplied by the number of tokens consumed, and long-running agents consume enormous numbers of tokens as they loop through tools, retries, and multi-step plans. This is where Grok 4.6's second efficiency claim matters. Artificial Analysis reports that on long-horizon tasks, Grok 4.6 reached comparable results in about 53 turns and roughly 0.5 billion input tokens on average, versus Claude Opus 5's roughly 103 turns and 2.0 billion input tokens on the same work.
The same long task, fewer steps and tokens
On long-horizon agentic work, Artificial Analysis reports Grok 4.6 reaching comparable results with markedly fewer turns and input tokens than Claude Opus 5. Lower consumption is where cost efficiency actually shows up on real agent runs.
Bars are scaled to each metric's larger value. Figures are averages reported by Artificial Analysis, not fixed per-task guarantees.
Stack those two effects together and the gap compounds. A model that is several times cheaper per token and also uses a fraction of the turns and tokens to finish the job is not marginally cheaper to run as an agent; it can be an order of magnitude cheaper on exactly the long, tool-heavy workloads that are becoming the dominant way AI gets used. That, not the Intelligence Index, is the substance behind the "leads on cost efficiency" line. It is also why xAI framed the release around long-running agents specifically: that is where its advantage is largest.
When the top four models sit within two points on intelligence, the leaderboard stops being the story. Cost per finished task becomes the story.
On a converged AI frontier
Why an incremental release still matters
Grok 4.6 is a point release, and it should be read as one. Grok 4.5, shipped a month earlier, was the large leap: a new architecture trained on real developer-usage data, a jump of more than fifteen points onto the frontier. Grok 4.6 refines that foundation rather than replacing it, with xAI describing added emphasis on self-testing and verification behavior, sturdier first attempts at visual and interactive work, and better stamina across long multi-step tasks, along with safeguards recalibrated to the expanded capabilities.
Those are agent-reliability improvements more than intelligence improvements, and that distinction reflects where the field's effort is going. As base intelligence converges, the marginal work moves to making models dependable over long horizons: catching their own mistakes, staying on task across dozens of steps, and not wasting tokens. A one-month cadence between a major release and a targeted follow-up is itself a signal of how fast this layer is now iterating.

Availability and the fine print
Grok 4.6 is available through xAI's API at console.x.ai, inside Cursor and Grok Build, through partners including OpenRouter, Vercel, and Cloudflare, and in the Grok iOS and Android apps, with a promotional period of doubled included usage in the first week. There is a faster variant priced at twice the standard rate, the usual latency-for-cost trade.
A few caveats keep the picture honest. Benchmark scores, however carefully constructed, are not the same as production performance on messy real tasks, and the efficiency figures are averages rather than guarantees; a given workload may consume more or fewer turns than the reported mean. The Intelligence Index reflects one firm's methodology and weighting, and different composites would reorder these closely spaced models. And "ties the frontier at a fifth of the price" is a claim about the specific configurations Artificial Analysis tested, not a universal statement that Grok 4.6 equals every competitor on every task. It clearly does not lead everywhere, as the per-benchmark spread shows.
The read-through
The larger significance of Grok 4.6 is what it says about the market rather than the model. Frontier-grade intelligence is becoming something several labs can deliver at once, which turns it from a differentiator into a baseline. Competition then moves to the dimensions that are still spread out: price, token and turn efficiency, latency, and reliability on long agentic runs. Grok 4.6 is a clean example of a lab choosing to compete on those dimensions rather than trying to claim the single highest score, and it continues the price pressure Grok 4.5 started a month earlier. For buyers, that pressure is good news; for the labs, it compresses the premium that a top score used to command.
For anyone building on these models, the practical consequence is that model selection is no longer a one-time decision to back the smartest system. When four models sit within two points on intelligence but diverge by four to five times on cost and by large margins on agentic efficiency, the right choice becomes task-dependent: a cheap, token-efficient model like Grok 4.6 for long, tool-heavy agent runs where cost compounds, and a different model where a specific benchmark or capability justifies the premium. Capturing that requires keeping the model layer swappable rather than hardwired, which is the design principle behind a model-agnostic platform like Metir AI: route each job to the model whose intelligence, price, and efficiency profile fits it, and re-route as the frontier keeps reshuffling. On a frontier this crowded and this fast-moving, the ability to switch is worth more than any single model's score.
Sources:
- Grok 4.6 | xAI News
- Grok 4.6 returns SpaceXAI to the intelligence frontier and leads on cost efficiency | Artificial Analysis
- SpaceXAI's Grok 4.6 matches OpenAI's best model and undercuts it on price | The Decoder
- SpaceXAI's Grok 4.6 model beats OpenAI's GPT-5.6 Sol in several AI benchmarks | Neowin
Image credits
Header and in-body image: the Berzelius AI supercomputer, an Atos and NVIDIA DGX-based system operated in Linkoping, Sweden, shown to illustrate the class of GPU supercomputers that train and serve frontier AI models. It is not xAI's own Colossus cluster. By Thor Balkhed / Erik And30 via Wikimedia Commons, licensed under CC BY-SA 4.0.