metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
Grok 4.7
xAI
AI Models
Benchmarks
Coding AI
Artificial Analysis

Grok 4.7: Where xAI's Benchmarks and Independent Testing Agree, and Where They Don't

xAI's Grok 4.7 claims a 38.0% Terminal-Bench score; Artificial Analysis measured 25.76%. Elsewhere the two sources match almost exactly. A neutral look at what the vendor says, what independent testing says, and where the two diverge.

Metir AI TeamSeptember 21, 202611 min read
Grok 4.7: Where xAI's Benchmarks and Independent Testing Agree, and Where They Don't

On September 21, 2026, xAI released Grok 4.7, describing it as "our most capable model for coding and knowledge work." The release page's subtitle makes the efficiency pitch explicit: "twice as fast, at half the price of comparable models." Grok 4.7 is served at the same $2 / $6 per million token price as Grok 4.6, unchanged.

Most model releases are reported from one side only, the vendor's own benchmark table. Grok 4.7 is unusual because the independent benchmarking firm Artificial Analysis (AA) published its own measurements of the same model on the same day, and on one widely used benchmark, the two sources land roughly 12 points apart. On several other measures they match almost exactly. That combination, real agreement in some places and a real, unexplained gap in another, is the more interesting story than either source read alone.

+2.1 ptsAA Intelligence Index gain44.31 to 46.45 vs Grok 4.6
38.0% vs 25.76%Terminal-Bench 4.0xAI claim vs AA measured
500KContext windowUnchanged from Grok 4.6
Sep 21, 2026Release dateSame price as Grok 4.6

What xAI changed

xAI's own description of the work behind Grok 4.7 is specific: "Grok 4.7 uses a new, larger base model compared to Grok 4.6. It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The model is better at verifying its own work and managing longer context. We also trained Grok 4.7 to natively understand the Grok Bot harness, making it better at conversational tasks and general knowledge work." Reasoning effort in the API comes in four tiers, low, medium, high, and xhigh, and it cannot be turned off entirely.

What xAI's own benchmarks show

xAI's headline comparison table sets Grok 4.7 at xhigh effort against Grok 4.6 at high effort, alongside GPT-5.6 Sol Max and Claude Fable 5.1 Max.

BenchmarkGrok 4.7Grok 4.6GPT-5.6 Sol MaxClaude Fable 5.1 Max
CursorBench 4.046.3%40.4%41.7%51.8%
DeepSWE v1.171.0%*65.2%72.7%70.0%
EEBench64.0%53.0%39.4%56.4%
AA Briefcase v1.11657154614871678
Terminal-Bench 4.038.0%20.3%37.3%57.9%
Harvey Legal Agent Benchmark19.6%15.8%2.5%6.7%
HealthBench Professional56.7%48.5%60.5%62.1%
Input / output per 1M tokens$2 / $6$4 / $20$10 / $50

*xAI reports DeepSWE v1.1 for Grok 4.7 at high effort rather than xhigh; all other Grok 4.7 figures in this table are at xhigh.

On this table Grok 4.7 does not lead everywhere. Anthropic's Fable 5.1 Max scores higher on CursorBench, DeepSWE, HealthBench, and Terminal-Bench, and OpenAI's GPT-5.6 Sol Max edges it on DeepSWE and HealthBench. Grok 4.7 does lead clearly on EEBench and on the Harvey legal agent benchmark, and it undercuts both rivals substantially on price. On GDPval, a measure of real-world professional task quality, xAI's own reported Elo has Fable 5.1 Max at 1735, Grok 4.7 at 1695, Grok 4.6 at 1605, and GPT-6 Astra at 1542.

xAI logoxAI
Anthropic logoAnthropic
OpenAI logoOpenAI
Grok 4.7 is being compared against Anthropic's and OpenAI's current flagship coding models.

Where the independent numbers agree with xAI's

Artificial Analysis published its own measurement of Grok 4.7 the same day, and two of its figures are worth flagging specifically because they match xAI's quoted numbers almost exactly: AA's own AA Briefcase v1.1 score of 1657.2 and its GDPval-AA Elo of 1695.21 correspond to the 1657 and 1695 xAI cites in its own table. That is not a coincidence, xAI is quoting AA's own measurements for those two benchmarks. It is a point in xAI's favor: where the vendor cites an independent source directly, the numbers check out.

A real gain over Grok 4.6, independently measured

Artificial Analysis Intelligence Index v4.3.2 at the Grok 4.7 launch. This is a genuine independent gain over Grok 4.6, and it is nearly identical at the high and xhigh effort tiers. Data from Artificial Analysis. Higher is better.

Grok 4.7 gains roughly +2.1 points over Grok 4.6 (44.31 to 46.45). AA has not yet published GPQA, LiveCodeBench, AIME-25, MMLU-Pro or SWE-bench Verified scores for Grok 4.7.

AA's own composite, the AA Intelligence Index v4.3.2, puts Grok 4.7 at 46.45 (xhigh) and 46.33 (high), both clear gains over Grok 4.6's 44.31 and well ahead of GPT-5.6 Sol's 33.47 on the same index. Note the gap between Grok 4.7's own two effort tiers: dropping from xhigh to high costs almost nothing on the index, about 0.12 points, for roughly 18% fewer output tokens. That is a genuine, independently confirmed improvement over Grok 4.6.

“

Where xAI cites Artificial Analysis directly, the numbers match almost exactly. On Terminal-Bench, where it does not, they diverge by about 12 points.

Reading Grok 4.7's two benchmark tables side by side

Where they diverge

Terminal-Bench 4.0 is the clearest disagreement. xAI reports 38.0% for Grok 4.7. AA measured 25.76% at xhigh effort and 24.75% at high, roughly 12 to 13 points below xAI's figure. That gap is unusual because it is not present on Grok 4.6: xAI's reported 20.3% and AA's measured 21.2% are within a point of each other on the prior model.

Where xAI's numbers and Artificial Analysis's numbers disagree

Terminal-Bench 4.0 pass rate, xAI's published figure vs Artificial Analysis's independent measurement. The two sources agree closely on Grok 4.6 and diverge by about 12 points on Grok 4.7. Data from xAI and Artificial Analysis.

AA has not published an explanation for the gap. A different test harness or scaffold is the likely cause; treat the 38.0% figure as unverified until AA or another party reproduces it.

Neither xAI nor AA has published an explanation for the Grok 4.7 gap. A different test harness or agent scaffold, which affects how a model is allowed to interact with the terminal, is a plausible cause on a benchmark like this, but that is inference, not a sourced explanation. Until AA or a third party reproduces xAI's number, the 38.0% figure should be treated as unverified, while the 25.76% AA figure stands as the independently checkable one.

Two other AA measurements sit in tension with xAI's own description of the release. AA-LCR, AA's long-context reasoning score, fell from 0.8033 on Grok 4.6 (high) to 0.7667 on Grok 4.7 (xhigh), a notable move against xAI's stated goal of a model "better at managing longer context." And on AA-Omniscience, Grok 4.7's hallucination rate rose to 29.34% (accuracy 47.45%) from Grok 4.6's 24.0%. Neither figure is catastrophic on its own, and composite scores can move for reasons unrelated to a single stated goal, but both cut against the specific claims in xAI's release notes rather than confirming them.

It is also worth being clear about what AA has not yet measured. As of this writing, AA has not published Grok 4.7 scores for GPQA, LiveCodeBench, AIME-25, MMLU-Pro, or SWE-bench Verified, and it has not yet published its own launch article on the model. The independent picture here is one day old and still filling in.

The verbosity caveat on "half the price"

xAI's price-per-token for Grok 4.7 output is unchanged from Grok 4.6, and both are well below GPT-5.6 Sol Max's $20 and Fable 5.1 Max's $50 per million output tokens. But price per token is not the same as cost per finished task, and AA's data adds a real qualification here.

A verbosity cost the headline price does not show

Average output tokens Artificial Analysis measured per task. At the default xhigh effort, roughly 58,544 of the 80,561 tokens are reasoning tokens. Cost per task is price times tokens, so a more verbose model can cost more to run even at an unchanged per-token price. Data from Artificial Analysis.

Across AA's full evaluation suite, Grok 4.7 generated about 240M tokens against a median of 92M for the models AA has tested, a score of 4 out of 4 on verbosity.

AA scores Grok 4.7 4 out of 4 on verbosity, noting that its evaluation suite generated about 240 million tokens running Grok 4.7, against a median of 92 million tokens across the models it has tested. At the task level, AA measured an average of 80,561 output tokens per task at xhigh effort (58,544 of them reasoning tokens), against 65,901 at high effort. Cost per task is price per token multiplied by tokens per task, so a model that is several times cheaper per token can still cost more per finished task than the sticker price implies if it also generates several times more tokens to get there. None of this makes the per-token price claim false. It does mean the "half the price" headline describes list price, not necessarily the bill for a comparable piece of finished work, and a fair comparison needs both numbers.

A second pricing detail worth noting: Grok 4.7's $2 / $6 rate applies below 200,000 prompt tokens. At or above that threshold, pricing steps up to $4 / $1 (cached) / $12 per million tokens, applied to the entire request rather than just the portion over the threshold.

Specs, availability, and what did not change

Dense server cabling and rack-mounted compute nodes inside an AI and high-performance computing data center
Rack-mounted compute nodes at the Datarmor Center, an Ifremer AI and HPC facility in Brest, France, shown to illustrate the class of server hardware behind frontier model releases. It is not xAI's own infrastructure. Photo by Gregory Rocher (IFREMER Centre Bretagne) via Wikimedia Commons, CC BY 4.0.

The context window is 500,000 tokens, unchanged from Grok 4.6. Modalities are text and image in, text out. Tools include function calling, web search, X search, and code execution. Knowledge cutoff is May 2026. Grok 4.7 (API model ID "grok-4.7") is available through the xAI API, Grok Build, Cursor, the OpenRouter, Vercel, and Cloudflare gateways, and Google Cloud Vertex AI; it is not available on Azure AI Foundry. A separate "Grok 4.7 Fast" variant runs on faster infrastructure at double the token rate, but it is offered only inside Cursor and Grok Build, not as a public API tier.

On safety, xAI says Grok 4.7 "was built with an entirely new safeguard stack," reporting 3.3% of risky prompts allowed on HackerBench, which is xAI's own benchmark rather than an independently run one, and a LatchBio biosafety score of 62.4%.

The read-through

Grok 4.7 is a real, independently confirmed gain in general intelligence over Grok 4.6, and where xAI cites Artificial Analysis's own numbers directly, those numbers hold up exactly. That combination gives the rest of xAI's table more credibility than a vendor benchmark table typically earns on its own. But the Terminal-Bench gap, the long-context score moving the wrong direction, the rising hallucination rate, and the verbosity behind the price claim are all real qualifications that a reader should weigh alongside the headline numbers, not against them. None of this makes Grok 4.7 a bad model; it makes it a model whose story is more layered than either "trust the vendor" or "trust the independent number" alone would suggest.

Grok 4.7 is available in Metir today across all three reasoning tiers, fast, reasoning, and high, at the same price point as Grok 4.6, and existing conversations pinned to Grok 4.6 move across automatically. The practical lesson from a release where vendor and independent numbers disagree this visibly is not which side to believe, it is that being able to switch models per task without rebuilding a workflow is a real hedge against exactly this kind of uncertainty. That is the portability Metir AI is built around: route a task to whichever model's measured behavior actually fits it, and change that choice as the independent evidence catches up.

Sources:

  • Grok 4.7 | xAI News
  • Grok 4.7 | Artificial Analysis
  • Grok 4.7 | xAI Developer Documentation

Image credits

Header and in-body images: the Datarmor Center, Ifremer's high-performance computing and artificial intelligence facility in Brest, France, shown to illustrate the class of NVIDIA GPU-based server infrastructure that trains and serves frontier AI models. It is not xAI's own facility. Photos by Gregory Rocher (IFREMER Centre Bretagne) via Wikimedia Commons, licensed under CC BY 4.0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational