metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
Anthropic
Claude
AI for Science
Physics
AI Agents
Research

Claude Computed a Nine-Loop Physics Amplitude: What It Means

Anthropic reported that Claude computed a nine-loop scattering amplitude in N=4 super Yang-Mills, beating the human record, for a few hundred dollars of compute. A neutral explainer.

Metir AI TeamSeptember 27, 20269 min read
Claude Computed a Nine-Loop Physics Amplitude: What It Means

On September 25, 2026, Anthropic published a research write-up reporting that two of its physicists, Liam Fitzpatrick and Siddharth Mishra-Sharma, used Claude to compute the six-particle scattering amplitude in planar N=4 super Yang-Mills at nine loops. The previous record, eight loops, was set by Lance Dixon in 2023. Dixon, now a professor at SLAC National Accelerator Laboratory, independently verified the new result. The striking numbers are not just the loop count. The computation ran nearly autonomously and cost on the order of a few hundred to a couple of thousand dollars. This piece explains what was actually computed, why it is hard, and what the result does and does not demonstrate.

9 loopsNew recordbeats the 8-loop mark from 2023
~$100Bootstrap-method compute96 CPUs for one week
$1,000 to $2,000Full computationacross two methods
4 to 6 hrsBetween human check-insnear-autonomous run

What is a scattering amplitude, in plain terms

Particle physics predicts what happens when particles collide by computing a "scattering amplitude," a quantity whose square gives the probability of a given outcome. These amplitudes are calculated as a series of increasingly precise corrections. The first term is the simplest picture of the interaction. Each additional "loop" adds a layer of quantum corrections, where particles briefly appear and vanish inside the process, and each layer makes the prediction more accurate.

The catch is that every extra loop is dramatically harder to compute than the one before. As Dixon put it, "Every loop order is harder than the previous one, computationally, even after finding lots of tricks." He had believed a direct nine-loop amplitude calculation was too hard to do directly. That is the wall the new result climbed over.

“

Every loop order is harder than the previous one, computationally, even after finding lots of tricks.

Lance Dixon, SLAC, who held the prior eight-loop record

The specific problem, and its honest limits

The theory here, planar N=4 super Yang-Mills, is not the physics of the real world. It is what amplitude researchers call a toy model: a mathematically clean, highly symmetric system used to develop and test calculation techniques that may later transfer to theories describing actual particles. Matt von Hippel, the former physicist and science writer who issued the public challenge to AI companies in August 2026, describes it exactly that way.

That framing matters for reading the result correctly. Claude did not discover new physics or invent a new method. It executed and extended known techniques, the amplitude "bootstrap" and a form-factor approach, to an order no one had reached, and it did so on a problem chosen precisely because it is a proving ground. The achievement is one of computation and endurance, not conceptual breakthrough. That is a meaningful thing to be, and it is a specific thing to be.

The ATLAS detector inside its experimental hall at CERN's Large Hadron Collider
The ATLAS experiment at CERN's Large Hadron Collider. Scattering-amplitude theory underpins the predictions tested at colliders like this, though the nine-loop computation itself is pure theory, run on CPUs rather than a collider. Photo by SimonWaldherr via Wikimedia Commons, CC BY-SA 4.0.

The two numbers that actually surprised people

Two figures make this more than a leaderboard update.

The first is autonomy. The researchers did not hand-hold the model through the calculation. They effectively told it to "keep working on this until I tell you to stop," checking in only every four to six hours. Dixon noted that this kind of work is very fragile, the sort where a single error can invalidate weeks of computation and require painstaking debugging. Sustaining a long, brittle, multi-stage calculation with light supervision is a different capability from answering a single hard question, and it is the one this result actually exercises.

The second is cost. The bootstrap calculation used about 96 CPUs for a week, for roughly $100 of compute. Running the full computation across both methods came to something like $1,000 to $2,000. Von Hippel's challenge deliberately restricted solvers to "computer resources an academic has access to," and the result stayed inside that limit. A task that might once have implied a national-lab allocation landed within the budget of an individual researcher.

From open challenge to verified frontier, in weeks

How the nine-loop result came together, and why the human check and the independent confirmation are part of the story rather than footnotes to it.

Aug 2026
The public challenge
Science writer Matt von Hippel challenges AI companies to solve a big open amplitudes problem using only compute an academic can access.
Late Aug to Sept
Near-autonomous run
Anthropic physicists run Claude on the hexagon amplitude with light steering: "keep working until I tell you to stop," checking in every 4 to 6 hours.
Result
Nine loops, ~$100 to $2,000
Claude reaches nine loops, past the 2023 eight-loop record, using the bootstrap and form-factor methods on about 96 CPUs for a week.
Verification
Human expert checks it
Lance Dixon (SLAC), who set the prior record, independently verifies the amplitude.
Within ~2 weeks
Independent confirmation
Song He’s group at the Chinese Academy of Sciences reaches partial nine-loop results using GPT-6 inside human-built frameworks.

The theory is a symmetric toy model used to test methods; the techniques were known, extended to a new loop order rather than newly invented.

Why the independent confirmation matters

Extraordinary computational claims deserve independent checks, and this one got two. Dixon, the human expert who held the prior record, verified the amplitude. Separately, Song He's group at the Chinese Academy of Sciences reached partial nine-loop results within about two weeks of Claude's completion, using GPT-6 assistance inside human-built frameworks.

The convergence is the interesting part. Two teams, on different continents, using models from different labs and different degrees of automation, arrived at the same frontier at nearly the same time. That pattern argues the result reflects a genuine shift in what these systems can do on this class of problem, rather than a single vendor's carefully staged demo. It also quietly makes a point about the tools: the capability was not unique to one model.

What it signals for AI in research

The useful way to read this is narrow and concrete. Frontier models are becoming able to carry out long, technical, error-intolerant computations with limited human steering, at a cost low enough that the bottleneck shifts from compute budget to the researcher's ability to frame the problem and verify the answer. That is a real change in how certain kinds of theoretical work could get done, and it lands squarely in domains, like amplitude computation, that reward relentless, exact manipulation over intuition.

It is also worth keeping the caveats in the same frame. This was a toy model, chosen for its tractability. The methods were known. A human expert still did the verification, and verification, not generation, remains the guarantee of correctness. None of that diminishes the result. It locates it: a demonstration that AI systems can now do a specific, demanding kind of scientific labor, checked by people, not a claim that they are doing science unsupervised.

“

The bottleneck shifts from compute budget to the researcher's ability to frame the problem and verify the answer.

For teams thinking about AI in their own technical work, the portable lesson is that the value increasingly sits in the workflow around the model, framing, orchestration, and verification, more than in any single model choice. Two different frontier models reached the same result here. A model-agnostic approach, the kind Metir AI takes by routing across providers, fits that reality: the frontier is a moving target reached by more than one system, and the durable advantage is in how you put it to work rather than which logo is on it.

Sources:

  • Claude computes a nine-loop amplitude in N=4 super-Yang-Mills | Anthropic
  • Anthropic Says Claude Computed a Nine-Loop Particle Physics Amplitude | Unite.AI
  • Anthropic's Claude Computes Nine-Loop Yang-Mills Amplitude | AI Weekly
  • Claude beats Dixon's record and computes a nine-loop physics amplitude | Pasquale Pillitteri
  • Anthropic's Claude solves nine-loop amplitude challenge in theoretical physics | Crypto Briefing

Image credits

Header and in-body image: the ATLAS detector at CERN's Large Hadron Collider, by SimonWaldherr, via Wikimedia Commons, CC BY-SA 4.0. Illustrative of experimental particle physics; the nine-loop amplitude is a theoretical computation, not an experiment run at this facility.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational