metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
Epoch AI
AI Economics
Inference Costs
AI Models
Industry Analysis

AI Is Getting 13x Cheaper Every Year: What Epoch AI's Numbers Mean

Epoch AI finds the cost of a fixed level of AI performance has fallen about 47% per quarter since 2023, roughly 13x a year. Here is what drives it and what it does not mean.

Metir AI TeamSeptember 23, 20269 min read
AI Is Getting 13x Cheaper Every Year: What Epoch AI's Numbers Mean

On September 23, 2026, the independent research group Epoch AI published a finding that puts a number on something practitioners have felt for two years but rarely measured directly: the cost of reaching a fixed level of AI performance has been falling by about 47% per quarter since 2023, or roughly 13x per year. That rate, Epoch AI says, is faster than any other general-purpose technology it has tracked, including DNA sequencing, computing power itself, and the historical decline in the cost of electricity.

OpenAI logoOpenAI
Anthropic logoAnthropic
Google logoGoogle
DeepSeek logoDeepSeek
Qwen logoQwen
NVIDIA logoNVIDIA
The decline spans every major lab: closed frontier providers cutting API prices and open-weight labs pricing aggressively against them, both riding the same underlying curve.

What Epoch AI actually measured

The study, titled "The plunging price of thought," is not a claim about any single model getting cheaper. It is a claim about the cost of a fixed capability getting cheaper as better, smaller, or more efficient models arrive to deliver that same capability. Epoch AI built the estimate from five benchmarks spanning mathematics, hard sciences, and games of skill: AIME (OTIS Mock), Chess Puzzles, FrontierMath Tiers 1-3, GPQA Diamond, and Mystery Game Puzzles. For each one, the researchers tracked what it cost, in API dollars, to reach a given score across every model released since 2023, then measured how that cost fell over time.

The average across all five is 47% per quarter. The rate is not uniform: math-heavy tasks like FrontierMath fall faster, at 50 to 52% per quarter (16 to 19x a year), while game-based puzzles fall slower, at 39 to 43% per quarter (7 to 10x a year). Two concrete examples from the same report make the abstraction tangible. OpenAI's o3, released January 31, 2025, was the first model to score above 75% on GPQA Diamond, a graduate-level science exam, at an average cost of about 30 cents per question. Under 18 months later, GPT-5.6 Luna matched that score for four hundredths of a cent, a 725-fold drop. On FrontierMath Tiers 1-3, o3 was also the first model to exceed 25% accuracy, at roughly 55 cents per attempt; this summer GPT-5.6 Luna reached the same mark for about 0.15 cents, a 377-fold drop in the same window.

47%Cost decline per quarterFixed AI performance level, since 2023
13xCost decline per yearEpoch AI's headline annualized rate
377xFrontierMath T1-3 cost drop~$0.55 to ~$0.0015 per attempt
725xGPQA Diamond cost drop~$0.30 to ~$0.0004 per question

A 47%-per-quarter decline, plotted on a log scale

Illustrative curve applying Epoch AI's measured average rate (cost of a fixed level of performance falling about 47% per quarter since 2023, roughly 13x per year) quarter over quarter, indexed to 100 in Q1 2023. This models the stated rate; it is not Epoch's own point-by-point cost series.

Log-scale y-axis. A straight downward line on this scale is a constant percentage decline per quarter, the signature of exponential improvement.

Two curves, not one

The easiest mistake to make with this number is to read it as "AI is getting cheaper," full stop. It is more precise, and more useful, to say the cost of any given capability is getting cheaper, while the cost of the capability nobody has reached yet keeps rising. Epoch AI's own separate research on frontier training costs finds the largest training runs have grown roughly 2 to 3x a year for the better part of a decade, with the biggest 2026 runs reportedly costing several hundred million dollars and industry watchers expecting the largest runs to cross a billion dollars within the next year or two. Those are two different curves moving in opposite directions from two different starting points, and both are true at once.

“

The frontier keeps getting more expensive to push forward. Everything behind the frontier keeps getting cheaper to reach. Both statements describe the same industry.

Reading Epoch AI's cost-of-performance data alongside its frontier-training-cost data

Two curves, not one: the frontier and any fixed point behind it

Illustrative indexed curves, both set to 100 in 2023. The cost of reaching a fixed capability falls at Epoch AI's measured rate of about 13x per year. The cost of the largest frontier training runs is shown rising at roughly 2.5x per year, within the 2-3x annual growth Epoch AI has separately measured. The two lines share an index, not a cost unit, so this shows shape, not a common price.

Same starting index, opposite directions: what got you 75% on a benchmark last year gets radically cheaper, while the very top of the frontier keeps costing more to reach.

This is a familiar shape from other technologies: it resembles a Jevons-style pattern where cheaper unit costs invite far more total consumption rather than less total spending, and it explains why total industry spending on compute and training can rise even as the price of any specific task falls. A lab spending more to push the absolute frontier is not in tension with last year's frontier becoming commodity-priced. Both are the normal behavior of a technology whose capability curve is being pushed from the top while the cost curve for everything already achieved collapses from below.

Rows of server racks with perforated cabinet doors in a commercial data center aisle
A data center server-rack aisle. Illustrative of the compute infrastructure whose amortized capital and energy cost underlies every dollar-per-task figure in this piece; it does not depict any specific lab's facility. Photo by PiDatacenters, via Wikimedia Commons, CC BY-SA 4.0.

The mechanisms behind the decline

No single cause explains a 47%-per-quarter drop; it is the sum of several trends compounding together. Algorithmic efficiency is one: Epoch AI's separate work on language model progress estimates that the compute needed to reach a fixed level of pretraining performance has been halving roughly every seven to eight months, worth about 3x a year, independent of any change in hardware. Better post-training and distillation is another: smaller, cheaper models increasingly match the benchmark scores that only much larger models could hit a year or two earlier, because labs have gotten better at compressing capability learned by a large teacher model into a small, fast, cheap-to-serve student model. Sparse mixture-of-experts architectures push in the same direction by activating only a small fraction of a model's total parameters per token, cutting inference compute without cutting capability.

Hardware price-performance and competitive pressure round it out. Each new generation of accelerators lowers the cost of a unit of inference, and 2026 has been a year of unusually intense price competition: OpenAI cut GPT-6 Sol and Luna API prices by roughly half compared with the prior GPT-5.6 generation, describing the cut as permanent rather than promotional, while Chinese open-weight labs have continued releasing frontier-adjacent models at prices well below Western closed models, forcing the whole market to reprice. None of these forces is sufficient alone. Together, they are what a 13x annual decline looks like from the inside.

Close-up of NVIDIA H100 GPU accelerator modules with NVLink bridges installed in a server chassis
NVIDIA H100 accelerator modules. Illustrative hardware photo, not tied to any lab's specific infrastructure; each generation of accelerators is one input among several (alongside algorithmic efficiency, distillation, and competition) into the cost decline described here. Photo by geekerwan, via Wikimedia Commons, CC BY 3.0.

What it means for budgeting, moats, and lock-in

For an enterprise budgeting AI spend, the practical takeaway is that a cost estimate has a shelf life measured in months, not years. A task that costs a dollar to automate today may cost a fraction of a cent to automate by the time a multi-year contract around it expires, and a build-versus-buy decision anchored to this quarter's model pricing can look badly outdated within two or three quarters. That argues for treating model choice as a decision to revisit routinely rather than a platform bet to make once.

It also complicates the idea of a durable moat built on model access alone. Epoch AI's own data on FrontierMath and GPQA suggests the premium a state-of-the-art model can charge decays quickly once competitors close in, with the fastest cost declines occurring right at the state of the art, exactly where a first mover would hope to hold pricing power longest. When the cheapest model capable of a given task changes every few months, and the previous quarter's leader is routinely undercut by 300x or more within eighteen months, committing a product's entire cost structure to one vendor is a bet that this quarter's price will still be the market's price next year. A model-agnostic approach, routing each task to whichever provider is currently strongest and cheapest for it, captures the savings this curve keeps generating instead of locking a business into one point on it; that portability is the premise behind platforms like Metir, which keep work movable across providers rather than tied to a single one.

The caveats worth keeping

Epoch AI's own report is careful about the limits of the measurement, and they are worth repeating rather than glossing over. The estimate is built from five specific benchmarks; it is a reasonable proxy for reasoning-and-knowledge-style tasks but not a universal price index for every kind of AI work, and real-world agentic or multimodal tasks may not fall at the same rate. Measuring "cost of a fixed level of performance" also depends on which benchmark defines that performance level, and different benchmarks in the same study show meaningfully different rates, from 39% to 52% per quarter. And the aggregate obscures unevenness within a single benchmark's history: Epoch AI finds costs fall fastest, around 66% per quarter, right after a new state-of-the-art model debuts, then slow to about 32% per quarter roughly two years on, a pattern consistent with early competitive undercutting rather than a perfectly smooth exponential. None of that undermines the headline finding, but it does mean a specific number from this study should be read as a well-evidenced estimate of a real trend, not a precise constant to extrapolate blindly for every workload.

The honest summary is that both halves of this story are true simultaneously. AI capability that already exists keeps getting radically cheaper, at a pace with no real precedent in the history of general-purpose technology. The frontier that has not yet been reached keeps getting more expensive to push toward. Enterprises, and the vendors selling to them, are making decisions inside both curves at once, and confusing one for the other is the most common way to get the economics of this industry wrong.

Sources:

  • The plunging price of thought | Epoch AI
  • Epoch AI on X: cost of a fixed AI performance level down ~47%/quarter since 2023
  • The Price of Intelligence is Falling Rapidly | Marginal Revolution
  • Algorithmic progress in language models | Epoch AI
  • How much does it cost to train frontier AI models? | Epoch AI
  • The rising costs of training frontier AI models | arXiv
  • OpenAI releases GPT-6 Sol and Luna models, slashing API costs 50% or more | VentureBeat

Image credits

Hero image: a commercial data center server-rack aisle, photographed by PiDatacenters, via Wikimedia Commons, licensed under CC BY-SA 4.0. It illustrates the compute infrastructure behind AI inference pricing and does not depict any specific lab's facility. In-body photograph: a close-up of NVIDIA H100 GPU accelerator modules by geekerwan, via Wikimedia Commons, licensed under CC BY 3.0. It is an illustrative hardware photo and is not tied to any specific lab's infrastructure.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational