For most of 2025 and early 2026 the story of AI model pricing had one direction: down. Each new model generation cost less per token than the last, each lab undercut the others, and the working assumption across the industry was that intelligence was becoming relentlessly cheaper. In August 2026 that assumption stopped being simple. In the same handful of weeks, OpenAI cut the price of a high-volume model by roughly 80 percent, while DeepSeek, the lab that built its reputation on being the cheapest option, raised prices on parts of its V4 family by more than tenfold.
Both moves are real, both are rational, and they point in opposite directions. Understanding why is a short course in how AI inference is actually priced, and why "AI keeps getting cheaper" was always a simplification of something more mechanical happening underneath.
What each side actually did
On July 30, OpenAI cut the input price of its GPT-5.6 Luna model by about 80 percent, from roughly $1.00 to $0.20 per million tokens, and made Luna the free default in ChatGPT with unlimited chats. Anthropic, separately, has positioned its Claude Opus 5 at roughly half the price of its higher-end Fable 5. These are the moves of established labs using price to win and keep high-volume users, especially against lower-cost competitors gaining traction among cost-conscious buyers.
DeepSeek went the other way. Alongside its V4-Pro launch it announced that V4 family prices would rise, effective August 16 at 16:00 UTC. On the V4-Pro model, peak-hour output pricing moves from a flat $0.87 to $3.96 per million tokens, and reporting described some V4 workloads rising by more than ten times. The V4 Flash tier rose about 93 percent, from $0.14 to $0.27 per million. DeepSeek also introduced peak and off-peak billing, with off-peak rates at about half the peak price.
The same weeks, opposite directions
Percent change in list price per 1M tokens. A frontier lab cutting to win high-volume users, and the low-cost disruptor raising prices as demand outruns capacity, are happening at once.
DeepSeek also introduced peak and off-peak billing, with off-peak rates at roughly half the peak price, a capacity-management tool more than a simple rate card.
The reversal is stark precisely because of who is doing it. DeepSeek spent the previous year as the industry's price hawk, running promotional discounts to pressure Western labs. Raising prices now, even substantially, is a reversal of that strategy, and the company framed it around demand straining its capacity.
The mechanism: capacity, not cost
The key to the apparent contradiction is that per-token list price is not a measure of how cheap the model is to build. It is a lever that sits on top of two different things: how much a lab wants to attract usage, and how much serving capacity it has to go around. Those two forces can push in opposite directions at the same lab, let alone across different ones.
Per-token price is not a measure of how cheap the model is. It is a lever sitting on top of demand strategy and serving capacity, which can push opposite ways at once.
On inference pricing
A lab with ample capacity and a strategic reason to grow, like OpenAI wanting a billion people on a free default model, can cut prices to drive volume, betting on scale and downstream monetization. A lab whose demand has outrun the GPUs it can actually serve, like DeepSeek describing capacity strain, has the opposite problem: too much usage at the current price, degrading service for everyone. Raising prices, and adding peak and off-peak tiers, is how you ration scarce capacity and push flexible workloads to quieter hours. It is the same logic that prices electricity higher at peak demand. The introduction of time-of-day pricing is the tell: this is capacity management, not a simple rate change.
Seen this way, the two moves are not a contradiction at all. They are two labs in different positions on the same underlying constraint, which is compute. The deflation narrative was really a story about that constraint loosening as capacity grew faster than demand. When demand for a specific model outruns its capacity, the arrow flips, even for the cheapest provider in the market.

What it means for anyone paying for tokens
For teams building on these APIs, the practical takeaways are concrete. First, a headline price is a snapshot, not a trend. DeepSeek was the cheapest option and remains inexpensive relative to Western frontier models even after the increase, but a workload sized around the promotional price just got materially more expensive overnight. Second, time-of-day pricing changes how you should schedule non-urgent work: batch jobs that can wait for off-peak hours now cost meaningfully less than the same work run at peak. Third, and most important, the direction of any single provider's pricing is no longer predictable, because it depends on that provider's private capacity situation as much as on the march of efficiency.
Even after DeepSeek's increase, the spread across the market is enormous. Frontier output pricing ranges from a few dollars per million tokens at the low end to $30 for OpenAI's fastest tier and $50 for Anthropic's Fable 5. A model that is a fiftieth of the price of another on paper may or may not be cheaper for a given task once you account for how many tokens each consumes to finish the work. List price is the start of the cost question, not the end of it.
The strategic read
This is the practical case for not hardwiring an application to a single model or provider. When the cheapest option can raise prices tenfold on two weeks' notice, and a frontier lab can cut its flagship by 80 percent in the same window, the ability to move workloads to wherever the current price-performance is best becomes a real operational advantage rather than a theoretical nicety. A team locked into one provider absorbs that provider's pricing swings in full. A team that can route each workload to the model that fits it, and re-route when prices move, turns volatility from a risk into an opportunity.
Anthropic
DeepSeek
Moonshot AIThat flexibility is the design principle behind model-agnostic platforms such as Metir AI, where the model is a swappable component chosen per task and per budget rather than a permanent commitment. The point is not that any one model is best. It is that the "best" answer now changes week to week and provider to provider, and the pricing inversion of August 2026 is a clean demonstration of why. The era of assuming AI only ever gets cheaper, in a straight line, from every provider at once, is over. What replaces it is a market where price is set by capacity, moves in both directions, and rewards whoever kept their options open.
Sources:
- DeepSeek raises some V4 prices by more than 10x as AI demand strains capacity | InfoWorld
- DeepSeek increases prices for AI services by multiple times | Fortune
- DeepSeek's New V4 Pro AI Model Quadruples Prices but Remains Far Cheaper Than Western AI | Android Headlines
- DeepSeek Price Hike 2026: What API Teams Must Do Now | Eden AI
- OpenAI and Anthropic lower prices as Chinese competitors gain users | Yahoo Finance
Image credits
Header image: rows of servers in a data center, representing the inference capacity that determines AI pricing. By BalticServers.com via Wikimedia Commons, licensed under CC BY-SA 3.0. In-body image: an aisle of server racks. By Christopher Bowns via Wikimedia Commons, licensed under CC BY-SA 2.0.
