metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
US-China AI Gap
LiveBench
DeepSeek V4.1 Flash
AI Benchmarks 2026
Open Weight AI Models

US-China AI Gap Narrows to 3% on LiveBench After DeepSeek V4.1

Bloomberg Intelligence says the US-China AI gap narrowed to about 3% on LiveBench after DeepSeek V4.1 Flash scored 81.1 versus 83.4. What the number means.

Metir AI TeamOctober 5, 20267 min read
US-China AI Gap Narrows to 3% on LiveBench After DeepSeek V4.1

The US-China AI gap on LiveBench has narrowed to about 3%, according to a note from Bloomberg Intelligence senior analyst Robert Lea reported on October 4 and 5, 2026. After DeepSeek released V4.1 Flash in September, the model scored 81.1 on the benchmark, against 83.4 for the leading Anthropic model, a spread of 2.3 points. Bloomberg Intelligence describes it as the smallest US-China top-model gap it has tracked.

The headline is simple, but a single benchmark number deserves careful reading. This post lays out what was reported, what LiveBench actually measures, and why a 3% gap on one leaderboard is a useful signal rather than a verdict.

81.1DeepSeek V4.1 Flash on LiveBenchMax Effort setting
83.4Top Anthropic model on LiveBenchClaude Fable 5.1, Max Effort
~3%US lead, October 2026Down from about 9% in May and about 15% earlier in 2026
77.3 vs 66.1Agentic coding sub-scoreDeepSeek vs Anthropic, Oct 4 snapshot
DeepSeek logoDeepSeek
Anthropic logoAnthropic
The two models at the center of the Bloomberg Intelligence comparison.

What Bloomberg Intelligence reported

According to coverage of Lea's note by Implicator and AI Weekly, the gap between the best US model and the best Chinese model on LiveBench was roughly 15% earlier in 2026, about 9% in May, and about 3% after V4.1 Flash. On the October 4 leaderboard snapshot, DeepSeek V4.1 Flash at Max Effort scored 81.1 and Anthropic's Claude Fable 5.1 at Max Effort scored 83.4. The 2.3-point difference is about 2.8% of the higher score, which rounds to the 3% in the headline.

The US lead on LiveBench, three snapshots

Approximate gap between the top US model and the top Chinese model, as a share of the US score. Source: Bloomberg Intelligence, as reported by Implicator and AI Weekly. Figures are rounded.

The three points on that chart are the only gap figures in the reporting, so the line between them is a trend, not a measured series. We could not read the Bloomberg article directly (it sits behind a paywall), so the figures here come from secondary coverage that cites the note.

Lea attributes the catch-up to Chinese labs improving their technical capabilities and optimizing models for domestic hardware, and says this raises questions about how effective US chip export controls are. The coverage also notes caveats: the benchmark results do not establish national AI leadership, and they do not show whether export restrictions are working.

DeepSeek V4.1 Flash: the release behind the number

DeepSeek released V4.1 Flash on September 10, 2026, as an open-weight model under an MIT license. In our earlier deep dive on its Causal Encoder-Decoder design, we covered the architecture, in which input tokens activate 8 billion parameters and output tokens activate 16 billion within a 552 billion-parameter backbone. DeepSeek also said it would route V4-Pro requests to V4.1 Flash at Flash pricing.

“

A 3% gap on one leaderboard is a signal worth tracking, not a ruling on who leads in AI.

Metir analysis

What LiveBench measures, and why it is hard to game

LiveBench was introduced in a paper accepted as an ICLR 2025 Spotlight, led by Colin White with a large group of co-authors. According to the paper, it tests models across math, coding, reasoning, language, instruction following and data analysis. Two design choices set it apart:

  • Contamination limiting. Test questions are drawn from recently released sources such as new math competitions, arXiv papers, news articles and datasets, so a model is less likely to have seen them in training. The paper says questions are added and updated monthly.
  • Objective scoring. Answers are graded automatically against ground-truth values rather than by human or LLM judges, which removes judge bias.

Those properties make LiveBench a better snapshot of fresh capability than static tests that leak into training data. They do not make it a complete measure.

Why one benchmark cannot settle the question

View over the western part of Hangzhou, the Chinese city where DeepSeek is headquartered
A view over the western part of Hangzhou, the city where DeepSeek is headquartered. The photo shows the city, not DeepSeek's offices. Photo by CatOnMars, CC BY-SA 4.0.

Several limits apply when reading the 3% figure:

  • It is one leaderboard. LiveBench covers a defined set of task categories. Performance on long-horizon enterprise work, safety behavior, multilingual use or tool reliability can look different elsewhere.
  • It compares two entries at a settings level. Both scores are Max Effort runs. Models rank differently at lower reasoning budgets, and the comparison says nothing about cost or latency per answer.
  • The top is narrow, the field is not. AI Weekly notes only three of the top 15 LiveBench entries are Chinese, so the convergence is concentrated at the frontier rather than across the board.
  • Sub-scores diverge from the average. On agentic coding, DeepSeek's 77.3 sat ahead of Anthropic's 66.1 on the same snapshot, so a single average hides where each model is stronger.
  • Leaderboards move. The October 4 snapshot is a point in time, and scores shift as new models and question sets arrive.
  • Rounding matters. A gap that is "about 3%" depends on whether you divide by the higher score, and a 2.3-point spread on a 100-point scale can look larger or smaller depending on the framing.

The economics behind the story

Benchmark parity is one input. Lea's note also looks at the business side: according to AI Weekly, China hosts more than 1,100 large language models, and Lea does not expect profitability in the sector before 2030. He is quoted as saying that putting the sector on a sustainable footing "will require a cooling of competitive pressures, an industry shakeout, and a more rational approach to pricing." In other words, closing a capability gap and building a profitable AI business are different problems. For more on how Chinese open models are showing up in Western company workflows, see our look at Chinese open models and US enterprise adoption.

What to take from it

  • Treat the 3% as a direction of travel: 15%, then 9%, then 3% across the reported snapshots.
  • Check sub-benchmarks that match your own workload, since averages blur real differences.
  • Expect the gap to keep moving as new releases land on both sides.
  • Keep your own evaluations. The practical way to test a claim like this is to run the same prompts through several models. Tools like Metir, which put models from multiple labs side by side, make that comparison cheap.

Whether the gap keeps shrinking, reopens, or just stays hard to measure, a benchmark lead of a few points is small relative to the variance you will see on your own tasks.

Sources:

  • US Lead in AI Over China Narrows After DeepSeek Gains, BI Says - Bloomberg (paywalled, headline and framing confirmed via search)
  • DeepSeek Cuts US AI Benchmark Lead to About 3% - Implicator
  • DeepSeek V4.1 Flash narrows US-China AI gap to 3% on LiveBench - AI Weekly
  • DeepSeek Narrows AI Gap With US to Just 3 Percent, Bloomberg Says - Startup Fortune
  • LiveBench: A Challenging, Contamination-Limited LLM Benchmark - arXiv
  • DeepSeek API changelog: V4.1-Flash release

Image credits

Header image: Hangzhou skyline across West Lake by Windmemories via Wikimedia Commons, licensed under CC BY-SA 4.0. In-body image: Skyline of the western part of Hangzhou by CatOnMars via Wikimedia Commons, licensed under CC BY-SA 4.0. Both photos show the city where DeepSeek is headquartered and do not depict the company's offices or infrastructure.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational