metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
ARC-AGI
François Chollet
AGI
LLMs
Benchmarks
Ndea

Chollet Says OpenAI Set Back AGI: ARC-AGI Explained

François Chollet says OpenAI set AGI back 5 to 10 years. We explain ARC-AGI-1, 2 and 3, the latest ARC Prize scores, the human gap and both sides of the debate.

Metir AI TeamOctober 5, 20268 min read
Chollet Says OpenAI Set Back AGI: ARC-AGI Explained

François Chollet, the creator of Keras and co-founder of the ARC Prize, is back in the headlines for a claim about OpenAI: that it "basically set back progress to AGI by five to 10 years." Yahoo Tech republished a TechRadar "quote of the day" on October 3, 2026, and the line spread quickly. The argument is that pouring money and attention into large language models came at the expense of other architectures. This explainer traces where the quote comes from, what ARC-AGI measures, where ARC-AGI-1, 2 and 3 scores stand against humans according to the ARC Prize, and what each side of the debate can point to.

OpenAI logoOpenAI
Anthropic logoAnthropic
Google logoGoogle
Labs whose models appear on the ARC Prize leaderboard discussed below.

Where the quote comes from

The remark is not new. The October 2026 write-up credits the Dwarkesh Podcast, and the original exchange is the June 11, 2024 episode with Chollet and Mike Knoop. There Chollet says OpenAI "basically set back progress towards AGI by quite a few years, probably like 5-10 years," and gives two mechanisms: the closing of frontier research publishing, and the hype around LLMs that, in his words, "sucked the oxygen out of the room." The TechRadar piece syndicated by Yahoo frames it as funding concentrated on LLMs "at the expense of other avenues or architectures."

Two cautions follow. First, the claim is a two-year-old opinion resurfaced, so it predates the reasoning-model results below. Second, none of the sources we fetched quantify the funding diverted or name which alternative architectures lost out, so "5 to 10 years" is a judgment, not a measurement.

“

Skill is not intelligence. That's the fundamental confusion that people run into.

François Chollet, Dwarkesh Podcast, June 2024

Chollet's definition of intelligence

The benchmark follows from a definition. In On the Measure of Intelligence (arXiv, November 2019), Chollet defines intelligence as skill-acquisition efficiency: how well a system turns limited experience and priors into new skills, rather than how well it performs on a task it was trained for. He argues that task-specific scores are misleading because skill is "heavily modulated by prior knowledge and experience," and he introduced the Abstraction and Reasoning Corpus (ARC) to test generalization with explicitly stated human-like priors.

ARC-AGI-1, 2 and 3 in plain terms

800ARC-AGI-1 tasksIntroduced 2019
120ARC-AGI-2 eval tasks per setEach solved by 2+ humans
0.51%Frontier AI at ARC-AGI-3 launchMarch 25, 2026
100%Human score on ARC-AGI-3Per ARC Prize
VersionFormatWhat it stresses
ARC-AGI-1 (2019)800 grid puzzles, about three examples eachInferring a rule from few examples
ARC-AGI-2 (2025)Harder grid tasks, calibrated with 400+ human testersSymbolic interpretation, compositional reasoning, contextual rules
ARC-AGI-3 (March 2026)Interactive turn-based games with no instructionsExploration, goal discovery, learning across levels

ARC Prize says ARC-AGI-1 stayed unsolved from 2019 to late 2024 despite a 50,000x scale-up of base LLM pretraining, until OpenAI's o3-preview scored 75% at low compute and 87% at higher compute in December 2024. ARC-AGI-2 was built to stress-test reasoning systems, with ARC Prize arguing that "log-linear scaling is insufficient" to beat it, and ARC-AGI-3, announced March 25, 2026, moved from static puzzles to agents that must discover the rules themselves.

Current scores and the human gap

The numbers below come from the ARC Prize's own leaderboard data (the leaderboard loads it from evaluations.json, which we read on October 5, 2026), covering semi-private sets.

Best lab-model score on ARC-AGI-1 and ARC-AGI-2

Semi-private evaluation sets, percent correct. Months without a new best for a benchmark are left blank and the line connects across them.

Source: ARC Prize leaderboard data. Human panel scores on the same sets: 98% (ARC-AGI-1) and 100% (ARC-AGI-2). Community refinement systems are excluded.

  • ARC-AGI-1: the top displayed score is 98.5%, shared by several models including Claude Fable 5 and GPT-6 Astra. The human panel entry is 98%.
  • ARC-AGI-2: the top score is 95.0% for GPT-6 Astra (Max) at about $1.12 per task, against a 100% human panel entry at $17 per task. Derived: that is a 5.0 point gap, with the best model about 15 times cheaper per task (17 divided by 1.12).
  • ARC-AGI-3: under the Standard harness, GPT-6 Astra (Max) scored 62.7% for $26,098 in run cost. Under the Provider Adapter harness, which preserves reasoning state and compacts long conversations, Astra (High) scored 99.9% for $18,817. Derived: against a 100% human line, the Standard gap is 37.3 points and the adapter gap is 0.1 points.

The ARC-AGI-3 harness gap is covered in more depth in our GPT-6 Astra analysis. ARC-AGI-3 scoring is relative action efficiency: ARC Prize says Astra used fewer actions than the human baseline (the median of roughly 500 participants) on 96.0% of levels. ARC Prize also states that "saturating the benchmark would not represent proof of achieving AGI," citing its bounded scope and deterministic mechanics.

François Chollet holding a microphone while speaking at an event
François Chollet speaking at an event; the photo dates from 2019 per Wikimedia Commons metadata and has been cropped. Photo: Kaicarver, CC BY-SA 4.0.

The case that scaling helped

Critics of Chollet's framing can cite the table above. The ARC Prize itself credits test-time adaptation methods pioneered by ARC Prize 2024 entrants and OpenAI for ending the five-year stall on ARC-AGI-1, and the ARC Prize 2025 analysis found the top verified commercial model at 37.6% on ARC-AGI-2 in late 2025. Our chart shows lab models moving from 37.6% to 95.0% on ARC-AGI-2 in under a year. A supporter of LLM scaling would say reasoning models, trained at large cost, are exactly what produced those gains, and that a benchmark designed to resist them has been nearly saturated.

The case that the benchmark itself shows the limits

Chollet's side has answers too. The ARC Prize 2025 results say the dominant theme was refinement loops, with the line "From an information theory perspective, refinement is intelligence," and the top Kaggle entry (24.03% on the ARC-AGI-2 private set) combined test-time training with small recursive-network components rather than a frontier LLM. ARC Prize also states that efficiency is part of the test, and reports cost per task alongside accuracy, so a high score at high compute is read differently from one at low compute. The harness dependence on ARC-AGI-3 (62.7% versus 99.9% for the same model) suggests scores depend heavily on scaffolding, a point both camps can use. Chollet has also said LLMs are useful, calling them good for automation while insisting that "automation is not the same as intelligence."

Ndea and the program synthesis bet

Chollet's alternative is concrete. In January 2025 he and Knoop launched Ndea, a lab focused on deep learning-guided program synthesis, which combines learned intuition with systems that generate explicit programs. The lab's stated view is that program synthesis today resembles deep learning in 2012, and Chollet wrote: "We believe we have a small but real chance of achieving a breakthrough." Ndea has not, in the sources we reviewed, published ARC leaderboard results, so the thesis remains untested at frontier scale.

What to watch

  • Whether ARC Prize reports more harness-independent results for ARC-AGI-3, since the gap between harnesses is larger than the gap to humans.
  • The ARC Prize 2026 competition tracks, which carry a prize pool of more than $2 million across ARC-AGI-3 and ARC-AGI-2 according to ARC Prize.
  • Any published results from alternative architectures such as Ndea's, which would be the direct test of the "oxygen" argument.

For teams choosing among models in the meantime, a model-agnostic workspace like Metir lets you compare them on your own tasks rather than relying on one benchmark.

Sources:

  • Yahoo Tech, Quote of the day by ARC Prize co-founder François Chollet (Oct 3, 2026)
  • TechRadar original of the same article
  • Dwarkesh Podcast, François Chollet and Mike Knoop (June 11, 2024)
  • On the Measure of Intelligence, arXiv 1911.01547
  • ARC Prize, ARC-AGI-1, ARC-AGI-2, ARC-AGI-3
  • ARC Prize, Announcing ARC-AGI-3
  • ARC Prize, GPT-6 Astra on ARC-AGI-3
  • ARC Prize 2025 Results and Analysis
  • ARC Prize leaderboard and evaluations data
  • The Decoder, Ndea launch (Jan 16, 2025)

Image credits

  • François Chollet speaking at an event: photo by Kaicarver, File:François Chollet.jpg on Wikimedia Commons, licensed CC BY-SA 4.0. Cropped from the original.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational