François Chollet, the creator of Keras and co-founder of the ARC Prize, is back in the headlines for a claim about OpenAI: that it "basically set back progress to AGI by five to 10 years." Yahoo Tech republished a TechRadar "quote of the day" on October 3, 2026, and the line spread quickly. The argument is that pouring money and attention into large language models came at the expense of other architectures. This explainer traces where the quote comes from, what ARC-AGI measures, where ARC-AGI-1, 2 and 3 scores stand against humans according to the ARC Prize, and what each side of the debate can point to.
AnthropicWhere the quote comes from
The remark is not new. The October 2026 write-up credits the Dwarkesh Podcast, and the original exchange is the June 11, 2024 episode with Chollet and Mike Knoop. There Chollet says OpenAI "basically set back progress towards AGI by quite a few years, probably like 5-10 years," and gives two mechanisms: the closing of frontier research publishing, and the hype around LLMs that, in his words, "sucked the oxygen out of the room." The TechRadar piece syndicated by Yahoo frames it as funding concentrated on LLMs "at the expense of other avenues or architectures."
Two cautions follow. First, the claim is a two-year-old opinion resurfaced, so it predates the reasoning-model results below. Second, none of the sources we fetched quantify the funding diverted or name which alternative architectures lost out, so "5 to 10 years" is a judgment, not a measurement.
Skill is not intelligence. That's the fundamental confusion that people run into.
François Chollet, Dwarkesh Podcast, June 2024
Chollet's definition of intelligence
The benchmark follows from a definition. In On the Measure of Intelligence (arXiv, November 2019), Chollet defines intelligence as skill-acquisition efficiency: how well a system turns limited experience and priors into new skills, rather than how well it performs on a task it was trained for. He argues that task-specific scores are misleading because skill is "heavily modulated by prior knowledge and experience," and he introduced the Abstraction and Reasoning Corpus (ARC) to test generalization with explicitly stated human-like priors.
ARC-AGI-1, 2 and 3 in plain terms
| Version | Format | What it stresses |
|---|---|---|
| ARC-AGI-1 (2019) | 800 grid puzzles, about three examples each | Inferring a rule from few examples |
| ARC-AGI-2 (2025) | Harder grid tasks, calibrated with 400+ human testers | Symbolic interpretation, compositional reasoning, contextual rules |
| ARC-AGI-3 (March 2026) | Interactive turn-based games with no instructions | Exploration, goal discovery, learning across levels |
ARC Prize says ARC-AGI-1 stayed unsolved from 2019 to late 2024 despite a 50,000x scale-up of base LLM pretraining, until OpenAI's o3-preview scored 75% at low compute and 87% at higher compute in December 2024. ARC-AGI-2 was built to stress-test reasoning systems, with ARC Prize arguing that "log-linear scaling is insufficient" to beat it, and ARC-AGI-3, announced March 25, 2026, moved from static puzzles to agents that must discover the rules themselves.
Current scores and the human gap
The numbers below come from the ARC Prize's own leaderboard data (the leaderboard loads it from evaluations.json, which we read on October 5, 2026), covering semi-private sets.
Best lab-model score on ARC-AGI-1 and ARC-AGI-2
Semi-private evaluation sets, percent correct. Months without a new best for a benchmark are left blank and the line connects across them.
Source: ARC Prize leaderboard data. Human panel scores on the same sets: 98% (ARC-AGI-1) and 100% (ARC-AGI-2). Community refinement systems are excluded.
- ARC-AGI-1: the top displayed score is 98.5%, shared by several models including Claude Fable 5 and GPT-6 Astra. The human panel entry is 98%.
- ARC-AGI-2: the top score is 95.0% for GPT-6 Astra (Max) at about $1.12 per task, against a 100% human panel entry at $17 per task. Derived: that is a 5.0 point gap, with the best model about 15 times cheaper per task (17 divided by 1.12).
- ARC-AGI-3: under the Standard harness, GPT-6 Astra (Max) scored 62.7% for $26,098 in run cost. Under the Provider Adapter harness, which preserves reasoning state and compacts long conversations, Astra (High) scored 99.9% for $18,817. Derived: against a 100% human line, the Standard gap is 37.3 points and the adapter gap is 0.1 points.
The ARC-AGI-3 harness gap is covered in more depth in our GPT-6 Astra analysis. ARC-AGI-3 scoring is relative action efficiency: ARC Prize says Astra used fewer actions than the human baseline (the median of roughly 500 participants) on 96.0% of levels. ARC Prize also states that "saturating the benchmark would not represent proof of achieving AGI," citing its bounded scope and deterministic mechanics.

The case that scaling helped
Critics of Chollet's framing can cite the table above. The ARC Prize itself credits test-time adaptation methods pioneered by ARC Prize 2024 entrants and OpenAI for ending the five-year stall on ARC-AGI-1, and the ARC Prize 2025 analysis found the top verified commercial model at 37.6% on ARC-AGI-2 in late 2025. Our chart shows lab models moving from 37.6% to 95.0% on ARC-AGI-2 in under a year. A supporter of LLM scaling would say reasoning models, trained at large cost, are exactly what produced those gains, and that a benchmark designed to resist them has been nearly saturated.
The case that the benchmark itself shows the limits
Chollet's side has answers too. The ARC Prize 2025 results say the dominant theme was refinement loops, with the line "From an information theory perspective, refinement is intelligence," and the top Kaggle entry (24.03% on the ARC-AGI-2 private set) combined test-time training with small recursive-network components rather than a frontier LLM. ARC Prize also states that efficiency is part of the test, and reports cost per task alongside accuracy, so a high score at high compute is read differently from one at low compute. The harness dependence on ARC-AGI-3 (62.7% versus 99.9% for the same model) suggests scores depend heavily on scaffolding, a point both camps can use. Chollet has also said LLMs are useful, calling them good for automation while insisting that "automation is not the same as intelligence."
Ndea and the program synthesis bet
Chollet's alternative is concrete. In January 2025 he and Knoop launched Ndea, a lab focused on deep learning-guided program synthesis, which combines learned intuition with systems that generate explicit programs. The lab's stated view is that program synthesis today resembles deep learning in 2012, and Chollet wrote: "We believe we have a small but real chance of achieving a breakthrough." Ndea has not, in the sources we reviewed, published ARC leaderboard results, so the thesis remains untested at frontier scale.
What to watch
- Whether ARC Prize reports more harness-independent results for ARC-AGI-3, since the gap between harnesses is larger than the gap to humans.
- The ARC Prize 2026 competition tracks, which carry a prize pool of more than $2 million across ARC-AGI-3 and ARC-AGI-2 according to ARC Prize.
- Any published results from alternative architectures such as Ndea's, which would be the direct test of the "oxygen" argument.
For teams choosing among models in the meantime, a model-agnostic workspace like Metir lets you compare them on your own tasks rather than relying on one benchmark.
Sources:
- Yahoo Tech, Quote of the day by ARC Prize co-founder François Chollet (Oct 3, 2026)
- TechRadar original of the same article
- Dwarkesh Podcast, François Chollet and Mike Knoop (June 11, 2024)
- On the Measure of Intelligence, arXiv 1911.01547
- ARC Prize, ARC-AGI-1, ARC-AGI-2, ARC-AGI-3
- ARC Prize, Announcing ARC-AGI-3
- ARC Prize, GPT-6 Astra on ARC-AGI-3
- ARC Prize 2025 Results and Analysis
- ARC Prize leaderboard and evaluations data
- The Decoder, Ndea launch (Jan 16, 2025)
Image credits
- François Chollet speaking at an event: photo by Kaicarver, File:François Chollet.jpg on Wikimedia Commons, licensed CC BY-SA 4.0. Cropped from the original.