metir
metir
Docs
Download on App StoreGet it on Google PlayLog inSign up
Back to Blog
ARC-AGI-3
ARC Prize
Kaggle
Open-Weight Models
AI Agents

ARC-AGI-3 Kaggle Scores Surge as Harnesses Take Over

ARC-AGI-3 Kaggle scores have jumped past 55% using small open-weight models in harnesses. Here is what the benchmark tests, the prize terms and the caveats.

Metir AI TeamOctober 5, 20267 min read
ARC-AGI-3 Kaggle Scores Surge as Harnesses Take Over

The ARC-AGI-3 Kaggle leaderboard has moved fast. As we read it in early October 2026, Tufa Labs sits first at 55.89 and Yi-Chia Chen second at 48.59, in a competition where entries must run offline on small models under fixed compute. The rise has been driven less by bigger models than by the harness wrapped around them, which is why the ARC-AGI-3 Kaggle results matter for anyone building AI agents. This post explains what the benchmark measures, what ARC Prize officially states about the 2026 competition, and why the leaderboard needs careful reading.

Qwen logoQwen
Gemma logoGemma
Open-weight model families named in the ARC Prize Milestone #1 results.

What we could and could not confirm

Reports circulating on Reddit say the top Kaggle score climbed from roughly 7% to roughly 56% in about 30 days. We could confirm the current top two scores, which appear on the Kaggle leaderboard (the page renders client-side, so we read it through a search snapshot and the figures may have moved since). We could not confirm the "7% to 56% in 30 days" trajectory from any official source, so we do not treat it as established. ARC Prize also posted on X that Yi-Chia Chen had taken first place with a 28.34% high score at an earlier point, which is consistent with steep progress but does not by itself fix the timeline.

55.89Top Kaggle scoreTufa Labs
48.59Second placeYi-Chia Chen
$850KARC-AGI-3 track prize pool
Nov 2, 2026Submission deadline

ARC-AGI-3 explained: interactive games, not static puzzles

ARC-AGI-1 and ARC-AGI-2 gave a model a few example grids and asked it to infer a rule. ARC-AGI-3, launched on March 25, 2026, is an interactive reasoning benchmark for AI agents: turn-based games with no instructions, where the agent has to explore, work out what the goal is, and carry what it learns across levels. For the broader lineage of the benchmark and Chollet's definition of intelligence, see our explainer on Chollet, OpenAI and ARC-AGI.

Because scoring depends on how efficiently an agent plays, a system cannot simply memorize answers. It has to act, observe the result and update, which is exactly where scaffolding matters.

ARC Prize 2026: format, prizes and deadlines

According to the ARC Prize 2026 page, the program offers $2M across three tracks (ARC-AGI-3, ARC-AGI-2 and a Paper Prize). The ARC-AGI-3 track carries $850K:

PrizeAmount
Grand prize$700K for the first eligible agent that scores 100%
Top score award$75K split among five places ($40K, $15K, $10K, $5K, $5K)
Milestone prizes$75K across two open-source deadlines (June 30 and September 30)

Key dates: submissions close on November 2, 2026, the paper deadline is November 8, and results are due on December 4. Rules require that solutions be open source under a permissive license such as CC0 or MIT-0, that there is no internet access during Kaggle evaluation, and that API-based systems such as GPT or Claude cannot be used. Those constraints are the reason small open-weight models dominate: nothing larger can be called.

“

Participants must open source their solutions before receiving official private evaluation scores.

ARC Prize 2026 rules

What the winning harnesses look like

The clearest public window into the approaches is the Milestone #1 results post, covering the period through June 30. Its three winners:

  • Tufa Labs ("The Duck"), $25K: a locally run Qwen 3.6 27B (FP8) that writes and runs Python in a live REPL, treating each game as an interactive coding problem. It combines rendered images, ASCII grids and segmentation tools, and uses what the team calls "infinite play via eviction" to manage the context window.
  • Reki, $10K: a vision-language policy on a locally served Gemma-4-31B that returns JSON actions each step, with a reflection memory refreshed about every 10 steps and click heuristics to avoid wasted actions.
  • Md Boktiar Mahbub Murad ("forge"), $2.5K: a similar image-plus-JSON design on Gemma-4-31B with a generator and arbiter that choose among candidate actions.
Portrait of François Chollet, co-founder of the ARC Prize
François Chollet, co-founder of the ARC Prize, who created the ARC benchmark. The photo shows Chollet only; it is not an image of the competition or its entries. Photo: Ramosset, CC BY-SA 4.0.

None of these is a frontier model. The common thread is engineering around a mid-sized model: structured perception, memory management, action filtering and code execution. The ARC Prize post does not publish per-submission scores or compute limits, so we cannot say how much each component contributes.

Why harness gains matter

Harness effects are not unique to the Kaggle track. In ARC Prize's own frontier testing, the same model can score very differently depending on the harness (see our GPT-6 Astra analysis). The Kaggle results point the same way: how a model perceives the game state, remembers earlier levels and chooses actions can matter as much as raw capability. For builders, that suggests agent quality is partly a systems problem, and that a model-agnostic workflow, such as the one in Metir, makes it easier to test the same task across models and scaffolds.

Caveats: public leaderboard, overfitting and what 56% means

  • Public versus private. The leaderboard is calculated on roughly half of the test data, and final results will use the other half, so rankings can change. That is the standard guard against tuning to what you can see.
  • Overfitting risk. Heavy iteration against a public score can encode quirks of the visible games. A jump in a short window, whatever its exact size, is a reason to wait for the private scores.
  • Not AGI. ARC Prize has stated that saturating the benchmark would not prove AGI, citing its bounded scope and deterministic mechanics. A score near 56 is also well short of the 100% needed for the $700K grand prize.
  • Single-source numbers. We read the leaderboard through a snapshot; check the live page before quoting figures.

What to watch before November 2

The final weeks will show whether the public gains survive on the held-out half, whether the September 30 Milestone #2 winners publish their notebooks, and whether other teams converge on the REPL-plus-vision pattern. Because every prize requires open-sourcing, the winning harnesses will be inspectable, which is arguably the competition's most durable output.

Sources:

  • Kaggle: ARC Prize 2026 - ARC-AGI-3 leaderboard
  • ARC Prize 2026 overview
  • ARC Prize 2026: ARC-AGI-3 track
  • ARC Prize 2026: ARC-AGI-3 Milestone Prize #1 results
  • ARC-AGI-3 benchmark page
  • ARC Prize, GPT-6 Astra on ARC-AGI-3
  • ARC Prize on X, 28.34% high score post

Image credits

  • François Chollet portrait (hero and in body): photo by Ramosset, File:Fchollet.jpg on Wikimedia Commons, licensed CC BY-SA 4.0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of service
  • Privacy policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational