metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
Anthropic
Claude
AI Research
Self-Improving AI
AI Safety

Claude Now Leads 26% of Anthropic's R&D: Reading the Number Carefully

Anthropic disclosed that Claude leads 26% of its own research and engineering work, up from under 1% in February, with about 30,000 agents running at any time. What 'leads' means, what the monitoring shows, and where the metric ends.

Metir AI TeamSeptember 18, 20268 min read
Claude Now Leads 26% of Anthropic's R&D: Reading the Number Carefully

On September 17, 2026, Anthropic published internal metrics it had not shared before: Claude now leads about 26 percent of the company's research and engineering work, up from under 1 percent in February. Roughly 30,000 agents run at any given time, and the company said more than 90 percent of its research and development happens with Claude as either a collaborator or a lead. It is a striking disclosure, and it is also easy to misread in either direction. The honest version is more interesting than either the hype or the dismissal.

The single most important word in the announcement is "leads," so start there.

26%R&D work led by ClaudeAugust 2026
<1%The same figurein February 2026
~30,000Agents runningat any given time
0.002%Actions blockedabout 1 in 47,000

What "leads" actually means

Anthropic drew a specific line between two modes. "Collaborates" means a human and Claude work on a task together, with the human driving. "Leads" means Claude can carry most of a task end-to-end from a high-level prompt, while a human supervisor reviews the result. The 26 percent figure is the "leads" share, and it climbed from essentially nothing to a quarter of the work in six months.

From under 1 percent to 26 percent in six months

Share of Anthropic's research and engineering work that Claude leads end-to-end under human oversight, as the company disclosed it.

"Leads" is a narrower claim than "does." It counts tasks a model can carry from a prompt to a result with a human reviewing the output, not autonomous research with no person in the loop.

That is a genuine jump, and it is worth not inflating. "Leads" is not "does autonomously with no human." There is a supervisor reviewing outputs, and the tasks that qualify are the ones a model can take from a prompt to a reviewable result, which skews toward well-scoped engineering and research work rather than open-ended judgment calls. Read precisely, the number says that a large and fast-growing slice of Anthropic's routine technical work can now be initiated and carried by a model under human review. That is a big claim on its own; it does not need to be stretched into "AI is running the lab."

Rows of servers with blue indicator lights in a data center
Thirty thousand concurrent agents is a compute story as much as a model story. The figure describes work initiated and carried by models across large fleets of machines, under human review. Photo: BalticServers.com, via Wikimedia Commons, CC BY-SA 3.0.

The recursive-improvement framing, handled with care

The reason this disclosure travels is that it touches the idea of AI systems helping build their own successors. If a company's frontier model is doing a rising share of the research that produces the next model, the loop that people describe when they talk about recursive self-improvement is, in a limited and supervised form, visible in a real org chart.

Anthropic logoAnthropic
Claude logoClaude
Anthropic is reporting on its own tools improving its own pipeline, which is why the self-reported nature of the metric matters.

It is worth keeping two things straight. First, "leads a quarter of R&D tasks under supervision" is a long way from "designs its own successor unsupervised," and Anthropic's own definitions make that clear. Second, the metric is self-reported, and the company that publishes it is also the one that benefits from the narrative that its models are exceptionally capable. That does not make the number wrong. It means the right posture is to treat it as a credible internal signal, disclosed with unusual specificity, rather than as an audited external fact.

The monitoring is the part worth studying

The most substantive detail is not the automation share. It is the oversight wrapped around it. Anthropic said every agent action passes through an online monitor before it runs, and 100 percent of executed actions are reviewed after the fact by a separate offline monitor. Across roughly 1 billion research and engineering decisions in August, about 1 in 47,000, or 0.002 percent, was blocked.

The oversight layer around the agents

The headline number is the automation. The part that governs it is the monitoring on either side of every action.

Agent proposes an actionOne of about 30,000 research and engineering agents running at any given time generates a step
↓
Online monitor screens itEvery action passes through an automated check before it is allowed to run
↓
Action executesApproved actions run; roughly 1 in 47,000 was blocked at this gate
↓
Offline monitor reviews it100 percent of executed actions are reviewed after the fact by a separate monitor

A 0.002 percent block rate can read as either reassuring or incomplete. It shows the monitor is active, and it says nothing about the actions it did not catch, which is the harder number to know.

That architecture is the interesting governance claim, and the 0.002 percent block rate cuts two ways. Read positively, it shows an active control plane catching a small, steady stream of actions before they execute. Read skeptically, a very low block rate tells you the monitor is permissive or the agents rarely try anything the monitor is tuned to stop, and it says nothing about the actions the monitor was not designed to catch. A block rate measures what the filter caught, never what it missed, and the second quantity is the one that matters most and is hardest to know.

“

A block rate is evidence the monitor is on. It is not evidence the monitor is complete. Those are different claims, and only one of them can be read off the number.

Metir analysis

What it signals for everyone else

Set aside the question of whether 26 percent is impressive and ask what the shape of the disclosure implies. Three things stand out. The work that automates first is well-scoped and reviewable, not open-ended. The scale is a fleet-of-agents story, which means orchestration and monitoring become as important as the underlying model. And the governance is being built as an explicit layer around the automation, not bolted on afterward.

For teams adopting AI for their own work, that is the transferable lesson. The value is showing up less in a single all-knowing model and more in the ability to run many scoped tasks across a fleet, with humans reviewing outputs and a monitoring layer in between. Building that layer so it is not tied to one provider, so the orchestration and oversight survive a model swap, is what lets an organization capture the productivity Anthropic is describing without betting everything on one vendor's roadmap. The automation is the headline. The oversight and portability around it are what make it usable.

Sources:

  • Anthropic Says Claude Drives 26% of Its Research and Development | Bloomberg
  • Anthropic says Claude leads 26% of its AI R&D work | Quartz
  • Anthropic Says Claude Leads 26% of Its AI R&D Work | Implicator.ai
  • Claude Now Leads 26% of Anthropic's R&D Work | Enterprise DNA

Image credits

BalticServers data center, by BalticServers.com, via Wikimedia Commons, licensed under CC BY-SA 3.0.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational