On September 17, 2026, Anthropic published internal metrics it had not shared before: Claude now leads about 26 percent of the company's research and engineering work, up from under 1 percent in February. Roughly 30,000 agents run at any given time, and the company said more than 90 percent of its research and development happens with Claude as either a collaborator or a lead. It is a striking disclosure, and it is also easy to misread in either direction. The honest version is more interesting than either the hype or the dismissal.
The single most important word in the announcement is "leads," so start there.
What "leads" actually means
Anthropic drew a specific line between two modes. "Collaborates" means a human and Claude work on a task together, with the human driving. "Leads" means Claude can carry most of a task end-to-end from a high-level prompt, while a human supervisor reviews the result. The 26 percent figure is the "leads" share, and it climbed from essentially nothing to a quarter of the work in six months.
From under 1 percent to 26 percent in six months
Share of Anthropic's research and engineering work that Claude leads end-to-end under human oversight, as the company disclosed it.
"Leads" is a narrower claim than "does." It counts tasks a model can carry from a prompt to a result with a human reviewing the output, not autonomous research with no person in the loop.
That is a genuine jump, and it is worth not inflating. "Leads" is not "does autonomously with no human." There is a supervisor reviewing outputs, and the tasks that qualify are the ones a model can take from a prompt to a reviewable result, which skews toward well-scoped engineering and research work rather than open-ended judgment calls. Read precisely, the number says that a large and fast-growing slice of Anthropic's routine technical work can now be initiated and carried by a model under human review. That is a big claim on its own; it does not need to be stretched into "AI is running the lab."

The recursive-improvement framing, handled with care
The reason this disclosure travels is that it touches the idea of AI systems helping build their own successors. If a company's frontier model is doing a rising share of the research that produces the next model, the loop that people describe when they talk about recursive self-improvement is, in a limited and supervised form, visible in a real org chart.
Anthropic
ClaudeIt is worth keeping two things straight. First, "leads a quarter of R&D tasks under supervision" is a long way from "designs its own successor unsupervised," and Anthropic's own definitions make that clear. Second, the metric is self-reported, and the company that publishes it is also the one that benefits from the narrative that its models are exceptionally capable. That does not make the number wrong. It means the right posture is to treat it as a credible internal signal, disclosed with unusual specificity, rather than as an audited external fact.
The monitoring is the part worth studying
The most substantive detail is not the automation share. It is the oversight wrapped around it. Anthropic said every agent action passes through an online monitor before it runs, and 100 percent of executed actions are reviewed after the fact by a separate offline monitor. Across roughly 1 billion research and engineering decisions in August, about 1 in 47,000, or 0.002 percent, was blocked.
The oversight layer around the agents
The headline number is the automation. The part that governs it is the monitoring on either side of every action.
A 0.002 percent block rate can read as either reassuring or incomplete. It shows the monitor is active, and it says nothing about the actions it did not catch, which is the harder number to know.
That architecture is the interesting governance claim, and the 0.002 percent block rate cuts two ways. Read positively, it shows an active control plane catching a small, steady stream of actions before they execute. Read skeptically, a very low block rate tells you the monitor is permissive or the agents rarely try anything the monitor is tuned to stop, and it says nothing about the actions the monitor was not designed to catch. A block rate measures what the filter caught, never what it missed, and the second quantity is the one that matters most and is hardest to know.
A block rate is evidence the monitor is on. It is not evidence the monitor is complete. Those are different claims, and only one of them can be read off the number.
Metir analysis
What it signals for everyone else
Set aside the question of whether 26 percent is impressive and ask what the shape of the disclosure implies. Three things stand out. The work that automates first is well-scoped and reviewable, not open-ended. The scale is a fleet-of-agents story, which means orchestration and monitoring become as important as the underlying model. And the governance is being built as an explicit layer around the automation, not bolted on afterward.
For teams adopting AI for their own work, that is the transferable lesson. The value is showing up less in a single all-knowing model and more in the ability to run many scoped tasks across a fleet, with humans reviewing outputs and a monitoring layer in between. Building that layer so it is not tied to one provider, so the orchestration and oversight survive a model swap, is what lets an organization capture the productivity Anthropic is describing without betting everything on one vendor's roadmap. The automation is the headline. The oversight and portability around it are what make it usable.
Sources:
- Anthropic Says Claude Drives 26% of Its Research and Development | Bloomberg
- Anthropic says Claude leads 26% of its AI R&D work | Quartz
- Anthropic Says Claude Leads 26% of Its AI R&D Work | Implicator.ai
- Claude Now Leads 26% of Anthropic's R&D Work | Enterprise DNA
Image credits
BalticServers data center, by BalticServers.com, via Wikimedia Commons, licensed under CC BY-SA 3.0.