metir
metir
Docs
Download on App StoreGet it on Google PlayLoginSign Up
Back to Blog
OpenAI
AI Agents
AI Research
Recursive Self-Improvement
Codex

OpenAI's Automated Research Intern: What the Data Says

OpenAI says coding agents now supply 3.1 agent-workdays per human workday. Here is what its automated research intern milestone does, and does not, prove.

Metir AI TeamSeptember 7, 20265 min read
OpenAI's Automated Research Intern: What the Data Says

On September 6, 2026, OpenAI said it had reached its internally defined automated research intern milestone. The company now has coding agents carrying out well-defined research tasks under human direction, including work it estimates would take a skilled researcher several days. That is a consequential claim, but it is not the same as an autonomous scientist choosing its own research agenda.

OpenAI logoOpenAI
OpenAI published its first detailed internal snapshot of agent-assisted AI research on September 6, 2026.
>$600/dayMedian researcher agent use at API pricesMid-August 2026
>$7,000/day90th-percentile researcher useMid-August 2026
3.1 to 1Agent-workdays per human workdayEight-hour normalization
>50%Successful 4-8 hour tasks needing interventionPrevious six months

What OpenAI's automated research intern actually does

OpenAI defines the milestone narrowly. Its system can execute a well-scoped research assignment under human supervision. People still decide what to study, judge which results matter, and choose whether to scale, pause, or deploy a system. OpenAI says high-level planning remains a minimal share of agent output, while build, technical support, and experiment-monitoring work have grown faster.

That distinction matters because research is not one task. Writing evaluation code, debugging infrastructure, running experiments, interpreting results, and selecting the next hypothesis require different kinds of judgment. OpenAI's evidence is strongest for execution. It is much weaker for independent research taste.

Agent runtime now exceeds human research time

Workdays of effort used across OpenAI's research organization per one human workday, as of mid-August 2026.

Method: OpenAI normalized agent runtime to standard eight-hour workdays. This measures runtime, not equal-value research output. Source: OpenAI.

The workday comparison is also an estimate of runtime, not a direct productivity multiple. By mid-August, OpenAI measured 3.1 eight-hour agent-workdays for each human workday across its research organization. Agents run concurrently, so the number shows the scale of delegated computation. It does not establish that one agent-hour creates the same value as one researcher-hour.

The evidence says acceleration, with large caveats

OpenAI reports that the median researcher was using more than $600 of agent inference per day at API prices by mid-August, while the 90th percentile exceeded $7,000. Experiments per active experimenter reached their highest level since tracking began in January 2025. Yet the company explicitly says compute capacity also grew, making it difficult to isolate how much of the increase came from better agents.

Reliability is another constraint. Success rates improved across several estimated task-duration buckets from January through July, but more than half of successful tasks estimated at four to eight human hours still needed at least one intervention. OpenAI also excludes uncertain outcomes and low-sample points from those charts. The disclosure is unusually concrete for a frontier lab, but it remains company-reported internal data without an independent replication set.

“

Three agent-workdays of runtime are evidence of scale, not proof of three times the scientific progress.

Metir analysis of OpenAI's methodology
The Pioneer Building in San Francisco, an office associated with OpenAI
The Pioneer Building in San Francisco, photographed in 2019 and described by Wikimedia Commons as housing OpenAI offices at the time. The photograph does not depict the research systems or the September 2026 disclosure. Photo by HaeB, CC BY-SA 4.0.

Why this is bigger than one lab

OpenAI is not alone in reporting this transition. Anthropic says more than 80 percent of the production code merged into its codebase was AI-authored by May 2026, while its typical engineer merged eight times as much code per day in the second quarter as in 2024. Anthropic also warns that lines of code overstate productivity and that major gaps remain in choosing goals.

Together, the disclosures suggest a common pattern: frontier labs are automating the execution layer of AI research before the direction-setting layer. That can still compound progress because each human can supervise more experiments, but bottlenecks move to review quality, compute, experimental design, and decisions about what not to pursue.

Safety controls can redirect work, not just slow it

OpenAI's own safety data illustrates the coordination problem. After new restrictions on its Astra model, Astra-class GPU allocation fell 59.2 percent in the following week. Allocation to other model classes rose 17.2 percent, offsetting about 85 percent of the decline. The company interprets that as researchers redirecting scarce compute rather than leaving it idle.

This makes governance harder than placing a brake on one model. Effective controls have to account for substitution across projects, while research environments need audit trails, access boundaries, and human approval for consequential actions. That lesson connects directly to OpenAI's recent misalignment incident reporting proposal and its Astra safety restrictions.

OpenAI's next stated target is an automated AI researcher by March 2028. The September milestone makes that roadmap more concrete, but not inevitable. What has been demonstrated is a rapidly expanding supervised research workforce made of agents. Whether those agents can reliably choose valuable research directions, and whether labs can keep them controlled as autonomy grows, remains the decisive open question.

Sources:

  • Research acceleration: The view inside OpenAI
  • An Alien Mind, by OpenAI chief scientist Jakub Pachocki
  • When AI builds itself | Anthropic Institute
  • AI Researchers' Views on Automating AI R&D and Intelligence Explosions
  • The Shift to Agentic AI: Evidence from Codex

Image credits

Hero and in-body photograph: the Pioneer Building in San Francisco, photographed by HaeB in 2019 and described by Wikimedia Commons as housing OpenAI offices at the time, via Wikimedia Commons, licensed under CC BY-SA 4.0. The photograph does not depict the research systems or the September 2026 disclosure.

Ready to experience AI that adapts to you?

metir brings together the world's best AI models in one seamless experience. Start for free today.

Get Started Free
metir

Agentic Operating System for Professionals buried in meetings, emails and docs.

© 2026 metir. All rights reserved.

Product

  • Features
  • Pricing
  • Research
  • Docs
  • Blog
  • Enterprise

Company

  • Docs
  • Support
  • Careers

Legal

  • Terms of Service
  • Privacy Policy

Personalisation is powerful. Privacy is non-negotiable.

Status: All systems operational