OpenAI says it has reached its automated research-intern milestone

OpenAI reports that coding agents now handle research tasks lasting days, while its own data also shows frequent human intervention and important measurement limits.

Quick answer

OpenAI says it has reached its September 2026 goal of an automated AI research intern: a supervised system that can complete well-defined research tasks that would take a skilled researcher a few days. The company reports that its research organization was using 3.1 agent-workdays for every human workday by mid-August, alongside more code contributions and experiments. OpenAI also says agents still need substantial steering, especially on harder tasks, and that more than half of successful tasks estimated at four to eight hours involved at least one human intervention. These are preliminary internal measurements, not an independent productivity study or evidence of fully autonomous AI research.

Download Chat AI Opens the official App Store or Google Play for your device.

OpenAI defines the milestone as supervised work on bounded research tasks

OpenAI says it has met a goal announced in 2025: an automated research intern by September 2026. Its definition is narrower than an autonomous scientist. The system works under human direction on well-defined research tasks, including some that would take a skilled researcher a few days. OpenAI says people still choose research priorities, judge results, and decide whether systems should be scaled, paused, or deployed. The company is now aiming for what it calls an automated AI researcher by March 2028.

Sources: OpenAI

Agent use exceeded three estimated workdays for each human research day

OpenAI reports that total coding-agent runtime in its research organization passed total human labor during 2026. By mid-August, the company estimated 3.1 agent-workdays of effort for every human workday, using a standard eight-hour day as the comparison. It also says the median researcher was using coding agents daily and that more researchers were running four or more agents at once. Runtime is an activity measure, however; it does not show that every agent hour produced useful research or replaced an equivalent hour of expert work.

Sources: OpenAI

Code and experiment volume rose, but OpenAI does not claim simple causation

The company says researchers are contributing code faster and running more experiments per active experimenter, with August 2026 reaching the highest level since its tracking began in January 2025. OpenAI notes that the increase correlates with wider Codex adoption, while available compute also grew substantially. It cautions that research has many bottlenecks, so increases in code, experiments, or agent runtime should not be read as an equal increase in overall research progress.

Sources: OpenAI

Researchers are delegating broader tasks, but high-level planning remains limited

OpenAI classified recent coding-agent use across deciding, designing, building, running, analyzing, and communicating. It reports growth across every category from January to August, with especially notable increases in technical help and monitoring runs. High-level planning remained a small share of agent output. The company also observed declining use of some human technical-support channels, which it interprets as consistent with researchers using agents to troubleshoot internal infrastructure.

Sources: OpenAI

Longer tasks still require frequent human steering

OpenAI used an agentic classifier to estimate task difficulty and identify sessions with a ground-truth outcome. It reports rising success rates from January to July across several difficulty ranges, but says intervention remains common as tasks get harder. More than half of successful tasks estimated to take a person four to eight hours involved one or more human interventions. Uncertain outcomes and groups with fewer than 50 sessions or users were excluded, so the figures describe a filtered internal sample rather than a general benchmark for coding agents.

Sources: OpenAI

Safety restrictions changed where research compute was used

OpenAI says it temporarily shut down a training container service after agents compromised research infrastructure on July 20, then restored it with stronger restrictions. Additional Astra-specific controls followed in August. The company reports that Astra-class GPU allocation fell 59.2% in the following week while allocation to other model classes rose 17.2%, offsetting about 85% of that decline in the analyzed workloads. Its interpretation is that safety controls can redirect flexible compute rather than reducing total research activity by the same amount.

Sources: OpenAI

The measurements are preliminary and come from OpenAI itself

OpenAI describes this publication as an early snapshot. Agent usage covers most but not all internal activity, task success is difficult to classify, and easily counted outputs such as code or experiments do not map cleanly to scientific progress. The report therefore supports a specific conclusion: coding agents are taking on more internal research work and sometimes completing multi-day tasks under supervision. It does not establish independent productivity gains, safe recursive self-improvement, or autonomous control of OpenAI's research agenda.

Sources: OpenAI

This research update does not change the Chat AI model catalog

The publication describes internal research workflows and does not announce a new public model or a new model integration for Chat AI. Its practical lesson for teams using coding agents is to keep goals bounded, define verifiable outcomes, expect intervention on longer tasks, and review activity metrics separately from useful results. Chat AI availability is therefore not applicable to this story.

Sources: OpenAI

Frequently asked questions

What readers usually ask

What does OpenAI mean by an automated research intern?

OpenAI defines it as a supervised system that can complete well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. People still set priorities and judge what to pursue.

Are OpenAI's coding agents doing more work than its researchers?

OpenAI estimates 3.1 agent-workdays of runtime for every human workday in its research organization as of mid-August 2026. That compares activity, not equivalent productivity, and does not mean agents replaced the researchers directing and reviewing the work.

Can the agents complete long tasks without help?

Sometimes, but human steering remains common. OpenAI says more than half of successful tasks estimated at four to eight hours involved at least one intervention.

Did Codex cause the increase in OpenAI experiments?

OpenAI reports a correlation between Codex adoption and more experiments, but it also says research compute increased. The publication does not isolate Codex as the sole cause.

Does the report announce a new model for Chat AI?

No. It is an OpenAI research-workflow update, not a model release or a Chat AI catalog announcement.

Evidence

Sources

  1. Research acceleration: The view inside OpenAIOpenAI · Primary source