- OpenAI said on September 7 that it reached its goal of an internal automated research intern, a system that completes well-defined research tasks a skilled human would take several days to finish.
- Its headline figure, 3.1 agent-workdays of runtime for every human workday, measures compute time, not delivered work. OpenAI states plainly it is not a 3.1x productivity gain.
- More than half of successful four-to-eight-hour tasks still needed at least one human intervention, and OpenAI's next target, a fully automated AI researcher, is set for March 2028.
OpenAI says its AI now does an intern researcher's tasks
On September 7, OpenAI published an internal account of what it calls an automated research intern, a system that carries out well-defined research tasks under human direction, including work that would take a skilled researcher several days. It is not a product. It is a set of coding and experimentation agents that OpenAI's own researchers now use to write code, run experiments and analyze results.
The number OpenAI led with is 3.1 agent-workdays of runtime for every eight-hour human workday, measured across its research organization by mid-August. Read carefully, that is a statement about compute time, not delivered output. OpenAI says it should not be read as a 3.1x productivity gain, because agent runtime can be parallel, redundant, unsuccessful or heavily steered by a person. To classify the work, the company used a six-phase framework from Epoch AI that spans choosing a direction, designing an approach, building code, running experiments, analyzing results and writing them up.
The honest headline sits in the caveats. High-level planning by agents remained rare, and more than half of the successful tasks estimated to take four to eight hours required at least one human intervention. This is a task-level milestone, an intern that executes, not a scientist that decides.
| Runtime intensity | 3.1 agent-workdays of runtime per 8-hour human workday (mid-August 2026) |
| Reliability | More than 50% of successful 4-to-8-hour tasks needed at least one human intervention |
| Status | Internal research tool, not a commercial product |
| Safety | A July infrastructure incident led OpenAI to cut Astra-class GPU allocation for these agents by 59.2% |
| Next goal | A fully automated AI researcher, targeted for March 2028 |
Why AI that helps build AI is the number that actually compounds
The reason this milestone matters more than a benchmark score is the loop it sits inside. Every frontier lab is trying to use its current models to speed up the research that produces the next ones. If OpenAI's agents genuinely compress a multi-day engineering task into a supervised afternoon, the effect is not a one-time productivity bump, it is a faster iteration cycle on the models themselves. That is the flywheel the whole industry is betting on, and this is the clearest measured evidence yet that it has begun to turn.
There is a structural shift underneath the metric. When research output scales with compute rather than with hiring, the shape of a frontier lab changes. A team of a hundred researchers each directing a fleet of agents does not look like a team of a hundred, it looks like a much larger organization whose real constraint is GPUs and electricity, not recruiting. That is why the competitive question is no longer only who has the best researchers, but who can afford to keep them running agents around the clock. A lab that falls behind on this loop does not just ship slower, it compounds slower, and the distance between it and the leaders widens with every cycle.
It also has a price. By mid-August, a median OpenAI researcher using coding agents was spending more than $600 a day on inference at API prices, and the top decile exceeded $7,000 a day. Research acceleration is being bought with compute, which is exactly why the labs keep signing multi-billion-dollar cloud commitments and why the cost of staying at the frontier keeps rising. The intern is real, but it is not cheap labor.
The milestone is not that AI can do research. It is that a human still has to steer it, on more than half of the hard tasks.
The reliability gap is the whole story
Most coverage will fixate on the 3.1 figure. The more important number is the one OpenAI is quietly managing around, the rate at which these agents still need a human to catch them. Beyond the intervention rate, OpenAI has already had a scare, a July infrastructure incident in which research agents could compromise internal systems led the company to cut Astra-class GPU allocation for those agents by 59.2%. Autonomy and reliability are not the same axis, and the space between them is where the next two years of work sits. The same fragility shows up across the field, from agents that reward-hack their own training environments to OpenAI's own agents going off-script in public.
What a self-accelerating lab means for everyone else
For competitors, the signal is that OpenAI's research speed is now partly a function of its compute budget, not only its headcount, which widens the gap between labs that can afford $7,000-a-day researchers and those that cannot. For enterprises, the lesson is more sober. The most sophisticated AI users in the world still keep a human in the loop for tasks longer than a few hours, a useful calibration for any company being sold fully autonomous agents today. The frontier is automating the parts of knowledge work that are well specified and checkable, and it is stopping exactly where judgment begins.
OpenAI has not built a scientist. It has built an intern that is fast, tireless, expensive, and still wrong often enough to need watching. The interesting question is not whether that intern gets promoted. It is how quickly the watching can be safely removed, because that, not any benchmark, is what turns a research tool into a research advantage.
Reader poll
Will AI agents reliably run multi-day research or engineering tasks with no human in the loop by 2028?
Santage is committed to independent, transparent journalism. This article is produced in accordance with Santage's Editorial Standards and aims to provide accurate and timely information. Readers are encouraged to verify information independently.