OpenAI: 3.1 agent-workdays per human workday
OpenAI says it hit its automated research intern goal and runs 3.1 agent-workdays per human workday at $600+ a day per researcher. Here is how to measure your own ratio.
OpenAI published its internal numbers on how much of its own research is now done by coding agents, and the headline figure is 3.1 agent-workdays of effort for every human workday. It also says it hit the goal it set last fall: an automated research intern by September 2026. The number worth copying is not the 3.1. It is that OpenAI bothered to measure it at all.
What actually happened
OpenAI defines a "research intern" narrowly: a system that carries out well-defined tasks under human direction, including work that would take a skilled researcher a few days. Not autonomy. Direction plus duration. The company says it is aiming for a fully automated AI researcher by March 2028.
The operating numbers are the interesting part. As of mid-August, the median OpenAI researcher was running more than $600 a day of inference on internal coding agents at API prices. The top decile was past $7,000 a day. That is per person, on tokens alone.
And OpenAI put a caveat on its own metric that most vendors would have quietly omitted: 3.1 agent-workdays is not a 3.1x productivity gain. Agent runtime can be parallel, redundant, unsuccessful, or heavily steered by the human sitting next to it. Chief scientist Jakub Pachocki separately said he expects the current pace to carry into recursive self-improvement — a claim about the future, not a measurement, and worth filing under forecast rather than fact.
Why this matters for your business
Two things transfer directly, and neither of them is "buy more agents."
First, the ratio is a metric you can actually run. Most small companies deploying AI have no denominator. They have a Claude subscription, a Copilot seat, and a feeling. OpenAI's framing — agent effort against human effort, tracked over months — is something you can approximate this quarter from your API dashboard and your team's hours. If the ratio is climbing and cycle time is not, you have a spend problem dressed up as an adoption win.
Second, $600 a day per seat is what unbounded agent use costs. That is a frontier lab optimizing for research speed with no budget constraint, so it is a ceiling, not a target. But it kills the assumption that agent cost is a rounding error next to salary. At $600 a day, five engineers is roughly $780K a year in tokens. Set per-project budget caps and route cheap work to cheap models before that curve finds you. We wrote about the same math when Writer showed a 41% cost drop from harness changes alone.
The honest read on the milestone itself: an intern that needs a defined task and a human reviewing the output is a good intern. It is not a replacement for the person assigning the work.
Key takeaways
- OpenAI reports 3.1 agent-workdays of effort per human workday across its research org as of mid-August 2026
- Median researcher inference spend is $600+ a day at API prices; the top decile exceeds $7,000 a day
- OpenAI explicitly warns the 3.1 figure is not a 3.1x productivity multiplier — runtime can be parallel, redundant, or failed
- Its "research intern" definition requires human direction on well-defined tasks; the automated researcher target is March 2028
- Track your own agent-effort-to-human-effort ratio against cycle time, and cap per-project token spend before it compounds
An agent you cannot measure is an invoice, not a system. We instrument agent workflows with per-project cost caps, model routing, and the metrics that tell you whether the spend moved anything. Run the numbers on your stack or book a working session.
Sources: OpenAI, Research acceleration: The view inside OpenAI, Help Net Security.
- #ai-agents
- #openai
- #coding-agents
- #automation
- #metrics
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
Thailand pauses 166 data centre projects: compute is political
Thailand's data centre policy board froze 49 builds and 117 pending approvals on September 4. AI capacity is a permitting question now, not just a price.
Read itAI agents ran a full ransomware breach in under 10 hours
Unit 42 documented an AI-driven ransomware intrusion that took under 10 hours instead of two weeks, used 50+ ATT&CK techniques, and left an 80-page audit.
Read it