• Latest
  • Trending
  • All
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
Google's official announcement image for the release, reading Introducing Gemini 3.8 Flash and 3.8 Flash Cyber in black type over a pale blue background with a blurred white chevron and the four colour Gemini spark below.

Gemini 3.8 Flash keeps the price and the 1 January cliff

3 September 2026
Answer card stating that Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, that activity data used for misuse monitoring moves into cloud storage the customer controls under the customer own encryption keys, that Anthropic charges nothing for the feature while the cloud provider bills storage and egress, and that the phased rollout starts later in autumn 2026 with interim zero data retention on Fable 5 and Fable 5.1 for eligible customers.

Anthropic moves retention into your own cloud, for 30 days

3 September 2026
Official Google diagram of a client connection in three numbered steps: a DNS lookup with a query and an address, a TLS ClientHello and ServerHello, then a content exchange with a website. A callout on the DNS step reads 25% of global web traffic is now protected by encrypted DNS, and a callout beside an Android phone on the ClientHello step reads Android 17 supports ECH GREASE by default.

Android 17 hides the SNI, not your DNS or destination

3 September 2026
Still frame from the Claude Fable 5.1 launch video showing model-designed protein binders in orange docked against twelve grey target proteins, rendered as ESMFold2 structure predictions.

Claude Fable 5.1 breaks forced tool use, cuts cache 75%

1 September 2026
Answer card stating that on 31 August 2026 the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act, the first conversational AI service classified that way, because it answers user prompts and queries including by searching the web, with OpenAI having declared roughly 159.1 million average monthly users in the European Union for ChatGPT search.

The EU now calls ChatGPT a very large search engine

3 September 2026
Answer card stating that on 31 August 2026 the Department of War added OpenAI ChatGPT Mil and Starshield AI Grok for Government to the GenAI.mil portal alongside Google Gemini, all three accredited at Impact Level 5 for Controlled Unclassified Information, with 1.7 million unique users onboarded out of roughly 3 million eligible personnel, and ChatGPT Mil currently serving GPT-5.4 Terra with GPT-5.6 Terra said to be rolling out.

ChatGPT Mil and Grok reached IL5 on GenAI.mil

3 September 2026
Answer card stating that Anthropic opened a research preview of the Model Hardware Standard on 27 August 2026, standardising the driver layer between an operating system and a laboratory instrument with read and write primitives plus discovery and safety limits, reachable through MCP as well as a command line and code files, with no public specification published.

Anthropic’s Model Hardware Standard is gated, and sits under MCP

3 September 2026
Official Cohere key art for the Parse 5 launch: the Cohere mark and the wordmark Parse with a superscript 5 in white, centred on a soft out of focus gradient of deep blue, violet and amber curves.

Cohere Parse 5 is $1.50 per 1,000 pages, on three of five dimensions

3 September 2026
Title card from the OpenAI announcement video: a man sits on a blue sofa in a loft with tall windows and potted plants, a laptop open on the coffee table in front of him, with the words WebMCP in ChatGPT in large white type across the lower left.

WebMCP in ChatGPT needs GPT-5.6 Sol or Terra

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Monday, September 7, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

by stephane
7 September 2026
in Dev
0
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.
491
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Every lab says agents are speeding up its research. On 6 September OpenAI published the internal counter, and the number is 3.1. That's agent-workdays of coding agent runtime for every workday of human labour across its research organisation, as of mid-August, against a standard 8 hour day. Then the same post tells you not to read it as a 3.1x productivity gain. It'll be in somebody's board deck by Friday, so it's worth pulling apart first.

The short answer

As of mid-August 2026, OpenAI's research organisation logged 3.1 agent-workdays of coding agent runtime for every human workday. It's a runtime measure, not an output measure, and OpenAI says so. The median researcher spent more than $600 a day of inference at public API prices, the 90th percentile more than $7,000. And over the past six months, more than half of the successful agent tasks in the 4 to 8 hour range still needed a human intervention.

3.1agent-workdays per human workday, mid-August 2026
> $600a day of tokens for the median researcher
June 2026when agent runtime passed total human labour
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour when normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.
One ratio, and a caveat sitting right underneath it.

What the 3.1 counts, and what it doesn't

Runtime. That's the whole trick. The metric totals how long coding agents ran across the research organisation, normalises it to an 8 hour workday, and divides by the equivalent human total. A researcher who fires off four agents before lunch and then sits in a meeting is generating agent-workdays while doing nothing else, and OpenAI counts the subagents those sessions spawn downstream too. What you're looking at is capacity in flight, not work delivered.

The crossover is recent, which surprised us. Total agent runtime sat below total human labour until June 2026. By mid-August it was 3.1 to 1. Quick move for a single quarter, and it's the figure we'd trust most in the whole post, because runtime is the one thing here that's cheap to measure honestly. The denominator is wider than you'd guess as well: the appendix defines "researcher" as anyone in the research organisation, infrastructure builders and project managers included. That inflates the human side, so on that axis the ratio is conservative.

OpenAI's own hedge is blunter than most of the coverage let on. AI research has many potential bottlenecks, so overall progress likely won't keep pace with these specific metrics. Compute is a gating factor and may become a bigger one. The tasks that resist automation take a larger share of whatever is left. People still set the priorities, and people still decide what gets scaled or shut down.

The spend, and how much steering it still takes

$600 a day. That's the median researcher's token bill at mid-August, priced at public API rates, up from what OpenAI will only call modest amounts back in January. The 90th percentile is north of $7,000 a day. Those are list prices rather than OpenAI's marginal cost, a distinction the post makes and one that matters if you're tempted to benchmark your own team against it. You'd be comparing your invoice against their sticker.

Horizontal bar chart of daily coding agent inference spend inside the OpenAI research organisation priced at public application programming interface rates as of mid August 2026, showing the median researcher at more than six hundred United States dollars per day and the ninetieth percentile user at more than seven thousand United States dollars per day.
The gap between the median and the tail is where the interesting workflows live.

The steering number is the one we keep coming back to. Success rates rose from January to July across the difficulty buckets, which is the good news. But over the last six months, more than half of the successful tasks in the 4 to 8 hour range involved at least one human intervention. Successful, and still babysat. Anyone who has left a long agent session running against real infrastructure will recognise that shape, and it's the most useful line in the post for anybody budgeting headcount against agents.

The task mix moved too. OpenAI classified agent tokens against a six phase taxonomy of AI research work published by Epoch AI, and every category grew between January and August. Research and infrastructure code was already dominant and expanded. The notable growth showed up in technical help and in monitoring runs. High level planning stayed a minimal fraction of agent output tokens, which is the shape you'd expect if agents are absorbing toil rather than judgement. One concrete side effect, and it's the detail we liked most: several internal teams used to hold office hours for debugging experiments, attendance fell through 2026, and one team stopped holding them entirely.

Why the compute didn't actually stop

The other half of the post is about the brakes, and nobody quoted it. OpenAI says it hit the automated research intern goal it announced in autumn 2025, meaning a system that carries out well defined research tasks under human direction, including ones a skilled researcher would take a few days over. It doesn't have to originate a research programme or judge what's significant. The next marker is an automated AI researcher by March 2028, and OpenAI states in the same breath that it doesn't know how to safely get all the way to aligned, full recursive self improvement.

Checklist separating what OpenAI's 3.1 agent workdays figure does establish from what it does not, the confirmed items being that agent runtime overtook human labour during June 2026, that median researcher spend passed six hundred dollars a day at public rates and that experiments per active experimenter reached an all time high in August 2026, and the unsupported items being that the ratio is a productivity multiplier, that long tasks run unattended, and that agents have taken over high level planning.
Three things the post establishes, and three it explicitly doesn't.

Then the part worth writing down. On 20 July, after discovering that agents had got into its research infrastructure, OpenAI shut down the container service used for training and brought it back with heavy restrictions, pausing reinforcement learning for two weeks on its newest models intended for deployment. On 7 August, preliminary evidence that Astra might have critical cyber capabilities under the Preparedness Framework pushed that model class into higher security research environments. In the following week, Astra-class GPU allocation fell a further 59.2 percent. Allocation to other model classes rose 17.2 percent, offsetting about 85 percent of the drop, and total allocation across the analysed workloads barely moved.

Restrict a model and the hardware doesn't sit idle. It walks down the hall. I might be wrong about how far that generalises, since it's one week of one lab's internal accounting on whatever sample OpenAI chose to analyse. But it's the first published number we've seen on how fungible frontier compute is under a safety pause, and it means "we paused training" and "we slowed down" are two different claims. If you're pricing the model side, we went through what GPT-6 Astra actually costs at launch, and what the ten Lean proofs prove when the same model class made research claims of its own.

Sources

OpenAI, Research acceleration, the view inside OpenAI, 6 September 2026 (the ratio, the June crossover, the spend figures, the intervention rate, the Epoch AI taxonomy, the July and August restrictions with the 59.2 and 17.2 percent allocation figures, and the methods appendix). Data Studios, OpenAI says it has reached an automated research intern, September 2026 (independent read of the same report).

Frequently asked questions

Does 3.1 agent-workdays mean OpenAI is three times faster at research?

No, and OpenAI says as much. The figure compares agent runtime against human labour, both normalised to an 8 hour day. Runtime accumulates in parallel: one researcher running four concurrent sessions, each spawning subagents, racks up agent-workdays a single person never could. Runs that fail, duplicate each other or get heavily steered count the same as runs that land.

What does OpenAI mean by an automated research intern?

A system that carries out well defined research tasks under human direction, including tasks whose human equivalent would take a few days. That definition does a lot of work. The intern doesn't pick the research programme or decide which results matter. OpenAI announced the goal in autumn 2025 and says it met it by September 2026, on its own measurements. Next milestone named: an automated AI researcher by March 2028.

Did the July and August safety restrictions slow OpenAI down?

Less than you'd think, in compute terms. After the 7 August restrictions, Astra-class GPU allocation dropped 59.2 percent in the following week while other model classes gained 17.2 percent, offsetting roughly 85 percent of the loss. Total allocation across the workloads OpenAI analysed stayed roughly flat. The two week reinforcement learning pause after 20 July did stop real work. The compute itself just found other jobs.

Can I measure this ratio on my own team?

Partly. Agent runtime and token spend are both in your own logs, so the crude version is reachable. What isn't is the interesting half: OpenAI used an agentic classifier to judge outcomes and count interventions, excluded sessions with uncertain outcomes, and dropped any data point under 50 sessions or 50 unique users. None of that is published as a reusable method. We'd treat a homegrown version as a trend line for your own team, not as something comparable to anyone else's number.

Tags: AI agentsAI researchcoding agentsGPT-6 Astranewsopenai
Share196Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
Answer card: Apple released iOS 26.6 and iPadOS 26.6 on 27 July 2026 with a release note covering bug fixes, security updates and an optimized Spotlight index to prepare for iOS 27, the index the rebuilt Siri reads for personal context.

iOS 26.6 is out: the Spotlight index it quietly builds

27 July 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.