• Latest
  • Trending
  • All
Answer card: on 30 July 2026 OpenAI cut GPT-5.6 Luna from $1 and $6 to $0.20 and $1.20 per million tokens, an 80 percent reduction on both input and output, cut Terra from $2.50 and $15 to $2 and $12, and left Sol unchanged at $5 and $30, twenty one days after the family went generally available.

GPT-5.6 price cut: Luna drops 80%, Sol stays at $5

3 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
Google's official announcement image for the release, reading Introducing Gemini 3.8 Flash and 3.8 Flash Cyber in black type over a pale blue background with a blurred white chevron and the four colour Gemini spark below.

Gemini 3.8 Flash keeps the price and the 1 January cliff

3 September 2026
Answer card stating that Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, that activity data used for misuse monitoring moves into cloud storage the customer controls under the customer own encryption keys, that Anthropic charges nothing for the feature while the cloud provider bills storage and egress, and that the phased rollout starts later in autumn 2026 with interim zero data retention on Fable 5 and Fable 5.1 for eligible customers.

Anthropic moves retention into your own cloud, for 30 days

3 September 2026
Official Google diagram of a client connection in three numbered steps: a DNS lookup with a query and an address, a TLS ClientHello and ServerHello, then a content exchange with a website. A callout on the DNS step reads 25% of global web traffic is now protected by encrypted DNS, and a callout beside an Android phone on the ClientHello step reads Android 17 supports ECH GREASE by default.

Android 17 hides the SNI, not your DNS or destination

3 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Tuesday, September 15, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

GPT-5.6 price cut: Luna drops 80%, Sol stays at $5

by stephane
3 September 2026
in Dev
0
Answer card: on 30 July 2026 OpenAI cut GPT-5.6 Luna from $1 and $6 to $0.20 and $1.20 per million tokens, an 80 percent reduction on both input and output, cut Terra from $2.50 and $15 to $2 and $12, and left Sol unchanged at $5 and $30, twenty one days after the family went generally available.
494
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Somebody on your team benchmarked GPT-5.6 three weeks ago and picked a tier. That estimate is now wrong. On 30 July OpenAI cut the API price of GPT-5.6 Luna by 80 percent and Terra by 20 percent, and left Sol exactly where it was at $5 in and $30 out per million tokens. Luna is the one to stare at. A dollar and six dollars became twenty cents and $1.20, and the reduction is the same 80 percent on input as on output, so a bill drops by the headline number whatever your token mix looks like. Terra gets the tidier version, a flat fifth off both legs. The thing nobody put in a headline is what this does to the distance between the rungs. It used to be five times from the cheapest tier to the top one. It's twenty-five now.

The short answer

On 30 July 2026 OpenAI repriced the two lower GPT-5.6 tiers. Luna goes from $1 / $6 to $0.20 / $1.20 per million tokens, Terra from $2.50 / $15 to $2 / $12, and Sol holds at $5 / $30. Each cut applies equally to input and output, so no token mix changes the arithmetic. The consequence worth acting on is the widened spread between tiers, which makes routing decisions matter far more than they did at launch.

$0.20Luna input, per million tokens
80%cut on both legs, not just input
25xnew gap from Luna up to Sol
Answer card showing the GPT-5.6 API price change of 30 July 2026: Luna from $1 and $6 down to $0.20 and $1.20 per million tokens for an 80 percent cut on both legs, Terra from $2.50 and $15 down to $2 and $12 for a 20 percent cut, and Sol unchanged at $5 and $30, twenty one days after the family reached general availability.
Two tiers moved. The one everybody quotes in benchmarks did not.

What actually changed

Three numbers, and they’re already live on the model pages. Luna: $0.20 per million input tokens, $1.20 per million output. Terra: $2 and $12. Sol: $5 and $30, same as it was on 9 July, same as GPT-5.5 charged before it.

The detail we’d flag first is that both cuts are uniform. Vendors love an asymmetric discount, usually a deep cut on input because input is cheap to serve and it makes the percentage look bigger than the saving. Not here. Luna’s input fell exactly 80 percent and so did its output. Terra fell exactly 20 percent on each. That means you don’t need to model your input-to-output ratio to know what happens to your invoice. It falls by the headline number. That’s rarer than it sounds, and it’s the reason this one is easy to act on.

Caching survives intact. Luna’s cached input reads at $0.02 per million on OpenAI’s own page, the usual 90 percent off, and the batch route still carries its own discount. So the floor under a well-cached high-volume workload just moved a long way down.

What didn’t change: Sol. If your agent loop runs on the frontier tier, congratulations, your bill on 31 July is identical to your bill on 29 July. We had a moment of thinking the headline applied to us before checking, and we suspect we won’t be the only ones.

The spread, not the discount

Here’s the part that actually reshapes architecture.

At general availability on 9 July, the ladder was gentle. Sol cost five times Luna on input and five times on output. We wrote at the time that Terra was the sensible default, and at $2.50 against $1 for Luna that was an easy call, because the quality difference bought more than the 2.5x price difference did.

Run the same division today. Sol against Luna is 25x on input and 25x on output. Terra against Luna is 10x. The tiers didn’t get closer together, they flew apart, and a routing mistake that used to cost you a small multiple now costs an order of magnitude.

Comparison chart of the price ratio between GPT-5.6 Sol and GPT-5.6 Luna, showing five times at general availability on 9 July 2026 against twenty five times from 30 July 2026, on both input and output tokens alike.
Same three tiers, very different geometry. This is what the percentage headline hides.

Which changes the calculus on cascades. A route-first design, where a cheap model handles the request and escalates only when it can’t, carries real engineering cost: a classifier, an escalation path, evaluation on both legs. At 2.5x savings that overhead rarely paid for itself and we mostly told people not to bother. At 10x it’s a different conversation, and on high-volume classification or extraction work it’s probably now the right answer.

Against a real bill

Say a modest production workload, 10 million input tokens and 2 million output tokens a day. Nothing exotic.

On Luna that used to be $10 plus $12, so $22 a day. Now it’s $2 plus $2.40, so $4.40. Over a month, roughly $660 becomes $132.

Terra on the same traffic was $25 plus $30, so $55 a day, and it’s now $20 plus $24, so $44. About $1,650 a month down to $1,320. Useful, not transformative.

Sol on that traffic is $50 plus $60, which is $110 a day, and it was $110 a day last week too.

So the honest summary: this is a volume play. If you’re spending real money at the bottom of the ladder, you just got a five-fold cut. If you’re spending real money at the top, you got a press release.

What we can verify, and what we can’t

The prices are checkable, and we checked them on the model pages rather than trusting the coverage. Worth doing yourself, because most third party pricing trackers were still quoting $1 / $6 for Luna hours after the announcement. Aggregator pages lag. Treat them as a starting point and confirm against the vendor.

The reasoning behind the cut is a different category of claim. OpenAI says it found roughly 20 percent lower end to end serving cost and better than 15 percent token generation efficiency, partly through work the model itself did on production code. Plausible, unaudited, and exactly the kind of number a company controls the definition of. Sam Altman has called costs “a huge issue” in recent remarks, which is at least a candid framing of why a discount lands three weeks into a launch.

The one to be most careful with: OpenAI’s claim that Luna matches models that were frontier-class a year ago at around six cents on the dollar per task, at nearly nine times the speed. Per task, not per token, which quietly folds in how many tokens the model spends thinking. It might well be true. It’s also unfalsifiable without their task set, so don’t put it in a budget.

Two-column checklist separating the confirmed facts of the 30 July 2026 GPT-5.6 repricing, including live Luna and Terra rates and 90 percent cached input discounts, from the unverified vendor claims about serving cost reductions and per-task cost comparisons.
Left column you can put in a spreadsheet. Right column you cannot.

Do you move

Not automatically. Nothing about any model’s quality changed on 30 July, only the invoice. Luna is exactly as good or as bad as it was when you last tested it, and if it flunked your accuracy bar in early July it still flunks it.

What changed is that the test is now worth rerunning, because the prize got five times bigger. Pull your evaluation set out, run Luna against it again, and see whether the tasks you rejected it for were genuinely beyond it or whether you rejected it on a cost-benefit calculation that no longer holds. That second category is where the money is. For the comparison against the previous generation, our GPT-5.6 against GPT-5.5 write-up still stands, with every price in it now stale for the bottom two tiers.

And a small note on planning. Two repricings inside a month, on a family that shipped in July, tells you the cost floor under this generation isn’t settled. We wouldn’t rebuild an architecture around a price that’s moved twice in three weeks. Build the routing layer so the tier is a config value, and let the next cut be somebody else’s migration.

Sources: the announcement and the rationale come from OpenAI’s own post, Advancing the price-performance frontier with GPT-5.6 (30 July 2026), with the live per-token rates and the $0.02 cached input figure taken from the GPT-5.6 Luna model page in the OpenAI API documentation. The old and new prices, the Fast mode rates and the general availability date are corroborated by Unite.AI. The serving cost and token efficiency figures, the Sam Altman quote and the enterprise spending context are reported by Yahoo Finance and CNBC.

What changed since

9 July 2026. OpenAI released GPT-5.6 to everyone after a US safety review, as three tiers rather than one model: Sol, Terra and Luna. The launch prices are the ones this price cut replaced.

30 July 2026. The cut described above.

Two later changes are worth knowing if you are picking a tier today. Sol reached ChatGPT before the API in August, and an Ultrafast preview appeared on 13 August running Sol at up to 750 tokens a second with no published price.

Frequently asked questions

How much does GPT-5.6 cost after the 30 July 2026 price cut?

Per million tokens: Luna is $0.20 input and $1.20 output, Terra is $2 input and $12 output, and Sol is unchanged at $5 input and $30 output. Cached input reads are 90 percent cheaper, which puts Luna cached input at $0.02 per million on OpenAI's own model page.

Which GPT-5.6 tiers got cheaper?

Only the two lower tiers. Luna fell 80 percent and Terra fell 20 percent, both uniformly across input and output tokens. Sol, the frontier tier, did not move at all, so anyone running most of their volume through Sol sees no change in their bill from this announcement.

Did the GPT-5.6 price cut change ChatGPT subscription prices?

No. This is an API pricing change, so it affects what you pay per token when you call the models yourself. ChatGPT consumer and business plan pricing was not part of the announcement.

Why did OpenAI cut GPT-5.6 prices three weeks after launch?

OpenAI attributes it to efficiency gains found after launch, citing roughly a 20 percent reduction in end to end serving cost and over 15 percent better token generation efficiency. The commercial backdrop is real too: reporting around the announcement points to enterprise buyers getting stricter about AI spend, and to pressure from cheap open weight models.

Should I switch from Terra to Luna now that Luna is 80 percent cheaper?

Only after you rerun your own evaluation. The price gap between Terra and Luna widened from 2.5x to 10x, which makes the test worth doing, but nothing about Luna's quality changed on 30 July. If Luna failed your accuracy bar three weeks ago, it still fails it, just more cheaply.

Tags: aigptllmnewsopenaipricing
Share198Tweet124
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Qwen 3.7 Max is API-only and cannot run locally yet; the open Qwen models (Qwen 3.6 27B, qwen3:8b to 32b) run offline via Ollama.

Qwen 3.7 local: what you can actually run offline

22 June 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.