• Latest
  • Trending
  • All
Answer card: Google shipped Gemini 3.6 Flash on July 21, 2026 at $1.50 per million input tokens and $7.50 per million output tokens, down from nine dollars output on 3.5 Flash, using 17 percent fewer output tokens and scoring higher on Google-reported benchmarks, while Gemini 3.5 Pro is still only in testing with partners.

Gemini 3.6 Flash is cheaper than 3.5 and scores higher

3 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
Google's official announcement image for the release, reading Introducing Gemini 3.8 Flash and 3.8 Flash Cyber in black type over a pale blue background with a blurred white chevron and the four colour Gemini spark below.

Gemini 3.8 Flash keeps the price and the 1 January cliff

3 September 2026
Answer card stating that Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, that activity data used for misuse monitoring moves into cloud storage the customer controls under the customer own encryption keys, that Anthropic charges nothing for the feature while the cloud provider bills storage and egress, and that the phased rollout starts later in autumn 2026 with interim zero data retention on Fable 5 and Fable 5.1 for eligible customers.

Anthropic moves retention into your own cloud, for 30 days

3 September 2026
Official Google diagram of a client connection in three numbered steps: a DNS lookup with a query and an address, a TLS ClientHello and ServerHello, then a content exchange with a website. A callout on the DNS step reads 25% of global web traffic is now protected by encrypted DNS, and a callout beside an Android phone on the ClientHello step reads Android 17 supports ECH GREASE by default.

Android 17 hides the SNI, not your DNS or destination

3 September 2026
Still frame from the Claude Fable 5.1 launch video showing model-designed protein binders in orange docked against twelve grey target proteins, rendered as ESMFold2 structure predictions.

Claude Fable 5.1 breaks forced tool use, cuts cache 75%

1 September 2026
Answer card stating that on 31 August 2026 the European Commission designated ChatGPT a Very Large Online Search Engine under the Digital Services Act, the first conversational AI service classified that way, because it answers user prompts and queries including by searching the web, with OpenAI having declared roughly 159.1 million average monthly users in the European Union for ChatGPT search.

The EU now calls ChatGPT a very large search engine

3 September 2026
Answer card stating that on 31 August 2026 the Department of War added OpenAI ChatGPT Mil and Starshield AI Grok for Government to the GenAI.mil portal alongside Google Gemini, all three accredited at Impact Level 5 for Controlled Unclassified Information, with 1.7 million unique users onboarded out of roughly 3 million eligible personnel, and ChatGPT Mil currently serving GPT-5.4 Terra with GPT-5.6 Terra said to be rolling out.

ChatGPT Mil and Grok reached IL5 on GenAI.mil

3 September 2026
Answer card stating that Anthropic opened a research preview of the Model Hardware Standard on 27 August 2026, standardising the driver layer between an operating system and a laboratory instrument with read and write primitives plus discovery and safety limits, reachable through MCP as well as a command line and code files, with no public specification published.

Anthropic’s Model Hardware Standard is gated, and sits under MCP

3 September 2026
Official Cohere key art for the Parse 5 launch: the Cohere mark and the wordmark Parse with a superscript 5 in white, centred on a soft out of focus gradient of deep blue, violet and amber curves.

Cohere Parse 5 is $1.50 per 1,000 pages, on three of five dimensions

3 September 2026
Title card from the OpenAI announcement video: a man sits on a blue sofa in a loft with tall windows and potted plants, a laptop open on the coffee table in front of him, with the words WebMCP in ChatGPT in large white type across the lower left.

WebMCP in ChatGPT needs GPT-5.6 Sol or Terra

3 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Tuesday, September 8, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Gemini 3.6 Flash is cheaper than 3.5 and scores higher

by stephane
3 September 2026
in Dev
0
Answer card: Google shipped Gemini 3.6 Flash on July 21, 2026 at $1.50 per million input tokens and $7.50 per million output tokens, down from nine dollars output on 3.5 Flash, using 17 percent fewer output tokens and scoring higher on Google-reported benchmarks, while Gemini 3.5 Pro is still only in testing with partners.
491
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

You opened the Gemini pricing page to check on 3.5 Pro, and instead there's a new Flash sitting where the old one used to be. Gemini 3.6 Flash shipped today, July 21, and the honest headline is simple: it's cheaper than the Flash it replaces, $1.50 in and $7.50 out per million tokens, down from nine dollars on the output side, while scoring higher on every benchmark Google published. We went and pulled the parts you can bank on apart from the marketing, because the price cut and the 17% drop in output tokens are real and verifiable, and the benchmark jumps are Google's own self-reported numbers. Oh, and the model everyone's actually waiting on, 3.5 Pro, still isn't out.

The short answer

Google shipped Gemini 3.6 Flash on July 21. It’s cheaper than the 3.5 Flash it replaces, $1.50 in and $7.50 out per million, and it uses about 17% fewer output tokens for the same job. The benchmark gains are real but they’re Google’s own numbers. Meanwhile the 3.5 Pro flagship everyone’s waiting on is still only in testing, and Gemini 4 pre-training has already started.

$7.50per 1M output, was $9
17%fewer output tokens
still MIAGemini 3.5 Pro
Answer card: Google shipped Gemini 3.6 Flash on July 21, 2026 at $1.50 per million input tokens and $7.50 per million output tokens, down from nine dollars output on 3.5 Flash, using 17 percent fewer output tokens and scoring higher on Google-reported benchmarks, while Gemini 3.5 Pro is still only in testing with partners.
The one-card version. Cheaper and lighter on tokens are the parts you can verify from your own invoice.

The actual win is the price

Let’s lead with the part that doesn’t need a benchmark to trust. Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output. The Flash it replaces, 3.5, charged $9 per million on the output side. So the new model is straightforwardly cheaper for the thing that usually dominates an agent bill, generation.

And there’s a second cut hiding behind the first. Google says 3.6 Flash produces about 17% fewer output tokens to do the same work, because it reasons in fewer steps and makes fewer tool calls. Lower rate, fewer tokens. If you run anything token-heavy, an agent loop, a batch summarizer, a code assistant that burns output all day, both of those land on the same line of your invoice and compound. That’s the headline for us, and it’s the one nobody has to take on faith.

Honestly, a Flash release that gets cheaper and claims to be smarter is a bit unusual. Usually you pay more for the upgrade. Here Google went the other way, which tells you where the pressure is: mid-tier models from OpenAI and xAI have been squeezing this exact price band, and Google clearly decided to compete on cost per useful token rather than on a bigger number.

The benchmarks look good. They’re also Google’s own.

Now the part to read with one eyebrow up. On every benchmark Google published, 3.6 Flash beats 3.5 Flash, sometimes by a lot.

Coding precision on DeepSWE went from 37% to 49%, which Google frames as fewer unwanted code edits and more production-ready output. The MLE Bench machine-learning research score jumped from 49.7% to 63.9%. Computer use on OSWorld-Verified moved from 78.4% to 83.0%. The GDPval-AA knowledge-work score climbed from 1349 to 1421. Independent of Google, Artificial Analysis put the model at 50 on its Intelligence Index, well over the roughly 31 median for models in this price tier.

Comparison bar chart of Gemini output pricing per million tokens: Gemini 3.5 Flash at nine dollars, the new Gemini 3.6 Flash at seven dollars fifty, and Gemini 3.5 Flash-Lite at two dollars fifty, showing 3.6 Flash undercutting the model it replaces.
Output price per million tokens. The new Flash is cheaper than the old one, and Flash-Lite sits well below both.

Here’s the caveat, and it matters. Those first four numbers are Google testing Google’s model on Google’s harness. That’s normal for a launch, and the gains are plausible given the price behavior. But vendor benchmarks are a best case, run on prompts the vendor picked. We’ve watched enough launches this month (the Qwen and Kimi model claims leaned hard on self-reported tables) to say the same thing every time: wait for someone with no stake to re-run these on the work you actually do. The Artificial Analysis 50 is the one figure here from outside the building, so weight it more.

Flash-Lite and the security model

Google didn’t just ship one model. Alongside 3.6 Flash came Gemini 3.5 Flash-Lite, the cheap-and-fast option: $0.30 per million input, $2.50 per million output, and Google clocks it at 350 output tokens per second. That’s the model you reach for when latency and throughput beat raw smarts, think classification, routing, high-volume extraction.

There’s also a third one, Gemini 3.5 Flash-Cyber, a security-focused model Google is pairing with its CodeMender agent. It isn’t a public release. Google says it’s going only to governments and trusted partners through a limited-access pilot, so for most of us it’s a name on a slide rather than something you can call. We’re noting it exists and moving on.

What this says about 3.5 Pro, and about Gemini 4

Step back and the naming gets weird. Google is now on Gemini 3.6 Flash, while Gemini 3.5 Pro, the actual frontier-tier model, still hasn’t shipped. On the same day it launched 3.6 Flash, Google repeated the same careful line: 3.5 Pro is testing with partners and will be broadly available when it’s ready. No date. Again.

Checklist splitting what is confirmed about Gemini 3.6 Flash from what to hold loosely: the price cut to 7.50 dollars output, the 17 percent token reduction, the app and API availability and the March 2026 knowledge cutoff are confirmed, while the coding and research benchmark gains are Google-reported and 3.5 Pro remains unreleased.
Two things you can verify from your own bill, four numbers that are Google's until someone else checks, and one flagship still missing.

So the Flash line keeps sprinting ahead while the Pro sits in the garage. And then, tucked into the same announcement, Google’s DeepMind team said it has “already started our most ambitious pre-training run yet, for Gemini 4.” Read that next to an unreleased 3.5 Pro and you get the honest subtext: Google is shipping the cheap, fast tier it can win on today, buying time on the flagship, and pointing at a future model to keep the narrative moving. If you need a frontier-grade Google model right now, you still can’t have one. If you need a cheap, fast Google model, this is a genuinely good week.

The honest read

Gemini 3.6 Flash is a real, useful update, and I don’t want the skepticism about benchmarks to bury that. The price cut is verifiable. The token efficiency, if it holds up in your workload, is money. The March 2026 knowledge cutoff is a nice bump over the old January 2025 one. For high-volume, cost-sensitive work, this is an easy model to test.

Just keep the two halves separate in your head. What Google can prove on your invoice: it’s cheaper and lighter on tokens. What Google is asserting: it’s smarter across the board. The first half you can act on today. The second, run it against your own prompts before you rewrite an eval around it. And if you came here for Pro, come back later. That story hasn’t changed since last week.

Sources: Google (official launch post), 9to5Google, OfficeChai and Artificial Analysis, July 21, 2026. Pricing ($1.50 input / $7.50 output per million for 3.6 Flash, $9 output for 3.5 Flash, $0.30/$2.50 for Flash-Lite), the 17% output-token reduction, the March 2026 knowledge cutoff and availability are from Google’s launch post and confirming coverage. The DeepSWE (49% vs 37%), MLE Bench (63.9% vs 49.7%), OSWorld-Verified (83.0% vs 78.4%) and GDPval-AA (1421 vs 1349) figures are Google-reported; the Intelligence Index score of 50 is from Artificial Analysis. Google states 3.5 Pro remains in partner testing with no release date and that Gemini 4 pre-training has begun.

Frequently asked questions

How much does Gemini 3.6 Flash cost?

Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens. That output price is a real cut from Gemini 3.5 Flash, which ran $9 per million output. On top of the lower rate, Google says 3.6 Flash uses about 17% fewer output tokens for the same work, so the effective bill drops further than the headline number suggests.

Is Gemini 3.6 Flash better than 3.5 Flash?

On Google's own published benchmarks, yes, on every one it listed. DeepSWE coding precision went from 37% to 49%, MLE Bench from 49.7% to 63.9%, OSWorld computer use from 78.4% to 83.0%, and the GDPval knowledge-work score from 1349 to 1421. Worth remembering these are Google's numbers on Google's test setup, not independent evaluations, so treat them as the vendor's best case until third parties re-run them.

Where can I use Gemini 3.6 Flash?

It's live in the Gemini app as of July 21, 2026. For developers, access is through the Gemini API in Google AI Studio and Android Studio, plus Google Antigravity and the Gemini Enterprise Agent Platform. Its knowledge cutoff moved forward to March 2026, up from January 2025 on the older model.

What is Gemini 3.5 Flash-Lite?

It's the smaller, faster sibling Google launched alongside 3.6 Flash. Flash-Lite is priced at $0.30 per million input and $2.50 per million output, and Google clocks it at 350 output tokens per second, aimed at high-throughput, low-latency jobs where you want the cheapest capable model rather than the smartest one.

Is Gemini 3.5 Pro out now?

No. Google reiterated on the same day that 3.5 Pro is still testing with partners and will go broadly available when it's ready, with no date. So the Flash line keeps advancing (now on 3.6) while the Pro flagship stays unreleased. Google also confirmed it has started pre-training Gemini 4.

Tags: aibig-techgeminigooglellmnews
Share196Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
Answer card: Apple released iOS 26.6 and iPadOS 26.6 on 27 July 2026 with a release note covering bug fixes, security updates and an optimized Spotlight index to prepare for iOS 27, the index the rebuilt Siri reads for personal context.

iOS 26.6 is out: the Spotlight index it quietly builds

27 July 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.