You opened the Gemini pricing page to check on 3.5 Pro, and instead there's a new Flash sitting where the old one used to be. Gemini 3.6 Flash shipped today, July 21, and the honest headline is simple: it's cheaper than the Flash it replaces, $1.50 in and $7.50 out per million tokens, down from nine dollars on the output side, while scoring higher on every benchmark Google published. We went and pulled the parts you can bank on apart from the marketing, because the price cut and the 17% drop in output tokens are real and verifiable, and the benchmark jumps are Google's own self-reported numbers. Oh, and the model everyone's actually waiting on, 3.5 Pro, still isn't out.
The short answer
Google shipped Gemini 3.6 Flash on July 21. It’s cheaper than the 3.5 Flash it replaces, $1.50 in and $7.50 out per million, and it uses about 17% fewer output tokens for the same job. The benchmark gains are real but they’re Google’s own numbers. Meanwhile the 3.5 Pro flagship everyone’s waiting on is still only in testing, and Gemini 4 pre-training has already started.
The actual win is the price
Let’s lead with the part that doesn’t need a benchmark to trust. Gemini 3.6 Flash costs $1.50 per million input tokens and $7.50 per million output. The Flash it replaces, 3.5, charged $9 per million on the output side. So the new model is straightforwardly cheaper for the thing that usually dominates an agent bill, generation.
And there’s a second cut hiding behind the first. Google says 3.6 Flash produces about 17% fewer output tokens to do the same work, because it reasons in fewer steps and makes fewer tool calls. Lower rate, fewer tokens. If you run anything token-heavy, an agent loop, a batch summarizer, a code assistant that burns output all day, both of those land on the same line of your invoice and compound. That’s the headline for us, and it’s the one nobody has to take on faith.
Honestly, a Flash release that gets cheaper and claims to be smarter is a bit unusual. Usually you pay more for the upgrade. Here Google went the other way, which tells you where the pressure is: mid-tier models from OpenAI and xAI have been squeezing this exact price band, and Google clearly decided to compete on cost per useful token rather than on a bigger number.
The benchmarks look good. They’re also Google’s own.
Now the part to read with one eyebrow up. On every benchmark Google published, 3.6 Flash beats 3.5 Flash, sometimes by a lot.
Coding precision on DeepSWE went from 37% to 49%, which Google frames as fewer unwanted code edits and more production-ready output. The MLE Bench machine-learning research score jumped from 49.7% to 63.9%. Computer use on OSWorld-Verified moved from 78.4% to 83.0%. The GDPval-AA knowledge-work score climbed from 1349 to 1421. Independent of Google, Artificial Analysis put the model at 50 on its Intelligence Index, well over the roughly 31 median for models in this price tier.
Here’s the caveat, and it matters. Those first four numbers are Google testing Google’s model on Google’s harness. That’s normal for a launch, and the gains are plausible given the price behavior. But vendor benchmarks are a best case, run on prompts the vendor picked. We’ve watched enough launches this month (the Qwen and Kimi model claims leaned hard on self-reported tables) to say the same thing every time: wait for someone with no stake to re-run these on the work you actually do. The Artificial Analysis 50 is the one figure here from outside the building, so weight it more.
Flash-Lite and the security model
Google didn’t just ship one model. Alongside 3.6 Flash came Gemini 3.5 Flash-Lite, the cheap-and-fast option: $0.30 per million input, $2.50 per million output, and Google clocks it at 350 output tokens per second. That’s the model you reach for when latency and throughput beat raw smarts, think classification, routing, high-volume extraction.
There’s also a third one, Gemini 3.5 Flash-Cyber, a security-focused model Google is pairing with its CodeMender agent. It isn’t a public release. Google says it’s going only to governments and trusted partners through a limited-access pilot, so for most of us it’s a name on a slide rather than something you can call. We’re noting it exists and moving on.
What this says about 3.5 Pro, and about Gemini 4
Step back and the naming gets weird. Google is now on Gemini 3.6 Flash, while Gemini 3.5 Pro, the actual frontier-tier model, still hasn’t shipped. On the same day it launched 3.6 Flash, Google repeated the same careful line: 3.5 Pro is testing with partners and will be broadly available when it’s ready. No date. Again.
So the Flash line keeps sprinting ahead while the Pro sits in the garage. And then, tucked into the same announcement, Google’s DeepMind team said it has “already started our most ambitious pre-training run yet, for Gemini 4.” Read that next to an unreleased 3.5 Pro and you get the honest subtext: Google is shipping the cheap, fast tier it can win on today, buying time on the flagship, and pointing at a future model to keep the narrative moving. If you need a frontier-grade Google model right now, you still can’t have one. If you need a cheap, fast Google model, this is a genuinely good week.
The honest read
Gemini 3.6 Flash is a real, useful update, and I don’t want the skepticism about benchmarks to bury that. The price cut is verifiable. The token efficiency, if it holds up in your workload, is money. The March 2026 knowledge cutoff is a nice bump over the old January 2025 one. For high-volume, cost-sensitive work, this is an easy model to test.
Just keep the two halves separate in your head. What Google can prove on your invoice: it’s cheaper and lighter on tokens. What Google is asserting: it’s smarter across the board. The first half you can act on today. The second, run it against your own prompts before you rewrite an eval around it. And if you came here for Pro, come back later. That story hasn’t changed since last week.
Sources: Google (official launch post), 9to5Google, OfficeChai and Artificial Analysis, July 21, 2026. Pricing ($1.50 input / $7.50 output per million for 3.6 Flash, $9 output for 3.5 Flash, $0.30/$2.50 for Flash-Lite), the 17% output-token reduction, the March 2026 knowledge cutoff and availability are from Google’s launch post and confirming coverage. The DeepSWE (49% vs 37%), MLE Bench (63.9% vs 49.7%), OSWorld-Verified (83.0% vs 78.4%) and GDPval-AA (1421 vs 1349) figures are Google-reported; the Intelligence Index score of 50 is from Artificial Analysis. Google states 3.5 Pro remains in partner testing with no release date and that Gemini 4 pre-training has begun.
Frequently asked questions
How much does Gemini 3.6 Flash cost?
Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens. That output price is a real cut from Gemini 3.5 Flash, which ran $9 per million output. On top of the lower rate, Google says 3.6 Flash uses about 17% fewer output tokens for the same work, so the effective bill drops further than the headline number suggests.
Is Gemini 3.6 Flash better than 3.5 Flash?
On Google's own published benchmarks, yes, on every one it listed. DeepSWE coding precision went from 37% to 49%, MLE Bench from 49.7% to 63.9%, OSWorld computer use from 78.4% to 83.0%, and the GDPval knowledge-work score from 1349 to 1421. Worth remembering these are Google's numbers on Google's test setup, not independent evaluations, so treat them as the vendor's best case until third parties re-run them.
Where can I use Gemini 3.6 Flash?
It's live in the Gemini app as of July 21, 2026. For developers, access is through the Gemini API in Google AI Studio and Android Studio, plus Google Antigravity and the Gemini Enterprise Agent Platform. Its knowledge cutoff moved forward to March 2026, up from January 2025 on the older model.
What is Gemini 3.5 Flash-Lite?
It's the smaller, faster sibling Google launched alongside 3.6 Flash. Flash-Lite is priced at $0.30 per million input and $2.50 per million output, and Google clocks it at 350 output tokens per second, aimed at high-throughput, low-latency jobs where you want the cheapest capable model rather than the smartest one.
Is Gemini 3.5 Pro out now?
No. Google reiterated on the same day that 3.5 Pro is still testing with partners and will go broadly available when it's ready, with no date. So the Flash line keeps advancing (now on 3.6) while the Pro flagship stays unreleased. Google also confirmed it has started pre-training Gemini 4.