DevNews

Gemini 3.5 Transcribe: $0.005 a minute, fifth on WER

On this page
  1. What actually shipped
  2. The benchmark says two different things
  3. What it costs, against the obvious alternatives
  4. The part nobody is calling a tradeoff
  5. Who should move
  6. Sources

Google published a chart on 26 August showing its new transcription model beating everyone else, and it is a perfectly honest chart. Then we opened the other benchmark, the one Google cites in the same post for its headline accuracy number, and Gemini 3.5 Transcribe sits fifth. Both things are true. The model went into public preview that day as gemini-3.5-transcribe for recorded audio and gemini-3.5-transcribe-live for streaming, at roughly half a cent a minute, and it is genuinely good. It is just not the most accurate speech model on the market, and the more interesting question is not accuracy at all. It is that this thing quietly edits what you said.

The short answer

Public preview since 26 August, in two flavours, one for files and one for live audio. Cheap, fast, multilingual. Not the most accurate model you can buy, and it rewrites your sentences on purpose.

$0.30per hour of recorded audio
2.6%word error rate, non streaming
5thon the board Google cites
Answer card stating that Gemini 3.5 Transcribe entered public preview on 26 August 2026 as two models, gemini-3.5-transcribe for recorded audio at about 0.005 dollars per minute and gemini-3.5-transcribe-live for streaming at about 0.009 dollars per minute, with a 2.6 percent non streaming word error rate that places it fifth on the Artificial Analysis leaderboard, and over 85 languages detected automatically.
Two model ids, two price points, one benchmark that needs reading twice. PNG

What actually shipped

Two model ids, and the split matters more than the marketing does.

gemini-3.5-transcribe takes recorded audio through the Interactions API. gemini-3.5-transcribe-live runs streaming through the Live API. Both went into public preview on 26 August 2026 in Google AI Studio and Google Antigravity, plus the Gemini Enterprise Agent Platform for the enterprise side. Chrome gets it for voice input into web fields at some unstated point, and Gemini Enterprise for Customer Experience is on the same soon list.

Some of it is already in front of real users. The Gemini app on macOS uses it for Speak to Window. Rambler on Gboard runs on it, in a subset of countries and languages Google has not enumerated.

Features worth knowing: automatic detection across 85 plus languages so you skip the language hint, custom vocabulary for the words your industry made up, and function calling so a transcript can hand work off to another Gemini model mid stream. Speaker separation tops out at three, and Google marks anything past three as experimental. That ceiling will disqualify it for plenty of meeting tooling.

The speed claim is the cleanest number in the announcement. Time to final transcription improves by 70 percent against Chirp 3, Google previous speech model. That is Google measuring itself against itself, which is the one comparison a vendor cannot really tilt.

The benchmark says two different things

Here is where it gets interesting.

Google published its own FLEURS chart with the announcement. Streaming speech recognition, top locales, lower is better.

Google official FLEURS benchmark chart titled Measuring streaming speech recognition accuracy across languages, lower is better, showing Gemini 3.5 Transcribe Live at 5.50 percent word error rate in blue, Google Cloud Chirp 3 at 7.32 percent, OpenAI GPT Live Transcribe at 8.97 percent, ElevenLabs Scribe v2 Realtime at 9.70 percent and Deepgram Nova-3 at 15.77 percent.

Image: Google, FLEURS chart published with the Gemini 3.5 Transcribe announcement, 26 August 2026.

Gemini Live lands at 5.50 percent against 7.32 for Chirp 3 and 15.77 for Deepgram Nova-3. Real gap, public multilingual test set, competition named rather than anonymised. Credit where it is due.

Now the other number. In the same post Google reports an average word error rate of 4.0 percent streaming and 2.6 percent non streaming, and attributes it to Artificial Analysis. So we went and read that board on 27 August. Gemini 3.5 Transcribe is fifth on it. Fun-Realtime-ASR-preview sits at 1.7 percent, ElevenLabs Scribe v2 at 2.2, then Microsoft MAI-Transcribe-1.5 and Smallest AI Pulse Pro tied at 2.4.

Checklist comparing two benchmark readings for Gemini 3.5 Transcribe: it wins Google own FLEURS streaming chart at 5.50 percent against Chirp 3 at 7.32 percent, but places fifth on the Artificial Analysis word error rate board at 2.6 percent behind Fun-Realtime-ASR-preview at 1.7 percent and ElevenLabs Scribe v2 at 2.2 percent, with the note that the FLEURS chart covers 28 locales in streaming mode while the Artificial Analysis figure is non streaming.
Same model, same week, two boards. Neither one is lying. PNG

Both readings are honest, and they measure different things. The FLEURS chart is streaming mode across 28 locales. The 2.6 percent is non streaming, on the Artificial Analysis audio mix. Pick the board that resembles your audio, not the one in the press post.

What we would not repeat is the line that Google now has the most accurate speech model going. It doesn’t, on the public board it quoted.

What it costs, against the obvious alternatives

Google prices it as a Gemini model, per million tokens, with a per minute equivalent printed beside it. Recorded audio comes to about $0.003 in and $0.002 out per minute, call it $0.005 a minute or $0.30 an hour. Streaming runs $0.005 in and $0.004 out, so around $0.009 a minute, $0.54 an hour. Free tier on both while they are in preview.

Bar chart of speech to text API cost in dollars per hour of recorded audio, lower is better, showing OpenAI gpt-4o-mini-transcribe at 0.18 dollars, ElevenLabs Scribe v2 at 0.22 dollars, Gemini 3.5 Transcribe at about 0.30 dollars, and OpenAI gpt-4o-transcribe and Whisper at 0.36 dollars, with a note that streaming rates run higher on every vendor.
Dollars per hour of audio. Cheap, not cheapest. PNG

So it undercuts OpenAI gpt-4o-transcribe and Whisper, both at $0.006 a minute. It does not undercut gpt-4o-mini-transcribe at $0.003, and it does not undercut ElevenLabs Scribe v2 at $0.22 an hour, which also happens to sit above it on that Artificial Analysis board. That combination is the awkward one for Google. Cheaper and more accurate, from a much smaller company.

Streaming is where the pricing looks best. Gemini Live at roughly $0.54 an hour against OpenAI live transcription at $0.017 a minute, which is $1.02. Half price. ElevenLabs Scribe v2 Realtime is still under it at $0.39, mind.

One habit worth carrying over: Google preview pricing is not a promise. We watched Gemini 3.7 Flash double its price on a published date earlier this year. Budget on the paid rate rather than the free tier, and read the page again when preview ends.

The part nobody is calling a tradeoff

Read Google description of what the model does with your words. It handles self corrections. It removes filler. It auto formats. Say “the meeting is at four, sorry, at five” and you get five.

For dictation that is wonderful. Honestly it is the reason Rambler works at all, and it is why this feels different from bolting Whisper onto a text box. The model isn’t writing down sounds, it is working out what you meant.

Now picture that behaviour on a deposition. Or a clinical intake. Or a research interview where the disfluency is the data, and where a self correction is exactly the thing the analyst cares about. A transcription model that improves your sentences has made an editorial decision on every line, and nothing in the output marks where it did.

I might be over worrying this. Google may well expose a verbatim switch before general availability, and preview is exactly when that kind of thing gets added. But it isn’t documented today, so if the transcript is ever going to be evidence, test that specific behaviour before you wire it in. Feed it a recording where somebody misspeaks and corrects themselves, then read what comes back.

Who should move

If you are already on Chirp 3, move when preview stabilises. Better error rate, 70 percent faster to final, same vendor and the same bill.

If you are running a voice agent and latency is the constraint, this deserves a bake off, and it slots into the same architecture we walked through for OpenAI gpt-realtime-2.1 voice agents. Half the streaming cost is not nothing at volume.

If you just want the lowest error rate on English recordings, the board says look elsewhere first. And if you need verbatim, none of the above matters until you have tested for it.

Preview means preview, anyway. Model ids move, prices land, behaviour shifts. Fine for a prototype this week. I wouldn’t migrate a production pipeline onto it yet.

Sources

Announcement, feature list, FLEURS chart and the 70 percent latency claim: Intelligent transcription with Gemini 3.5 Transcribe on the Google blog, 26 August 2026. Product surfaces and launch detail confirmed against 9to5Google. Prices for both model ids were read from the Gemini Developer API pricing page on 27 August 2026. Comparison prices come from the vendors own pages, OpenAI API pricing and ElevenLabs API pricing, read the same day. The word error rate ranking is the Artificial Analysis speech to text leaderboard, also read on 27 August 2026, and positions on it change.

Frequently asked questions

What does Gemini 3.5 Transcribe cost?

Google prices two models. For recorded audio, gemini-3.5-transcribe reads $2.00 per million audio input tokens, which the pricing page equates to about $0.003 per minute, with text output at $12.00 per million or roughly $0.002 per minute. That lands near $0.005 a minute, so about $0.30 an hour of audio. The streaming model, gemini-3.5-transcribe-live, is $3.50 in and $21.00 out per million, listed as $0.005 and $0.004 per minute, so around $0.009 a minute or $0.54 an hour. Both carry a free tier while the models sit in preview.

Is Gemini 3.5 Transcribe the most accurate speech-to-text model?

No, and Google does not quite claim that. On the Artificial Analysis word error rate board that Google cites for its own 2.6 percent figure, we read it in fifth place on 27 August 2026, behind Fun-Realtime-ASR-preview at 1.7 percent, ElevenLabs Scribe v2 at 2.2 percent, and both Microsoft MAI-Transcribe-1.5 and Smallest AI Pulse Pro at 2.4 percent. On Google own FLEURS streaming chart it wins comfortably at 5.50 percent. Different test sets, different answers.

Does it transcribe verbatim?

Deliberately not. Google describes smart transcription that handles self corrections, strips filler words and auto formats the result. Say a number wrong then correct yourself, and you get the corrected version rather than both. Excellent for dictation and meeting notes. If your use case is legal, medical, research interviews or subtitles that have to match the audio, a model that improves your sentences becomes a liability rather than a feature, so test that behaviour before committing.

How many speakers can it separate?

Up to three, with Google flagging anything beyond three as experimental. If you are transcribing a five person meeting and need reliable speaker labels on every turn, that ceiling is the thing to check first, ahead of the error rate.

Which languages does it handle?

Google says it detects and transcribes over 85 languages automatically, so you do not have to declare the language up front. The FLEURS chart in the announcement is scoped to a top locales subset of 28, which is not the same as claiming even quality across all 85. Test your own languages.