DevNews

Grok Voice Think Fast 2.0: grok-voice-latest costs 60% more

On this page
  1. The switch was announced, the default was not neutral
  2. The number that matters is 37.5%
  3. What you get for the money
  4. What we’d do this afternoon
  5. What isn’t published

Go and look at your xAI bill in September. If your voice agent connects with model=grok-voice-latest, which is the exact string sitting in xAI's own quickstart, it started answering from a different model today at $0.08 a minute instead of $0.05. That's the alias moving from Grok Voice Think Fast 1.0 to 2.0, announced on July 29 with a hard date of August 5, and today is August 5. The upgrade is real. Artificial Analysis has 2.0 at 82.9% on speech-to-speech where 1.0 sat at 75.7%, and time to first audio falls from 1.25 seconds to 0.70. But the rate went up 60% and the window to pin the old model closed this morning. So the question we'd actually ask a colleague is whether the calls got short enough to absorb it. We ran that number.

The short answer

xAI’s grok-voice-latest alias now resolves to Grok Voice Think Fast 2.0 instead of 1.0. The new model is genuinely stronger and noticeably faster off the mark, at 0.70 seconds to first audio against 1.25. It also costs $4.80 an hour instead of $3.00. Pinning grok-voice-think-fast-1.0 still works and still bills at the old rate. Net for you: if you use the alias, check your call volume today, because nothing in your code changed and your unit cost did.

Aug 5alias moved to 2.0
+60%$0.05 to $0.08 a minute
82.9%speech-to-speech, up from 75.7
Official xAI announcement artwork for Grok Voice Think Fast 2.0, the model name set over the dark xAI release graphic. Image: xAI, announcement artwork from the Grok Voice Think Fast 2.0 release post.

Most model upgrades ask you to change a string. This one asked you to change a string to avoid it.

Answer card: on 5 August 2026 xAI moved the grok-voice-latest alias from Grok Voice Think Fast 1.0 to 2.0, raising the audio rate from $0.05 to $0.08 a minute while speech-to-speech scores rise from 75.7 to 82.9 percent.
A better model, a faster first syllable, and a 60% jump in the rate that bills you. PNG

The switch was announced, the default was not neutral

Credit where it’s due: xAI did this in the open. The changelog entry went up on July 29 and said plainly that on August 5 the alias would move from grok-voice-think-fast-1.0 to grok-voice-think-fast-2.0, and that anyone wanting to stay on 1.0 should pin the explicit string first. A week of notice, a named date, no ambiguity.

The friction is that xAI’s own voice quickstart connects like this:

wss://api.x.ai/v1/realtime?model=grok-voice-latest

So the documented happy path is the alias. Anyone who followed the guide, shipped, and moved on is on the new model this morning, paying the new rate, having read nothing. That isn’t a trick. It’s just what “latest” means, and it’s the second time in a fortnight we’ve written about a model swapping out from under a stable identifier, after DeepSeek replaced V4-Flash in place behind the same name. The difference is that DeepSeek held the price. xAI didn’t.

The number that matters is 37.5%

Here’s the arithmetic nobody seems to be doing. Grok Voice bills per minute of audio, not per token. Think Fast 1.0 is $0.05 a minute, or $3.00 an hour. Think Fast 2.0 is $0.08 a minute, or $4.80 an hour. Both figures come straight off xAI’s model docs.

For your monthly bill to come out the same after the switch, your calls have to get shorter by the ratio between those rates. That’s 0.05 divided by 0.08, which is 0.625. Your average call needs to lose 37.5% of its length.

Now put the latency win against that. Time to first audio drops 0.55 seconds, from 1.25 to 0.70. Say a five minute call with twenty assistant turns, which is a generous turn count for a support call. That’s eleven seconds saved out of three hundred. Under 4%.

It’s a real improvement and callers will feel it. It does not come close to paying for the rate. To break even on latency alone you’d need something like two hundred turns inside a five minute call, which isn’t a phone call, it’s a stutter.

Bar chart comparing Grok Voice hourly rates and a thousand-hour monthly bill: Think Fast 1.0 at $3.00 an hour and $3,000, Think Fast 2.0 at $4.80 an hour and $4,800.
A thousand hours of calls a month is where the change stops being abstract. Rates are xAI's, the monthly totals are ours. PNG

At a thousand hours of calls a month, a mid-sized support line, that’s $3,000 becoming $4,800. An extra $1,800 a month, $21,600 a year, for a model you did not choose and a config you did not edit.

One more wrinkle worth naming. xAI says 2.0 uses roughly 60% fewer reasoning tokens than 1.0. That’s a genuine efficiency gain, but read where it lands: when billing is a flat rate per minute of audio, cheaper thinking shows up in xAI’s margin, not on your invoice. The only path from that efficiency to your bill runs through shorter calls, and we’ve just seen how much shorter they’d have to get. That reading is ours, not xAI’s.

What you get for the money

Enough grumbling about the rate, because the model did move. xAI’s announcement attributes its headline numbers to Artificial Analysis’ speech-to-speech benchmark, and they’re good.

Overall, 82.9% against 75.7% for 1.0. That’s a big jump for a point release, and it puts 2.0 ahead of gpt-realtime-2.1 at 79.1% and well clear of Gemini 3.1 Flash at 69.5%. The agentic score, the one that tracks whether the thing can actually call your tools without falling over mid-sentence, goes 52.1% to 56.5%, against 45.7% for GPT-Realtime-2.1. If you’re building a voice agent that books appointments rather than one that chats, that second number is the one to care about.

Transcription is where xAI pushes hardest: 1.5 to 2.0 times more accurate than Deepgram Nova 3 and ElevenLabs Scribe v2, 1.4 times better than its own 1.0, and around ten times better once you add background noise and telephony compression. Telephony compression is the honest detail there, because that’s the actual condition of a real support call and it’s where most speech stacks quietly fall apart.

Bar chart of Artificial Analysis speech-to-speech and agentic scores: Think Fast 2.0 at 82.9 and 56.5, Think Fast 1.0 at 75.7 and 52.1, GPT-Realtime-2.1 at 79.1 and 45.7, Gemini 3.1 Flash at 69.5 and 37.7.
The agentic column is the wider gap, and it's the one that matters if your agent calls tools. PNG

And the line nobody is quoting: on the conversational benchmark, Think Fast 2.0 scores 95.1% against 95.7% for GPT-Realtime-2.1. It loses that one. Small margin, and xAI published it anyway, which is more than a lot of announcements manage. But if your agent’s job is to hold a natural conversation rather than drive an API, the case for switching gets thinner fast.

Language coverage is the one place the reporting doesn’t line up. We’ve seen 24 languages in one write-up of the announcement and 25 or more in another, and we couldn’t resolve it against a first-party list. Two dozen or so is the safe read.

What we’d do this afternoon

Grep your codebase for grok-voice-latest. That’s the whole first step, and it takes a minute. If it’s there, you’re on 2.0 and your unit cost is up 60% as of today.

Then decide deliberately instead of by default. Pin grok-voice-think-fast-1.0 if your agent is mostly conversational, your callers aren’t complaining about lag, and volume is high enough that $1.80 an hour compounds into something you’d have to explain. Take 2.0 if you’re doing agentic work with tool calls, or if you’re transcribing noisy phone audio, because that’s where the gap is genuinely large.

Honestly, I’d take 2.0 for anything touching tools and pin 1.0 for the rest, and I might be wrong about that if your call volume is small enough that the whole argument is worth forty dollars a month. Whichever way you go, pin something. Running production voice on an alias means your unit economics are a xAI changelog entry away from changing again, and there’s no reason the next move has to be announced a week ahead.

Checklist of what xAI published about Grok Voice Think Fast 2.0 versus what it left unstated, covering pricing, benchmark provenance, language count and independent evaluation.
Published, and still open. The benchmark provenance is the row we'd watch. PNG

What isn’t published

No independent evaluation exists yet that we can find. Every number above traces back to xAI’s own announcement, including the ones credited to Artificial Analysis, and we could not pull a first-party benchmark page to confirm them. There’s no deprecation date for Think Fast 1.0, so today’s rate on the pinned model is current but not promised. There’s no context window published for either voice model, no per-language accuracy breakdown behind the “1.5 to 2.0 times” claim, and no statement about whether the old rate is grandfathered for existing volume commitments.

That last one is worth an email to your account contact if you’re spending real money here.

Sources: xAI’s model and pricing docs and voice guide, the Grok Voice Think Fast 2.0 announcement, the xAI changelog entry of July 29 as collected by Releasebot, plus benchmark reporting from TestingCatalog. Hourly rates, the 37.5% break-even and the thousand-hour totals are our arithmetic on xAI’s published per-minute prices.

Frequently asked questions

What happened to grok-voice-latest on August 5, 2026?

The alias stopped resolving to grok-voice-think-fast-1.0 and started resolving to grok-voice-think-fast-2.0. xAI announced the change on July 29 with August 5 named as the switchover date, and told developers to pin grok-voice-think-fast-1.0 explicitly before then if they wanted to stay put. If you never pinned it, your sessions are on 2.0 now with no code change on your side.

How much does Grok Voice Think Fast 2.0 cost?

xAI's model docs list $0.08 per minute of audio, which works out to $4.80 an hour, plus the same separate text input charge that applied before. Think Fast 1.0 is still listed at $0.05 per minute, or $3.00 an hour. That is a 60% increase on the audio line, which is the line that dominates a voice agent bill.

Is Grok Voice Think Fast 2.0 actually better than 1.0?

On the numbers xAI published, yes. Artificial Analysis speech-to-speech goes from 75.7% to 82.9%, the agentic score from 52.1% to 56.5%, and time to first audio from 1.25 seconds to 0.70. Transcription is reported as 1.4 times more accurate than 1.0 across roughly two dozen languages. All of it comes from xAI's own announcement, so treat it as vendor-reported until someone runs it independently.

Does the lower latency pay for the higher price?

Not on our arithmetic. Billing is per minute of audio, so to hold your bill flat at $0.08 instead of $0.05 your calls have to get 37.5% shorter. The latency win is 0.55 seconds per turn. On a five minute call with twenty assistant turns that is about eleven seconds, under 4%. The speed is worth having. It is not worth 60% on its own.

Can I still use Grok Voice Think Fast 1.0?

Yes. The model is still listed in xAI's docs at its old $0.05 per minute rate, so pinning the explicit grok-voice-think-fast-1.0 string keeps both the old behaviour and the old price. What you cannot do is get there by accident any more. The default alias now points at the newer, dearer model, so staying on 1.0 is an active choice you have to write down.