• Latest
  • Trending
  • All
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
Answer card: Sakana Fugu Max at $2 and $6 per million tokens, Fugu Ultra v2 unchanged at $5 and $30, and Sakana saying Ultra v2 scores without Fable 5 or GPT-6 Astra in its pool.

Fugu Max costs $2 and $6 while Fugu Ultra v2 runs without Fable 5

13 September 2026
Answer card stating that DeepSeek released DeepSeek-V4.1-Flash on 10 September 2026 as a 552 billion parameter mixture of experts model with a new causal encoder decoder architecture that activates 8 billion parameters on input and 16 billion on output, with native vision, a one million token context and MIT licensed weights, that the API model name is now deepseek-flash at 0.15 dollars per million input tokens and 0.60 dollars per million output tokens off peak, and that DeepSeek announced V4 Pro would be routed to V4.1-Flash from 14 September and reversed that on 11 September.

DeepSeek V4.1-Flash arrived, and the V4 Pro retirement lasted a day

12 September 2026
Answer card stating that Cognition released SWE-2 on 10 September 2026, a coding model post-trained from Kimi K3, scoring 50.0 percent on FrontierCode 1.1 Main against 50.9 percent for Claude Fable 5.1 and 27.3 percent on Terminal-Bench 4 against 55.8 percent, available only inside Devin.

SWE-2 trails Fable 5.1 by one point, and by 28 on Terminal-Bench 4

11 September 2026
Answer card for Meta Muse, free to 100 million tokens a week then $20 a month, launched 8 September 2026 for United States adults only, running in a dedicated per user virtual machine.

Does Meta Muse do enough to earn your inbox and a card on file?

9 September 2026
Answer card stating that the public download pages for the VMware Virtual Disk Development Kit on developer.broadcom.com began returning 404 errors on 25 August 2026 with no announcement or deprecation notice, that Broadcom support tells customers the kit is no longer available for use or download, and that release lines 7.0.3.1, 8.x and 9.x are all affected.

Broadcom pulled VDDK 8.0 and 9.0, and the 404 is the only notice

8 September 2026
Answer card stating that OpenAI published its research acceleration measurements on 6 September 2026, that as of mid August 2026 its research organisation logged 3.1 agent workdays of coding agent runtime for every workday of human labour normalised to a standard eight hour day, and that OpenAI states this should not be read as a 3.1 times productivity gain because it measures runtime rather than delivered output.

OpenAI’s 3.1 agent-workdays per human day is not a 3.1x gain

7 September 2026
Answer card stating that Mullvad announced on 3 September 2026 that it is shutting down its public encrypted domain name system servers on 2 November 2026 and sponsoring the Quad9 Foundation instead, with 194.242.2.2 and its five sibling addresses all going away, and virtual private network customers unaffected.

Mullvad’s DNS servers go dark on 2 November, and Quad9 blocks no ads

5 September 2026
OpenAI announcement image for GPT-6 Astra, a spiral galaxy of white, blue and amber points of light curling around a bright core on a near black star field.

GPT-6 Astra lists at $10 and $50, 2.5x what GPT-5.6 Sol costs

6 September 2026
Google's official announcement image for the release, reading Introducing Gemini 3.8 Flash and 3.8 Flash Cyber in black type over a pale blue background with a blurred white chevron and the four colour Gemini spark below.

Gemini 3.8 Flash keeps the price and the 1 January cliff

3 September 2026
Answer card stating that Anthropic announced Enterprise Frontier Safeguards on 1 September 2026, that activity data used for misuse monitoring moves into cloud storage the customer controls under the customer own encryption keys, that Anthropic charges nothing for the feature while the cloud provider bills storage and egress, and that the phased rollout starts later in autumn 2026 with interim zero data retention on Fable 5 and Fable 5.1 for eligible customers.

Anthropic moves retention into your own cloud, for 30 days

3 September 2026
  • About
  • Contact
  • Privacy
  • Legal
Thursday, September 17, 2026
  • Login
Packet Nebula
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About
No Result
View All Result
Packet Nebula
No Result
View All Result
Home Dev

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

by stephane
16 September 2026
in Dev
0
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.
491
SHARES
1.4k
VIEWS
Share on FacebookShare on Twitter

Every voice agent we've built on a Live API has the same ugly seam. The caller asks for something, the model has to call a tool, and the line goes quiet while it waits. On 15 September 2026 Google shipped two models aimed at that seam. Gemini 3.8 Live is the cheap conversational one. Gemini 3.8 Live Extended Thinking reasons and runs tools in the background while it keeps talking, and to make that work it refuses any function you declare the old way. We've read the announcement, the pricing page, the model card and the Live API thinking guide, because the announcement is about benchmark rows and the guide is where your code breaks.

The short answer

Both models are live in the Gemini API and Google AI Studio, at the same price: $3 per million audio tokens in and $12 out, which Google also quotes as $0.005 and $0.018 a minute. Text is $0.75 in and $4.50 out. The Extended Thinking model takes a thinkingLevel of low, medium or high, requires behavior: NON_BLOCKING on every function declaration, and errors on a synchronous tool. Limits are 131,072 tokens in and 65,536 out, with no context caching and no structured outputs. Artificial Analysis scores the thinking model 82.6 on its Speech to Speech Quality Index and measures 1.35 seconds to first audio. Every second of output carries a SynthID watermark you can't switch off.

$0.018per minute of audio out, both models, paid tier
82.6Speech to Speech Quality Index, Extended Thinking, per Artificial Analysis
NON_BLOCKINGthe only tool behaviour the thinking model accepts
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 in the Gemini API and Google AI Studio, that both cost 3 dollars per million audio input tokens and 12 dollars per million audio output tokens, about half a cent a minute in and 1.8 cents a minute out, that the Extended Thinking model reasons in the background while speaking and requires every declared function to be asynchronous, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.
Two model IDs, one price, and a tool contract that isn't backward compatible.

What Google shipped, and what it costs

The model IDs are gemini-3.8-live and gemini-3.8-live-extended-thinking, both stable, both dated September 2026 on the model card. Inputs are text, images, audio and video. Output is text and audio. The card lists the input limit at 131,072 tokens and output at 65,536, and it's honest about the gaps: caching not supported, code execution not supported, structured outputs not supported, Batch API not supported. Function calling is marked "Supported (Async only)" on the thinking model, which is the line this whole article hangs on. Search grounding is in, at 5,000 free requests a month shared across the 3.x family and $14 per thousand after that.

Pricing is the same for both. The Gemini API page lists audio at $3.00 per million tokens in and $12.00 out, and next to each figure it prints the per-minute equivalent: $0.005 in, $0.018 out. Image and video input is $1.00, or $0.002 a minute. Text is $0.75 in and $4.50 out. That last number is worth a second look. Gemini 3.8 Flash charges $0.75 and $3 for text today, so text out of a Live session costs half again what the same tokens cost from Flash. Not a lot in absolute terms. It adds up on a transcript-heavy agent.

What hasn't changed is the audio rate. Gemini 2.5 Flash Native Audio, the previous generation on the same page, is $3 in and $12 out too. Google's Live Translate model bills at 25 tokens a second of audio, and the 3.8 Live per-minute figures work out to the same meter: 1,500 tokens a minute times $12 per million is $0.018. So this is a capability release, not a price move. The 3.1 Flash Live Preview is still listed at the same rate, and we couldn't find a retirement date for it on 16 September. I'd expect one.

Horizontal bar chart of audio output prices in dollars per million tokens for speech to speech models, showing GPT-Realtime-2.1 at 64 dollars, Gemini 3.5 Live Translate at 21 dollars, GPT-Realtime-2.1 mini at 20 dollars, Gemini 3.8 Live and 3.8 Live Extended Thinking at 12 dollars and Gemini 2.5 Flash Native Audio at 12 dollars, with a note that Google bills audio at 25 tokens a second so 12 dollars per million works out to about 1.8 cents a minute.
Audio out, per million tokens. The new models inherit the old audio price and undercut the OpenAI flagship by about five to one.

Extended Thinking rewrites your tool contract

Here's the part the launch post glosses. On gemini-3.8-live-extended-thinking you set thinkingConfig.thinkingLevel to low, medium or high. There's no minimal, and the parameter is rejected outright on plain gemini-3.8-live. Fine so far. Then the guide says, in one sentence, that thinking models run tools asynchronously in the background while streaming verbal updates, and that synchronous blocking tools return an error. Every function declaration you send to the thinking model needs "behavior": "NON_BLOCKING". Not recommended. Required.

What you get for that is the behaviour Google is actually selling. The model starts a tool call, keeps talking with a filler ("checking flights to Seattle"), and folds the result in when it lands. The caller never hears the wait. We've faked this ourselves with canned phrases and a timer, and honestly the fake version is worse than it sounds, because the filler ends before the tool does and you're back to silence. Having the model own the gap is the right design. It just means your tool layer has to be genuinely asynchronous. Return immediately and deliver the result as a later functionResponse. And be ready for the model to ask a follow-up before the first answer is back.

The second change is quieter and will bite more people. Your client shouldn't key on turnComplete any more. The thinking model reports an interactionStatus field, IN_PROGRESS while it's reasoning or a tool is still out, IDLE when it's actually ready for the user. The guide is blunt: only return to idle when the status is IDLE. If you unmute the microphone on turnComplete, as every Live API sample since 2025 did, you'll get callers talking over a model that's mid-thought and tool results landing into a turn that's already moved on. That's a two-line fix once you know. It's a week of confused bug reports if you don't.

Checklist of what changes in Live API code when moving to Gemini 3.8 Live Extended Thinking, listing that thinkingLevel accepts low, medium and high but not minimal, that every function declaration needs behavior NON_BLOCKING or the call errors, that the client should wait for interactionStatus IDLE rather than turnComplete, that thinkingLevel is rejected by the plain gemini-3.8-live model, and that neither model supports context caching, structured outputs or the Batch API.
From the thinking guide and the model card. The first item is a feature. The other four are the migration.

Structured outputs being unsupported matters here too. If your agent extracts a booking reference or a postcode from the conversation, you can't ask the model for JSON on this endpoint. You do it in a tool, asynchronously, or you hand the transcript to a text model afterwards. The WebSocket endpoint itself hasn't moved: it's still BidiGenerateContent on generativelanguage.googleapis.com, and Simon Willison had a browser client talking to both models within hours of the release using nothing but the Web Audio API.

The benchmark numbers, and who measured them

Google's post leads with 82.6 on the Speech to Speech Quality Index, first place, and 68.6% on τ-Voice and 35.1% on Sierra's τ-Voice-banking for agentic task completion, plus 97.7% on Big Bench Audio. Those are the Extended Thinking figures. The plain 3.8 Live model gets second place on Speech Agent Arena, which is a human preference vote. We went to the Artificial Analysis leaderboard rather than take the post's word: it shows Extended Thinking at 82.6, plain 3.8 Live at 76.0, and the previous 3.1 Flash Live at 71.5, so about eleven points of headroom over the last generation for the thinking model. The same page measures time to first audio at 1.18 seconds for 3.8 Live and 1.35 for Extended Thinking.

That latency is the number I'd watch. A 1.35 second gap before the first syllable is fine in a support flow and noticeable in anything that's supposed to feel like a conversation. The whole point of the background reasoning is that the gap shouldn't grow when a tool is involved, and nobody outside Google has published a measurement of that yet. I might be wrong, but I'd guess the filler phrases paper over a good deal of it and the tail latencies on high are what you'll end up tuning. Run it on low first.

Two more things for the record. Both models auto-detect and switch between 97 languages mid-call, which is the same claim as 3.8 Flash's text side and a real step up from the 70-plus on Live Translate. And every second of generated audio carries a SynthID watermark. There's no opt-out on the paid tier that we could find, and no public detector either, the same shape as the text watermarking we covered in August. If you're transcribing the other side of the call, Gemini 3.5 Transcribe is the model Google points you at, and its per-minute price is the same half cent as the Live input side. For the OpenAI comparison, our GPT-Realtime-2.1 coverage from July has the $32 and $64 audio figures the chart above uses.

Sources

Google, Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking, 15 September 2026 (the two models, the 82.6, 68.6%, 35.1% and 97.7% figures, the Speech Agent Arena placing, the 97 languages, the availability list, the SynthID watermark). Google, Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe, 15 September 2026 (the $0.005 and $0.018 per-minute prices, the background reasoning and asynchronous function calling description, the partner list). Google AI for Developers, Gemini Developer API pricing, read 16 September 2026 (the $3, $12, $0.75, $4.50 and $1.00 per million figures, the 2.5 Flash Native Audio and Live Translate rows, the 25 tokens a second note, the grounding quota). Google AI for Developers, Gemini 3.8 Live Extended Thinking model card, read 16 September 2026 (the 131,072 and 65,536 limits, the capability table). Google AI for Developers, Thinking in the Live API, read 16 September 2026 (the thinkingLevel values, the NON_BLOCKING requirement, the interactionStatus field). Artificial Analysis, Speech to Speech leaderboard, read 16 September 2026 (the 82.6, 76.0 and 71.5 scores, the 1.18 and 1.35 second latencies). Simon Willison, Gemini Live audio, 15 September 2026 (the browser client and the WebSocket endpoint). The GPT-Realtime-2.1 prices are from our own July coverage.

Frequently asked questions

What's the difference between Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking?

Same price, same limits, same audio. The Extended Thinking model reasons in the background while it speaks and can run tools without pausing the conversation, and it exposes a thinkingLevel setting of low, medium or high. The plain 3.8 Live model has no thinking setting and is meant for high-volume conversational use. On the Artificial Analysis index they score 82.6 and 76.0 respectively.

How much does Gemini 3.8 Live cost per minute?

On the paid tier, $0.005 a minute of audio in and $0.018 a minute of audio out, which is $3 and $12 per million audio tokens. Text is $0.75 in and $4.50 out per million, and image or video input is $1.00 per million or $0.002 a minute. Both 3.8 Live models are priced identically, and the audio rate is unchanged from the older 2.5 Flash Native Audio model.

Why does my existing tool fail on gemini-3.8-live-extended-thinking?

Because the thinking model only accepts asynchronous tools. Every function declaration must carry behavior set to NON_BLOCKING, and the Live API docs state that a synchronous blocking tool returns an error. Your tool should acknowledge immediately and deliver its result as a later function response. Also switch your client from watching turnComplete to waiting for interactionStatus to read IDLE.

Does Gemini 3.8 Live support structured outputs or context caching?

No. The model card marks caching, structured outputs, code execution, URL context and the Batch API as not supported on the Extended Thinking model, and the plain Live model shares the same capability list. Search grounding and audio generation are supported. If you need JSON from the conversation, extract it inside an asynchronous tool or post-process the transcript with a text model.

Is the audio from Gemini 3.8 Live watermarked?

Yes. Google says all audio generated by both models is watermarked with SynthID, its imperceptible audio watermark. There's no opt-out on the API that we could find, and Google hasn't published a public detector for third parties. Treat every second of generated speech as attributable.

Tags: geminiGemini 3.8 LivegoogleLive APInewsspeech to speechvoice agents
Share196Tweet123
stephane

stephane

  • Trending
  • Comments
  • Latest
Answer card: Proton Lumo 2.0 is private by policy, not by locality. Saved history is locked so even Proton cannot read it, but the prompt is decrypted on a Proton EU server to answer it, then forgotten.

Proton Lumo 2.0 review: how private is it, really?

3 September 2026
The Agentic Coding section of the official Hy4 preview benchmark appendix published by Tencent, a table comparing Hy3 and Hy4 preview against DeepSeek V4 Pro 0813, Qwen 3.8 Max, GLM 5.3, Kimi K3, GPT 5.6 Sol and Claude Opus 5 across SWE-bench Multilingual, SWE-bench Pro, DeepSWE, three SWE Atlas tasks, SWE-Marathon, Terminal-Bench 2.1, NL2Repo-Bench, CyberGym, ProgramBench, PostTrainBench and Harbor-Index.

Tencent’s 770B Hy4 tops one benchmark row in 46

3 September 2026
Answer card: Apple released iOS 26.6 and iPadOS 26.6 on 27 July 2026 with a release note covering bug fixes, security updates and an optimized Spotlight index to prepare for iOS 27, the index the rebuilt Siri reads for personal context.

iOS 26.6 is out: the Spotlight index it quietly builds

27 July 2026
Answer card: JWTs are not encrypted, anyone can read them; the signature proves who issued the token, not who may read it.

Are JWTs encrypted? No, and the difference will bite you

0
Answer card: a random 8 character password falls in under 2 hours offline, while 16 random characters hold for 1.4 trillion years at the same speed.

How long does it take to crack a password in 2026?

0
Answer card: three DNS records decide if your mail lands or bounces; SPF lists allowed senders, DKIM signs messages, DMARC sets the failure policy.

SPF, DKIM and DMARC explained: the records your email needs

0
Answer card stating that Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 at 3 dollars per million audio input tokens and 12 dollars out, that the thinking model requires asynchronous tools, and that Artificial Analysis scores it 82.6 on its Speech to Speech Quality Index.

Gemini 3.8 Live Extended Thinking rejects any tool that blocks

16 September 2026
Answer card summarising the Atria Dawn Preview release: 744B GLM-5.2 base, MIT licence, 1.5 TB BF16 and 756 GB FP8 checkpoints, 256K context, top on five of sixteen benchmark rows and trailing on SWE-bench Pro.

Atria Dawn Preview is 744B under MIT, and the BF16 weighs 1.5 TB

15 September 2026
Answer card stating that OpenAI released the Agents API in public beta on 10 September 2026 with no separate fee, billed through model tokens, tool calls and hosted sandbox time, with a choice of OpenAI hosted, self hosted or partner sandboxes, US only data residency and no Zero Data Retention support.

OpenAI’s Agents API has no fee, no ZDR and a one hour sandbox clock

14 September 2026
  • About
  • Contact
  • Privacy
  • Legal

Copyright © 2026 Stephane Cardon.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • Articles
    • Security
    • Network
    • Dev
    • Sysadmin
    • SEO
    • Email & DNS
  • Tools
    • Network tools: free, fast, no signup
    • Security tools: free, fast, no signup
    • Developer tools: free, fast, no signup
    • Sysadmin tools: free, fast, no signup
    • SEO tools: free, fast, no signup
    • Email & DNS tools: free, fast, no signup
  • Download
  • About

Copyright © 2026 Stephane Cardon.