All News
googlegeminivoice-airealtimebenchmarks

Gemini 3.8 Live: Google's new voice models put the best scores on the pricier tier

Google's Gemini 3.8 Live and Live Extended Thinking land for voice agents. The flagship benchmark rows belong to the expensive variant, not the cheap one.

Vlad MakarovVlad Makarovreviewed and published
5 min read
Gemini 3.8 Live: Google's new voice models put the best scores on the pricier tier

Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15, 2026, describing both as "our most advanced live dialogue models yet." Tom Ouyang and Malini Jaganathan signed the launch post on behalf of the Gemini Audio team. The same day brought the models to Search Live, Gemini Live, the Gemini API and AI Studio, with enterprise access still behind a private preview.

Two models, two jobs

This is a split release, and the split matters more than the naming. Gemini 3.8 Live is "built for scale and cost efficiency," pairing conversational intelligence with fluid dialogue and visual grounding; it is the model behind the Search Live experience. Gemini 3.8 Live Extended Thinking is "built for high-complexity tasks, with increased intelligence and multi-step reasoning."

That arrangement lets Google sell one story to two audiences at once. Cost-conscious developers get the cheaper sibling, while the marketing numbers come from the variant most of them will not deploy by default. It is a reasonable product decision and a confusing way to read a headline.

Talking while thinking

The technical pitch leans on latency and overlap. Google says the models do near-real-time reasoning for voice agents, and that 3.8 Live processes visual inputs in near real time and "automatically detects and transitions between 97 supported languages mid-conversation." It executes tools and API calls in the background while continuing the conversation, so the model can acknowledge a request and keep chatting while the task finishes.

Extended Thinking attacks the same problem from the other side: it "reasons and speaks simultaneously," using early verbal cues like "Let me check that..." to acknowledge a prompt, and narrates progress through multi-step background tasks. Filler that once signalled a stall becomes a status report.

Google's demos run through an employee-onboarding agent that answers live questions with visual context, a model playing chess in near real time, raw sketches plus voice feedback turning into functional React components, and restaurant bookings coordinated through asynchronous function calls. Demos are chosen to work, so treat them as feasibility evidence rather than performance evidence.

The benchmark rows Google chose

Every figure below appears on Google's launch post. These are vendor-selected benchmark rows, and they mix third-party indices with runs Google controlled.

BenchmarkResultWho runs it
Artificial Analysis Speech to Speech Quality Index82.6, #1 overallThird party
Tau-Voice (agentic task completion)68.6%Vendor
Sierra Tau-Voice-banking35.1%Third party benchmark, vendor-reported
Big Bench Audio97.7%Vendor
Speech Agent ArenaSecond placeThird party

On ServiceNow's EVA-Bench, Google says the models "push the Pareto Frontier" for complex workflows, and its post carries a footnote that the run happened "on the Live API on Gemini Enterprise Agent Platform." Take the footnote at face value: a vendor scoring its own model on a third-party benchmark's hosted stack is still a vendor-run number, and nothing published lets an outsider reproduce it.

One misreading to avoid: "Gemini 3.8 Live scores 82.6" is wrong. The headline positions belong to Extended Thinking, the variant Google frames as the reasoning model. The base model is the one sold on cost efficiency, and Google did not attach an equivalent scorecard to it.

What the announcement leaves out

The release notes are thin on the things that would make the claims checkable. There are no parameter counts, no architecture details, no pricing in the announcement, and no published latency measurements behind live-dialogue claims of the sub-100-millisecond family. Independent verification is left to whoever has API access and a stopwatch.

Two omissions look intentional rather than accidental. Pricing is where the cost-efficiency argument would have to survive contact with a real bill, and it is absent. Latency is the entire premise of a live-dialogue model, and the only evidence offered is that the demos feel responsive.

To Google's credit, one disclosure is unusually explicit: all audio generated by its AI products carries a SynthID watermark, which the post describes as an "imperceptible watermark... woven directly into the audio output." The model card lives on the DeepMind model-cards page under Gemini 3.8 Audio, and developer details sit in the Gemini API documentation.

Where each variant ships

  • Developers: Gemini 3.8 Live and Extended Thinking in the Gemini API and Google AI Studio.
  • Enterprise: private preview in Gemini Enterprise, coming soon to Gemini Enterprise for Customer Experience and, for Extended Thinking, Google Workspace business customers.
  • Consumers: 3.8 Live in Search Live, Extended Thinking in Gemini Live, both in Workspace Docs for AI Pro and Ultra subscribers.
  • Free for subscribers: Gmail and Keep live features for all Google AI subscribers.

The integration list is long and commercially real: the Gemini Live API is used by Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel and Vision Agents, with Salesforce, Genspark and Lumeris named as launch partners praising latency, fluidity and tool calling. Partner quotes are endorsements from companies shipping on the stack, not measurements.

The bottom line

9to5Google's write-up frames the launch as the successor to Gemini 3.1 Flash Live from March 2026 and reads Extended Thinking as the engine for the recently launched Gmail Live, Docs Live and Keep Live. One reader on that article reports more concise answers, less latency, and that Google "has gotten rid of the audio cues between answers which feels much more natural." That is a single user's impression, not a test, and it is the only hands-on signal in the coverage so far.

The pattern here echoes Gemini 3.8 Flash two weeks earlier: a fast follow-up, a genuine capability advance on paper, and an evidence base that stops at the vendor's own charts. Voice agents are the one category where that gap is easy to close, because latency is measurable with a phone and a timer. Until someone publishes that measurement, the honest read is that Google has shipped two plausible live-dialogue models with one well-documented scorecard and one missing.

Related Articles

Scroll down

to load the next article