All News
googlegeminivoice-aitext-to-speechsynthid

Gemini 3.8 Flash TTS clones a voice from 30 seconds of audio. Six jurisdictions will not get that feature.

Google's new voice models clone a speaker from a 30-second sample, but replication will not run in Illinois, Texas, the EEA, the UK, Switzerland or India.

Vlad MakarovVlad Makarovreviewed and published
5 min read
Gemini 3.8 Flash TTS clones a voice from 30 seconds of audio. Six jurisdictions will not get that feature.

Google announced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS on September 23, 2026, in a launch post signed by Leland Rechis and Alan Cowen on behalf of the Gemini Audio team. One model is built for character design, the other for volume. The detail worth more attention than voice quality is what Google attached to the cloning tool: a consent check, a SynthID watermark, C2PA credentials, and a list of jurisdictions where the feature will not run.

Two models, one studio, and a library the post calls infinite

Gemini 3.8 Flash TTS is the flagship, pitched at "deep creative direction and character design." It generates entirely new voices from natural-language prompts for games, immersive audiobooks, podcasts and interactive media. Gemini 3.8 Flash-Lite TTS is the volume sibling, built for high-volume dubbing, audio content creation and expressive voice agents. Both ship first in the Gemini API and Google AI Studio, whose speech generation docs cover the new audio playground.

Generative voice design is the headline capability: customizing role, accent and voice characteristics across more than 100 languages and dialects by prompting. Google says creators can scale "from 30 original voices to an infinite library," and separately advertises 2,000+ production-ready voices with regional varieties such as Mexican Spanish, Quebec French and Scots English. The release sits inside a fast-moving Gemini Audio family, after 3.5 Live Translate, 3.5 Transcribe, and Gemini 3.8 Live and 3.8 Live Extended Thinking; the same team has also shipped Gemini 3.8 Flash and 3.8 Flash Cyber.

Gemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTS
Built forCharacter design, creative directionHigh-volume dubbing, voice agents
Developer accessGemini API, Google AI StudioGemini API, Google AI Studio
Consumer surfaceGemini NotebookGoogle Vids
EnterpriseComing soon via APIComing soon via API

Infinite is a marketing figure. The post's own concrete inventory is 2,000+ voices, and nothing published measures how a library behaves at either number. The two-model split also lets Google sell one capability to two budgets: the creative-direction story rides on Flash TTS, while the volume argument rests on Flash-Lite.

Cloning a voice from thirty seconds of audio

Voice replication recreates a vocal profile from a 30-second sample, either your own voice or one you have the rights to use. Google gates it with consent verification, requiring a verbal consent recording from the voice owner that matches the reference speaker before a voice can be created. Generated clips carry a SynthID watermark and C2PA credentials.

Thirty seconds is the number to sit with. Voice-verification systems in banking, call centres and identity checks assume a short sample cannot reproduce a speaker. A production feature that does exactly that, inside consumer-accessible tooling, weakens that assumption for everyone rather than only for the people who opt in. The voice cloning literature has carried the warning for years; what changed is the sample length and the price of access.

What Google's scoreboard does and does not show

The launch post does carry third-party numbers. Flash TTS took the top spot on Hume AI's Voice Design Benchmark at 71.4 and led its accent-modeling row at 60.8, and the two models ranked first and second on Hume's Overall Quality Index. On Voice Arena's blind human-preference leaderboard, Google says they placed at the top among competitors in Japanese, Brazilian Portuguese, Vietnamese, Modern Standard Arabic, Mexican Spanish and Hindi. All of those rows are vendor-selected and vendor-reported, and the post names no competitor models behind them.

Nothing in the announcement supplies a word-error rate, a speaker-similarity score for the clone, or a comparison against ElevenLabs or OpenAI's audio models. Voice remixing, which would let a user fine-tune timbre, pitch, pace and accent on a library voice, is listed as coming soon.

Can a single script stage two speakers?

Both models take direction at the line level, and the post lists the controls plainly: acting cues, pacing, dialect shifts and backchanneling can be written into the script or left to the model. Long-form generation is pitched at holding voice quality, pacing and character timbre steady across hours of continuous audio, which is the audiobook and podcast case, and Google says both models improve on Gemini 3.1 Flash TTS for long-form content and dual-speaker screenplay control.

The most unusual capability is native two-speaker scene staging: multi-turn dialogue directed from one script, with turn-taking handled by the model and the two voices kept apart. Google also supports scripted non-verbal bursts, written as <laughs>, <sigh> or <gasp>, plus active-listening interjections such as |mhm| and |yeah|. Comedic timing and reaction beats stop being a post-production problem.

Neither claim comes with a published measurement of drift over a long session, which is the one most likely to disappoint a paying customer.

SynthID and C2PA, shipped as an admission

Google says every audio clip generated by its Gemini Audio models is watermarked with SynthID, an imperceptible mark the post describes as woven directly into the output, and that replication carries C2PA credentials. Putting a detector next to a generator is a reasonable disclosure, and it is also an admission that a cloning capability has a known abuse path. Watermarking only helps if someone can find the mark, and the post publishes no independent detection rate for SynthID-marked audio.

Six jurisdictions that will not get voice replication

One footnote qualifies the whole rollout: voice replication through AI Studio is not available in Illinois, Texas, the European Economic Area, the United Kingdom, Switzerland, or India. Each of those six has a voice, biometric or data-protection statute on the books, and the excluded list is shorter than the spec it qualifies. That is a rollout map drawn by regulation rather than by engineering.

The footnote also scopes the exclusion to AI Studio, leaving the API path in those same markets unaddressed. Anyone reading the list as a global safety boundary is reading it too generously; the honest description is a product with a jurisdictional carve-out and an open question about where else it applies.

What would settle it

Three artifacts would move this from announcement to evidence. Published evaluations of clone fidelity and drift across long audio, on a harness someone outside Google can rerun. A working demonstration of the consent check against a non-consensual sample, showing it rejects a reconstructed voice rather than a mismatched one. And independent detection rates for SynthID-marked audio, measured by a party that did not build the watermark.

Until then, the launch is a capable voice family with a vendor scoreboard, an unusually explicit statement of its own abuse path, and a reception on Hacker News of 215 points and 112 comments as of 15:29 UTC on September 23.

Related Articles

Scroll down

to load the next article