The new voice studio
Google introduced Gemini 3.8 Flash TTS and Flash-Lite TTS for generated speech. Flash is positioned for detailed creative direction and character design; Flash-Lite is aimed at high-volume, lower-cost voice work. Rather than choosing only from a handful of fixed narrators, a creator can describe a voice’s role, accent and character in natural language.
Google says the system supports more than 100 languages and dialects and offers over 2,000 production-ready voices. For a consistent custom voice, Flash can use a 30-second sample of a voice the user owns or has rights to use. The company says consent verification, SynthID watermarking and C2PA credentials are built around that replication route.
Directing the performance
Both models let a writer specify how a line should sound: pacing, emotion, whispering or conversational backchannels such as a laugh or sigh. The system can stage two speakers from one script while maintaining distinct voices. Google also claims long-form consistency over hours of audio, a requirement for audiobooks and podcasts where voice drift can ruin continuity.
The primary announcement includes audio and video demonstrations, making this more than a static benchmark claim. Hearing a generated scene with turn-taking is a more direct test of whether the output sounds like a conversation than reading a quality score. Google also cites outside voice-design evaluations, but the end use still depends on a creator’s script and editing.
Where it is available
Google lists AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids among the places these voice capabilities appear. Availability can differ by product and plan, so a reader should check the specific tool before designing a workflow around it.
The practical leap is control over a performance, not just conversion of text to audio. A team can record one approved voice, reuse it consistently and adjust delivery scene by scene. The guardrail for replication is straightforward: only use a voice you own or have explicit permission to reproduce.
Source published 2026-09-23. Coverage is based on the maker’s announcement and demonstration.
