AI voiceovers can read a script. But getting them to actually perform it with the right accent, pacing, emotion, pauses, and even the occasional sigh is another story. To address this, Google has spent weeks growing Gemini 3.8 from the original Gemini 3.8 Flash into new Live and Live Extended Thinking models. And now it wants the same family to act the part, with Gemini 3.8 Flash Text-To-Speech (TTS) and Flash-Lite TTS rolling out.
The biggest change for users is control. In a blog post, Google says Gemini 3.8 Flash TTS can create an original voice from a natural-language description, with support for more than 100 languages and dialects. You also get over 2,000 production-ready voices, and voice replication can recreate a consistent voice from a 30-second sample, provided you have permission to use it.