
Google is introducing two speech synthesis models — Gemini 3.8 Flash TTS and Flash-Lite TTS. According to The Decoder, they support more than 100 languages.
Flash TTS allows creating a voice based on a text description. Both models support stage directions for individual lines and the generation of two-voice dialogues from a single script. A separate voice cloning function creates a voice profile from a 30-second sample.
The practical implication of this innovation is shifting part of the voiceover setup from technical parameters to ordinary descriptions and scripts. However, since the source consists of a single independent publication and metadata, the quality, access conditions, and safeguards against misuse remain unclear.
editorial commentary
Why it matters
A likely consequence is a simpler release of multilingual voiceovers and dialogue materials. The next verifiable signal will be Google's publication regarding access terms, documentation, and voice cloning rules. Significant uncertainty remains concerning feature quality and protection against misuse.