Google is introducing two speech synthesis models — Gemini 3.8 Flash TTS and Flash-Lite TTS. According to The Decoder, they support more than 100 languages.

Flash TTS allows creating a voice based on a text description. Both models support stage directions for individual lines and the generation of two-voice dialogues from a single script. A separate voice cloning function creates a voice profile from a 30-second sample.

The practical implication of this innovation is shifting part of the voiceover setup from technical parameters to ordinary descriptions and scripts. However, since the source consists of a single independent publication and metadata, the quality, access conditions, and safeguards against misuse remain unclear.