Google DeepMind has announced the release of its latest audio model, named Gemini 3.1 Flash TTS. According to the developers, this system represents the next generation of speech synthesis technology, focused on enhanced expressiveness.

A key feature of the model is the introduction of granular audio tags. This mechanism provides users with precise control over the generation process, allowing them to directly specify desired emotional and intonational characteristics of the created audio.

The presented technology is positioned as a tool for creating more lifelike and controllable artificial speech. The implementation of such tags aims to eliminate limitations of previous systems, where fine-tuning voice was difficult.