Gradium AI released a new default speech synthesis model. According to MarkTechPost, it achieved a pass rate of 81,0% in the evaluation of 500 complex sentences across five languages and showed a median time to first audio of 216 milliseconds on Coval.

The evaluation was conducted by humans, and the set of sentences is published on Hugging Face under CC BY 4.0. These data indicate an attempt to measure both speech quality on difficult examples and the speed of speech onset.

The practical interest lies in the potential application of the model where it is important for the user to hear results quickly rather than wait for full generation.