According to data from an independent report by The Decoder, the Qwen Audio 3.0 TTS Plus text-to-speech model, developed by Alibaba, has taken first place in the Speech Arena leaderboard from the organization Artificial Analysis. The system supports operation in 16 languages and allows users to control pronunciation style using natural language commands or special tags, such as [angry].

Despite high output quality, the model's operating speed is only 16 characters per second. This is significantly slower than the performance of major market competitors—systems Sonic 3.5 and Simba 3.2—which demonstrate higher productivity during audio generation.

The data regarding the model's leadership is based exclusively on the meta-description of the publication in The Decoder, without access to the full text of the report or primary testing data. The information confirms the fact of occupying the top spot in the ranking but does not reveal details of the evaluation methodology or specific numerical indicators of voice quality.