
Kyutai released Voice of Reason — two open-weight models for speech-to-speech transformation. According to MarkTechPost, they are built on GLM-4-Voice-9B and are intended to solve spoken mathematical problems.
In the description, it states that after supervised fine-tuning and reinforcement learning, the accuracy on the spoken GSM8K increased from 27,3% to 77,1%. The claimed scheme includes no transcription step and no textual language model in the processing chain.
Both checkpoints are published on Hugging Face and, according to the same release, run on a single H100. This information is provided by a single independent source and is based on its metadata, rather than the full text of the publication available here.
editorial commentary
Why it matters
A probable consequence is increased interest in speech models that solve tasks directly from audio, without a required textual intermediary. The next verifiable signal will be Kyutai's initial publication with the methodology and independent results. Significant uncertainty arises because only a synopsis from one source is currently available.