Kyutai released Voice of Reason — two open-weight models for speech-to-speech transformation. According to MarkTechPost, they are built on GLM-4-Voice-9B and are intended to solve spoken mathematical problems.

In the description, it states that after supervised fine-tuning and reinforcement learning, the accuracy on the spoken GSM8K increased from 27,3% to 77,1%. The claimed scheme includes no transcription step and no textual language model in the processing chain.

Both checkpoints are published on Hugging Face and, according to the same release, run on a single H100. This information is provided by a single independent source and is based on its metadata, rather than the full text of the publication available here.