MarkTechPost published a benchmark of inference APIs for voice and real-time agents. At the center of the analysis is Time To First Token (TTFT), which teams often use when selecting APIs.

The material examines four layers of the voice stack: large language models, speech recognition, speech synthesis, and speech-to-speech transformation. The description states that the results were cross-checked against primary sources 30 August 2026 year, and the provenance of each figure was indicated separately.

The source package does not contain the numeric results themselves, the list of participants, or a final ranking. Therefore it confirms the methodological scope of the benchmark but does not permit naming the fastest API.