Alibaba Qwen3 TTS 1.7b
Alibaba Qwen3 TTS 1.7b is an open-weights text-to-speech (TTS) API from Alibaba (Qwen), served by Baseten, measured by one independent benchmark, Coval.
254.8 ms P50 time-to-first-audio on Coval; ranked 18th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 44.5/100, ranked 18th of 112 in the TTS for Voice Agents 2026 index (Contender). Strongest criterion: real-time speed, 40 of 60 points.
Rime Mist v3 is a proprietary text-to-speech (TTS) API from Rime, measured by two independent benchmarks: Coval and Speko.
Overall score
44.5 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Rime Mist v3 ranks 18th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 44.5/100.
On Coval's 30-day window ending October 1, 2026 (11,799 runs), Rime Mist v3 has a P50 (median) time-to-first-audio (TTFA) of 254.8 ms, 13th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 364.3 ms (11th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 26.1 ms (3rd lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 127.5 ms.
Coval measures an accuracy (word error rate) of 5.3% over 11,779 samples on the same window, 24th lowest of the 31 models Coval measures.
Speko rates its naturalness at 1,428 Elo from blind A/B votes (36th highest of the 41 models Speko rates).
Speko lists a cost of about $30 per 1M characters.
Alibaba Qwen3 TTS 1.7b is an open-weights text-to-speech (TTS) API from Alibaba (Qwen), served by Baseten, measured by one independent benchmark, Coval.
SpaceXAI Grok TTS is a proprietary text-to-speech (TTS) API from xAI, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Cartesia Sonic 3.6 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.