Cartesia Sonic 3.6
Cartesia Sonic 3.6 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
103.5 ms P50 time-to-first-audio on Coval; ranked 17th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 45.3/100, ranked 17th of 112 in the TTS for Voice Agents 2026 index (Contender). Strongest criterion: real-time speed, 43.9 of 60 points.
Alibaba Qwen3 TTS 1.7b is an open-weights text-to-speech (TTS) API from Alibaba (Qwen), served by Baseten, measured by one independent benchmark, Coval.
Overall score
45.3 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Alibaba Qwen3 TTS 1.7b ranks 17th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 45.3/100.
On Coval's 30-day window ending October 1, 2026 (781 runs), Alibaba Qwen3 TTS 1.7b has a P50 (median) time-to-first-audio (TTFA) of 103.5 ms, 5th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 165.4 ms (4th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 37.5 ms (5th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 72.8 ms.
Coval measures an accuracy (word error rate) of 5.4% over 781 samples on the same window, 25th lowest of the 31 models Coval measures.
Neither Artificial Analysis nor Speko rates its voice quality, so it counts 0 of 20.
No price is published by Artificial Analysis or Speko for this model; price counts 0 of 5.
Coval measures this model through Baseten's API, so the latency above is Baseten's. Coval's model registry names Alibaba (Qwen) as the creator and Baseten as the host.
Cartesia Sonic 3.6 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Rime Mist v3 is a proprietary text-to-speech (TTS) API from Rime, measured by two independent benchmarks: Coval and Speko.
Fluxions vui is a proprietary text-to-speech (TTS) API from Fluxions, measured by one independent benchmark, Coval.