Fluxions vui
Fluxions vui is a proprietary text-to-speech (TTS) API from Fluxions, measured by one independent benchmark, Coval.
347.9 ms P50 time-to-first-audio on Coval; ranked 16th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 45.9/100, ranked 16th of 112 in the TTS for Voice Agents 2026 index (Contender). Strongest criterion: voice quality, 17.4 of 20 points.
Cartesia Sonic 3.6 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Overall score
45.9 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Cartesia Sonic 3.6 ranks 16th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 45.9/100.
On Coval's 30-day window ending October 1, 2026 (11,856 runs), Cartesia Sonic 3.6 has a P50 (median) time-to-first-audio (TTFA) of 347.9 ms, 21st lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 840 ms (22nd lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 117.6 ms (15th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 235.7 ms. Speko measures a P50 TTFA of 120 ms from US East (n=30), dated August 28, 2026: 10th lowest of the 29 models Speko measures.
Coval measures an accuracy (word error rate) of 5.2% over 11,836 samples on the same window, 20th lowest of the 31 models Coval measures. Speko rates its pronunciation robustness on numbers, dates and currency at 0.21 on a 0-to-1 scale (20th highest of the 28 models Speko rates).
In the Artificial Analysis Speech Arena (capture of October 1, 2026), Cartesia Sonic 3.6 has an Elo of 1,273 (±16, 95% confidence) from 2,093 blind comparisons, 2nd highest of the 92 models Artificial Analysis rates. Speko rates its naturalness at 1,578 Elo from blind A/B votes (11th highest of the 41 models Speko rates).
Artificial Analysis lists Cartesia Sonic 3.6 at $49 per 1M characters (54th lowest of the 78 models Artificial Analysis lists a price for). Speko lists a cost of about $50 per 1M characters.
Artificial Analysis changed its Speech Arena Elo between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (1,275, then 1,273): the score uses the average, 1,274. The value shown is the latest. Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.5, then 0.21): the score uses the average, 0.355. The value shown is the latest.
Tested in Spanish, German, French, Portuguese, Japanese, Chinese, Arabic, Hindi and Vietnamese by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.
Fluxions vui is a proprietary text-to-speech (TTS) API from Fluxions, measured by one independent benchmark, Coval.
Alibaba Qwen3 TTS 1.7b is an open-weights text-to-speech (TTS) API from Alibaba (Qwen), served by Baseten, measured by one independent benchmark, Coval.
Rime Coda is a proprietary text-to-speech (TTS) API from Rime, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.