Cartesia Sonic 3
Cartesia Sonic 3 is a proprietary text-to-speech (TTS) API from Cartesia, measured by two independent benchmarks: Artificial Analysis and Speko.
294 ms P50 time-to-first-audio on Speko; ranked 39th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 11.4/100, ranked 39th of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: voice quality, 10 of 20 points.
MiniMax Speech 2.8 HD is a proprietary text-to-speech (TTS) API from MiniMax, measured by two independent benchmarks: Artificial Analysis and Speko.
Overall score
11.4 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
MiniMax Speech 2.8 HD ranks 39th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 11.4/100.
Speko measures a P50 TTFA of 294 ms from US East (n=30), dated July 3, 2026: 19th lowest of the 29 models Speko measures.
Speko rates its pronunciation robustness on numbers, dates and currency at 0.15 on a 0-to-1 scale (26th highest of the 28 models Speko rates).
In the Artificial Analysis Speech Arena (capture of October 1, 2026), MiniMax Speech 2.8 HD has an Elo of 1,170 (±11, 95% confidence) from 4,560 blind comparisons, 17th highest of the 92 models Artificial Analysis rates. Speko rates its naturalness at 1,431 Elo from blind A/B votes (34th highest of the 41 models Speko rates).
Artificial Analysis lists MiniMax Speech 2.8 HD at $100 per 1M characters (69th lowest of the 78 models Artificial Analysis lists a price for). Speko lists a cost of about $100 per 1M characters.
Artificial Analysis changed its Speech Arena Elo between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (1,171, then 1,170): the score uses the average, 1,170.5. The value shown is the latest. Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0, then 0.15): the score uses the average, 0.075. The value shown is the latest.
Tested in Spanish, German, French, Portuguese, Japanese, Chinese, Arabic, Hindi, Vietnamese, Filipino and Norwegian by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.
Cartesia Sonic 3 is a proprietary text-to-speech (TTS) API from Cartesia, measured by two independent benchmarks: Artificial Analysis and Speko.
Alibaba Qwen-Audio-3.0-TTS-Plus is a proprietary text-to-speech (TTS) API from Alibaba (Qwen), measured by one independent benchmark, Artificial Analysis.
Bland Speech v3 is a proprietary text-to-speech (TTS) API from Bland AI, measured by two independent benchmarks: Artificial Analysis and Speko.