Cartesia Sonic 3.5
Cartesia Sonic 3.5 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
247.2 ms P50 time-to-first-audio on Coval; ranked 9th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 57.2/100, ranked 9th of 112 in the TTS for Voice Agents 2026 index (Strong). Strongest criterion: accuracy, 10.6 of 15 points.
Soniox TTS RT v1 is a proprietary text-to-speech (TTS) API from Soniox, measured by two independent benchmarks: Coval and Speko.
Overall score
57.2 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Soniox TTS RT v1 ranks 9th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 57.2/100.
On Coval's 30-day window ending October 1, 2026 (10,106 runs), Soniox TTS RT v1 has a P50 (median) time-to-first-audio (TTFA) of 247.2 ms, 12th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 339.8 ms (9th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 104.8 ms (13th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 179.1 ms. Speko measures a P50 TTFA of 362 ms from US East (n=30), dated July 3, 2026: 21st lowest of the 29 models Speko measures.
Coval measures an accuracy (word error rate) of 4% over 10,086 samples on the same window, 3rd lowest of the 31 models Coval measures. Speko rates its pronunciation robustness on numbers, dates and currency at 0.35 on a 0-to-1 scale (11th highest of the 28 models Speko rates).
Speko rates its naturalness at 1,569 Elo from blind A/B votes (14th highest of the 41 models Speko rates).
Speko lists a cost of about $13 per 1M characters.
Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.49, then 0.35): the score uses the average, 0.42. The value shown is the latest.
Tested in Spanish, German and French by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.
Cartesia Sonic 3.5 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
ElevenLabs Eleven v3 Conversational is a proprietary text-to-speech (TTS) API from ElevenLabs, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Palabra TTS v1 is a proprietary text-to-speech (TTS) API from Palabra, measured by two independent benchmarks: Coval and Speko.