Palabra TTS v1
Palabra TTS v1 is a proprietary text-to-speech (TTS) API from Palabra, measured by two independent benchmarks: Coval and Speko.
280 ms P50 time-to-first-audio on Coval; ranked 8th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 58.5/100, ranked 8th of 112 in the TTS for Voice Agents 2026 index (Strong). Strongest criterion: voice quality, 15.8 of 20 points.
Cartesia Sonic 3.5 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Overall score
58.5 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Cartesia Sonic 3.5 ranks 8th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 58.5/100.
On Coval's 30-day window ending October 1, 2026 (11,855 runs), Cartesia Sonic 3.5 has a P50 (median) time-to-first-audio (TTFA) of 280 ms, 15th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 421.8 ms (13th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 45.6 ms (7th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 152.1 ms. Speko measures a P50 TTFA of 121 ms from US East (n=30), dated July 3, 2026: 11th lowest of the 29 models Speko measures.
Coval measures an accuracy (word error rate) of 5.8% over 11,835 samples on the same window, 28th lowest of the 31 models Coval measures. Speko rates its pronunciation robustness on numbers, dates and currency at 0.18 on a 0-to-1 scale (24th highest of the 28 models Speko rates).
In the Artificial Analysis Speech Arena (capture of October 1, 2026), Cartesia Sonic 3.5 has an Elo of 1,187 (±12, 95% confidence) from 3,128 blind comparisons, 14th highest of the 92 models Artificial Analysis rates. Speko rates its naturalness at 1,574 Elo from blind A/B votes (12th highest of the 41 models Speko rates).
Artificial Analysis lists Cartesia Sonic 3.5 at $49 per 1M characters (54th lowest of the 78 models Artificial Analysis lists a price for). Speko lists a cost of about $50 per 1M characters.
Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.55, then 0.18): the score uses the average, 0.365. The value shown is the latest.
Tested in Spanish, German, French, Portuguese, Japanese, Chinese, Arabic, Hindi, Vietnamese, Korean, Filipino, Norwegian, Tamil, Telugu and Thai by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.
Palabra TTS v1 is a proprietary text-to-speech (TTS) API from Palabra, measured by two independent benchmarks: Coval and Speko.
Soniox TTS RT v1 is a proprietary text-to-speech (TTS) API from Soniox, measured by two independent benchmarks: Coval and Speko.
Gradium TTS (Aug 2026) is a proprietary text-to-speech (TTS) API from Gradium, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.