Gradium TTS (Aug 2026)
Gradium TTS (Aug 2026) is a proprietary text-to-speech (TTS) API from Gradium, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
105 ms P50 time-to-first-audio on Coval; ranked 7th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 58.6/100, ranked 7th of 112 in the TTS for Voice Agents 2026 index (Strong). Strongest criterion: real-time speed, 49.9 of 60 points.
Palabra TTS v1 is a proprietary text-to-speech (TTS) API from Palabra, measured by two independent benchmarks: Coval and Speko.
Overall score
58.6 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Palabra TTS v1 ranks 7th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 58.6/100.
On Coval's 30-day window ending October 1, 2026 (11,682 runs), Palabra TTS v1 has a P50 (median) time-to-first-audio (TTFA) of 105 ms, 6th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 299.8 ms (7th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 42.4 ms (6th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 111.9 ms. Speko measures a P50 TTFA of 72 ms from US East (n=30), dated July 3, 2026: 2nd lowest of the 29 models Speko measures.
Coval measures an accuracy (word error rate) of 5.6% over 11,662 samples on the same window, 27th lowest of the 31 models Coval measures. Speko rates its pronunciation robustness on numbers, dates and currency at 0.16 on a 0-to-1 scale (25th highest of the 28 models Speko rates).
Speko rates its naturalness at 1,520 Elo from blind A/B votes (22nd highest of the 41 models Speko rates).
Speko lists a cost of about $30 per 1M characters.
Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.3, then 0.16): the score uses the average, 0.23. The value shown is the latest.
Gradium TTS (Aug 2026) is a proprietary text-to-speech (TTS) API from Gradium, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Cartesia Sonic 3.5 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Inworld Realtime TTS-2 Flash is a proprietary text-to-speech (TTS) API from Inworld, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.