Aller au contenu

Top rated for

Best TTS for Voice agents

6 text-to-speech (TTS) models, ranked from the highest benchmark score to the lowest, out of 100.

Data checked on September 30, 2026 · updated after each benchmark capture

6 text-to-speech (TTS) models start speaking within 200 ms on 9 turns out of 10: their P90 time-to-first-audio on Coval's continuous 30-day measurement is 200 ms or less. That is the typical gap between turns in human conversation (Levinson & Torreira, 2015).

Why 200 ms at the P90

A voice agent is judged on every turn, not the average one. The P90 is the time-to-first-audio that 90% of turns stay under; a model that meets 200 ms at the P90 leaves the rest of the agent pipeline (speech recognition, language model) room to answer like a person would. Models Coval does not measure are not listed.

Frequently asked questions

What is the best TTS for voice agents in 2026?
Among the 6 models that meet the 200 ms P90 threshold, Gradium TTS Beta has the highest overall benchmark score, 76.8 out of 100, with a P50 time-to-first-audio of 47.2 ms on Coval.
Why is a fast TTS model missing from this page?
Either Coval does not measure it, or its P90 time-to-first-audio over 30 days is above 200 ms. Its own page shows which, with the measured values.