Gradium TTS Beta
Gradium TTS Beta is a proprietary text-to-speech (TTS) API from Gradium, measured by two independent benchmarks: Coval and Speko.
Top rated for
6 text-to-speech (TTS) models, ranked from the highest benchmark score to the lowest, out of 100.
Data checked on September 30, 2026 · updated after each benchmark capture
6 text-to-speech (TTS) models start speaking within 200 ms on 9 turns out of 10: their P90 time-to-first-audio on Coval's continuous 30-day measurement is 200 ms or less. That is the typical gap between turns in human conversation (Levinson & Torreira, 2015).
A voice agent is judged on every turn, not the average one. The P90 is the time-to-first-audio that 90% of turns stay under; a model that meets 200 ms at the P90 leaves the rest of the agent pipeline (speech recognition, language model) room to answer like a person would. Models Coval does not measure are not listed.
Gradium TTS Beta is a proprietary text-to-speech (TTS) API from Gradium, measured by two independent benchmarks: Coval and Speko.
Alibaba Qwen3 TTS Fast is an open-weights text-to-speech (TTS) API from Alibaba (Qwen), served by Nari Labs, measured by two independent benchmarks: Coval and Speko.
Inworld Realtime TTS-2 Flash is a proprietary text-to-speech (TTS) API from Inworld, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Palabra TTS v1 is a proprietary text-to-speech (TTS) API from Palabra, measured by two independent benchmarks: Coval and Speko.
Fluxions vui is a proprietary text-to-speech (TTS) API from Fluxions, measured by one independent benchmark, Coval.
Alibaba Qwen3 TTS 1.7b is an open-weights text-to-speech (TTS) API from Alibaba (Qwen), served by Baseten, measured by one independent benchmark, Coval.