Kokoro 82M v1.0
Kokoro 82M v1.0 is an open-weights text-to-speech (TTS) API from Kokoro, measured by one independent benchmark, Artificial Analysis.
Speech Arena Elo 1,079 on Artificial Analysis, latency not measured; ranked 61st (tied) of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 7.1/100, ranked 61st (tied) of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: price, 1.6 of 5 points.
Mistral Voxtral TTS is an open-weights text-to-speech (TTS) API from Mistral, measured by one independent benchmark, Artificial Analysis.
Overall score
7.1 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Mistral Voxtral TTS ranks 61st of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 7.1/100.
Neither Coval nor Speko measures Mistral Voxtral TTS's latency, so it scores 0 of 60 on real-time speed and does not appear in the real-time ranking.
No accuracy measurement: Coval does not measure its word error rate and Speko does not rate its pronunciation robustness, so accuracy counts 0 of 15.
In the Artificial Analysis Speech Arena (capture of October 1, 2026), Mistral Voxtral TTS has an Elo of 1,079 (±13, 95% confidence) from 2,094 blind comparisons, 42nd highest of the 92 models Artificial Analysis rates.
Artificial Analysis lists Mistral Voxtral TTS at $16 per 1M characters (28th lowest of the 78 models Artificial Analysis lists a price for).
Artificial Analysis changed its Speech Arena Elo between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (1,080, then 1,079): the score uses the average, 1,079.5. The value shown is the latest.
Tested in Spanish, German, French, Portuguese and Hindi by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.
Kokoro 82M v1.0 is an open-weights text-to-speech (TTS) API from Kokoro, measured by one independent benchmark, Artificial Analysis.
Resemble AI Chatterbox HD is a proprietary text-to-speech (TTS) API from Resemble AI, measured by one independent benchmark, Artificial Analysis.
MiniMax Speech-02-HD is a proprietary text-to-speech (TTS) API from MiniMax, measured by one independent benchmark, Artificial Analysis.