Fish Audio S2.1 Pro
Fish Audio S2.1 Pro is a proprietary text-to-speech (TTS) API from Fish Audio, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
1,031 ms P50 time-to-first-audio on Speko; ranked 24th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 29.7/100, ranked 24th of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: voice quality, 19.3 of 20 points.
Google Gemini 3.8 Flash TTS is a proprietary text-to-speech (TTS) API from Google, measured by two independent benchmarks: Artificial Analysis and Speko.
Overall score
29.7 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Google Gemini 3.8 Flash TTS ranks 24th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 29.7/100.
Speko measures a P50 TTFA of 1,031 ms from US East (n=30), dated September 24, 2026: 29th lowest of the 29 models Speko measures.
Speko rates its pronunciation robustness on numbers, dates and currency at 0.7 on a 0-to-1 scale (4th highest of the 28 models Speko rates).
In the Artificial Analysis Speech Arena (capture of October 1, 2026), Google Gemini 3.8 Flash TTS has an Elo of 1,272 (±16, 95% confidence) from 2,330 blind comparisons, 3rd highest of the 92 models Artificial Analysis rates. Speko rates its naturalness at 1,642 Elo from blind A/B votes (3rd highest of the 41 models Speko rates).
Artificial Analysis lists Google Gemini 3.8 Flash TTS at $16.5 per 1M characters (31st lowest of the 78 models Artificial Analysis lists a price for). Speko lists a cost of about $15 per 1M characters.
Artificial Analysis changed its Speech Arena Elo between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (1,268, then 1,272): the score uses the average, 1,270. The value shown is the latest. Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.95, then 0.7): the score uses the average, 0.825. The value shown is the latest.
Tested in Spanish, German, French, Portuguese, Japanese, Chinese, Arabic, Hindi and Vietnamese by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.
Fish Audio S2.1 Pro is a proprietary text-to-speech (TTS) API from Fish Audio, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Google Gemini 3.8 Flash-Lite TTS is a proprietary text-to-speech (TTS) API from Google, measured by two independent benchmarks: Artificial Analysis and Speko.
Deepdub Phantom Z 3.4 conversational is a proprietary text-to-speech (TTS) API from Deepdub, measured by one independent benchmark, Coval.