Aller au contenu

Fish Audio S2.1 Pro Fish Audio · text-to-speech (TTS) API

Listed

300.1 ms P50 time-to-first-audio on Coval; ranked 23rd of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 33.2/100, ranked 23rd of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: voice quality, 7.7 of 20 points.

Fish Audio S2.1 Pro is a proprietary text-to-speech (TTS) API from Fish Audio, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.

Overall score

33.2 /100

23 rd of 112

Score breakdown

  • Real-time speed 18.6 /60
  • Voice quality 7.7 /20
  • Accuracy 5.1 /15
  • Price 1.8 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Fish Audio S2.1 Pro Compare with Google Gemini 3.8 Flash TTS

Fish Audio S2.1 Pro ranks 23rd of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 33.2/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,855 runs), Fish Audio S2.1 Pro has a P50 (median) time-to-first-audio (TTFA) of 300.1 ms, 18th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 885.9 ms (24th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 193.9 ms (21st lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 320.2 ms.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 4.6% over 11,835 samples on the same window, 10th lowest of the 31 models Coval measures.

Voice quality

In the Artificial Analysis Speech Arena (capture of October 1, 2026), Fish Audio S2.1 Pro has an Elo of 1,137 (±13, 95% confidence) from 2,291 blind comparisons, 22nd highest of the 92 models Artificial Analysis rates.

Price

Artificial Analysis lists Fish Audio S2.1 Pro at $15 per 1M characters (19th lowest of the 78 models Artificial Analysis lists a price for).

Scoring note

Artificial Analysis changed its Speech Arena Elo between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (1,139, then 1,137): the score uses the average, 1,138. The value shown is the latest.

Languages tested

Tested in Spanish, German, French, Portuguese, Japanese, Chinese, Arabic, Hindi and Vietnamese by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
339 ms
P50 / median time-to-first-audio (TTFA)
300.1 ms
P90 TTFA
402.6 ms
P95 TTFA
621.4 ms
P99 TTFA (slowest turns)
885.9 ms
Latency consistency (standard deviation)
193.9 ms
Leading silence
18.8 ms
Time to first byte (network arrival)
320.2 ms
Accuracy: word error rate (WER)
4.6%
WER basis
per-clip mean
Latency runs
11,855 runs
Accuracy runs
11,835 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
s2.1-pro
Coval model page
benchmarks.coval.ai

Speko

Speko model
s2.1-pro
Speko model page
benchmarks.speko.ai

Artificial Analysis Speech Arena

Speech Arena Elo (blind listener preference)
1,137
Elo 95% interval (±)
13
Arena rank
22
Arena rank range
21-27
Arena appearances
2,291
Voices tested
8
Price per 1M characters
15 USD
Released
Jun 2026
Artificial Analysis model
Fish Audio S2.1 Pro
Artificial Analysis leaderboard
artificialanalysis.ai

Not collected

Similar TTS models