Aller au contenu

Fish Audio S2.1 Pro Free Fish Audio · text-to-speech (TTS) API

Listed

405.6 ms P50 time-to-first-audio on Coval; ranked 67th of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 6.1/100, ranked 67th of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: accuracy, 5.1 of 15 points.

Fish Audio S2.1 Pro Free is a proprietary text-to-speech (TTS) API from Fish Audio, measured by one independent benchmark, Coval.

Overall score

6.1 /100

67 th of 112

Score breakdown

  • Real-time speed 1 /60
  • Voice quality 0 /20
  • Accuracy 5.1 /15
  • Price 0 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Fish Audio S2.1 Pro Free Compare with StepFun Step Audio EditX (Mar 2026)

Fish Audio S2.1 Pro Free ranks 67th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 6.1/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,831 runs), Fish Audio S2.1 Pro Free has a P50 (median) time-to-first-audio (TTFA) of 405.6 ms, 26th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 7,919.1 ms (31st lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 2,855.9 ms (31st lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 895.3 ms.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 4.6% over 11,811 samples on the same window, 10th lowest of the 31 models Coval measures.

Voice quality

Neither Artificial Analysis nor Speko rates its voice quality, so it counts 0 of 20.

Price

No price is published by Artificial Analysis or Speko for this model; price counts 0 of 5.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
913.8 ms
P50 / median time-to-first-audio (TTFA)
405.6 ms
P90 TTFA
1,097.3 ms
P95 TTFA
1,637.1 ms
P99 TTFA (slowest turns)
7,919.1 ms
Latency consistency (standard deviation)
2,855.9 ms
Leading silence
18.5 ms
Time to first byte (network arrival)
895.3 ms
Accuracy: word error rate (WER)
4.6%
WER basis
per-clip mean
Latency runs
11,831 runs
Accuracy runs
11,811 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
s2.1-pro-free
Coval model page
benchmarks.coval.ai

Not collected

Similar TTS models