StepFun Step Audio EditX (Mar 2026)
StepFun Step Audio EditX (Mar 2026) is an open-weights text-to-speech (TTS) API from StepFun, measured by one independent benchmark, Artificial Analysis.
405.6 ms P50 time-to-first-audio on Coval; ranked 67th of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 6.1/100, ranked 67th of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: accuracy, 5.1 of 15 points.
Fish Audio S2.1 Pro Free is a proprietary text-to-speech (TTS) API from Fish Audio, measured by one independent benchmark, Coval.
Overall score
6.1 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Fish Audio S2.1 Pro Free ranks 67th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 6.1/100.
On Coval's 30-day window ending October 1, 2026 (11,831 runs), Fish Audio S2.1 Pro Free has a P50 (median) time-to-first-audio (TTFA) of 405.6 ms, 26th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 7,919.1 ms (31st lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 2,855.9 ms (31st lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 895.3 ms.
Coval measures an accuracy (word error rate) of 4.6% over 11,811 samples on the same window, 10th lowest of the 31 models Coval measures.
Neither Artificial Analysis nor Speko rates its voice quality, so it counts 0 of 20.
No price is published by Artificial Analysis or Speko for this model; price counts 0 of 5.
StepFun Step Audio EditX (Mar 2026) is an open-weights text-to-speech (TTS) API from StepFun, measured by one independent benchmark, Artificial Analysis.
Google Gemini 2.5 Flash TTS (Dec 2025) is a proprietary text-to-speech (TTS) API from Google, measured by one independent benchmark, Artificial Analysis.
MiniMax Speech-02-Turbo is a proprietary text-to-speech (TTS) API from MiniMax, measured by one independent benchmark, Artificial Analysis.