Deepdub Phantom Z 3.4 conversational
Deepdub Phantom Z 3.4 conversational is a proprietary text-to-speech (TTS) API from Deepdub, measured by one independent benchmark, Coval.
300.1 ms P50 time-to-first-audio on Coval; ranked 23rd of 112 in the TTS for Voice Agents 2026 index.
Benchmark summary: 33.2/100, ranked 23rd of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: voice quality, 7.7 of 20 points.
Fish Audio S2.1 Pro is a proprietary text-to-speech (TTS) API from Fish Audio, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Overall score
33.2 /100
Data checked on September 30, 2026
Updated after each benchmark capture
Every point comes from a public benchmark: sources · methodology
Fish Audio S2.1 Pro ranks 23rd of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 33.2/100.
On Coval's 30-day window ending October 1, 2026 (11,855 runs), Fish Audio S2.1 Pro has a P50 (median) time-to-first-audio (TTFA) of 300.1 ms, 18th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 885.9 ms (24th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 193.9 ms (21st lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 320.2 ms.
Coval measures an accuracy (word error rate) of 4.6% over 11,835 samples on the same window, 10th lowest of the 31 models Coval measures.
In the Artificial Analysis Speech Arena (capture of October 1, 2026), Fish Audio S2.1 Pro has an Elo of 1,137 (±13, 95% confidence) from 2,291 blind comparisons, 22nd highest of the 92 models Artificial Analysis rates.
Artificial Analysis lists Fish Audio S2.1 Pro at $15 per 1M characters (19th lowest of the 78 models Artificial Analysis lists a price for).
Artificial Analysis changed its Speech Arena Elo between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (1,139, then 1,137): the score uses the average, 1,138. The value shown is the latest.
Tested in Spanish, German, French, Portuguese, Japanese, Chinese, Arabic, Hindi and Vietnamese by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.
Deepdub Phantom Z 3.4 conversational is a proprietary text-to-speech (TTS) API from Deepdub, measured by one independent benchmark, Coval.
Google Gemini 3.8 Flash TTS is a proprietary text-to-speech (TTS) API from Google, measured by two independent benchmarks: Artificial Analysis and Speko.
Smallest.ai Lightning V3.1 Pro (Jul 2026) is a proprietary text-to-speech (TTS) API from Smallest.ai, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.