Aller au contenu

Deepdub Phantom Z 3.4 conversational Deepdub · text-to-speech (TTS) API

Listed

261.3 ms P50 time-to-first-audio on Coval; ranked 22nd of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 33.9/100, ranked 22nd of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: real-time speed, 33.3 of 60 points.

Deepdub Phantom Z 3.4 conversational is a proprietary text-to-speech (TTS) API from Deepdub, measured by one independent benchmark, Coval.

Overall score

33.9 /100

22 nd of 112

Score breakdown

  • Real-time speed 33.3 /60
  • Voice quality 0 /20
  • Accuracy 0.6 /15
  • Price 0 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Deepdub Phantom Z 3.4 conversational Compare with Fish Audio S2.1 Pro

Deepdub Phantom Z 3.4 conversational ranks 22nd of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 33.9/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,585 runs), Deepdub Phantom Z 3.4 conversational has a P50 (median) time-to-first-audio (TTFA) of 261.3 ms, 14th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 427.4 ms (14th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 80.8 ms (12th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 241.2 ms.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 5.8% over 11,565 samples on the same window, 28th lowest of the 31 models Coval measures.

Voice quality

Neither Artificial Analysis nor Speko rates its voice quality, so it counts 0 of 20.

Price

No price is published by Artificial Analysis or Speko for this model; price counts 0 of 5.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
271.2 ms
P50 / median time-to-first-audio (TTFA)
261.3 ms
P90 TTFA
318.2 ms
P95 TTFA
343.1 ms
P99 TTFA (slowest turns)
427.4 ms
Latency consistency (standard deviation)
80.8 ms
Leading silence
30 ms
Time to first byte (network arrival)
241.2 ms
Accuracy: word error rate (WER)
5.8%
WER basis
per-clip mean
Latency runs
11,585 runs
Accuracy runs
11,565 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
dd-etts-3.3
Coval model page
benchmarks.coval.ai

Not collected

Similar TTS models