Aller au contenu

Alibaba Qwen3 TTS Flash Realtime Alibaba (Qwen) · text-to-speech (TTS) API

Listed

708.8 ms P50 time-to-first-audio on Coval; ranked 84th of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 4.3/100, ranked 84th of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: real-time speed, 4.3 of 60 points.

Alibaba Qwen3 TTS Flash Realtime is a proprietary text-to-speech (TTS) API from Alibaba (Qwen), measured by one independent benchmark, Coval.

Overall score

4.3 /100

84 th of 112

Score breakdown

  • Real-time speed 4.3 /60
  • Voice quality 0 /20
  • Accuracy 0 /15
  • Price 0 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Alibaba Qwen3 TTS Flash Realtime Compare with ElevenLabs Turbo v2

Alibaba Qwen3 TTS Flash Realtime ranks 84th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 4.3/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,470 runs), Alibaba Qwen3 TTS Flash Realtime has a P50 (median) time-to-first-audio (TTFA) of 708.8 ms, 31st lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 4,208.7 ms (29th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 566.7 ms (28th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 646.2 ms.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 8.7% over 11,450 samples on the same window, 31st lowest of the 31 models Coval measures.

Voice quality

Neither Artificial Analysis nor Speko rates its voice quality, so it counts 0 of 20.

Price

No price is published by Artificial Analysis or Speko for this model; price counts 0 of 5.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
742.7 ms
P50 / median time-to-first-audio (TTFA)
708.8 ms
P90 TTFA
783.8 ms
P95 TTFA
843.8 ms
P99 TTFA (slowest turns)
4,208.7 ms
Latency consistency (standard deviation)
566.7 ms
Leading silence
96.5 ms
Time to first byte (network arrival)
646.2 ms
Accuracy: word error rate (WER)
8.7%
WER basis
per-clip mean
Latency runs
11,470 runs
Accuracy runs
11,450 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
qwen3-tts-flash-realtime
Coval model page
benchmarks.coval.ai

Not collected

Similar TTS models