Aller au contenu

Alibaba Qwen3 TTS Fast Alibaba (Qwen) · text-to-speech (TTS) API

Leader

64.6 ms P50 time-to-first-audio on Coval; ranked 2nd (tied) of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 73.7/100, ranked 2nd (tied) of 112 in the TTS for Voice Agents 2026 index (Leader). Strongest criterion: real-time speed, 57.7 of 60 points.

Alibaba Qwen3 TTS Fast is an open-weights text-to-speech (TTS) API from Alibaba (Qwen), served by Nari Labs, measured by two independent benchmarks: Coval and Speko.

Overall score

73.7 /100

2 nd of 112

Score breakdown

  • Real-time speed 57.7 /60
  • Voice quality 5.3 /20
  • Accuracy 8.4 /15
  • Price 2.3 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Alibaba Qwen3 TTS Fast Compare with Gradium TTS Beta

Alibaba Qwen3 TTS Fast ranks 2nd of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 73.7/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (7,083 runs), Alibaba Qwen3 TTS Fast has a P50 (median) time-to-first-audio (TTFA) of 64.6 ms, 3rd lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 125.3 ms (2nd lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 14.6 ms (the lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 60.9 ms. Speko measures a P50 TTFA of 64 ms from US East (n=30), dated September 12, 2026: the lowest of the 29 models Speko measures.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 3.7% over 7,076 samples on the same window, 2nd lowest of the 31 models Coval measures. Speko rates its pronunciation robustness on numbers, dates and currency at 0.3 on a 0-to-1 scale (14th highest of the 28 models Speko rates).

Voice quality

Speko rates its naturalness at 1,536 Elo from blind A/B votes (20th highest of the 41 models Speko rates).

Price

Speko lists a cost of about $10 per 1M characters.

Provider and host

Coval and Speko measure this model through Nari Labs' API, so the latency above is Nari Labs'. Coval's model registry names Alibaba (Qwen) as the creator and Nari Labs as the host.

Scoring note

Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.15, then 0.3): the score uses the average, 0.225. The value shown is the latest.

Benchmark measurements

Hosting

Served by
Nari Labs

Latency and accuracy — Coval (30-day window)

Mean TTFA
70.3 ms
P50 / median time-to-first-audio (TTFA)
64.6 ms
P90 TTFA
87.8 ms
P95 TTFA
100.2 ms
P99 TTFA (slowest turns)
125.3 ms
Latency consistency (standard deviation)
14.6 ms
Leading silence
9.3 ms
Time to first byte (network arrival)
60.9 ms
Accuracy: word error rate (WER)
3.7%
WER basis
pooled
Latency runs
7,083 runs
Accuracy runs
7,076 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
qwen3-tts-fast
Coval model page
benchmarks.coval.ai

Speko

P50 TTFA, US East
64 ms
Naturalness (Elo)
1,536
Pronunciation robustness (0 to 1)
0.3
Voice drift (lower is steadier)
0
Deterministic output
✓
Price per 1M characters
10 USD
Measured on
2026-09-12
Speko model
Qwen3-TTS Fast
Speko model page
benchmarks.speko.ai

Not collected

Similar TTS models