Aller au contenu

Alibaba Qwen3 TTS 1.7b Alibaba (Qwen) · text-to-speech (TTS) API

Contender

103.5 ms P50 time-to-first-audio on Coval; ranked 17th of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 45.3/100, ranked 17th of 112 in the TTS for Voice Agents 2026 index (Contender). Strongest criterion: real-time speed, 43.9 of 60 points.

Alibaba Qwen3 TTS 1.7b is an open-weights text-to-speech (TTS) API from Alibaba (Qwen), served by Baseten, measured by one independent benchmark, Coval.

Overall score

45.3 /100

17 th of 112

Score breakdown

  • Real-time speed 43.9 /60
  • Voice quality 0 /20
  • Accuracy 1.4 /15
  • Price 0 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Alibaba Qwen3 TTS 1.7b Compare with Rime Mist v3

Alibaba Qwen3 TTS 1.7b ranks 17th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 45.3/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (781 runs), Alibaba Qwen3 TTS 1.7b has a P50 (median) time-to-first-audio (TTFA) of 103.5 ms, 5th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 165.4 ms (4th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 37.5 ms (5th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 72.8 ms.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 5.4% over 781 samples on the same window, 25th lowest of the 31 models Coval measures.

Voice quality

Neither Artificial Analysis nor Speko rates its voice quality, so it counts 0 of 20.

Price

No price is published by Artificial Analysis or Speko for this model; price counts 0 of 5.

Provider and host

Coval measures this model through Baseten's API, so the latency above is Baseten's. Coval's model registry names Alibaba (Qwen) as the creator and Baseten as the host.

Benchmark measurements

Hosting

Served by
Baseten

Latency and accuracy — Coval (30-day window)

Mean TTFA
106.5 ms
P50 / median time-to-first-audio (TTFA)
103.5 ms
P90 TTFA
127.9 ms
P95 TTFA
136.1 ms
P99 TTFA (slowest turns)
165.4 ms
Latency consistency (standard deviation)
37.5 ms
Leading silence
33.7 ms
Time to first byte (network arrival)
72.8 ms
Accuracy: word error rate (WER)
5.4%
WER basis
per-clip mean
Latency runs
781 runs
Accuracy runs
781 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
qwen3-tts-1.7b
Coval model page
benchmarks.coval.ai

Not collected

Similar TTS models