Aller au contenu

Real-time TTS · speed only · 2026

Best real-time TTS API, ranked on measured speed

This ranking looks at one thing only: how fast each text-to-speech API starts speaking. It orders the 44 models whose latency is measured by Coval or Speko on P50 and P99 time-to-first-audio (TTFA) and latency consistency, nothing else. Voice quality, accuracy and price are not counted here; for the realtime choice that weighs all four, see the overall TTS benchmark.

Coval 30-day window ending October 1, 2026 · Speko US East P50 · updated after each benchmark capture

# Model Speed /60 P50 TTFA (Coval) P99 TTFA (Coval) Consistency (std dev) P50 TTFA (Speko)
1 Gradium TTS Beta Gradium 58.7 47.2 ms 108.8 ms 16.6 ms 100 ms
2 Alibaba Qwen3 TTS Fast Alibaba (Qwen) 57.7 64.6 ms 125.3 ms 14.6 ms 64 ms
3 Palabra TTS v1 Palabra 49.9 105 ms 299.8 ms 42.4 ms 72 ms
4 ElevenLabs Flash v2.5 ElevenLabs 45.7 184.8 ms 332.4 ms 33.2 ms —
5 Fluxions vui Fluxions 44.6 50.2 ms 128.1 ms 241.6 ms —
6 Alibaba Qwen3 TTS 1.7b Alibaba (Qwen) 43.9 103.5 ms 165.4 ms 37.5 ms —
7 Inworld Realtime TTS-2 Flash Inworld 43.5 69.4 ms 175.2 ms 373.3 ms 82 ms
8 Soniox TTS Real-Time v2 Soniox 40.1 245.6 ms 343.4 ms 49.8 ms 381 ms
9 Rime Mist v3 Rime 40 254.8 ms 364.3 ms 26.1 ms —
10 Gradium TTS (Aug 2026) Gradium 39.7 213.5 ms 382.5 ms 70.2 ms 272 ms
11 Soniox TTS RT v1 Soniox 37.9 247.2 ms 339.8 ms 104.8 ms 362 ms
12 Inworld Realtime TTS-2 Inworld 37.8 163.1 ms 293.7 ms 567.8 ms 116 ms
13 Cartesia Sonic 3.5 Cartesia 37.1 280 ms 421.8 ms 45.6 ms 121 ms
14 Deepgram Flux TTS Deepgram 36.4 193.9 ms 550.3 ms 131 ms 106 ms
15 Deepdub Phantom Z 3.4 conversational Deepdub 33.3 261.3 ms 427.4 ms 80.8 ms —
16 Rime Coda Rime 32.4 294.2 ms 438.1 ms 46 ms —
17 ElevenLabs Eleven v3 Conversational ElevenLabs 29.9 323.1 ms 497.7 ms 52.9 ms 266 ms
18 Deepgram Aura 2 Deepgram 27.9 287.9 ms 576 ms 140.8 ms 125 ms
19 Cartesia Sonic 3.6 Cartesia 21.5 347.9 ms 840 ms 117.6 ms 120 ms
20 Smallest.ai Lightning V3.1 Pro (Jul 2026) Smallest.ai 21.4 337.1 ms 793.8 ms 117.4 ms —
21 SpeechifyAI Simba 3.2 Speechify 19.6 362.7 ms 921.5 ms 388.2 ms 85 ms
22 SpaceXAI Grok TTS xAI 19.3 386.6 ms 542.4 ms 163.1 ms 272 ms
23 Fish Audio S2.1 Pro Fish Audio 18.6 300.1 ms 885.9 ms 193.9 ms —
24 Fish Audio OpenAudio S1 Fish Audio 17.1 364.1 ms 789.6 ms 163.5 ms —
25 Murf AI Falcon 2 Murf AI 14.8 557.5 ms 860.9 ms 141.3 ms —
26 SpeechifyAI Simba 3.0 Speechify 14.3 390.6 ms 906.3 ms 498.6 ms —
27 Google Chirp 3: HD Google 10 515.2 ms 1282.2 ms 225.8 ms —
28 Deepgram Aura 2 En Deepgram 6.9 520.2 ms 1696.9 ms 326.6 ms —
29 Alibaba Qwen3 TTS Flash Realtime Alibaba (Qwen) 4.3 708.8 ms 4208.7 ms 566.7 ms —
30 Nari Qwen3-TTS Alibaba (Qwen) 2.7 — — — 75 ms
31 Maya 2 Native Maya Research 2.3 — — — 100 ms
32 Smallest.ai Lightning v3.1 Smallest.ai 1.6 — — — 173 ms
33 ElevenLabs Eleven v4 Turbo ElevenLabs 1.5 — — — 205 ms
34 Rime arcanav3 Rime 1.4 — — — 238 ms
35 OpenAI GPT-4o mini TTS OpenAI 1.3 567.4 ms 5896.6 ms 1575.6 ms 691 ms
36 Fish Audio S2.1 Pro Free Fish Audio 1 405.6 ms 7919.1 ms 2855.9 ms —
36 MiniMax Speech 2.8 HD MiniMax 1 — — — 294 ms
38 Bland Speech v3 Bland AI 0.9 — — — 303 ms
39 Hume AI Octave 2 Hume AI 0.6 — — — 448 ms
40 Alibaba Qwen3 TTS Flash Alibaba (Qwen) 0.5 — — — 472 ms
41 Google Gemini 3.8 Flash-Lite TTS Google 0.4 — — — 497 ms
42 ElevenLabs Eleven v4 ElevenLabs 0.2 — — — 844 ms
43 Google Gemini 3.1 Flash TTS Google 0.1 — — — 978 ms
44 Google Gemini 3.8 Flash TTS Google 0 — — — 1031 ms

Frequently asked questions

Which TTS API should be used for real-time applications?
On measured speed alone, Gradium TTS Beta ranks first: 58.7 of 60 speed points, with a P50 time-to-first-audio of 47.2 ms and a P99 of 108.8 ms on Coval's 30-day window ending October 1, 2026. Speed is not the whole choice: the overall ranking also weighs voice quality, accuracy and price.
Which TTS API has the lowest P99 latency?
Gradium TTS Beta: 108.8 ms P99 time-to-first-audio on Coval's 30-day window. The P99 is the slowest 1% of turns, the ones a caller notices.
How is real-time speed measured here?
With two independent benchmarks. Coval measures time-to-first-audio continuously over 30 days (thousands of runs per model): we use its mean, P99 and standard deviation. Speko publishes a P50 measured from US East (n=30 per model). Each source weighs by the square root of the number of runs it publishes, so Coval carries most of the speed score.
Why are some TTS APIs missing from this page?
This page lists only the 44 models whose time-to-first-audio Coval or Speko measure. The 68 others are in the overall ranking, where their speed counts zero and their page says why.