TTS API comparison 2026
Deepgram Flux TTS vs Google Gemini 3.8 Flash TTS Deepgram vs Google · text-to-speech for voice agents
Deepgram Flux TTS ranks 13th of 112 with 49.6/100; Google Gemini 3.8 Flash TTS ranks 24th with 29.7/100: a gap of 19.9 points on the same public rules, built on Coval, Speko and Artificial Analysis.
Data checked on September 30, 2026 · Coval 30-day window ending October 1, 2026 · updated after each benchmark capture
Side-by-side benchmark
| Measure | Deepgram Flux TTS | Google Gemini 3.8 Flash TTS |
|---|---|---|
| Overall score | 49.6/100 | 29.7/100 |
| Rank | 13th of 112 | 24th of 112 |
| Real-time speed /60 | 36.4 | 0 |
| Voice quality /20 | 5.8 | 19.3 |
| Accuracy /15 | 6.2 | 6.9 |
| Price /5 | 1.2 | 3.5 |
| P50 / median time-to-first-audio Coval, 30-day window | 193.9 ms | not measured |
| P99 time-to-first-audio (slowest turns) Coval, 30-day window | 550.3 ms | not measured |
| Latency consistency (standard deviation) Coval, 30-day window | 131 ms | not measured |
| P50 time-to-first-audio, US East Speko | 106 ms | 1031 ms |
| Accuracy: word error rate (WER) Coval, 30-day window | 4.9% | not measured |
| Pronunciation robustness (0 to 1) Speko | 0.2 | 0.7 |
| Speech Arena Elo (blind listener preference) Artificial Analysis | not measured | 1272 |
| Naturalness (Elo) Speko | 1550 | 1642 |
| Price per 1M characters Artificial Analysis | not measured | 16.5 USD |
Choose Deepgram Flux TTS if…
- → Ranks 13th of 112 in the overall benchmark, against 24th for Google Gemini 3.8 Flash TTS
- → Better P50 time-to-first-audio, US East: 106 ms vs 1031 ms (Speko)
Choose Google Gemini 3.8 Flash TTS if…
- → Better pronunciation robustness (0 to 1): 0.7 vs 0.2 (Speko)
- → Better naturalness (Elo): 1642 vs 1550 (Speko)
Frequently asked questions
- What is the difference between Deepgram Flux TTS and Google Gemini 3.8 Flash TTS?
- In the 2026 TTS API benchmark, Deepgram Flux TTS ranks 13th of 112 with 49.6/100 and Google Gemini 3.8 Flash TTS ranks 24th with 29.7/100, a gap of 19.9 points on the same rules.
- Which is faster for real-time voice agents, Deepgram Flux TTS or Google Gemini 3.8 Flash TTS?
- On Coval's 30-day window ending October 1, 2026, the P50 time-to-first-audio is 193.9 ms for Deepgram Flux TTS and not measured for Google Gemini 3.8 Flash TTS; the P99 is 550.3 ms and not measured.
- On which criteria does Google Gemini 3.8 Flash TTS beat Deepgram Flux TTS?
- Google Gemini 3.8 Flash TTS scores higher on Voice quality (19.3 vs 5.8 out of 20), Accuracy (6.9 vs 6.2 out of 15), Price (3.5 vs 1.2 out of 5).
- How are these two text-to-speech APIs scored?
- Both are scored by the same public rules: 4 criteria, 100 points (Real-time speed 60, Voice quality 20, Accuracy 15, Price 5), computed from three independent benchmarks: Coval, Speko and Artificial Analysis. A measurement a benchmark does not publish counts zero, and the model page says why.
- Can these numbers be checked?
- Yes. Each model page lists every measurement with its source and date, and links to the benchmark page it comes from. The rules are on the methodology page.