Aller au contenu

Deepgram Aura 2 En Deepgram · text-to-speech (TTS) API

Listed

520.2 ms P50 time-to-first-audio on Coval; ranked 34th (tied) of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 14.4/100, ranked 34th (tied) of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: accuracy, 7.5 of 15 points.

Deepgram Aura 2 En is a proprietary text-to-speech (TTS) API from Deepgram, served by Cloudflare, measured by one independent benchmark, Coval.

Overall score

14.4 /100

34 th of 112

Score breakdown

  • Real-time speed 6.9 /60
  • Voice quality 0 /20
  • Accuracy 7.5 /15
  • Price 0 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Deepgram Aura 2 En Compare with Nari Qwen3-TTS

Deepgram Aura 2 En ranks 34th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 14.4/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (1,116 runs), Deepgram Aura 2 En has a P50 (median) time-to-first-audio (TTFA) of 520.2 ms, 28th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 1,696.9 ms (28th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 326.6 ms (24th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 360.6 ms.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 2.5% over 1,116 samples on the same window, the lowest of the 31 models Coval measures.

Voice quality

Neither Artificial Analysis nor Speko rates its voice quality, so it counts 0 of 20.

Price

No price is published by Artificial Analysis or Speko for this model; price counts 0 of 5.

Provider and host

Coval measures this model through Cloudflare's API, so the latency above is Cloudflare's. Coval's model registry names Deepgram as the creator and Cloudflare as the host.

Benchmark measurements

Hosting

Served by
Cloudflare

Latency and accuracy — Coval (30-day window)

Mean TTFA
557.9 ms
P50 / median time-to-first-audio (TTFA)
520.2 ms
P90 TTFA
829 ms
P95 TTFA
1,059.1 ms
P99 TTFA (slowest turns)
1,696.9 ms
Latency consistency (standard deviation)
326.6 ms
Leading silence
197.3 ms
Time to first byte (network arrival)
360.6 ms
Accuracy: word error rate (WER)
2.5%
WER basis
pooled
Latency runs
1,116 runs
Accuracy runs
1,116 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
aura-2-en
Coval model page
benchmarks.coval.ai

Not collected

Similar TTS models