Aller au contenu

ElevenLabs Flash v2.5 ElevenLabs · text-to-speech (TTS) API

Strong

184.8 ms P50 time-to-first-audio on Coval; ranked 11th of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 56.1/100, ranked 11th of 112 in the TTS for Voice Agents 2026 index (Strong). Strongest criterion: real-time speed, 45.7 of 60 points.

ElevenLabs Flash v2.5 is a proprietary text-to-speech (TTS) API from ElevenLabs, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.

Overall score

56.1 /100

11 th of 112

Score breakdown

  • Real-time speed 45.7 /60
  • Voice quality 8.7 /20
  • Accuracy 0.2 /15
  • Price 1.5 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs ElevenLabs Flash v2.5 Compare with SpeechifyAI Simba 3.2

ElevenLabs Flash v2.5 ranks 11th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 56.1/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,842 runs), ElevenLabs Flash v2.5 has a P50 (median) time-to-first-audio (TTFA) of 184.8 ms, 8th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 332.4 ms (8th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 33.2 ms (4th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 148.2 ms.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 6.5% over 11,822 samples on the same window, 30th lowest of the 31 models Coval measures.

Voice quality

In the Artificial Analysis Speech Arena (capture of October 1, 2026), ElevenLabs Flash v2.5 has an Elo of 1,076 (±11, 95% confidence) from 5,643 blind comparisons, 45th highest of the 92 models Artificial Analysis rates. Speko rates its naturalness at 1,494 Elo from blind A/B votes (27th highest of the 41 models Speko rates).

Price

Artificial Analysis lists ElevenLabs Flash v2.5 at $50 per 1M characters (57th lowest of the 78 models Artificial Analysis lists a price for). Speko lists a cost of about $50 per 1M characters.

Languages tested

Tested in Spanish, German, French, Portuguese, Japanese, Chinese, Arabic, Hindi and Vietnamese by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
188.7 ms
P50 / median time-to-first-audio (TTFA)
184.8 ms
P90 TTFA
215.6 ms
P95 TTFA
228.6 ms
P99 TTFA (slowest turns)
332.4 ms
Latency consistency (standard deviation)
33.2 ms
Leading silence
40.5 ms
Time to first byte (network arrival)
148.2 ms
Accuracy: word error rate (WER)
6.5%
WER basis
per-clip mean
Latency runs
11,842 runs
Accuracy runs
11,822 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
eleven_flash_v2_5
Coval model page
benchmarks.coval.ai

Speko

Naturalness (Elo)
1,494
Price per 1M characters
50 USD
Measured on
2026-09-08
Speko model
eleven_flash_v2_5
Speko model page
benchmarks.speko.ai

Artificial Analysis Speech Arena

Speech Arena Elo (blind listener preference)
1,076
Elo 95% interval (±)
11
Arena rank
45
Arena rank range
41-48
Arena appearances
5,643
Voices tested
7
Price per 1M characters
50 USD
Released
Dec 2024
Artificial Analysis model
Flash v2.5
Artificial Analysis leaderboard
artificialanalysis.ai

Not collected

Similar TTS models