Aller au contenu

Methodology

How the TTS benchmark is built: three independent sources, 4 criteria, no estimates

TTS for Voice Agents ranks 112 text-to-speech (TTS) models for real-time voice agents. It measures nothing itself: every number comes from one of three independent benchmarks, Coval, Speko and Artificial Analysis, and every model page links to the source and gives the date.

What it adds is a set of rules, applied the same way to every model: which measurements count, how much each weighs, and what happens when a benchmark does not measure a model. Nothing is estimated: a measurement a benchmark does not publish counts zero, and the model page says why.

Publisher: Rank-ly, company being registered. Contact: contact@rank-ly.com.

A missing model, an outdated value? Submit a model or contact us.

The three sources

Coval

31 models measured

Time-to-first-audio and word error rate, measured continuously from us-east-1 every 30 minutes. We use the 30-day aggregate ending October 1, 2026: mean, P50, P90, P95, P99 and standard deviation of TTFA, 781 to 11856 runs per model.

Speko

43 models measured

P50 time-to-first-audio from US East (n=30 per model), naturalness Elo from blind A/B votes, pronunciation robustness on numbers, dates and currency, voice drift and published cost. Also its multilingual boards, for the languages tested.

Artificial Analysis

92 models measured

Speech Arena Elo from blind listener preference, with its 95% interval and number of appearances; published API price per 1M characters; open-weights flag; listening categories (customer service, assistants, entertainment, knowledge sharing); controlled-voice boards by language.

The score

How the 100 points are split

Real-time speed

60 pts

Time to first audio. Coval (30-day mean 50%, P99 25%, standard deviation 25%) and Speko (US East P50), each source weighted by the square root of the number of runs it publishes for the model (Coval: thousands per model; Speko: n=30). Not measured = 0.

Voice quality

20 pts

Blind listener preference: Artificial Analysis Speech Arena Elo and Speko naturalness, equal weights (Speko publishes no vote count). Values a source changed between two captures without a new measurement date are averaged across captures.

Accuracy

15 pts

Coval word error rate (30 days) and Speko robustness on numbers, dates and currency, equal weights. Undated recalculations are averaged across captures.

Price

5 pts

Published price per million characters: Artificial Analysis and Speko, equal weights. Lower is better.

The rules

  1. Rank, not raw value. Each measurement becomes a score from 0 to 1 by its percentile rank among the models measured: latencies range from under 50 ms to several seconds, and a linear scale would flatten everything but the extremes. Ties share the average rank.
  2. Not measured counts zero. A model a benchmark does not measure gets zero on that measurement, and its page lists it under "Not collected" with the reason.
  3. Sources weigh by sample size. Within a criterion, each source weighs by the square root of the number of runs it publishes for the model, the standard way to combine measurements of unequal precision. Coval publishes thousands of latency runs per model and Speko 30, so Coval carries most of the speed score. Where a source publishes no run count (Speko naturalness and robustness), the sources stay at equal weights.
  4. Undated recalculations are averaged. When a benchmark changes a value between two captures without dating a new measurement, the score uses the average of the captures, so a silent recalculation cannot flip the ranking from one day to the next. The model page shows the latest published value.

Labels

Score tiers

The score ranks; the tier groups. It follows mechanically from the score, with no editorial override.

Frequently asked questions

Is this TTS benchmark independent?
It is built only on measurements published by three independent benchmarks, Coval, Speko and Artificial Analysis, and links to each source page. It adds no measurement of its own; what it adds is the scoring rules on this page, applied the same way to every model.
What happens when a benchmark does not measure a model?
The missing measurement counts zero, and the model page lists it under "Not collected" with the reason: not measured by the benchmark, or listed without that value. Nothing is estimated.
How often is the benchmark updated?
The three sources are captured again and every score is recomputed after each benchmark capture. Each model page shows the date its data was checked; the latest check was on September 30, 2026.
What does the benchmark not measure?
Developer experience (SDKs, documentation), SSML support, word timestamps and voice design: none of the three benchmarks measures them per model, so they are not scored.

Get benchmark updates

An occasional email when the sources are captured again or the scores change.

Benchmark updates only. Unsubscribe at any time. See our privacy policy.