Aller au contenu

Google Chirp 3: HD Google · text-to-speech (TTS) API

Listed

515.2 ms P50 time-to-first-audio on Coval; ranked 32nd of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 18.5/100, ranked 32nd of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: accuracy, 3.4 of 15 points.

Google Chirp 3: HD is a proprietary text-to-speech (TTS) API from Google, measured by two independent benchmarks: Artificial Analysis and Coval.

Overall score

18.5 /100

32 nd of 112

Score breakdown

  • Real-time speed 10 /60
  • Voice quality 4 /20
  • Accuracy 3.4 /15
  • Price 1.1 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Google Chirp 3: HD Compare with Nari Qwen3-TTS

Google Chirp 3: HD ranks 32nd of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 18.5/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,855 runs), Google Chirp 3: HD has a P50 (median) time-to-first-audio (TTFA) of 515.2 ms, 27th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 1,282.2 ms (27th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 225.8 ms (22nd lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 189.6 ms.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 5% over 11,835 samples on the same window, 17th lowest of the 31 models Coval measures.

Voice quality

In the Artificial Analysis Speech Arena (capture of October 1, 2026), Google Chirp 3: HD has an Elo of 1,053 (±12, 95% confidence) from 3,296 blind comparisons, 56th highest of the 92 models Artificial Analysis rates.

Price

Artificial Analysis lists Google Chirp 3: HD at $30 per 1M characters (43rd lowest of the 78 models Artificial Analysis lists a price for).

Languages tested

Tested in Japanese, Hindi, Tamil, Telugu and Thai by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
542.5 ms
P50 / median time-to-first-audio (TTFA)
515.2 ms
P90 TTFA
814.3 ms
P95 TTFA
899.2 ms
P99 TTFA (slowest turns)
1,282.2 ms
Latency consistency (standard deviation)
225.8 ms
Leading silence
352.9 ms
Time to first byte (network arrival)
189.6 ms
Accuracy: word error rate (WER)
5%
WER basis
per-clip mean
Latency runs
11,855 runs
Accuracy runs
11,835 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
chirp-3-hd
Coval model page
benchmarks.coval.ai

Artificial Analysis Speech Arena

Speech Arena Elo (blind listener preference)
1,053
Elo 95% interval (±)
12
Arena rank
57
Arena rank range
49-62
Arena appearances
3,296
Voices tested
8
Price per 1M characters
30 USD
Released
Mar 2025
Artificial Analysis model
Chirp 3: HD
Artificial Analysis leaderboard
artificialanalysis.ai

Not collected

Similar TTS models