Aller au contenu

OpenAI GPT-4o mini TTS OpenAI · text-to-speech (TTS) API

Listed

567.4 ms P50 time-to-first-audio on Coval; ranked 34th (tied) of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 14.4/100, ranked 34th (tied) of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: accuracy, 10.4 of 15 points.

OpenAI GPT-4o mini TTS is a proprietary text-to-speech (TTS) API from OpenAI, measured by two independent benchmarks: Coval and Speko.

Overall score

14.4 /100

34 th of 112

Score breakdown

  • Real-time speed 1.3 /60
  • Voice quality 1 /20
  • Accuracy 10.4 /15
  • Price 1.7 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs OpenAI GPT-4o mini TTS Compare with Nari Qwen3-TTS

OpenAI GPT-4o mini TTS ranks 34th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 14.4/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,780 runs), OpenAI GPT-4o mini TTS has a P50 (median) time-to-first-audio (TTFA) of 567.4 ms, 30th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 5,896.6 ms (30th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 1,575.6 ms (30th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 962.7 ms. Speko measures a P50 TTFA of 691 ms from US East (n=30), dated July 3, 2026: 26th lowest of the 29 models Speko measures.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 4.7% over 11,801 samples on the same window, 12th lowest of the 31 models Coval measures. Speko rates its pronunciation robustness on numbers, dates and currency at 0.44 on a 0-to-1 scale (10th highest of the 28 models Speko rates).

Voice quality

Speko rates its naturalness at 1,424 Elo from blind A/B votes (37th highest of the 41 models Speko rates).

Price

Speko lists a cost of about $20 per 1M characters.

Scoring note

Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.93, then 0.44): the score uses the average, 0.685. The value shown is the latest.

Languages tested

Tested in Spanish, German, French, Japanese, Chinese, Korean, Filipino, Norwegian and Thai by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
1,097.3 ms
P50 / median time-to-first-audio (TTFA)
567.4 ms
P90 TTFA
3,459 ms
P95 TTFA
4,259.5 ms
P99 TTFA (slowest turns)
5,896.6 ms
Latency consistency (standard deviation)
1,575.6 ms
Leading silence
128.1 ms
Time to first byte (network arrival)
962.7 ms
Accuracy: word error rate (WER)
4.7%
WER basis
per-clip mean
Latency runs
11,780 runs
Accuracy runs
11,801 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
gpt-4o-mini-tts
Coval model page
benchmarks.coval.ai

Speko

P50 TTFA, US East
691 ms
Naturalness (Elo)
1,424
Head-to-head win rate
37%
Pronunciation robustness (0 to 1)
0.44
Voice drift (lower is steadier)
22
Price per 1M characters
20 USD
Measured on
2026-07-03
Speko model
gpt-4o-mini-tts
Speko model page
benchmarks.speko.ai

Not collected

Similar TTS models