Aller au contenu

MiniMax Speech 2.8 HD MiniMax · text-to-speech (TTS) API

Listed

294 ms P50 time-to-first-audio on Speko; ranked 39th of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 11.4/100, ranked 39th of 112 in the TTS for Voice Agents 2026 index (Listed). Strongest criterion: voice quality, 10 of 20 points.

MiniMax Speech 2.8 HD is a proprietary text-to-speech (TTS) API from MiniMax, measured by two independent benchmarks: Artificial Analysis and Speko.

Overall score

11.4 /100

39 th of 112

Score breakdown

  • Real-time speed 1 /60
  • Voice quality 10 /20
  • Accuracy 0.1 /15
  • Price 0.3 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs MiniMax Speech 2.8 HD Compare with Alibaba Qwen-Audio-3.0-TTS-Plus

MiniMax Speech 2.8 HD ranks 39th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 11.4/100.

Real-time latency

Speko measures a P50 TTFA of 294 ms from US East (n=30), dated July 3, 2026: 19th lowest of the 29 models Speko measures.

Accuracy (word error rate)

Speko rates its pronunciation robustness on numbers, dates and currency at 0.15 on a 0-to-1 scale (26th highest of the 28 models Speko rates).

Voice quality

In the Artificial Analysis Speech Arena (capture of October 1, 2026), MiniMax Speech 2.8 HD has an Elo of 1,170 (±11, 95% confidence) from 4,560 blind comparisons, 17th highest of the 92 models Artificial Analysis rates. Speko rates its naturalness at 1,431 Elo from blind A/B votes (34th highest of the 41 models Speko rates).

Price

Artificial Analysis lists MiniMax Speech 2.8 HD at $100 per 1M characters (69th lowest of the 78 models Artificial Analysis lists a price for). Speko lists a cost of about $100 per 1M characters.

Scoring note

Artificial Analysis changed its Speech Arena Elo between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (1,171, then 1,170): the score uses the average, 1,170.5. The value shown is the latest. Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0, then 0.15): the score uses the average, 0.075. The value shown is the latest.

Languages tested

Tested in Spanish, German, French, Portuguese, Japanese, Chinese, Arabic, Hindi, Vietnamese, Filipino and Norwegian by Artificial Analysis or Speko. A language missing here was not tested, which does not mean it is unsupported.

Benchmark measurements

Speko

P50 TTFA, US East
294 ms
Naturalness (Elo)
1,431
Head-to-head win rate
36%
Pronunciation robustness (0 to 1)
0.15
Voice drift (lower is steadier)
12
Price per 1M characters
100 USD
Measured on
2026-07-03
Speko model
speech-2.8-hd
Speko model page
benchmarks.speko.ai

Artificial Analysis Speech Arena

Speech Arena Elo (blind listener preference)
1,170
Elo 95% interval (±)
11
Arena rank
18
Arena rank range
15-18
Arena appearances
4,560
Voices tested
8
Price per 1M characters
100 USD
Released
Feb 2026
Artificial Analysis model
Speech 2.8 HD
Artificial Analysis leaderboard
artificialanalysis.ai

Not collected

Similar TTS models