Aller au contenu

SpeechifyAI Simba 3.2 Speechify · text-to-speech (TTS) API

Contender

362.7 ms P50 time-to-first-audio on Coval; ranked 12th of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 51.6/100, ranked 12th of 112 in the TTS for Voice Agents 2026 index (Contender). Strongest criterion: price, 4.6 of 5 points.

SpeechifyAI Simba 3.2 is a proprietary text-to-speech (TTS) API from Speechify, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.

Overall score

51.6 /100

12 th of 112

Score breakdown

  • Real-time speed 19.6 /60
  • Voice quality 16.4 /20
  • Accuracy 11 /15
  • Price 4.6 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs SpeechifyAI Simba 3.2 Compare with Deepgram Flux TTS

SpeechifyAI Simba 3.2 ranks 12th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 51.6/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,833 runs), SpeechifyAI Simba 3.2 has a P50 (median) time-to-first-audio (TTFA) of 362.7 ms, 22nd lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 921.5 ms (26th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 388.2 ms (26th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 298.8 ms. Speko measures a P50 TTFA of 85 ms from US East (n=30), dated July 3, 2026: 5th lowest of the 29 models Speko measures.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 4.3% over 11,826 samples on the same window, 5th lowest of the 31 models Coval measures. Speko rates its pronunciation robustness on numbers, dates and currency at 0.21 on a 0-to-1 scale (20th highest of the 28 models Speko rates).

Voice quality

In the Artificial Analysis Speech Arena (capture of October 1, 2026), SpeechifyAI Simba 3.2 has an Elo of 1,239 (±13, 95% confidence) from 2,669 blind comparisons, 7th highest of the 92 models Artificial Analysis rates. Speko rates its naturalness at 1,573 Elo from blind A/B votes (13th highest of the 41 models Speko rates).

Price

Artificial Analysis lists SpeechifyAI Simba 3.2 at $6.6 per 1M characters (6th lowest of the 78 models Artificial Analysis lists a price for). Speko lists a cost of about $10 per 1M characters.

Scoring note

Artificial Analysis changed its Speech Arena Elo between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (1,241, then 1,239): the score uses the average, 1,240. The value shown is the latest. Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.7, then 0.21): the score uses the average, 0.455. The value shown is the latest.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
323.6 ms
P50 / median time-to-first-audio (TTFA)
362.7 ms
P90 TTFA
480.7 ms
P95 TTFA
631.6 ms
P99 TTFA (slowest turns)
921.5 ms
Latency consistency (standard deviation)
388.2 ms
Leading silence
24.9 ms
Time to first byte (network arrival)
298.8 ms
Accuracy: word error rate (WER)
4.3%
WER basis
per-clip mean
Latency runs
11,833 runs
Accuracy runs
11,826 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
simba-3.2
Coval model page
benchmarks.coval.ai

Speko

P50 TTFA, US East
85 ms
Naturalness (Elo)
1,573
Head-to-head win rate
63%
Pronunciation robustness (0 to 1)
0.21
Voice drift (lower is steadier)
7
Price per 1M characters
10 USD
Measured on
2026-07-03
Speko model
simba-3.2
Speko model page
benchmarks.speko.ai

Artificial Analysis Speech Arena

Speech Arena Elo (blind listener preference)
1,239
Elo 95% interval (±)
13
Arena rank
7
Arena rank range
5-8
Arena appearances
2,669
Voices tested
8
Price per 1M characters
6.6 USD
Released
Jul 2026
Artificial Analysis model
Simba 3.2
Artificial Analysis leaderboard
artificialanalysis.ai

Similar TTS models