Aller au contenu

Palabra TTS v1 Palabra · text-to-speech (TTS) API

Strong

105 ms P50 time-to-first-audio on Coval; ranked 7th of 112 in the TTS for Voice Agents 2026 index.

Benchmark summary: 58.6/100, ranked 7th of 112 in the TTS for Voice Agents 2026 index (Strong). Strongest criterion: real-time speed, 49.9 of 60 points.

Palabra TTS v1 is a proprietary text-to-speech (TTS) API from Palabra, measured by two independent benchmarks: Coval and Speko.

Overall score

58.6 /100

7 th of 112

Score breakdown

  • Real-time speed 49.9 /60
  • Voice quality 4.8 /20
  • Accuracy 2.4 /15
  • Price 1.5 /5

Data checked on September 30, 2026

Updated after each benchmark capture

Every point comes from a public benchmark: sources · methodology

Gradium TTS Beta vs Palabra TTS v1 Compare with Cartesia Sonic 3.5

Palabra TTS v1 ranks 7th of 112 in the 2026 text-to-speech (TTS) API benchmark, with a score of 58.6/100.

Real-time latency

On Coval's 30-day window ending October 1, 2026 (11,682 runs), Palabra TTS v1 has a P50 (median) time-to-first-audio (TTFA) of 105 ms, 6th lowest of the 31 models Coval measures. Its P99, the slowest 1% of turns, is 299.8 ms (7th lowest of the 31 models Coval measures). Latency consistency: a standard deviation of 42.4 ms (6th lowest of the 31 models Coval measures). Mean time to first byte (TTFB) is 111.9 ms. Speko measures a P50 TTFA of 72 ms from US East (n=30), dated July 3, 2026: 2nd lowest of the 29 models Speko measures.

Accuracy (word error rate)

Coval measures an accuracy (word error rate) of 5.6% over 11,662 samples on the same window, 27th lowest of the 31 models Coval measures. Speko rates its pronunciation robustness on numbers, dates and currency at 0.16 on a 0-to-1 scale (25th highest of the 28 models Speko rates).

Voice quality

Speko rates its naturalness at 1,520 Elo from blind A/B votes (22nd highest of the 41 models Speko rates).

Price

Speko lists a cost of about $30 per 1M characters.

Scoring note

Speko changed its pronunciation robustness between the captures of September 30, 2026 and October 1, 2026 without a new measurement date (0.3, then 0.16): the score uses the average, 0.23. The value shown is the latest.

Benchmark measurements

Latency and accuracy — Coval (30-day window)

Mean TTFA
114.6 ms
P50 / median time-to-first-audio (TTFA)
105 ms
P90 TTFA
130.4 ms
P95 TTFA
147 ms
P99 TTFA (slowest turns)
299.8 ms
Latency consistency (standard deviation)
42.4 ms
Leading silence
2.7 ms
Time to first byte (network arrival)
111.9 ms
Accuracy: word error rate (WER)
5.6%
WER basis
per-clip mean
Latency runs
11,682 runs
Accuracy runs
11,662 runs
Window
30d
Coval snapshot
2026-10-01T10:11:47.318752Z
Coval model ID
palabra-tts-v1
Coval model page
benchmarks.coval.ai

Speko

P50 TTFA, US East
72 ms
Naturalness (Elo)
1,520
Pronunciation robustness (0 to 1)
0.16
Voice drift (lower is steadier)
11
Price per 1M characters
30 USD
Measured on
2026-07-03
Speko model
palabra-tts-v1
Speko model page
benchmarks.speko.ai

Not collected

Similar TTS models