Soniox TTS Real-Time v2
Soniox TTS Real-Time v2 is a proprietary text-to-speech (TTS) API from Soniox, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Language tested
29 text-to-speech (TTS) models, ranked from the highest benchmark score to the lowest, out of 100.
Data checked on September 30, 2026 · updated after each benchmark capture
29 of the 112 text-to-speech (TTS) models in this benchmark were tested in Chinese by Artificial Analysis (controlled-voice arena) or Speko (multilingual board). A model missing from this page was not tested in Chinese: that does not mean it cannot speak Chinese.
A model is listed here only when an independent benchmark measured its output in Chinese. The ranking below is the overall benchmark score (latency, accuracy, voice quality and price), not a Chinese-only score; the model pages give the measurements and their dates.
Soniox TTS Real-Time v2 is a proprietary text-to-speech (TTS) API from Soniox, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Inworld Realtime TTS-2 is a proprietary text-to-speech (TTS) API from Inworld, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Inworld Realtime TTS-2 Flash is a proprietary text-to-speech (TTS) API from Inworld, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Cartesia Sonic 3.5 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
ElevenLabs Eleven v3 Conversational is a proprietary text-to-speech (TTS) API from ElevenLabs, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
ElevenLabs Flash v2.5 is a proprietary text-to-speech (TTS) API from ElevenLabs, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Cartesia Sonic 3.6 is a proprietary text-to-speech (TTS) API from Cartesia, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
SpaceXAI Grok TTS is a proprietary text-to-speech (TTS) API from xAI, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Fish Audio S2.1 Pro is a proprietary text-to-speech (TTS) API from Fish Audio, measured by three independent benchmarks: Artificial Analysis, Coval and Speko.
Google Gemini 3.8 Flash TTS is a proprietary text-to-speech (TTS) API from Google, measured by two independent benchmarks: Artificial Analysis and Speko.
Google Gemini 3.8 Flash-Lite TTS is a proprietary text-to-speech (TTS) API from Google, measured by two independent benchmarks: Artificial Analysis and Speko.
Fish Audio OpenAudio S1 is a proprietary text-to-speech (TTS) API from Fish Audio, measured by two independent benchmarks: Artificial Analysis and Coval.
ElevenLabs Eleven v4 is a proprietary text-to-speech (TTS) API from ElevenLabs, measured by two independent benchmarks: Artificial Analysis and Speko.
Google Gemini 3.1 Flash TTS is a proprietary text-to-speech (TTS) API from Google, measured by two independent benchmarks: Artificial Analysis and Speko.
OpenAI GPT-4o mini TTS is a proprietary text-to-speech (TTS) API from OpenAI, measured by two independent benchmarks: Coval and Speko.
Cartesia Sonic 3 is a proprietary text-to-speech (TTS) API from Cartesia, measured by two independent benchmarks: Artificial Analysis and Speko.
MiniMax Speech 2.8 HD is a proprietary text-to-speech (TTS) API from MiniMax, measured by two independent benchmarks: Artificial Analysis and Speko.
Alibaba Qwen-Audio-3.0-TTS-Plus is a proprietary text-to-speech (TTS) API from Alibaba (Qwen), measured by one independent benchmark, Artificial Analysis.
ElevenLabs Turbo v2.5 is a proprietary text-to-speech (TTS) API from ElevenLabs, measured by two independent benchmarks: Artificial Analysis and Speko.
BreezeBlue Breeze TTS 2 is an open-weights text-to-speech (TTS) API from BreezeBlue, measured by one independent benchmark, Artificial Analysis.
MiniMax Speech 2.6 HD is a proprietary text-to-speech (TTS) API from MiniMax, measured by two independent benchmarks: Artificial Analysis and Speko.
ElevenLabs Multilingual v2 is a proprietary text-to-speech (TTS) API from ElevenLabs, measured by two independent benchmarks: Artificial Analysis and Speko.
MiniMax Speech 2.8 Turbo is a proprietary text-to-speech (TTS) API from MiniMax, measured by two independent benchmarks: Artificial Analysis and Speko.
Fish Audio S2 Pro is an open-weights text-to-speech (TTS) API from Fish Audio, measured by one independent benchmark, Artificial Analysis.
ElevenLabs Eleven v3 is a proprietary text-to-speech (TTS) API from ElevenLabs, measured by one independent benchmark, Artificial Analysis.
Fish Audio OpenAudio S1 Mini is an open-weights text-to-speech (TTS) API from Fish Audio, measured by one independent benchmark, Artificial Analysis.
Boson AI Higgs Audio V3 TTS is an open-weights text-to-speech (TTS) API from Boson AI, measured by one independent benchmark, Artificial Analysis.
OpenVoice v2 is an open-weights text-to-speech (TTS) API from OpenVoice, measured by one independent benchmark, Artificial Analysis.
Coqui XTTS v2 is an open-weights text-to-speech (TTS) API from Coqui, measured by one independent benchmark, Artificial Analysis.