Request a new model

Tell us which model you want measured. We review every request.

Text-to-Speech

Naturalness

higher is better
field averagetied with bestQwenQwenqwen3-tts-flash1310HumeHumeoctave-21377MiniMaxMiniMaxspeech-2.6-turbo~1408OpenAIgpt-4o-mini-tts1424RimeRimemistv3~1428RimeRimearcanav31429MiniMaxMiniMaxspeech-2.8-hd1431SmallestSmallestlightning_v3.1_pro~1447MiniMaxMiniMaxspeech-2.6-hd~1450HumeHumeoctave-1~1455ElevenLabseleven_turbo_v2_5~1466ElevenLabseleven_multilingual_v2~1472ElevenLabseleven_turbo_v2~1492ElevenLabseleven_flash_v2_5~1494xAI Grokgrok-tts1495InworldInworldtts-2-flash~1501ElevenLabseleven_flash_v2~1507PalabraPalabratts-v1~1520RimeRimecoda~1535SmallestSmallestlightning_v3.11544Deepgramflux~1550Cartesiasonic-3~1551InworldInworldtts-21561Fish AudioFish Audios2.1-pro~1566Sonioxtts-rt-v11569BlandBlandspeech~1569SpeechifySpeechifysimba-3.21573Cartesiasonic-3.51574Cartesiasonic-3.6~1578Deepgramaura-21584GradiumTTS~1585ElevenLabseleven_v3_conversational1590GeminiGemini3.1-flash-tts-preview1591Sonioxtts-rt-v2~1606
1300140015001600

Time to first audio

Region
InworldInworldtts-2-flash50ms · 57ms
InworldInworldtts-284ms · 96ms
PalabraPalabratts-v186ms · 96ms
Maya2 Native94ms · 102ms
Deepgramflux118ms · 153ms
RimeRimemistv3120ms · 139ms
Cartesiasonic-3.6130ms · 264ms
Cartesiasonic-3.5132ms · 258ms
Sonioxtts-rt-v2134ms · 204ms
Deepgramaura-2147ms · 156ms
RimeRimecoda148ms · 194ms
Cartesiasonic-3150ms · 177ms
BlandBlandspeech192ms · 262ms
Fish AudioFish Audios2.1-pro198ms · 233ms
GradiumTTS217ms · 351ms
xAI Grokgrok-tts258ms · 312ms
RimeRimearcanav3260ms · 271ms
HumeHumeoctave-2435ms · 505ms
QwenQwenqwen3-tts-flash472ms · 493ms
SpeechifySpeechifysimba-3.2474ms · 485ms
OpenAIgpt-4o-mini-tts518ms · 929ms
HumeHumeoctave-1589ms · 903ms
MiniMaxMiniMaxspeech-2.8-hd1.35s · 1.61s
MiniMaxMiniMaxspeech-2.6-hd1.88s · 2.25s
0ms300ms600ms900ms1.20s

p50 p90

Robustness

higher is better

Whether the voice says the right words for a number, date, currency amount or operator — a different question from drift, which is whether it sounds like itself across repeats.

Cost

lower is better

What a vendor charges to synthesize a million characters of text, at its published list rate.

FAQ

Which TTS voice sounds the most natural?
No single voice wins it. In the blind human A/B arena the top 8 systems have overlapping 95% intervals - from Gemini 3.1 Flash TTS at Elo 1591 down to Smallest lightning_v3.1 at 1544, against a field mean of 1500 - so picking a winner inside that band would be reading vote noise.
What is the fastest TTS API for a real-time agent?
Palabra palabra-tts-v1 reached first audio in 72ms p50, the fastest synth measured; Inworld inworld-tts-2 is the fastest fully arena-rated voice at 116ms.
How is TTS naturalness measured on this board?
Each score is an arena Elo from blind human A/B votes on phone-agent lines, fit with Bradley-Terry and published with a bootstrap 95% interval, field mean 1500. Read the interval and not the point: the top 8 voices all overlap, while OpenAI gpt-4o-mini-tts at Elo 1424 sits clear below that band.
Which TTS API is cheapest per million characters?
Maya Maya 2 Native (~$4.0) and Speechify simba-3.2 ($10.0) list the lowest per-character rates on the board. On naturalness: Maya Maya 2 Native carries no naturalness measurement on this board, and Speechify simba-3.2 scored Elo 1573, inside the top arena tie band.
Which TTS voice reads numbers, dates and currency correctly?
Robustness is a 0-1 defect score on numbers, dates, currency and operators. Inworld inworld-tts-2 leads at 0.93.