Text-to-Speech
Naturalness
higher is betterfield averagetied with bestQwenqwen3-tts-flash1310Humeoctave-21377OpenAIgpt-4o-mini-tts1424Rimearcanav31429MiniMaxspeech-2.8-hd1431xAI Grokgrok-tts1495Inworldtts-2-flash~1501Palabratts-v1~1520Smallestlightning_v3.11544Deepgramflux~1550Inworldtts-21561Fish Audios2.1-pro~1566Sonioxtts-rt-v11569Blandspeech~1569Speechifysimba-3.21573Cartesiasonic-3.51574Cartesiasonic-3.6~1578Deepgramaura-21584Gradiumdefault~1585ElevenLabseleven_v3_conversational1590Gemini3.1-flash-tts-preview1591Sonioxtts-rt-v2~1606
1300140015001600
Time to first audio
Region0ms300ms600ms900ms1.20s
p50 p90
Robustness
higher is betterWhether the voice says the right words for a number, date, currency amount or operator — a different question from drift, which is whether it sounds like itself across repeats.
0.000.250.500.751.00
Cost
lower is betterWhat a vendor charges to synthesize a million characters of text, at its published list rate.
$0$25$50$75$100
Maya2 NativeSpeechifysimba-3.2Qwenqwen3-tts-flashSonioxtts-rt-v1Sonioxtts-rt-v2Inworldtts-2-flashFish Audios2.1-proxAI Grokgrok-ttsBlandspeechOpenAIgpt-4o-mini-ttsInworldtts-2Smallestlightning_v3.1Deepgramaura-2Palabratts-v1Gemini3.1-flash-tts-previewRimearcanav3DeepgramfluxCartesiasonic-3.5Cartesiasonic-3.6GradiumdefaultElevenLabseleven_v3_conversationalMiniMaxspeech-2.8-hdHumeoctave-2