Speech-to-Text · Multilingual
Spanish
Language
WER
Share of words the transcript got wrong.
AssemblyAIUniversal-3.5 Pro3.9%
ElevenLabsScribe v2 Realtime4.4%
Sonioxstt-rt-v55.2%
AlibabaQwen3-ASR Realtime6.6%
OpenAIGPT-4o-mini Transcribe7.2%
DeepgramNova-37.6%
GradiumASR8.5%
OpenAIGPT-4o Transcribe8.7%
GoogleChirp 39.3%
Smallest AIPulse10.3%
CartesiaInk-Whisper10.9%
DeepgramNova-212.2%
0.0%5.0%10.0%15.0%20.0%
Finalize latency
lower is betterTime from the caller stopping to the finalized transcript, not the first partial — the agent cannot reply until it lands.
DeepgramNova-3272ms · 533ms
DeepgramNova-2288ms · 416ms
Sonioxstt-rt-v5478ms · 1.13s
AssemblyAIUniversal-3.5 Pro650ms · 1.67s
Smallest AIPulse876ms · 983ms
OpenAIGPT-4o-mini Transcribe898ms · 996ms
OpenAIGPT-4o Transcribe910ms · 1.15s
AlibabaQwen3-ASR Realtime936ms · 1.09s
GradiumASR1.06s · 1.25s
GoogleChirp 31.12s · 1.50s
ElevenLabsScribe v2 Realtime2.00s · 2.25s
0ms300ms600ms900ms1.20s
p50 p90
Time to first token
lower is betterHow long before the first word of the transcript arrives.
AssemblyAIUniversal-3.5 Pro1.77s · 2.28s
DeepgramNova-22.12s · 3.17s
DeepgramNova-32.12s · 3.03s
ElevenLabsScribe v2 Realtime2.21s · 3.16s
Sonioxstt-rt-v52.22s · 2.91s
Smallest AIPulse2.39s · 3.60s
GradiumASR3.00s · 3.51s
GoogleChirp 36.23s · 7.41s
OpenAIGPT-4o Transcribe6.96s · 10.7s
OpenAIGPT-4o-mini Transcribe7.33s · 10.7s
AlibabaQwen3-ASR Realtime7.42s · 8.46s
0ms2.00s4.00s6.00s8.00s
p50 p90
Cost
lower is better$ per 1,000 minutes of audio · vendor list price, the same in every language