Request a new model

Tell us which model you want measured. We review every request.

Speech-to-Text

WER

Finalize latency

lower is better

Time from the caller stopping to the finalized transcript, not the first partial — the agent cannot reply until it lands.

Sonioxstt-rt-v578ms · 96ms
CartesiaInk-2102ms · 151ms
DeepgramNova-3106ms · 138ms
InworldInworldRealtime STT-1139ms · 158ms
Smallest AISmallest AIPulse178ms · 196ms
xAIGrok STT305ms · 356ms
Gradium ASR334ms · 406ms
DeepgramFlux406ms · 1.19s
AlibabaQwen3-ASR424ms · 606ms
GoogleChirp 3581ms · 882ms
GGladiaSolaria-1596ms · 1.00s
ModulateModulateVelma 21.11s · 1.42s
0ms300ms600ms900ms1.20s

p50 p90

Time to first token

lower is better
Sonioxstt-rt-v51.04s · 1.68s
DeepgramFlux1.11s · 1.59s
DeepgramNova-31.12s · 1.58s
CartesiaInk-21.16s · 2.06s
InworldInworldRealtime STT-11.21s · 1.81s
xAIGrok STT1.34s · 1.39s
Smallest AISmallest AIPulse1.39s · 2.40s
ModulateModulateVelma 21.48s · 1.50s
GGladiaSolaria-11.49s · 2.10s
Gradium ASR1.83s · 2.72s
GoogleChirp 35.59s · 6.21s
AlibabaQwen3-ASR6.60s · 7.31s
0ms2.00s4.00s6.00s8.00s

p50 p90

Cost

lower is better

What a vendor charges to transcribe a thousand minutes of audio, at its published streaming rate.