Request a new model

Tell us which model you want measured. We review every request.

Speech-to-Text

WER

Finalize latency

Region

Time from the caller stopping to the finalized transcript, not the first partial — the agent cannot reply until it lands.

DeepgramNova-3222ms · 576ms
Sonioxstt-rt-v5345ms · 1.33s
CartesiaInk-2556ms · 960ms
InworldInworldRealtime STT-1561ms · 1.09s
DeepgramFlux706ms · 1.29s
Smallest AISmallest AIPulse793ms · 1.08s
OpenAIGPT Transcribe820ms · 1.12s
GGladiaSolaria-1837ms · 3.09s
AlibabaQwen3-ASR873ms · 1.17s
Gradium ASR1.13s · 1.34s
xAIGrok STT1.13s · 1.35s
ModulateModulateVelma 21.21s · 1.41s
0ms300ms600ms900ms1.20s

p50 p90

Time to first token

Region
DeepgramFlux792ms · 1.49s
DeepgramNova-31.02s · 2.03s
xAIGrok STT1.04s · 2.01s
InworldInworldRealtime STT-11.13s · 1.93s
ModulateModulateVelma 21.48s · 1.53s
GGladiaSolaria-11.49s · 2.61s
Sonioxstt-rt-v51.53s · 2.03s
CartesiaInk-21.76s · 2.37s
Smallest AISmallest AIPulse2.27s · 2.29s
Gradium ASR2.44s · 2.95s
OpenAIGPT Transcribe6.37s · 10.6s
AlibabaQwen3-ASR6.79s · 7.44s
0ms2.00s4.00s6.00s8.00s

p50 p90

Cost

lower is better

What a vendor charges to transcribe a thousand minutes of audio, at its published streaming rate.