Request a new model

Tell us which model you want measured. We review every request.

Speech-to-Text

WER

Route

Finalize latency

Region

Time from the caller stopping to the finalized transcript, not the first partial — the agent cannot reply until it lands.

DeepgramNova-3189ms · 770ms
Sonioxstt-rt-v5340ms · 1.33s
CartesiaInk-2587ms · 989ms
InworldInworldRealtime STT-1670ms · 1.16s
DeepgramFlux712ms · 1.31s
OpenAIGPT Transcribe824ms · 1.16s
Smallest AISmallest AIPulse839ms · 1.08s
AlibabaQwen3-ASR889ms · 1.21s
GladiaSolaria-1910ms · 3.10s
xAIGrok STT1.13s · 1.36s
GradiumASR1.13s · 1.34s
ModulateModulateVelma 21.21s · 1.41s
0ms300ms600ms900ms1.20s

p50 p90

Time to first token

Region
DeepgramFlux811ms · 1.49s
DeepgramNova-3976ms · 1.98s
xAIGrok STT1.03s · 2.01s
InworldInworldRealtime STT-11.32s · 1.92s
ModulateModulateVelma 21.48s · 1.49s
Sonioxstt-rt-v51.53s · 2.04s
GladiaSolaria-11.61s · 3.65s
CartesiaInk-21.78s · 2.40s
Smallest AISmallest AIPulse2.27s · 2.32s
GradiumASR2.44s · 2.95s
OpenAIGPT Transcribe6.51s · 10.7s
AlibabaQwen3-ASR6.81s · 7.50s
0ms2.00s4.00s6.00s8.00s

p50 p90

Cost

lower is better

What a vendor charges to transcribe a thousand minutes of audio, at its published streaming rate.