Request a new model

Tell us which model you want measured. We review every request.

Speech-to-Text

WER

Route

Finalize latency

Region

Time from the caller stopping to the finalized transcript, not the first partial — the agent cannot reply until it lands.

DeepgramNova-3222ms · 576ms
Sonioxstt-rt-v5345ms · 1.33s
CartesiaInk-2556ms · 960ms
InworldInworldRealtime STT-1561ms · 1.09s
DeepgramFlux706ms · 1.29s
Smallest AISmallest AIPulse793ms · 1.08s
OpenAIGPT Transcribe820ms · 1.12s
GladiaSolaria-1837ms · 3.09s
AlibabaQwen3-ASR873ms · 1.17s
GradiumASR1.13s · 1.34s
xAIGrok STT1.13s · 1.35s
ModulateModulateVelma 21.21s · 1.41s
0ms300ms600ms900ms1.20s

p50 p90

Time to first token

Region
DeepgramFlux792ms · 1.49s
DeepgramNova-31.02s · 2.03s
xAIGrok STT1.04s · 2.01s
InworldInworldRealtime STT-11.13s · 1.93s
ModulateModulateVelma 21.48s · 1.49s
GladiaSolaria-11.49s · 2.61s
Sonioxstt-rt-v51.53s · 2.03s
CartesiaInk-21.76s · 2.37s
Smallest AISmallest AIPulse2.27s · 2.29s
GradiumASR2.44s · 2.95s
OpenAIGPT Transcribe6.37s · 10.6s
AlibabaQwen3-ASR6.79s · 7.44s
0ms2.00s4.00s6.00s8.00s

p50 p90

Cost

lower is better

What a vendor charges to transcribe a thousand minutes of audio, at its published streaming rate.

FAQ

Which STT API has the lowest latency for phone calls?
AssemblyAI Universal-3.5 Pro finalizes a turn in 66ms (p50, end of speech to final transcript, us-east4), the fastest measured on the board. Soniox stt-rt-v5 is next at 78ms.
What is the most accurate speech-to-text API on the streaming path?
AssemblyAI Universal-3.5 Pro measured 2.0% WER over the live streaming socket, the lowest streaming error rate on the board, and it is the rare model that gives up nothing against its own 2.0% batch WER.
Why is streaming WER higher than batch WER?
Batch uploads the whole clip; streaming transcribes it live, and most models lose accuracy on that path. Deepgram Nova-3 measured 12.0% WER batch against 12.9% streaming on the same clips.
What is the cheapest streaming STT API?
Modulate Velma 2 lists at $0.001 per minute of audio (5.4% streaming WER) and Soniox stt-rt-v5 at $0.002 (7.3% streaming WER), the two lowest rates on the board.
What does end-of-turn latency mean for a voice agent?
It is the wait between the caller finishing and the final transcript arriving, and the agent cannot reply until it lands. Measured p50s on the board span 66ms (AssemblyAI Universal-3.5 Pro) to 1.22s (Google Gemini 3.5 Transcribe Live).