Request a new model

Tell us which model you want measured. We review every request.

Speech-to-Text

WER

Route

Finalize latency

Region

Time from the caller stopping to the finalized transcript, not the first partial — the agent cannot reply until it lands.

DeepgramNova-3189ms · 770ms
Sonioxstt-rt-v5340ms · 1.33s
CartesiaInk-2587ms · 989ms
InworldInworldRealtime STT-1670ms · 1.16s
DeepgramFlux712ms · 1.31s
OpenAIGPT Transcribe824ms · 1.16s
Smallest AISmallest AIPulse839ms · 1.08s
AlibabaQwen3-ASR889ms · 1.21s
GladiaSolaria-1910ms · 3.10s
xAIGrok STT1.13s · 1.36s
GradiumASR1.13s · 1.34s
ModulateModulateVelma 21.21s · 1.41s
0ms300ms600ms900ms1.20s

p50 p90

Time to first token

Region
DeepgramFlux811ms · 1.49s
DeepgramNova-3976ms · 1.98s
xAIGrok STT1.03s · 2.01s
InworldInworldRealtime STT-11.32s · 1.92s
ModulateModulateVelma 21.48s · 1.49s
Sonioxstt-rt-v51.53s · 2.04s
GladiaSolaria-11.61s · 3.65s
CartesiaInk-21.78s · 2.40s
Smallest AISmallest AIPulse2.27s · 2.32s
GradiumASR2.44s · 2.95s
OpenAIGPT Transcribe6.51s · 10.7s
AlibabaQwen3-ASR6.81s · 7.50s
0ms2.00s4.00s6.00s8.00s

p50 p90

Cost

lower is better

What a vendor charges to transcribe a thousand minutes of audio, at its published streaming rate.

FAQ

Which STT API has the lowest latency for phone calls?
AssemblyAI Universal-3.5 Pro finalizes a turn in 66ms (p50, end of speech to final transcript, us-east4), the fastest measured on the board. Soniox stt-rt-v5 is next at 78ms.
What is the most accurate speech-to-text API on the streaming path?
AssemblyAI Universal-3.5 Pro measured 2.0% WER over the live streaming socket, the lowest streaming error rate on the board, and it is the rare model that gives up nothing against its own 2.0% batch WER.
Why is streaming WER higher than batch WER?
Batch uploads the whole clip; streaming transcribes it live, and most models lose accuracy on that path. Deepgram Nova-3 measured 12.0% WER batch against 12.9% streaming on the same clips.
What is the cheapest streaming STT API?
Modulate Velma 2 lists at $0.001 per minute of audio (5.4% streaming WER) and Soniox stt-rt-v5 at $0.002 (7.3% streaming WER), the two lowest rates on the board.
What does end-of-turn latency mean for a voice agent?
It is the wait between the caller finishing and the final transcript arriving, and the agent cannot reply until it lands. Measured p50s on the board span 66ms (AssemblyAI Universal-3.5 Pro) to 1.22s (Google Gemini 3.5 Transcribe Live).