Request a new model

Tell us which model you want measured. We review every request.

Speech-to-Text

Benchmarks

X
Y

WER

AzureAzureMAI-Transcribe-2 Streaming6.2%
InworldInworldRealtime STT-17.0%
AssemblyAIAssemblyAIUniversal-3.5 Pro7.2%
AssemblyAIAssemblyAIUniversal-3.6 Pro7.2%
NariNariQwen3-ASR7.6%
NariNariQwen3-ASR Fast7.6%
ElevenLabsScribe v2 Realtime8.2%
ModulateModulateVelma 28.4%
MetaMetaMuse Voice Transcribe8.5%
OpenAIGPT Live Transcribe8.8%
SpeechmaticsSpeechmaticsEnhanced9.0%
SpeechmaticsSpeechmaticsLinden 19.5%
Sonioxstt-rt-v59.5%
xAIGrok Voice Transcribe 1.09.7%
Smallest AISmallest AIPulse10.0%
xAIGrok Voice Transcribe 2.010.2%
OpenAIGPT Transcribe10.3%
OpenAIGPT Realtime Whisper10.4%
DeepgramNova-311.0%
SpeechmaticsSpeechmaticsStandard11.3%
OpenAIGPT-4o mini Transcribe11.4%
DeepgramNova-3 Medical12.3%
GladiaSolaria-112.3%
CartesiaInk-212.6%
DeepgramFlux12.8%
DeepgramFlux Multilingual14.1%
AssemblyAIAssemblyAIUniversal-Streaming English14.3%
GradiumASR18.4%
OpenAIGPT-4o Transcribe19.3%
AssemblyAIAssemblyAIUniversal-Streaming Multilingual21.6%
DeepgramBase Phonecall23.5%
0.0%10.0%20.0%30.0%40.0%

End of turn

lower is better

Time from the last spoken word to the final transcript, decided by the model itself.

MetaMetaMuse Voice Transcribe301ms · 584ms
Sonioxstt-rt-v5489ms · 2.40s
AssemblyAIAssemblyAIUniversal-3.6 Pro589ms · 1.58s
AssemblyAIAssemblyAIUniversal-3.5 Pro606ms · 1.62s
DeepgramFlux Multilingual654ms · 1.36s
NariNariQwen3-ASR Fast684ms · 965ms
DeepgramNova-3771ms · 1.21s
DeepgramFlux773ms · 1.48s
ElevenLabsScribe v2 Realtime793ms · 1.06s
OpenAIGPT-4o mini Transcribe926ms · 1.35s
Smallest AISmallest AIPulse958ms · 1.45s
GladiaSolaria-11.09s · 1.59s
GradiumASR1.22s · 1.35s
ModulateModulateVelma 21.25s · 1.54s
OpenAIGPT-4o Transcribe1.28s · 1.70s
xAIGrok Voice Transcribe 2.01.35s · 1.75s
xAIGrok Voice Transcribe 1.01.40s · 1.95s
SpeechmaticsSpeechmaticsStandard1.45s · 1.81s
InworldInworldRealtime STT-11.90s · 2.71s
0ms500ms1.00s1.50s2.00s2.50s

p50 p90

Time to first token

lower is better

Time from the caller starting to speak to the first word of the live transcript.

AzureAzureMAI-Transcribe-2 Streaming241ms · 342ms
DeepgramFlux Multilingual506ms · 1.65s
DeepgramFlux641ms · 1.19s
GladiaSolaria-1671ms · 1.20s
SpeechmaticsSpeechmaticsStandard690ms · 1.05s
MetaMetaMuse Voice Transcribe728ms · 1.16s
AssemblyAIAssemblyAIUniversal-3.6 Pro775ms · 974ms
AssemblyAIAssemblyAIUniversal-3.5 Pro775ms · 973ms
Sonioxstt-rt-v5786ms · 1.27s
DeepgramNova-3885ms · 1.45s
InworldInworldRealtime STT-1942ms · 1.02s
OpenAIGPT Realtime Whisper987ms · 1.27s
OpenAIGPT Live Transcribe988ms · 1.13s
xAIGrok Voice Transcribe 1.0989ms · 1.10s
xAIGrok Voice Transcribe 2.0995ms · 1.18s
Smallest AISmallest AIPulse1.04s · 2.01s
CartesiaInk-21.19s · 1.61s
ModulateModulateVelma 21.38s · 1.47s
NariNariQwen3-ASR Fast1.65s · 1.74s
GradiumASR1.65s · 2.43s
ElevenLabsScribe v2 Realtime2.02s · 2.19s
OpenAIGPT-4o mini Transcribe2.91s · 6.12s
OpenAIGPT-4o Transcribe4.72s · 8.46s
0ms1.00s2.00s3.00s4.00s5.00s

p50 p90

Live words

lower is better

Time from a word being spoken to it appearing in the live transcript for good.

AzureAzureMAI-Transcribe-2 Streaming114ms · 515ms
DeepgramFlux Multilingual180ms · 715ms
DeepgramFlux190ms · 586ms
SpeechmaticsSpeechmaticsStandard403ms · 1.07s
MetaMetaMuse Voice Transcribe486ms · 716ms
Sonioxstt-rt-v5496ms · 739ms
InworldInworldRealtime STT-1583ms · 1.86s
DeepgramNova-3583ms · 1.72s
AssemblyAIAssemblyAIUniversal-3.5 Pro803ms · 2.52s
AssemblyAIAssemblyAIUniversal-3.6 Pro815ms · 2.71s
ModulateModulateVelma 2896ms · 2.35s
NariNariQwen3-ASR Fast901ms · 2.04s
ElevenLabsScribe v2 Realtime945ms · 1.97s
Smallest AISmallest AIPulse1.17s · 1.98s
CartesiaInk-21.22s · 1.71s
GradiumASR1.48s · 1.66s
GladiaSolaria-11.50s · 3.98s
xAIGrok Voice Transcribe 2.01.89s · 10.22s
xAIGrok Voice Transcribe 1.02.08s · 8.91s
OpenAIGPT-4o mini Transcribe2.47s · 5.29s
OpenAIGPT-4o Transcribe3.43s · 7.18s
OpenAIGPT Realtime Whisper3.96s · 11.28s
OpenAIGPT Live Transcribe4.02s · 11.34s
0ms1.00s2.00s3.00s4.00s5.00s

p50 p90

Rewrites

lower is better

Words the live transcript takes back, per word in the final transcript.

GradiumASR0%
CartesiaInk-20%
MetaMetaMuse Voice Transcribe0%
Smallest AISmallest AIPulse0%
OpenAIGPT-4o Transcribe3%
OpenAIGPT-4o mini Transcribe5%
OpenAIGPT Realtime Whisper6%
OpenAIGPT Live Transcribe10%
Sonioxstt-rt-v513%
NariNariQwen3-ASR Fast22%
ElevenLabsScribe v2 Realtime29%
InworldInworldRealtime STT-130%
DeepgramNova-333%
ModulateModulateVelma 233%
SpeechmaticsSpeechmaticsStandard35%
AssemblyAIAssemblyAIUniversal-3.6 Pro35%
AssemblyAIAssemblyAIUniversal-3.5 Pro36%
AzureAzureMAI-Transcribe-2 Streaming37%
DeepgramFlux40%
DeepgramFlux Multilingual41%
xAIGrok Voice Transcribe 2.064%
xAIGrok Voice Transcribe 1.083%
GladiaSolaria-1156%
0%40%80%120%160%

Cost

lower is better

Published list price per thousand minutes; batch-only models show their batch rate.