Speech-to-Text
Benchmarks
XYWER
AzureMAI-Transcribe-2 Streaming6.2%
InworldRealtime STT-17.0%
AssemblyAIUniversal-3.5 Pro7.2%
AssemblyAIUniversal-3.6 Pro7.2%
NariQwen3-ASR7.6%
NariQwen3-ASR Fast7.6%
ElevenLabsScribe v2 Realtime8.2%
ModulateVelma 28.4%
MetaMuse Voice Transcribe8.5%
OpenAIGPT Live Transcribe8.8%
SpeechmaticsEnhanced9.0%
SpeechmaticsLinden 19.5%
Sonioxstt-rt-v59.5%
xAIGrok Voice Transcribe 1.09.7%
Smallest AIPulse10.0%
xAIGrok Voice Transcribe 2.010.2%
OpenAIGPT Transcribe10.3%
OpenAIGPT Realtime Whisper10.4%
DeepgramNova-311.0%
SpeechmaticsStandard11.3%
OpenAIGPT-4o mini Transcribe11.4%
DeepgramNova-3 Medical12.3%
GladiaSolaria-112.3%
CartesiaInk-212.6%
DeepgramFlux12.8%
DeepgramFlux Multilingual14.1%
AssemblyAIUniversal-Streaming English14.3%
GradiumASR18.4%
OpenAIGPT-4o Transcribe19.3%
AssemblyAIUniversal-Streaming Multilingual21.6%
DeepgramBase Phonecall23.5%
0.0%10.0%20.0%30.0%40.0%
End of turn
lower is betterTime from the last spoken word to the final transcript, decided by the model itself.
MetaMuse Voice Transcribe301ms · 584ms
Sonioxstt-rt-v5489ms · 2.40s
AssemblyAIUniversal-3.6 Pro589ms · 1.58s
AssemblyAIUniversal-3.5 Pro606ms · 1.62s
DeepgramFlux Multilingual654ms · 1.36s
NariQwen3-ASR Fast684ms · 965ms
DeepgramNova-3771ms · 1.21s
DeepgramFlux773ms · 1.48s
ElevenLabsScribe v2 Realtime793ms · 1.06s
OpenAIGPT-4o mini Transcribe926ms · 1.35s
Smallest AIPulse958ms · 1.45s
GladiaSolaria-11.09s · 1.59s
GradiumASR1.22s · 1.35s
ModulateVelma 21.25s · 1.54s
OpenAIGPT-4o Transcribe1.28s · 1.70s
xAIGrok Voice Transcribe 2.01.35s · 1.75s
xAIGrok Voice Transcribe 1.01.40s · 1.95s
SpeechmaticsStandard1.45s · 1.81s
InworldRealtime STT-11.90s · 2.71s
0ms500ms1.00s1.50s2.00s2.50s
p50 p90
Time to first token
lower is betterTime from the caller starting to speak to the first word of the live transcript.
AzureMAI-Transcribe-2 Streaming241ms · 342ms
DeepgramFlux Multilingual506ms · 1.65s
DeepgramFlux641ms · 1.19s
GladiaSolaria-1671ms · 1.20s
SpeechmaticsStandard690ms · 1.05s
MetaMuse Voice Transcribe728ms · 1.16s
AssemblyAIUniversal-3.6 Pro775ms · 974ms
AssemblyAIUniversal-3.5 Pro775ms · 973ms
Sonioxstt-rt-v5786ms · 1.27s
DeepgramNova-3885ms · 1.45s
InworldRealtime STT-1942ms · 1.02s
OpenAIGPT Realtime Whisper987ms · 1.27s
OpenAIGPT Live Transcribe988ms · 1.13s
xAIGrok Voice Transcribe 1.0989ms · 1.10s
xAIGrok Voice Transcribe 2.0995ms · 1.18s
Smallest AIPulse1.04s · 2.01s
CartesiaInk-21.19s · 1.61s
ModulateVelma 21.38s · 1.47s
NariQwen3-ASR Fast1.65s · 1.74s
GradiumASR1.65s · 2.43s
ElevenLabsScribe v2 Realtime2.02s · 2.19s
OpenAIGPT-4o mini Transcribe2.91s · 6.12s
OpenAIGPT-4o Transcribe4.72s · 8.46s
0ms1.00s2.00s3.00s4.00s5.00s
p50 p90
Live words
lower is betterTime from a word being spoken to it appearing in the live transcript for good.
AzureMAI-Transcribe-2 Streaming114ms · 515ms
DeepgramFlux Multilingual180ms · 715ms
DeepgramFlux190ms · 586ms
SpeechmaticsStandard403ms · 1.07s
MetaMuse Voice Transcribe486ms · 716ms
Sonioxstt-rt-v5496ms · 739ms
InworldRealtime STT-1583ms · 1.86s
DeepgramNova-3583ms · 1.72s
AssemblyAIUniversal-3.5 Pro803ms · 2.52s
AssemblyAIUniversal-3.6 Pro815ms · 2.71s
ModulateVelma 2896ms · 2.35s
NariQwen3-ASR Fast901ms · 2.04s
ElevenLabsScribe v2 Realtime945ms · 1.97s
Smallest AIPulse1.17s · 1.98s
CartesiaInk-21.22s · 1.71s
GradiumASR1.48s · 1.66s
GladiaSolaria-11.50s · 3.98s
xAIGrok Voice Transcribe 2.01.89s · 10.22s
xAIGrok Voice Transcribe 1.02.08s · 8.91s
OpenAIGPT-4o mini Transcribe2.47s · 5.29s
OpenAIGPT-4o Transcribe3.43s · 7.18s
OpenAIGPT Realtime Whisper3.96s · 11.28s
OpenAIGPT Live Transcribe4.02s · 11.34s
0ms1.00s2.00s3.00s4.00s5.00s
p50 p90
Rewrites
lower is betterWords the live transcript takes back, per word in the final transcript.
GradiumASR0%
CartesiaInk-20%
MetaMuse Voice Transcribe0%
Smallest AIPulse0%
OpenAIGPT-4o Transcribe3%
OpenAIGPT-4o mini Transcribe5%
OpenAIGPT Realtime Whisper6%
OpenAIGPT Live Transcribe10%
Sonioxstt-rt-v513%
NariQwen3-ASR Fast22%
ElevenLabsScribe v2 Realtime29%
InworldRealtime STT-130%
DeepgramNova-333%
ModulateVelma 233%
SpeechmaticsStandard35%
AssemblyAIUniversal-3.6 Pro35%
AssemblyAIUniversal-3.5 Pro36%
AzureMAI-Transcribe-2 Streaming37%
DeepgramFlux40%
DeepgramFlux Multilingual41%
xAIGrok Voice Transcribe 2.064%
xAIGrok Voice Transcribe 1.083%
GladiaSolaria-1156%
0%40%80%120%160%
Cost
lower is betterPublished list price per thousand minutes; batch-only models show their batch rate.
ModulateVelma 2NariQwen3-ASRAzureMAI-Transcribe-2NariQwen3-ASR FastSonioxstt-rt-v5AlibabaQwen3-ASR FlashInworldRealtime STT-1AssemblyAIUniversal-Streaming EnglishAssemblyAIUniversal-Streaming MultilingualSpeechmaticsLinden 1MetaMuse Voice TranscribeOpenAIGPT-4o mini TranscribexAIGrok Voice Transcribe 1.0xAIGrok Voice Transcribe 2.0ElevenLabsScribe v2Smallest AIPulseSpeechmaticsStandardOpenAIGPT TranscribeDeepgramNova-3OpenAIGPT-4o TranscribeFish AudioTranscribe 1 ProElevenLabsScribe v2 RealtimeDeepgramFluxSpeechmaticsEnhancedAssemblyAIUniversal-3.6 ProAssemblyAIUniversal-3.5 ProDeepgramFlux MultilingualCartesiaInk-2GradiumASRGladiaSolaria-1OpenAIGPT Live TranscribeOpenAIGPT Realtime Whisper