← Speech-to-Text 
#10 / 11
xAI
Grok STT
* Conditional
Only 11/20 clips finalized on native VAD — needs a forced finalize.
WER · batch
4.8%
Loudness-normalized best-case; degrades on raw input.
WER · stream
10.9%
Raw audio, n=31: 15/50 very-quiet FLEURS clips excluded where the model returns no transcript (loudness-sensitive; batch 4.8% is loudness-normalized).
End-of-turn
452ms
n=11/20 · native
9/20 clips never finalized on native VAD alone; needs a forced finalize.
Cost / min
—
Share this result
Embed the live badge
Markdown
[](https://benchmarks.speko.ai/stt/xai-grok-stt)
HTML
<a href="https://benchmarks.speko.ai/stt/xai-grok-stt"><img src="https://benchmarks.speko.ai/badge/stt/xai-grok-stt.svg" alt="Speko stt rank"></a>
URL
https://benchmarks.speko.ai/stt/xai-grok-stt