← Speech-to-Text
#10 / 11

xAI

Grok STT

* Conditional

Only 11/20 clips finalized on native VAD — needs a forced finalize.

WER · batch

4.8%

Loudness-normalized best-case; degrades on raw input.

WER · stream

10.9%

Raw audio, n=31: 15/50 very-quiet FLEURS clips excluded where the model returns no transcript (loudness-sensitive; batch 4.8% is loudness-normalized).

End-of-turn

452ms

n=11/20 · native

9/20 clips never finalized on native VAD alone; needs a forced finalize.

Cost / min

Share this result

Embed the live badge

Speko stt rank
Markdown
[![Speko stt rank](https://benchmarks.speko.ai/badge/stt/xai-grok-stt.svg)](https://benchmarks.speko.ai/stt/xai-grok-stt)
HTML
<a href="https://benchmarks.speko.ai/stt/xai-grok-stt"><img src="https://benchmarks.speko.ai/badge/stt/xai-grok-stt.svg" alt="Speko stt rank"></a>
URL
https://benchmarks.speko.ai/stt/xai-grok-stt