Best STT APIs 2026

Measured on live streaming calls through one gateway; latest capture 2026-08-08. Speech-to-text has no single overall winner, so each axis below is measured and ranked separately.

Fastest end-of-turn

AssemblyAI Universal-3.5 Pro turns end of caller speech into a final transcript in 66ms at the median, the fastest finalize we measured on the live streaming socket (2026-07-03). It holds a 2.0% streaming word error rate at $0.0075 per minute. Soniox stt-rt-v5 (78ms) and Cartesia Ink-2 (102ms) are close behind; both trade accuracy for the speed.

Model Finalize p50 WER stream Cost / min
AssemblyAI Universal-3.5 Pro 66ms2.0%$0.0075
Soniox stt-rt-v5 78ms7.3%$0.002
Cartesia Ink-2 * 102ms9.9%$0.0090
Deepgram Nova-3 106ms12.9%$0.0048
Inworld Realtime STT-1 * 139ms3.6%$0.0025

Most accurate

AssemblyAI Universal-3.5 Pro posts the lowest streaming word error rate we measured: 2.0% on read English over the live socket, against 2.0% on the batch path for the same clips (2026-07-03). The same model also posts the fastest finalize on this board, so accuracy costs nothing in speed. ElevenLabs Scribe v2 Realtime is next at 3.4%.

Model WER stream WER batch Finalize p50
AssemblyAI Universal-3.5 Pro 2.0%2.0%66ms
ElevenLabs Scribe v2 Realtime * 3.4%-233ms
Inworld Realtime STT-1 * 3.6%3.3%139ms
Alibaba Qwen3-ASR * 4.0%2.8%424ms
OpenAI GPT Live Transcribe 4.5%-1124ms

Best at code-switching

AssemblyAI universal-3-5-pro keeps both languages alive on every tested pair (4 / 4, measured 2026-07-31), with 9.8% mixed-token error on Mandarin-English and 3.5% on English-Spanish over the live stream. Soniox stt-rt-v5 also covers 4 / 4 pairs, at 11.8% on Mandarin-English.

Model Pairs handled EN-ZH EN-ES EN-DE EN-FR
Soniox stt-rt-v5 4 / 411.8%4.9%5.1%8.8%
AssemblyAI universal-3-5-pro 4 / 49.8%3.5%3.7%2.6%
ElevenLabs scribe_v2_realtime * 3 / 475.0%3.8%3.1%7.8%

Single-language accuracy across 9 languages is a separate study: see Multilingual STT.

Best value

Modulate Velma 2 lists the lowest streaming rate on the board at $0.001 per minute, with 5.4% streaming WER; the tradeoff is finalize latency, at 1.11s p50 (2026-08-08). Caveat: Batch measured vendor-direct; streaming re-scored through the gateway. Inworld Realtime STT-1 at $0.0025 posts the lowest error rate of the cheap lane: 3.6% streaming WER and a 139ms finalize. Caveat: Finalizes eagerly — ~1 in 5 turns commit before speech ends.

Model Cost / min WER stream Finalize p50
Modulate Velma 2 * $0.0015.4%1.11s
Soniox stt-rt-v5 $0.0027.3%78ms
Inworld Realtime STT-1 * $0.00253.6%139ms
OpenAI GPT-4o-mini Transcribe $0.0036.4%460ms
xAI Grok STT $0.003310.9%305ms

How we measured

Method: how Speko benchmarks STT. Full boards, column definitions and caveats: the Speech-to-Text board, Multilingual STT, streaming WER by production condition, and the raw dataset at /data.json.