Best STT APIs 2026
Measured on live streaming calls through one gateway; latest capture 2026-08-08. Speech-to-text has no single overall winner, so each axis below is measured and ranked separately.
Fastest end-of-turn
AssemblyAI Universal-3.5 Pro turns end of caller speech into a final transcript in 66ms at the median, the fastest finalize we measured on the live streaming socket (2026-07-03). It holds a 2.0% streaming word error rate at $0.0075 per minute. Soniox stt-rt-v5 (78ms) and Cartesia Ink-2 (102ms) are close behind; both trade accuracy for the speed.
| Model | Finalize p50 | WER stream | Cost / min |
|---|---|---|---|
| AssemblyAI Universal-3.5 Pro | 66ms | 2.0% | $0.0075 |
| Soniox stt-rt-v5 | 78ms | 7.3% | $0.002 |
| Cartesia Ink-2 * | 102ms | 9.9% | $0.0090 |
| Deepgram Nova-3 | 106ms | 12.9% | $0.0048 |
| Inworld Realtime STT-1 * | 139ms | 3.6% | $0.0025 |
Most accurate
AssemblyAI Universal-3.5 Pro posts the lowest streaming word error rate we measured: 2.0% on read English over the live socket, against 2.0% on the batch path for the same clips (2026-07-03). The same model also posts the fastest finalize on this board, so accuracy costs nothing in speed. ElevenLabs Scribe v2 Realtime is next at 3.4%.
| Model | WER stream | WER batch | Finalize p50 |
|---|---|---|---|
| AssemblyAI Universal-3.5 Pro | 2.0% | 2.0% | 66ms |
| ElevenLabs Scribe v2 Realtime * | 3.4% | - | 233ms |
| Inworld Realtime STT-1 * | 3.6% | 3.3% | 139ms |
| Alibaba Qwen3-ASR * | 4.0% | 2.8% | 424ms |
| OpenAI GPT Live Transcribe | 4.5% | - | 1124ms |
Best at code-switching
AssemblyAI universal-3-5-pro keeps both languages alive on every tested pair (4 / 4, measured 2026-07-31), with 9.8% mixed-token error on Mandarin-English and 3.5% on English-Spanish over the live stream. Soniox stt-rt-v5 also covers 4 / 4 pairs, at 11.8% on Mandarin-English.
| Model | Pairs handled | EN-ZH | EN-ES | EN-DE | EN-FR |
|---|---|---|---|---|---|
| Soniox stt-rt-v5 | 4 / 4 | 11.8% | 4.9% | 5.1% | 8.8% |
| AssemblyAI universal-3-5-pro | 4 / 4 | 9.8% | 3.5% | 3.7% | 2.6% |
| ElevenLabs scribe_v2_realtime * | 3 / 4 | 75.0% | 3.8% | 3.1% | 7.8% |
Single-language accuracy across 9 languages is a separate study: see Multilingual STT.
Best value
Modulate Velma 2 lists the lowest streaming rate on the board at $0.001 per minute, with 5.4% streaming WER; the tradeoff is finalize latency, at 1.11s p50 (2026-08-08). Caveat: Batch measured vendor-direct; streaming re-scored through the gateway. Inworld Realtime STT-1 at $0.0025 posts the lowest error rate of the cheap lane: 3.6% streaming WER and a 139ms finalize. Caveat: Finalizes eagerly — ~1 in 5 turns commit before speech ends.
| Model | Cost / min | WER stream | Finalize p50 |
|---|---|---|---|
| Modulate Velma 2 * | $0.001 | 5.4% | 1.11s |
| Soniox stt-rt-v5 | $0.002 | 7.3% | 78ms |
| Inworld Realtime STT-1 * | $0.0025 | 3.6% | 139ms |
| OpenAI GPT-4o-mini Transcribe | $0.003 | 6.4% | 460ms |
| xAI Grok STT | $0.0033 | 10.9% | 305ms |
How we measured
Method: how Speko benchmarks STT. Full boards, column definitions and caveats: the Speech-to-Text board, Multilingual STT, streaming WER by production condition, and the raw dataset at /data.json.