Best STT APIs 2026
Measured vendor-direct; latest capture 2026-09-30. Speech-to-text has no single overall winner, so each axis below is measured and ranked separately.
Fastest end-of-turn
Nari Qwen3-ASR turns end of caller speech into a final transcript in 25ms at the median, the fastest finalize we measured on the live streaming socket (2026-09-26). Speed is all this ranks: it holds a 3.88% streaming word error rate, and 12.02% on accented conversational English at $0.001 per minute. Nari Qwen3-ASR Fast (33ms) and Deepgram Base Phonecall (171ms) are close behind.
| Model | Finalize p50 | WER stream | Cost / min |
|---|---|---|---|
| Nari Qwen3-ASR | 25ms | 3.88% | $0.001 |
| Nari Qwen3-ASR Fast | 33ms | 3.83% | $0.002 |
| Deepgram Base Phonecall | 171ms | 27.03% | — |
| Deepgram Nova-3 | 189ms | 10.8% | $0.0048 |
| Deepgram Nova-3 Medical | 190ms | 13.12% | — |
Most accurate
AssemblyAI Universal-3.6 Pro posts the lowest streaming word error rate we measured: 3.07% on read English over the live socket, against — on the batch path for the same clips (2026-09-30). AssemblyAI Universal-3.5 Pro is next at 3.73%.
| Model | WER stream | WER batch | Finalize p50 |
|---|---|---|---|
| AssemblyAI Universal-3.6 Pro | 3.07% | — | 505ms |
| AssemblyAI Universal-3.5 Pro | 3.73% | 2.0% | 408ms |
| Nari Qwen3-ASR Fast | 3.83% | 3.3% | 33ms |
| Nari Qwen3-ASR | 3.88% | 3.4% | 25ms |
| Alibaba Qwen3-ASR Realtime | 4.0% | - | 889ms |
Best at code-switching
AssemblyAI universal-3-5-pro keeps both languages alive on every tested pair (4 / 4, measured 2026-07-31), with 9.8% mixed-token error on Mandarin-English and 3.5% on English-Spanish over the live stream. Soniox stt-rt-v5 also covers 4 / 4 pairs, at 11.8% on Mandarin-English.
| Model | Pairs handled | EN-ZH | EN-ES | EN-DE | EN-FR |
|---|---|---|---|---|---|
| Soniox stt-rt-v5 | 4 / 4 | 11.8% | 4.9% | 5.1% | 8.8% |
| AssemblyAI universal-3-5-pro | 4 / 4 | 9.8% | 3.5% | 3.7% | 2.6% |
| ElevenLabs scribe_v2_realtime | 3 / 4 | 75.0% | 3.8% | 3.1% | 7.8% |
Single-language accuracy across 14 languages is a separate study: see Multilingual STT.
Best value
Modulate Velma 2 lists the lowest streaming rate on the board at $0.00083 per minute, with 5.35% streaming WER; the tradeoff is finalize latency, at 1.21s p50 (2026-08-08). Nari Qwen3-ASR Fast at $0.002 posts the lowest error rate of the cheap lane: 3.83% streaming WER and a 33ms finalize.
| Model | Cost / min | WER stream | Finalize p50 |
|---|---|---|---|
| Modulate Velma 2 | $0.00083 | 5.35% | 1.21s |
| Nari Qwen3-ASR | $0.001 | 3.88% | 25ms |
| Soniox stt-rt-v5 | $0.002 | 6.31% | 340ms |
| Nari Qwen3-ASR Fast | $0.002 | 3.83% | 33ms |
| Alibaba Qwen3-ASR | $0.0021 | - | — |
How we measured
Method: how Speko benchmarks STT. Full boards, column definitions and caveats: the Speech-to-Text board, Multilingual STT, streaming WER by production condition, and the raw dataset at /data.json.