Best LLMs for Voice Agents 2026
Behaviour measured on multi-step tool scenarios over a live Speko route, first-token latency from one US region; latest capture 2026-10-08. First-token latency, task completion, fabrication and cost for voice-agent LLMs.
The LLM board publishes no composite score and no single ranking: first-token latency, task completion, fabrication, dead air and cost are each measured and ranked on their own, and the model that wins one axis is routinely mid-table on another. The sections below name a leader per axis.
Fastest first token
Cerebras gpt-oss-120b returns a first token in 139ms at the median (captured 2026-10-04) at $0.75 per million output tokens.
| Model | TTFT p50 | Task done | Cost / 1M tok |
|---|---|---|---|
| Cerebras gpt-oss-120b | 139ms | 82.02% | $0.75 |
| Baseten GLM-4.7 | 148ms | 80.88% | $2.20 |
| Together DeepSeek-V4.1-Flash | 157ms | 84.21% | $1.20 |
| Baseten Nemotron-3-Ultra | 245ms | 86.84% | $2.40 |
| Cerebras Qwen3.8-27B | 246ms | 86.14% | $1.49 |
| OpenAI gpt-5.4-mini | 390ms | 91.75% | $4.50 |
Most reliable
19 of the 27 measured models never invented a fact they were not given. Baseten DeepSeek-V4-Flash-0731 is the strongest of those on task completion while holding a real-time first-token measurement: 93.16% of the multi-step tool tasks, 36.8% dead air, 558ms p50 (captured 2026-10-04).
| Model | Task done | Fabrication | Dead-air | TTFT p50 |
|---|---|---|---|---|
| Gemini gemini-3.5-flash-lite | 94.74% | 10% | 63.1% | 500ms |
| Baseten DeepSeek-V4-Flash-0731 | 93.16% | 0% | 36.8% | 558ms |
| OpenAI gpt-4.1-mini | 92.63% | 3% | 61.4% | 480ms |
| OpenAI gpt-5.4-mini | 91.75% | 0% | 62.2% | 390ms |
| xAI Grok-4.3 | 90.88% | 0% | 56.9% | 2.38s |
| OpenAI gpt-4.1 | 89.12% | 10% | 61.5% | 462ms |
Best multilingual
Top scorer across the 13 languages: OpenAI gpt-4.1-mini (top in 6 of 13, never below 78.07% of scenarios, in the caller's language throughout) and Baseten DeepSeek-V4-Flash-0731 (top in 4 of 13, never below 88.6% of scenarios, in the caller's language throughout). Outside that pair, Gemini gemini-3.5-flash-lite has the highest floor: its weakest language still completes 85.09% of scenarios, in the caller's language throughout.
| Language | Top model | Task done | Adherence |
|---|---|---|---|
| Spanish | Baseten DeepSeek-V4-Flash-0731 | 91.23% | 100% |
| Arabic | OpenAI gpt-5.6-luna | 92.98% | 100% |
| Filipino | Gemini gemini-3.5-flash-lite | 94.74% | 100% |
| German | Gemini gemini-3.5-flash-lite | 94.74% | 100% |
| French | Baseten DeepSeek-V4-Flash-0731 | 92.98% | 100% |
| Norwegian | OpenAI gpt-4.1-mini | 93.86% | 100% |
| Hindi | OpenAI gpt-4.1-mini | 95.91% | 100% |
| Tamil | OpenAI gpt-4.1-mini | 98.25% | 100% |
| Telugu | Baseten DeepSeek-V4-Flash-0731 | 94.74% | 100% |
| Chinese | OpenAI gpt-4.1-mini | 95.91% | 100% |
| Korean | OpenAI gpt-4.1-mini | 92.11% | 100% |
| Thai | Baseten DeepSeek-V4-Flash-0731 | 89.47% | 100% |
| Japanese | OpenAI gpt-4.1-mini | 90.35% | 100% |
Numeric fidelity is its own failure axis: in the Arabic study, Alibaba qwen-turbo spoke tool-returned values back correctly 43% of the time. Full studies: Multilingual LLM.
Best value
Baseten DeepSeek-V4-Flash-0731 is the cheapest model on the board that never invented a fact it was not given: $0.26 per million output tokens, 558ms to first token, 93.16% task completion, with 36.8% dead air as the tradeoff (captured 2026-10-04). Grounding is not the only way to invent: check its tool-argument column before shipping it, which is a separate measurement and a separate failure. Alibaba qwen-turbo lists at $0.20 but completes 30.61% of tasks and fabricates on 90% of out-of-policy questions.
| Model | Cost / 1M tok | Task done | TTFT p50 |
|---|---|---|---|
| Alibaba qwen-turbo | $0.20 | 30.61% | 430ms |
| Baseten DeepSeek-V4-Flash-0731 | $0.26 | 93.16% | 558ms |
| OpenAI gpt-5-nano | $0.40 | 77.02% | 2.41s |
| Crusoe gemma-4-31b-it | $0.40 | — | — |
| Anthropic Claude Haiku 5.5 | $0.50 | 82.46% | 579ms |
How we measured
Full board and column definitions: LLM board; multilingual task and numeric-fidelity studies: Multilingual LLM; raw dataset: /data.json.