Best LLMs for Voice Agents 2026
Behaviour measured on multi-step tool scenarios over a live gateway, first-token latency from one US vantage; latest capture 2026-09-03. First-token latency, task completion, fabrication and cost for voice-agent LLMs.
The LLM board publishes no composite score and no single ranking: first-token latency, task completion, fabrication, dead air and cost are each measured and ranked on their own, and the model that wins one axis is routinely mid-table on another. The sections below name a leader per axis.
Fastest first token
Baseten inkling-small returns a first token in 177ms at the median (captured 2026-08-17) at $1.20 per million output tokens. Caveat: Fastest first token measured here; repeated a state-changing refund on 5 of 5 H3 runs.
| Model | TTFT p50 | Task done | Cost / 1M tok |
|---|---|---|---|
| Baseten inkling-small * | 177ms | 57.72% | $1.20 |
| Cerebras gemma-4-31b * | 192ms | 77.72% | $1.49 |
| Cerebras gpt-oss-120b * | 195ms | 85.00% | $0.75 |
| Baseten inkling * | 211ms | 71.14% | $4.05 |
| Baseten GLM-4.7 * | 275ms | 72.46% | $2.20 |
| Baseten Nemotron-3-Ultra * | 300ms | 72.54% | $2.40 |
Most reliable
11 of the 23 measured models never invented a fact they were not given. OpenAI gpt-5.6-luna is the strongest of those on task completion while holding a real-time first-token measurement: 94.21% of the multi-step tool tasks, 69.7% dead air, 659ms p50 (captured 2026-08-05). Cerebras Qwen3.8-27B and Gemini gemini-3.8-flash also never fabricated but have no real-time first-token measurement, so they stay off the ranked lane.
| Model | Task done | Fabrication | Dead-air | TTFT p50 |
|---|---|---|---|---|
| OpenAI gpt-5.6-luna * | 94.21% | 0% | 69.7% | 659ms |
| xAI Grok-4.3 * | 92.46% | 0% | 54.8% | 2.15s |
| OpenAI gpt-4.1 * | 91.40% | 10% | 58.9% | 640ms |
| OpenAI gpt-5-mini | 89.12% | 0% | 67.1% | 746ms |
| Cerebras Qwen3.8-27B * | 87.72% | 0% | 33.2% | — |
| OpenAI gpt-5.6-terra * | 87.37% | 0% | 66.3% | 701ms |
Best multilingual
Top scorer across the 11 languages: OpenAI gpt-5.6-luna (top in 4 of 11, never below 83.33% of scenarios) and OpenAI gpt-4.1-mini (top in 3 of 11, never below 73.83% of scenarios, in the caller's language throughout). Outside that pair, OpenAI gpt-4.1 has the highest floor: its weakest language still completes 87.13% of scenarios, in the caller's language throughout.
| Language | Top model | Task done | Adherence |
|---|---|---|---|
| Spanish | xAI Grok-4.3 | 91.23% | 100% |
| Arabic | OpenAI gpt-5.6-luna | 96.49% | 100% |
| Filipino | OpenAI gpt-5.6-luna | 89.47% | 100% |
| German | OpenAI gpt-4.1-mini | 92.54% | 100% |
| French | OpenAI gpt-4.1 | 92.11% | 100% |
| Norwegian | OpenAI gpt-4.1 | 92.98% | 100% |
| Hindi | OpenAI gpt-4.1-mini | 92.98% | 100% |
| Tamil | OpenAI gpt-5.6-luna | 93.86% | 100% |
| Telugu | OpenAI gpt-5.6-terra | 93.57% | 97% |
| Chinese | OpenAI gpt-5.6-luna | 90.35% | 97% |
| Korean | Cerebras gpt-oss-120b | 87.72% | 100% |
Numeric fidelity is its own failure axis: in the Arabic study, Alibaba qwen-turbo spoke tool-returned values back correctly 43% of the time (said $70.45 for a $75.50 total on every run). Full studies: Multilingual LLM.
Best value
Together Llama-3.3-70B is the cheapest model on the board that never invented a fact it was not given: $1.04 per million output tokens, 698ms to first token, 64.21% task completion, with 74.3% dead air as the tradeoff (captured 2026-08-25). Grounding is not the only way to invent: check its tool-argument column before shipping it, which is a separate measurement and a separate failure. Alibaba qwen-turbo lists at $0.20 but completes 6.58% of tasks and fabricates on 87% of out-of-policy questions.
| Model | Cost / 1M tok | Task done | TTFT p50 |
|---|---|---|---|
| Alibaba qwen-turbo * | $0.20 | 6.58% | 483ms |
| Baseten DeepSeek-V4-Flash-0731 * | $0.26 | 80.18% | 361ms |
| OpenAI gpt-5-nano | $0.40 | 32.37% | 525ms |
| Baseten gpt-oss-120b * | $0.50 | 86.32% | 393ms |
| Cerebras gpt-oss-120b * | $0.75 | 85.00% | 195ms |
How we measured
Full board and column definitions: LLM board; multilingual task and numeric-fidelity studies: Multilingual LLM; raw dataset: /data.json.