Best LLMs for Voice Agents 2026
Behaviour measured on multi-step tool scenarios over a live gateway, first-token latency from one US vantage; latest capture 2026-08-18. First-token latency, task completion, fabrication and cost for voice-agent LLMs.
The LLM board publishes no composite score and no single ranking: first-token latency, task completion, fabrication, dead air and cost are each measured and ranked on their own, and the model that wins one axis is routinely mid-table on another. The sections below name a leader per axis.
Fastest first token
Baseten inkling-small returns a first token in 177ms at the median (captured 2026-08-17) at $1.20 per million output tokens. Caveat: Fastest first token measured here; repeated a state-changing refund on 5 of 5 H3 runs.
| Model | TTFT p50 | Task done | Cost / 1M tok |
|---|---|---|---|
| Baseten inkling-small * | 177ms | 53% | $1.20 |
| Cerebras gemma-4-31b * | 192ms | 73% | $1.49 |
| Cerebras gpt-oss-120b * | 195ms | 80% | $0.75 |
| Baseten inkling * | 211ms | 53% | $4.05 |
| Baseten GLM-4.7 * | 275ms | 70% | $2.20 |
| Baseten Nemotron-3-Ultra * | 300ms | 67% | $2.40 |
Most reliable
No model on the current board pairs perfect task completion with zero fabrication and zero dead air; the closest profiles are below.
| Model | Task done | Fabrication | Dead-air | TTFT p50 |
|---|---|---|---|---|
| xAI Grok-4.3 * | 100% | 0% | 46% | 2.15s |
| OpenAI gpt-5.6-terra * | 100% | 0% | 55% | 701ms |
| OpenAI gpt-5.6-luna * | 100% | 0% | 56% | 659ms |
| Alibaba Qwen3.8-27B * | 93% | 0% | 7% | — |
| Anthropic Claude Haiku 4.5 * | 90% | 0% | 2% | 532ms |
| Together DeepSeek-V4-Pro * | 83% | 0% | 0% | — |
Best multilingual
Top scorer across the 10 languages: OpenAI gpt-5.6-terra (top in 8 of 10, never below 89% of scenarios, in the caller's language throughout) and OpenAI gpt-5.6-luna (top in 6 of 10, never below 83% of scenarios). Outside that pair, OpenAI gpt-4.1 has the highest floor: its weakest language still completes 78% of scenarios, in the caller's language throughout. Adherence is the real separator: xAI Grok-4.3, OpenAI gpt-4.1-mini, Cerebras gpt-oss-120b, Together Llama-3.3-70B, Anthropic Claude Haiku 4.5 answer the caller in English in at least two of the tested languages (xAI Grok-4.3 completes 83% of Arabic tasks while speaking no Arabic).
| Language | Top model | Task done | Adherence |
|---|---|---|---|
| Spanish | OpenAI gpt-5.6-luna | 100% | 100% |
| Arabic | Anthropic Claude Sonnet 5 | 100% | 100% |
| Filipino | OpenAI gpt-5.6-luna | 100% | 100% |
| German | OpenAI gpt-5.6-luna | 100% | 100% |
| French | OpenAI gpt-5.6-terra | 100% | 100% |
| Norwegian | Anthropic Claude Sonnet 5 | 100% | 100% |
| Hindi | Anthropic Claude Sonnet 5 | 100% | 100% |
| Tamil | Anthropic Claude Sonnet 5 | 100% | 100% |
| Telugu | OpenAI gpt-5.6-terra | 100% | 100% |
| Chinese | OpenAI gpt-5.6-luna | 100% | 100% |
Numeric fidelity is its own failure axis: in the Arabic study, Alibaba qwen-turbo spoke tool-returned values back correctly 43% of the time (said $70.45 for a $75.50 total on every run). Full studies: Multilingual LLM.
Best value
OpenAI gpt-5.6-luna is the cheapest model on the board that still completes every task without fabricating: $1.20 per million output tokens, 659ms to first token, with 56% dead air as the tradeoff (captured 2026-08-05). Alibaba qwen-turbo lists at $0.20 but completes 33% of tasks and fabricates on 38% of out-of-policy questions.
| Model | Cost / 1M tok | Task done | TTFT p50 |
|---|---|---|---|
| Alibaba qwen-turbo * | $0.20 | 33% | 483ms |
| Baseten DeepSeek-V4-Flash-0731 * | $0.26 | 67% | 361ms |
| OpenAI gpt-5-nano | $0.40 | 23% | 525ms |
| Baseten gpt-oss-120b * | $0.50 | 83% | 393ms |
| Cerebras gpt-oss-120b * | $0.75 | 80% | 195ms |
How we measured
Full board and column definitions: LLM board; multilingual task and numeric-fidelity studies: Multilingual LLM; raw dataset: /data.json.