Best LLMs for Voice Agents 2026

Behaviour measured on multi-step tool scenarios over a live gateway, first-token latency from one US vantage; latest capture 2026-08-18. First-token latency, task completion, fabrication and cost for voice-agent LLMs.

The LLM board publishes no composite score and no single ranking: first-token latency, task completion, fabrication, dead air and cost are each measured and ranked on their own, and the model that wins one axis is routinely mid-table on another. The sections below name a leader per axis.

Fastest first token

Baseten inkling-small returns a first token in 177ms at the median (captured 2026-08-17) at $1.20 per million output tokens. Caveat: Fastest first token measured here; repeated a state-changing refund on 5 of 5 H3 runs.

Model TTFT p50 Task done Cost / 1M tok
Baseten inkling-small * 177ms53%$1.20
Cerebras gemma-4-31b * 192ms73%$1.49
Cerebras gpt-oss-120b * 195ms80%$0.75
Baseten inkling * 211ms53%$4.05
Baseten GLM-4.7 * 275ms70%$2.20
Baseten Nemotron-3-Ultra * 300ms67%$2.40

Most reliable

No model on the current board pairs perfect task completion with zero fabrication and zero dead air; the closest profiles are below.

Model Task done Fabrication Dead-air TTFT p50
xAI Grok-4.3 * 100%0%46%2.15s
OpenAI gpt-5.6-terra * 100%0%55%701ms
OpenAI gpt-5.6-luna * 100%0%56%659ms
Alibaba Qwen3.8-27B * 93%0%7%
Anthropic Claude Haiku 4.5 * 90%0%2%532ms
Together DeepSeek-V4-Pro * 83%0%0%

Best multilingual

Top scorer across the 10 languages: OpenAI gpt-5.6-terra (top in 8 of 10, never below 89% of scenarios, in the caller's language throughout) and OpenAI gpt-5.6-luna (top in 6 of 10, never below 83% of scenarios). Outside that pair, OpenAI gpt-4.1 has the highest floor: its weakest language still completes 78% of scenarios, in the caller's language throughout. Adherence is the real separator: xAI Grok-4.3, OpenAI gpt-4.1-mini, Cerebras gpt-oss-120b, Together Llama-3.3-70B, Anthropic Claude Haiku 4.5 answer the caller in English in at least two of the tested languages (xAI Grok-4.3 completes 83% of Arabic tasks while speaking no Arabic).

Language Top model Task done Adherence
Spanish OpenAI gpt-5.6-luna100%100%
Arabic Anthropic Claude Sonnet 5100%100%
Filipino OpenAI gpt-5.6-luna100%100%
German OpenAI gpt-5.6-luna100%100%
French OpenAI gpt-5.6-terra100%100%
Norwegian Anthropic Claude Sonnet 5100%100%
Hindi Anthropic Claude Sonnet 5100%100%
Tamil Anthropic Claude Sonnet 5100%100%
Telugu OpenAI gpt-5.6-terra100%100%
Chinese OpenAI gpt-5.6-luna100%100%

Numeric fidelity is its own failure axis: in the Arabic study, Alibaba qwen-turbo spoke tool-returned values back correctly 43% of the time (said $70.45 for a $75.50 total on every run). Full studies: Multilingual LLM.

Best value

OpenAI gpt-5.6-luna is the cheapest model on the board that still completes every task without fabricating: $1.20 per million output tokens, 659ms to first token, with 56% dead air as the tradeoff (captured 2026-08-05). Alibaba qwen-turbo lists at $0.20 but completes 33% of tasks and fabricates on 38% of out-of-policy questions.

Model Cost / 1M tok Task done TTFT p50
Alibaba qwen-turbo * $0.2033%483ms
Baseten DeepSeek-V4-Flash-0731 * $0.2667%361ms
OpenAI gpt-5-nano $0.4023%525ms
Baseten gpt-oss-120b * $0.5083%393ms
Cerebras gpt-oss-120b * $0.7580%195ms

How we measured

Full board and column definitions: LLM board; multilingual task and numeric-fidelity studies: Multilingual LLM; raw dataset: /data.json.