Best LLMs for Voice Agents 2026

Behaviour measured on multi-step tool scenarios over a live gateway, first-token latency from one US vantage; latest capture 2026-09-03. First-token latency, task completion, fabrication and cost for voice-agent LLMs.

The LLM board publishes no composite score and no single ranking: first-token latency, task completion, fabrication, dead air and cost are each measured and ranked on their own, and the model that wins one axis is routinely mid-table on another. The sections below name a leader per axis.

Fastest first token

Baseten inkling-small returns a first token in 177ms at the median (captured 2026-08-17) at $1.20 per million output tokens. Caveat: Fastest first token measured here; repeated a state-changing refund on 5 of 5 H3 runs.

Model TTFT p50 Task done Cost / 1M tok
Baseten inkling-small * 177ms57.72%$1.20
Cerebras gemma-4-31b * 192ms77.72%$1.49
Cerebras gpt-oss-120b * 195ms85.00%$0.75
Baseten inkling * 211ms71.14%$4.05
Baseten GLM-4.7 * 275ms72.46%$2.20
Baseten Nemotron-3-Ultra * 300ms72.54%$2.40

Most reliable

11 of the 23 measured models never invented a fact they were not given. OpenAI gpt-5.6-luna is the strongest of those on task completion while holding a real-time first-token measurement: 94.21% of the multi-step tool tasks, 69.7% dead air, 659ms p50 (captured 2026-08-05). Cerebras Qwen3.8-27B and Gemini gemini-3.8-flash also never fabricated but have no real-time first-token measurement, so they stay off the ranked lane.

Model Task done Fabrication Dead-air TTFT p50
OpenAI gpt-5.6-luna * 94.21%0%69.7%659ms
xAI Grok-4.3 * 92.46%0%54.8%2.15s
OpenAI gpt-4.1 * 91.40%10%58.9%640ms
OpenAI gpt-5-mini 89.12%0%67.1%746ms
Cerebras Qwen3.8-27B * 87.72%0%33.2%
OpenAI gpt-5.6-terra * 87.37%0%66.3%701ms

Best multilingual

Top scorer across the 11 languages: OpenAI gpt-5.6-luna (top in 4 of 11, never below 83.33% of scenarios) and OpenAI gpt-4.1-mini (top in 3 of 11, never below 73.83% of scenarios, in the caller's language throughout). Outside that pair, OpenAI gpt-4.1 has the highest floor: its weakest language still completes 87.13% of scenarios, in the caller's language throughout.

Language Top model Task done Adherence
Spanish xAI Grok-4.391.23%100%
Arabic OpenAI gpt-5.6-luna96.49%100%
Filipino OpenAI gpt-5.6-luna89.47%100%
German OpenAI gpt-4.1-mini92.54%100%
French OpenAI gpt-4.192.11%100%
Norwegian OpenAI gpt-4.192.98%100%
Hindi OpenAI gpt-4.1-mini92.98%100%
Tamil OpenAI gpt-5.6-luna93.86%100%
Telugu OpenAI gpt-5.6-terra93.57%97%
Chinese OpenAI gpt-5.6-luna90.35%97%
Korean Cerebras gpt-oss-120b87.72%100%

Numeric fidelity is its own failure axis: in the Arabic study, Alibaba qwen-turbo spoke tool-returned values back correctly 43% of the time (said $70.45 for a $75.50 total on every run). Full studies: Multilingual LLM.

Best value

Together Llama-3.3-70B is the cheapest model on the board that never invented a fact it was not given: $1.04 per million output tokens, 698ms to first token, 64.21% task completion, with 74.3% dead air as the tradeoff (captured 2026-08-25). Grounding is not the only way to invent: check its tool-argument column before shipping it, which is a separate measurement and a separate failure. Alibaba qwen-turbo lists at $0.20 but completes 6.58% of tasks and fabricates on 87% of out-of-policy questions.

Model Cost / 1M tok Task done TTFT p50
Alibaba qwen-turbo * $0.206.58%483ms
Baseten DeepSeek-V4-Flash-0731 * $0.2680.18%361ms
OpenAI gpt-5-nano $0.4032.37%525ms
Baseten gpt-oss-120b * $0.5086.32%393ms
Cerebras gpt-oss-120b * $0.7585.00%195ms

How we measured

Full board and column definitions: LLM board; multilingual task and numeric-fidelity studies: Multilingual LLM; raw dataset: /data.json.