Best LLMs for Voice Agents 2026

Behaviour measured on multi-step tool scenarios over a live Speko route, first-token latency from one US region; latest capture 2026-10-08. First-token latency, task completion, fabrication and cost for voice-agent LLMs.

The LLM board publishes no composite score and no single ranking: first-token latency, task completion, fabrication, dead air and cost are each measured and ranked on their own, and the model that wins one axis is routinely mid-table on another. The sections below name a leader per axis.

Fastest first token

Cerebras gpt-oss-120b returns a first token in 139ms at the median (captured 2026-10-04) at $0.75 per million output tokens.

Model TTFT p50 Task done Cost / 1M tok
Cerebras gpt-oss-120b 139ms82.02%$0.75
Baseten GLM-4.7 148ms80.88%$2.20
Together DeepSeek-V4.1-Flash 157ms84.21%$1.20
Baseten Nemotron-3-Ultra 245ms86.84%$2.40
Cerebras Qwen3.8-27B 246ms86.14%$1.49
OpenAI gpt-5.4-mini 390ms91.75%$4.50

Most reliable

19 of the 27 measured models never invented a fact they were not given. Baseten DeepSeek-V4-Flash-0731 is the strongest of those on task completion while holding a real-time first-token measurement: 93.16% of the multi-step tool tasks, 36.8% dead air, 558ms p50 (captured 2026-10-04).

Model Task done Fabrication Dead-air TTFT p50
Gemini gemini-3.5-flash-lite 94.74%10%63.1%500ms
Baseten DeepSeek-V4-Flash-0731 93.16%0%36.8%558ms
OpenAI gpt-4.1-mini 92.63%3%61.4%480ms
OpenAI gpt-5.4-mini 91.75%0%62.2%390ms
xAI Grok-4.3 90.88%0%56.9%2.38s
OpenAI gpt-4.1 89.12%10%61.5%462ms

Best multilingual

Top scorer across the 13 languages: OpenAI gpt-4.1-mini (top in 6 of 13, never below 78.07% of scenarios, in the caller's language throughout) and Baseten DeepSeek-V4-Flash-0731 (top in 4 of 13, never below 88.6% of scenarios, in the caller's language throughout). Outside that pair, Gemini gemini-3.5-flash-lite has the highest floor: its weakest language still completes 85.09% of scenarios, in the caller's language throughout.

Language Top model Task done Adherence
Spanish Baseten DeepSeek-V4-Flash-073191.23%100%
Arabic OpenAI gpt-5.6-luna92.98%100%
Filipino Gemini gemini-3.5-flash-lite94.74%100%
German Gemini gemini-3.5-flash-lite94.74%100%
French Baseten DeepSeek-V4-Flash-073192.98%100%
Norwegian OpenAI gpt-4.1-mini93.86%100%
Hindi OpenAI gpt-4.1-mini95.91%100%
Tamil OpenAI gpt-4.1-mini98.25%100%
Telugu Baseten DeepSeek-V4-Flash-073194.74%100%
Chinese OpenAI gpt-4.1-mini95.91%100%
Korean OpenAI gpt-4.1-mini92.11%100%
Thai Baseten DeepSeek-V4-Flash-073189.47%100%
Japanese OpenAI gpt-4.1-mini90.35%100%

Numeric fidelity is its own failure axis: in the Arabic study, Alibaba qwen-turbo spoke tool-returned values back correctly 43% of the time. Full studies: Multilingual LLM.

Best value

Baseten DeepSeek-V4-Flash-0731 is the cheapest model on the board that never invented a fact it was not given: $0.26 per million output tokens, 558ms to first token, 93.16% task completion, with 36.8% dead air as the tradeoff (captured 2026-10-04). Grounding is not the only way to invent: check its tool-argument column before shipping it, which is a separate measurement and a separate failure. Alibaba qwen-turbo lists at $0.20 but completes 30.61% of tasks and fabricates on 90% of out-of-policy questions.

Model Cost / 1M tok Task done TTFT p50
Alibaba qwen-turbo $0.2030.61%430ms
Baseten DeepSeek-V4-Flash-0731 $0.2693.16%558ms
OpenAI gpt-5-nano $0.4077.02%2.41s
Crusoe gemma-4-31b-it $0.40——
Anthropic Claude Haiku 5.5 $0.5082.46%579ms

How we measured

Full board and column definitions: LLM board; multilingual task and numeric-fidelity studies: Multilingual LLM; raw dataset: /data.json.