Request a new model

Tell us which model you want measured. We review every request.

LLM

Task done

higher is better

Did the agent finish the job without doing damage — look the order up, refund the right amount, and leave alone what it must not touch?

94.21%
92.46%
91.40%
89.12%
Cerebras87.72%
87.37%
87.19%
Baseten86.32%
Cerebras85.00%
Anthropic82.46%
Anthropic81.72%
Baseten80.18%
Cerebras77.72%
Gemini77.37%
74.74%
Baseten72.54%
Baseten72.46%
Gemini72.02%
Baseten71.14%
Together64.21%
Baseten57.72%
32.37%
6.58%
0%25%50%75%100%

Stalled

lower is better

The agent says it will do something, then goes quiet without doing it. The caller has to ask again.

Fabrication

lower is better

The agent states a fact it was never given — an opening time, a delivery window, a policy. It sounds confident and it is invented.

Refusal quality

higher is better

When the agent does not know something, does it say what it is missing and offer a way forward, or just stop? Scored on the same probes as Fabrication.

Silence

lower is better

How often the caller hears nothing at all — every turn counted, not just the ones where a tool runs. The small number is the narrower measure: the share of tool calls the agent made in complete silence, which for 10 of the 22 measured models is all of them.

AnthropicClaude Haiku 4.52%3% of tool calls4 of 248 turns
BasetenGLM-4.72%4% of tool calls6 of 260 turns
BasetenDeepSeek-V4-Flash-07315%7% of tool calls16 of 314 turns
qwen-turbo9%100% of tool calls ≥16%10 of 110 turns
CerebrasQwen3.8-27B33%49% of tool calls98 of 295 turns
gpt-5-nano39%100% of tool calls ≥80%65 of 165 turns
AnthropicClaude Sonnet 540%59% of tool calls117 of 294 turns
gpt-4.1-mini40%66% of tool calls102 of 255 turns
Cerebrasgemma-4-31b53%95% of tool calls122 of 229 turns
Grok-4.355%83% of tool calls155 of 283 turns
BasetenNemotron-3-Ultra56%94% of tool calls139 of 248 turns
gpt-4.159%94% of tool calls159 of 270 turns
Baseteninkling66%94% of tool calls203 of 306 turns
gpt-5.6-terra66%100% of tool calls ≥82%195 of 294 turns
gpt-5-mini67%100% of tool calls ≥82%198 of 295 turns
Cerebrasgpt-oss-120b68%100% of tool calls ≥82%204 of 299 turns
Geminigemini-3.5-flash-lite69%100% of tool calls ≥82%197 of 287 turns
Geminigemini-3.8-flash69%100% of tool calls ≥82%214 of 310 turns
gpt-5.6-luna70%100% of tool calls ≥82%230 of 330 turns
Basetengpt-oss-120b71%100% of tool calls ≥82%227 of 319 turns
gpt-572%100% of tool calls ≥82%216 of 301 turns
Baseteninkling-small72%95% of tool calls244 of 339 turns
TogetherLlama-3.3-70B74%100% of tool calls ≥82%260 of 350 turns
0%25%50%75%100%

Tool-arg fabrication

lower is better

The prompt orders a lookup the agent has no id for. It calls the tool anyway and fills the argument in — blank, a placeholder, or a number it made up.

AnthropicTogether6 models0 of 300%
gpt-5.6-luna30 of 30100%
gpt-5-mini30 of 30100%
gpt-530 of 30100%
gpt-5-nano30 of 30100%
TogetherLlama-3.3-70B30 of 30100%
AnthropicClaude Haiku 4.523 of 3077%
Cerebrasgpt-oss-120b15 of 3050%
Cerebrasgemma-4-31b10 of 3033%
gpt-4.1-mini8 of 3027%
gpt-5.6-terra2 of 307%

Placeholder spoken

lower is better

A template variable the platform never filled in. The agent reads it out loud and the caller hears the braces.

Time to first token

Region

How long the caller waits after they stop speaking before the agent starts to answer.

CerebrasCerebrasgemma-4-31b78ms · 118ms
CerebrasCerebrasgpt-oss-120b100ms · 177ms
BasetenBaseteninkling-small100ms · 109ms
CerebrasCerebrasQwen3.8-27B103ms · 148ms
BasetenBaseteninkling109ms · 132ms
BasetenBasetengpt-oss-120b224ms · 343ms
TogetherTogetherLlama-3.3-70B267ms · 554ms
BasetenBasetenGLM-4.7348ms · 545ms
Alibabaqwen-turbo383ms · 419ms
OpenAIgpt-4.1407ms · 562ms
GeminiGemini3.5-flash-lite408ms · 502ms
OpenAIgpt-5.6-luna463ms · 623ms
OpenAIgpt-4.1-mini514ms · 636ms
OpenAIgpt-5-mini556ms · 826ms
OpenAIgpt-5.6-terra562ms · 824ms
OpenAIgpt-5595ms · 770ms
OpenAIgpt-5-nano671ms · 841ms
xAIGrok-4.32.05s · 2.66s
0ms500ms1.00s1.50s2.00s2.50s

p50 p90

Cost

lower is better

What the model charges for a million tokens of its own output, at each vendor's published list price.

FAQ

What is the best LLM for a real-time voice agent?
The board ranks each parameter on its own and publishes no single winner, and on this run the parameters disagree. OpenAI gpt-5.6-luna each completed 94.21% of the multi-step tool tasks while staying silent on most of their tool calls, whereas Anthropic Claude Haiku 4.5 completed 82.46% and goes silent on only 2.7%.
Which LLM has the lowest first-token latency?
Baseten inkling-small measured 177ms TTFT p50 from us-east4, the lowest on the board. Cerebras gpt-oss-120b measured 195ms but stays silent on 100.0% of tool calls.
How fast does an LLM need to respond for voice?
Every millisecond before the first token is silence to the caller, so the board measures first-token latency as its own parameter. Measured p50s range from 177ms (Baseten inkling-small) to 2.15s (xAI Grok-4.3), which is too slow to lead a live turn.
Which LLMs fabricate answers in a voice agent?
Fabrication is inventing a policy or fact the agent was never given. OpenAI gpt-4.1-mini measured 57% on out-of-policy questions and Alibaba qwen-turbo 87%, while Anthropic Claude Haiku 4.5 measured 0%.
What is dead air in a voice agent?
A turn where the caller hears silence instead of speech. Anthropic Claude Haiku 4.5 measured 1.6% of turns against 68.2% for Cerebras gpt-oss-120b, but a quiet agent is not the same as a working one: Alibaba qwen-turbo speaks on almost every turn and still finished only 6.58% of the tool tasks, because it narrates instead of acting.