LLM

Multilingual

The English board’s 19 scenarios and grader, run per language at 3 iterations.

Request a new language

Tell us which language you want added to this benchmark.

Caller speaks

Spanish

Task done

higher is better

Did the agent actually finish what the caller asked for — look the order up, issue the refund, book the appointment?

91.23%
90.35%
88.89%
85.96%
85.09%
Gemini83.77%
Cerebras83.33%
Cerebras81.58%
Anthropic79.39%
79.24%
Baseten76.17%
Anthropic76.02%
Baseten75.73%
Baseten74.56%
73.98%
Baseten69.01%
Baseten63.16%
31.87%
Together31.58%
12.43%
0%25%50%75%100%

Answered in the language (nothing asked)

higher is better

With no line in the prompt naming a language, how much of what the agent says comes back in the caller's language rather than English.

AnthropicAnthropicClaude Sonnet 5100%≥82%
BasetenBasetenGLM-4.7100%≥82%
OpenAIgpt-5100%≥82%
OpenAIgpt-5-mini100%≥82%
OpenAIgpt-5.6-luna100%≥82%
OpenAIgpt-5.6-terra100%≥82%
BasetenBaseteninkling100%≥82%
OpenAIgpt-4.198%
CerebrasCerebrasgemma-4-31b95%
OpenAIgpt-5-nano93%
AnthropicAnthropicClaude Haiku 4.591%
Alibabaqwen-turbo90%
BasetenBaseteninkling-small86%
GeminiGemini3.5-flash-lite72%
BasetenBasetenNemotron-3-Ultra46%
TogetherTogetherLlama-3.3-70B19%
CerebrasCerebrasgpt-oss-120b16%
OpenAIgpt-4.1-mini13%
xAIGrok-4.312%
BasetenBasetenDeepSeek-V4-Flash-073111%
0%25%50%75%100%

Stalled

lower is better

The agent says it will do something, then goes quiet without doing it. The caller has to ask again.

Silence

lower is better

How often the caller hears nothing at all — every turn counted. The small number is the narrower measure: the share of tool calls made in complete silence, which is pinned at 100% for about half the rows in every language and so cannot rank them.

AnthropicClaude Haiku 4.51%1% of tool calls1 of 146 turns
BasetenGLM-4.71%1% of tool calls1 of 156 turns
BasetenDeepSeek-V4-Flash-073115%21% of tool calls27 of 186 turns
qwen-turbo15%79% of tool calls11 of 74 turns
gpt-5-nano37%100% of tool calls ≥78%35 of 95 turns
AnthropicClaude Sonnet 539%57% of tool calls64 of 166 turns
gpt-4.1-mini41%69% of tool calls61 of 148 turns
Cerebrasgemma-4-31b55%99% of tool calls76 of 137 turns
Grok-4.357%92% of tool calls90 of 158 turns
BasetenNemotron-3-Ultra57%100% of tool calls ≥82%79 of 139 turns
gpt-4.160%97% of tool calls93 of 156 turns
gpt-5.6-terra64%100% of tool calls ≥82%107 of 167 turns
gpt-5-mini65%100% of tool calls ≥82%110 of 170 turns
Baseteninkling-small66%92% of tool calls132 of 201 turns
Cerebrasgpt-oss-120b69%100% of tool calls ≥82%128 of 185 turns
gpt-5.6-luna70%100% of tool calls ≥82%140 of 199 turns
Geminigemini-3.5-flash-lite70%100% of tool calls ≥82%124 of 178 turns
Baseteninkling71%96% of tool calls141 of 199 turns
gpt-573%100% of tool calls ≥82%137 of 187 turns
TogetherLlama-3.3-70B94%100% of tool calls ≥82%240 of 256 turns
0%25%50%75%100%

Fabrication

lower is better

The agent states a fact it was never given — an opening time, a delivery window, a policy. It sounds confident and it is invented. Measured on the run without the language line, the only pass that ran these probes.

AnthropicCerebras7 models0 of 3031% · 10 independent0%
qwen-turbo21 of 3070%
BasetenGLM-4.716 of 3053%
gpt-4.1-mini15 of 3050%
gpt-512 of 3040%
Cerebrasgpt-oss-120b9 of 3030%
gpt-5-nano6 of 3020%
gpt-4.15 of 3017%
gpt-5-mini5 of 3017%
TogetherLlama-3.3-70B1 of 303%

Says numbers back correctly

higher is better

A tool returns $75.50 or a date; does the agent say that exact value out loud? Judge-scored on the 2026-07 model set, so it covers three of these languages — three vendors independently turn $75.50 into $57.50 in Arabic, because Arabic reads units before tens.

CerebrasCerebrasgemma-4-31b100%
OpenAIgpt-4.1100%
OpenAIgpt-4.1-mini100%
OpenAIgpt-5100%
OpenAIgpt-5-mini100%
CerebrasCerebrasgpt-oss-120b100%
xAIGrok-4.3100%
TogetherTogetherLlama-3.3-70B100%
Alibabaqwen-turbo86%
0%25%50%75%100%