LLM
Multilingual
The English board’s 19 scenarios and grader, run per language at 3 iterations.
Spanish
Task done
higher is betterDid the agent actually finish what the caller asked for — look the order up, issue the refund, book the appointment?
Answered in the language (nothing asked)
higher is betterWith no line in the prompt naming a language, how much of what the agent says comes back in the caller's language rather than English.
Stalled
lower is betterThe agent says it will do something, then goes quiet without doing it. The caller has to ask again.
Silence
lower is betterHow often the caller hears nothing at all — every turn counted. The small number is the narrower measure: the share of tool calls made in complete silence, which is pinned at 100% for about half the rows in every language and so cannot rank them.
Fabrication
lower is betterThe agent states a fact it was never given — an opening time, a delivery window, a policy. It sounds confident and it is invented. Measured on the run without the language line, the only pass that ran these probes.
Says numbers back correctly
higher is betterA tool returns $75.50 or a date; does the agent say that exact value out loud? Judge-scored on the 2026-07 model set, so it covers three of these languages — three vendors independently turn $75.50 into $57.50 in Arabic, because Arabic reads units before tens.