LLM

Multilingual

Same scenarios and grader as the English LLM board, per language.

Request a new language

Tell us which language you want added to this benchmark.

Task completion

Same six scenarios and the same deterministic grader as the English board. English lives on the LLM board — one number, one place. The sub-line is the share of replies actually spoken in the caller's language; models that answer in English are ranked last whatever they score.

#Model
1gpt-4.1OpenAI100%100%89%
2gpt-4.1-miniOpenAI100%100%100%51% in-language
3Cgemma-4-31bCerebras100%89%94%60% in-language
4Cgpt-oss-120bCerebras100%52% in-language94% answered in English78% answered in English
5TDeepSeek-V4-ProTogether89%94%94%
6gpt-5OpenAI83% 83% 78% 52% in-language
7gpt-5-miniOpenAI56%61%72%84% in-language
8qwen-turboAlibaba33% 80% in-language33% 90% in-language33% 55% in-language
9Grok-4.3xAI100% answered in English100% answered in English94% answered in English
10TLlama-3.3-70BTogether67% answered in English67% answered in English67% answered in English
Numeric fidelity

Different question: when a tool returns a date or an amount, does the model say that value back correctly? Judge-scored — the smaller number is the judge band.

#Model
1gpt-4.1-miniOpenAI100%100%81% 76–8195%
2gpt-5OpenAI100%100%90%90–95100%
3gpt-5-miniOpenAI100%100%71% 90%
4Cgpt-oss-120bCerebras100%100%81% 81–8686%
5Cgemma-4-31bCerebras100%100%100%95%
6TDeepSeek-V4-ProTogether100%90%90–100
7TLlama-3.3-70BTogether100%100%90% 86–90100%95–100
8Grok-4.3xAI100%100%100%100%
9gpt-4.1OpenAI95%100%94–10095%100%
10qwen-turboAlibaba86% 86%43% 67% 52–67