LLM
Multilingual
Same scenarios and grader as the English LLM board, per language.
Task completion
Same six scenarios and the same deterministic grader as the English board. English lives on the LLM board — one number, one place. The sub-line is the share of replies actually spoken in the caller's language; models that answer in English are ranked last whatever they score.
| # | Model | |||
|---|---|---|---|---|
| 1 | gpt-4.1OpenAI | 100% | 100% | 89% |
| 2 | gpt-4.1-miniOpenAI | 100% | 100% | 100%51% in-language |
| 3 | gemma-4-31bCerebras | 100% | 89% | 94%60% in-language |
| 4 | gpt-oss-120bCerebras | 100%52% in-language | 94% ⚠answered in English | 78% ⚠answered in English |
| 5 | DeepSeek-V4-ProTogether | 89% | 94% | 94% |
| 6 | gpt-5OpenAI | 83% ⚠ | 83% ⚠ | 78% ⚠52% in-language |
| 7 | gpt-5-miniOpenAI | 56% | 61% | 72%84% in-language |
| 8 | qwen-turboAlibaba | 33% ⚠80% in-language | 33% ⚠90% in-language | 33% ⚠55% in-language |
| 9 | Grok-4.3xAI | 100% ⚠answered in English | 100% ⚠answered in English | 94% ⚠answered in English |
| 10 | Llama-3.3-70BTogether | 67% ⚠answered in English | 67% ⚠answered in English | 67% ⚠answered in English |
Numeric fidelity
Different question: when a tool returns a date or an amount, does the model say that value back correctly? Judge-scored — the smaller number is the judge band.
| # | Model | ||||
|---|---|---|---|---|---|
| 1 | gpt-4.1-miniOpenAI | 100% | 100% | 81% ⚠76–81 | 95% |
| 2 | gpt-5OpenAI | 100% | 100% | 90%90–95 | 100% |
| 3 | gpt-5-miniOpenAI | 100% | 100% | 71% ⚠ | 90% |
| 4 | gpt-oss-120bCerebras | 100% | 100% | 81% ⚠81–86 | 86% |
| 5 | gemma-4-31bCerebras | 100% | 100% | 100% | 95% |
| 6 | DeepSeek-V4-ProTogether | 100% | 90%90–100 | — | — |
| 7 | Llama-3.3-70BTogether | 100% | 100% | 90% ⚠86–90 | 100%95–100 |
| 8 | Grok-4.3xAI | 100% | 100% | 100% | 100% |
| 9 | gpt-4.1OpenAI | 95% | 100%94–100 | 95% | 100% |
| 10 | qwen-turboAlibaba | 86% ⚠ | 86% | 43% ⚠ | 67% ⚠52–67 |