Open models

Multi-step tool tasks completed end-to-end (chains, compound, clarify), % of runs.Task doneInvented a policy or fact it was never given, % of out-of-policy questions.FabricationTurns where the caller hears silence (no spoken words), % of turns.Dead-airTime-to-first-token, p50 (us-east4). Whisker spans p50→p90.TTFT p50
05Rank 5 of 12 on the full LLM board — not re-ranked for this page.
Cerebras
gpt-oss-120b
94%
10%
51%
195ms195–260 p50–p90
01Rank 1 of 12 on the full LLM board — not re-ranked for this page.
Cerebras
gemma-4-31b
89% †
0%
12%
192ms192–232 p50–p90
10Rank 10 of 12 on the full LLM board — not re-ranked for this page.
Together
Llama-3.3-70B
56%
0%
29%
698ms698–860 p50–p90
Together
DeepSeek-V4-Pro
100%
0%
0%
No real-time first-token measurement; exceeds the ~800ms voice latency budget.