Open models

Mean per-run score over 19 scripted scenarios: 1 completed, 0 acted dangerously, part-credit if it stopped early but safely.Task doneInvented a policy or fact it was never given, % of out-of-policy questions.FabricationTurns where the caller hears silence (no spoken words), % of turns.Dead-airTime-to-first-token, p50 (us-east4). Whisker spans p50→p90.TTFT p50
Cerebras
gpt-oss-120b
85.00%70–96% 95% CI bootstrap over items · n=95 · measured 2026-08-25
30%9 of 30 probe runs
68.2%204 of 299 turns
195ms195–260 p50–p90
Cerebras
gemma-4-31b
77.72%62–92% 95% CI bootstrap over items · n=95 · measured 2026-08-25
0%0 of 10 probes (30 probe runs) · ≤31%
53.3%122 of 229 turns
192ms192–232 p50–p90
Together
Llama-3.3-70B
64.21%43–84% 95% CI bootstrap over items · n=95 · measured 2026-08-25
0%0 of 10 probes (30 probe runs) · ≤31%
74.3%260 of 350 turns
698ms698–860 p50–p90
Together
DeepSeek-V4-Pro
no measurement on the current corpus — transport unavailable
fabrication probes not run for this model
not measured on the current corpus
No real-time first-token measurement; exceeds the ~800ms voice latency budget.

Data: CC BY 4.0, attribution "Speko Benchmarks".