← LLM 
Baseten
gpt-oss-120b
reasoning off (gateway default)
Measured 2026-08-17
* Conditional
Silent on every tool call; fabricates on 27% of out-of-policy questions.
Score
75
Task done
86.32%
72–98% 95% CI bootstrap over items · n=95 · measured 2026-08-25
Refusal quality
80.0%
names the gap 67% · offers a route 83% · 10 probes x 3 iterations · 3 silent turns scored zero
Fabrication
27%
8 of 30 probe runs
Dead-air
71.2%
227 of 319 turns
Tool silence
100.0%
all of 19 items (227 tool-call turns) · ≥82%
Stalled
8.9%
8 of 90 runs
TTFT p50
393ms
393–536 p50–p90
Cost / 1M tok
$0.50
Share this result
Embed the live badge
Markdown
[](https://benchmarks.speko.ai/llm/baseten-gpt-oss-120b)
HTML
<a href="https://benchmarks.speko.ai/llm/baseten-gpt-oss-120b"><img src="https://benchmarks.speko.ai/badge/llm/baseten-gpt-oss-120b.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/baseten-gpt-oss-120b