← LLM

Baseten

gpt-oss-120b

reasoning off (gateway default)

Measured 2026-08-17

* Conditional

Silent on every tool call; fabricates on 27% of out-of-policy questions.

Score

75

Task done

86.32%

72–98% 95% CI bootstrap over items · n=95 · measured 2026-08-25

Refusal quality

80.0%

names the gap 67% · offers a route 83% · 10 probes x 3 iterations · 3 silent turns scored zero

Fabrication

27%

8 of 30 probe runs

Dead-air

71.2%

227 of 319 turns

Tool silence

100.0%

all of 19 items (227 tool-call turns) · ≥82%

Stalled

8.9%

8 of 90 runs

TTFT p50

393ms

393–536 p50–p90

Cost / 1M tok

$0.50

Share this result

Embed the live badge

Speko llm rank
Markdown
[![Speko llm rank](https://benchmarks.speko.ai/badge/llm/baseten-gpt-oss-120b.svg)](https://benchmarks.speko.ai/llm/baseten-gpt-oss-120b)
HTML
<a href="https://benchmarks.speko.ai/llm/baseten-gpt-oss-120b"><img src="https://benchmarks.speko.ai/badge/llm/baseten-gpt-oss-120b.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/baseten-gpt-oss-120b