← LLM 
OpenAI
gpt-5
reasoning
Measured 2026-08-25
Score
67
Task done
74.74%
60–88% 95% CI bootstrap over items · n=95 · measured 2026-08-25
Refusal quality
84.4%
names the gap 77% · offers a route 77% · 10 probes x 3 iterations
Fabrication
10%
3 of 30 probe runs
Dead-air
71.8%
216 of 301 turns
Tool silence
100.0%
all of 19 items (216 tool-call turns) · ≥82%
Stalled
25.6%
23 of 90 runs
TTFT p50
635ms
635–679 p50–p90
Cost / 1M tok
$10.00
Share this result
Embed the live badge
Markdown
[](https://benchmarks.speko.ai/llm/openai-gpt-5)
HTML
<a href="https://benchmarks.speko.ai/llm/openai-gpt-5"><img src="https://benchmarks.speko.ai/badge/llm/openai-gpt-5.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/openai-gpt-5