← LLM

OpenAI

gpt-5

reasoning

Measured 2026-08-25

Score

67

Task done

74.74%

60–88% 95% CI bootstrap over items · n=95 · measured 2026-08-25

Refusal quality

84.4%

names the gap 77% · offers a route 77% · 10 probes x 3 iterations

Fabrication

10%

3 of 30 probe runs

Dead-air

71.8%

216 of 301 turns

Tool silence

100.0%

all of 19 items (216 tool-call turns) · ≥82%

Stalled

25.6%

23 of 90 runs

TTFT p50

635ms

635–679 p50–p90

Cost / 1M tok

$10.00

Share this result

Embed the live badge

Speko llm rank
Markdown
[![Speko llm rank](https://benchmarks.speko.ai/badge/llm/openai-gpt-5.svg)](https://benchmarks.speko.ai/llm/openai-gpt-5)
HTML
<a href="https://benchmarks.speko.ai/llm/openai-gpt-5"><img src="https://benchmarks.speko.ai/badge/llm/openai-gpt-5.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/openai-gpt-5