← LLM 
OpenAI
gpt-4.1
Measured 2026-08-25
* Conditional
Fast enough for a live turn; priciest of the group. TTFT measured 2026-07-27.
Score
74
Task done
91.40%
78–100% 95% CI bootstrap over items · n=95 · measured 2026-08-25
Refusal quality
78.9%
names the gap 70% · offers a route 80% · 10 probes x 3 iterations · 3 silent turns scored zero
Fabrication
10%
3 of 30 probe runs
Dead-air
58.9%
159 of 270 turns
Tool silence
93.5%
159 of 170 tool-call turns
Stalled
14.4%
13 of 90 runs
TTFT p50
640ms
640–775 p50–p90
Cost / 1M tok
$8.00
Share this result
Embed the live badge
Markdown
[](https://benchmarks.speko.ai/llm/openai-gpt-4-1)
HTML
<a href="https://benchmarks.speko.ai/llm/openai-gpt-4-1"><img src="https://benchmarks.speko.ai/badge/llm/openai-gpt-4-1.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/openai-gpt-4-1