← LLM

OpenAI

gpt-4.1

Measured 2026-08-25

* Conditional

Fast enough for a live turn; priciest of the group. TTFT measured 2026-07-27.

Score

74

Task done

91.40%

78–100% 95% CI bootstrap over items · n=95 · measured 2026-08-25

Refusal quality

78.9%

names the gap 70% · offers a route 80% · 10 probes x 3 iterations · 3 silent turns scored zero

Fabrication

10%

3 of 30 probe runs

Dead-air

58.9%

159 of 270 turns

Tool silence

93.5%

159 of 170 tool-call turns

Stalled

14.4%

13 of 90 runs

TTFT p50

640ms

640–775 p50–p90

Cost / 1M tok

$8.00

Share this result

Embed the live badge

Speko llm rank
Markdown
[![Speko llm rank](https://benchmarks.speko.ai/badge/llm/openai-gpt-4-1.svg)](https://benchmarks.speko.ai/llm/openai-gpt-4-1)
HTML
<a href="https://benchmarks.speko.ai/llm/openai-gpt-4-1"><img src="https://benchmarks.speko.ai/badge/llm/openai-gpt-4-1.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/openai-gpt-4-1