← LLM

OpenAI

gpt-4.1-mini

Measured 2026-08-25

* Conditional

Fabricates on 57% of out-of-policy questions — the worst rate on the board bar qwen-turbo. TTFT measured 2026-07-27.

Score

70

Task done

87.19%

78–95% 95% CI bootstrap over items · n=95 · measured 2026-08-25

Refusal quality

83.3%

names the gap 80% · offers a route 80% · 10 probes x 3 iterations · 3 silent turns scored zero

Fabrication

57%

17 of 30 probe runs

Dead-air

40.0%

102 of 255 turns

Tool silence

65.8%

102 of 155 tool-call turns

Stalled

16.7%

15 of 90 runs

TTFT p50

708ms

708–1094 p50–p90

Cost / 1M tok

$1.60

Share this result

Embed the live badge

Speko llm rank
Markdown
[![Speko llm rank](https://benchmarks.speko.ai/badge/llm/openai-gpt-4-1-mini.svg)](https://benchmarks.speko.ai/llm/openai-gpt-4-1-mini)
HTML
<a href="https://benchmarks.speko.ai/llm/openai-gpt-4-1-mini"><img src="https://benchmarks.speko.ai/badge/llm/openai-gpt-4-1-mini.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/openai-gpt-4-1-mini