← LLM

Baseten

DeepSeek-V4-Flash-0731

reasoning off (gateway default)

Measured 2026-08-17

* Conditional

Cheapest row on the board; stalls mid-task on 27.8% of runs.

Score

89

Task done

80.18%

71–90% 95% CI bootstrap over items · n=95 · measured 2026-08-25

Refusal quality

72.2%

names the gap 57% · offers a route 67% · 10 probes x 3 iterations

Fabrication

10%

3 of 30 probe runs

Dead-air

5.1%

16 of 314 turns

Tool silence

7.4%

16 of 217 tool-call turns

Stalled

27.8%

25 of 90 runs

TTFT p50

361ms

361–432 p50–p90

Cost / 1M tok

$0.26

Share this result

Embed the live badge

Speko llm rank
Markdown
[![Speko llm rank](https://benchmarks.speko.ai/badge/llm/baseten-deepseek-v4-flash-0731.svg)](https://benchmarks.speko.ai/llm/baseten-deepseek-v4-flash-0731)
HTML
<a href="https://benchmarks.speko.ai/llm/baseten-deepseek-v4-flash-0731"><img src="https://benchmarks.speko.ai/badge/llm/baseten-deepseek-v4-flash-0731.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/baseten-deepseek-v4-flash-0731