← LLM 
Baseten
DeepSeek-V4-Flash-0731
reasoning off (gateway default)
Measured 2026-08-17
* Conditional
Cheapest row on the board; stalls mid-task on 27.8% of runs.
Score
89
Task done
80.18%
71–90% 95% CI bootstrap over items · n=95 · measured 2026-08-25
Refusal quality
72.2%
names the gap 57% · offers a route 67% · 10 probes x 3 iterations
Fabrication
10%
3 of 30 probe runs
Dead-air
5.1%
16 of 314 turns
Tool silence
7.4%
16 of 217 tool-call turns
Stalled
27.8%
25 of 90 runs
TTFT p50
361ms
361–432 p50–p90
Cost / 1M tok
$0.26
Share this result
Embed the live badge
Markdown
[](https://benchmarks.speko.ai/llm/baseten-deepseek-v4-flash-0731)
HTML
<a href="https://benchmarks.speko.ai/llm/baseten-deepseek-v4-flash-0731"><img src="https://benchmarks.speko.ai/badge/llm/baseten-deepseek-v4-flash-0731.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/baseten-deepseek-v4-flash-0731