← LLM 
Baseten
Nemotron-3-Ultra
non-thinking at default
Measured 2026-08-18
* Conditional
Fabricates on 27% of out-of-policy questions and stalls on 54.4% of runs.
Score
76
Task done
72.54%
60–86% 95% CI bootstrap over items · n=95 · measured 2026-08-25
Refusal quality
75.6%
names the gap 83% · offers a route 60% · 10 probes x 3 iterations · 1 silent turn scored zero
Fabrication
27%
8 of 30 probe runs
Dead-air
56.0%
139 of 248 turns
Tool silence
93.9%
139 of 148 tool-call turns
Stalled
54.4%
49 of 90 runs
TTFT p50
300ms
300–370 p50–p90
Cost / 1M tok
$2.40
Share this result
Embed the live badge
Markdown
[](https://benchmarks.speko.ai/llm/baseten-nemotron-3-ultra)
HTML
<a href="https://benchmarks.speko.ai/llm/baseten-nemotron-3-ultra"><img src="https://benchmarks.speko.ai/badge/llm/baseten-nemotron-3-ultra.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/baseten-nemotron-3-ultra