← LLM

Baseten

Nemotron-3-Ultra

non-thinking at default

Measured 2026-08-18

* Conditional

Fabricates on 27% of out-of-policy questions and stalls on 54.4% of runs.

Score

76

Task done

72.54%

60–86% 95% CI bootstrap over items · n=95 · measured 2026-08-25

Refusal quality

75.6%

names the gap 83% · offers a route 60% · 10 probes x 3 iterations · 1 silent turn scored zero

Fabrication

27%

8 of 30 probe runs

Dead-air

56.0%

139 of 248 turns

Tool silence

93.9%

139 of 148 tool-call turns

Stalled

54.4%

49 of 90 runs

TTFT p50

300ms

300–370 p50–p90

Cost / 1M tok

$2.40

Share this result

Embed the live badge

Speko llm rank
Markdown
[![Speko llm rank](https://benchmarks.speko.ai/badge/llm/baseten-nemotron-3-ultra.svg)](https://benchmarks.speko.ai/llm/baseten-nemotron-3-ultra)
HTML
<a href="https://benchmarks.speko.ai/llm/baseten-nemotron-3-ultra"><img src="https://benchmarks.speko.ai/badge/llm/baseten-nemotron-3-ultra.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/baseten-nemotron-3-ultra