← LLM 
Together
Llama-3.3-70B
Measured 2026-08-25
* Conditional
Fails restraint — fires a tool with a made-up ID instead of asking for the missing one.
Score
63
Task done
64.21%
43–84% 95% CI bootstrap over items · n=95 · measured 2026-08-25
Refusal quality
20.0%
names the gap 33% · offers a route 13% · 10 probes x 3 iterations · 19 silent turns scored zero
Fabrication
0%
0 of 10 probes (30 probe runs) · ≤31%
Dead-air
74.3%
260 of 350 turns
Tool silence
100.0%
all of 19 items (260 tool-call turns) · ≥82%
Stalled
26.7%
24 of 90 runs
TTFT p50
698ms
698–860 p50–p90
Cost / 1M tok
$1.04
Share this result
Embed the live badge
Markdown
[](https://benchmarks.speko.ai/llm/together-llama-3-3-70b)
HTML
<a href="https://benchmarks.speko.ai/llm/together-llama-3-3-70b"><img src="https://benchmarks.speko.ai/badge/llm/together-llama-3-3-70b.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/together-llama-3-3-70b