← LLM

Gemini

gemini-3.8-flash

* Conditional

Measured 2026-09-03; forbids-and-duplicates a state-changing cancel on two scenarios, 5/5 iterations each. Fabricates nothing (0/30).

Score

No real-time first-token measurement, so a speed-weighted score is not computed — unranked rather than judged on partial data.

Task done

77.37%

61–92% 95% CI bootstrap over items · n=95 · measured 2026-09-03

Refusal quality

96.7%

names the gap 97% · offers a route 97% · 10 probes x 3 iterations

Fabrication

0%

0 of 10 probes (30 probe runs) · ≤31%

Dead-air

69.0%

214 of 310 turns

Tool silence

100.0%

all of 19 items (214 tool-call turns) · ≥82%

Stalled

28.9%

26 of 90 runs

TTFT p50

Bimodal: about half the calls answer near 600ms and the rest take 1-5s, so the median lands in the gap and shifts between sweeps (668ms, then 1168ms minutes later). Unresolved on both transports and in both regions.

Cost / 1M tok

$3.75

Vendor list, paid tier, per 1M output tokens; rises to $7.50 on 2027-01-01.

Share this result

Embed the live badge

Speko llm rank
Markdown
[![Speko llm rank](https://benchmarks.speko.ai/badge/llm/gemini-gemini-3-8-flash.svg)](https://benchmarks.speko.ai/llm/gemini-gemini-3-8-flash)
HTML
<a href="https://benchmarks.speko.ai/llm/gemini-gemini-3-8-flash"><img src="https://benchmarks.speko.ai/badge/llm/gemini-gemini-3-8-flash.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/gemini-gemini-3-8-flash