← LLM 
Alibaba
qwen-turbo
Measured 2026-08-25
* Conditional
Promises to act and then goes quiet on 94.4% of runs; fabricates on 87% of out-of-policy questions.
Score
56
Task done
6.58%
0–18% 95% CI bootstrap over items · n=95 · measured 2026-08-25
Refusal quality
11.1%
names the gap 0% · offers a route 10% · 10 probes x 3 iterations · 3 silent turns scored zero
Fabrication
87%
26 of 30 probe runs
Dead-air
9.1%
10 of 110 turns
Tool silence
100.0%
all of 2 items (10 tool-call turns) · ≥16%
only 10 tool-call turns — the model rarely calls a tool
Stalled
94.4%
85 of 90 runs
TTFT p50
483ms
483–497 p50–p90
Cost / 1M tok
$0.20
Share this result
Embed the live badge
Markdown
[](https://benchmarks.speko.ai/llm/alibaba-qwen-turbo)
HTML
<a href="https://benchmarks.speko.ai/llm/alibaba-qwen-turbo"><img src="https://benchmarks.speko.ai/badge/llm/alibaba-qwen-turbo.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/alibaba-qwen-turbo