← LLM

Alibaba

qwen-turbo

Measured 2026-08-25

* Conditional

Promises to act and then goes quiet on 94.4% of runs; fabricates on 87% of out-of-policy questions.

Score

56

Task done

6.58%

0–18% 95% CI bootstrap over items · n=95 · measured 2026-08-25

Refusal quality

11.1%

names the gap 0% · offers a route 10% · 10 probes x 3 iterations · 3 silent turns scored zero

Fabrication

87%

26 of 30 probe runs

Dead-air

9.1%

10 of 110 turns

Tool silence

100.0%

all of 2 items (10 tool-call turns) · ≥16%

only 10 tool-call turns — the model rarely calls a tool

Stalled

94.4%

85 of 90 runs

TTFT p50

483ms

483–497 p50–p90

Cost / 1M tok

$0.20

Share this result

Embed the live badge

Speko llm rank
Markdown
[![Speko llm rank](https://benchmarks.speko.ai/badge/llm/alibaba-qwen-turbo.svg)](https://benchmarks.speko.ai/llm/alibaba-qwen-turbo)
HTML
<a href="https://benchmarks.speko.ai/llm/alibaba-qwen-turbo"><img src="https://benchmarks.speko.ai/badge/llm/alibaba-qwen-turbo.svg" alt="Speko llm rank"></a>
URL
https://benchmarks.speko.ai/llm/alibaba-qwen-turbo