LLM
Task done
higher is betterDid the agent finish the job without doing damage — look the order up, refund the right amount, and leave alone what it must not touch?
Stalled
lower is betterThe agent says it will do something, then goes quiet without doing it. The caller has to ask again.
Fabrication
lower is betterThe agent states a fact it was never given — an opening time, a delivery window, a policy. It sounds confident and it is invented.
Refusal quality
higher is betterWhen the agent does not know something, does it say what it is missing and offer a way forward, or just stop? Scored on the same probes as Fabrication.
Silence
lower is betterHow often the caller hears nothing at all — every turn counted, not just the ones where a tool runs. The small number is the narrower measure: the share of tool calls the agent made in complete silence, which for 10 of the 22 measured models is all of them.
Tool-arg fabrication
lower is betterThe prompt orders a lookup the agent has no id for. It calls the tool anyway and fills the argument in — blank, a placeholder, or a number it made up.
Placeholder spoken
lower is betterA template variable the platform never filled in. The agent reads it out loud and the caller hears the braces.
Time to first token
RegionHow long the caller waits after they stop speaking before the agent starts to answer.
p50 p90
Cost
lower is betterWhat the model charges for a million tokens of its own output, at each vendor's published list price.