Open models
| Multi-step tool tasks completed end-to-end (chains, compound, clarify), % of runs.Task done | Invented a policy or fact it was never given, % of out-of-policy questions.Fabrication | Turns where the caller hears silence (no spoken words), % of turns.Dead-air | Time-to-first-token, p50 (us-east4). Whisker spans p50→p90.TTFT p50 | |
|---|---|---|---|---|
05Rank 5 of 12 on the full LLM board — not re-ranked for this page. Cerebras | 94% | 10% | 51% | 195ms195–260 p50–p90 |
01Rank 1 of 12 on the full LLM board — not re-ranked for this page. Cerebras | 89% † | 0% | 12% | 192ms192–232 p50–p90 |
10Rank 10 of 12 on the full LLM board — not re-ranked for this page. Together | 56% | 0% | 29% | 698ms698–860 p50–p90 |
— Together | 100% | 0% | 0% | —No real-time first-token measurement; exceeds the ~800ms voice latency budget. |
| Average Word Error Rate over 7 English short-form test sets. Lower is better. Sourced from the Hugging Face Open ASR Leaderboard, 2026-07-16 — not measured on our rig.WER avg | Seconds of audio transcribed per second of compute (the Open ASR Leaderboard's H200 run). Higher is faster.RTFx | |
|---|---|---|
01 Boson AI | 4.62% | 110× |
02 AutoArk AI | 4.64% | 484× |
03 OpenMOSS (Fudan) | 4.66% | 149× |
04 Boson AI | 4.73% | 139× |
4.90% | 547× | |
4.95% | 2079× | |
07 Alibaba (Qwen) | 4.99% | 796× |
09 NVIDIA | 5.06% | 861× |
5.10% | 661× | |
11 AutoArk AI | 5.14% | 672× |
12 Cohere | 5.20% | 916× |
13 SoundsgoodAI | 5.21% | 160× |
5.29% | 264× | |
15 NVIDIA | 5.39% | 6038× |
| Blind pairwise listening preference (which voice sounds more natural). Higher is better. Sourced from a public speech arena, 2026-07-16 — not measured on our rig.Elo | |
|---|---|
01 Fish Audio | 1116 |
02 StepFun | 1110 |
03 Mistral | 1078 |
04 NVIDIA | 1060 |
05 hexgrad | 1058 |
06 Maya Research | 1048 |
07 Resemble AI | 1012 |
09 Zyphra | 1000 |
10 MyShell | 957 |
12 Coqui | 913 |
13 yl4579 | 889 |
14 MetaVoice | 835 |
| Correct END-vs-WAIT decisions across 200 real clips, at each model’s best operating threshold (min false-cutoff s.t. END-recall ≥ 85%).EOT accuracy | Completed turns correctly ended — the agent didn’t leave the caller hanging.END-recall | WAIT clips wrongly ended = talked over the caller mid-thought. The costly live-call error.False-cutoff | Model forward pass, warm CPU. Text models additionally wait on the STT transcript; audio runs in parallel with STT.Inference | |
|---|---|---|---|---|
01Rank 1 of 5 on the full Turn-taking board — not re-ranked for this page. Smart Turn v3.2 | 94.0% | 89.0% | 1.0%One false cutoff in 100 incomplete clips. | 32msRuns in parallel with STT — no transcript wait.49 p90 |
02Rank 2 of 5 on the full Turn-taking board — not re-ranked for this page. LiveKit Intl | 87.0% | 86.0% | 12.0% | 29msWaits for the STT transcript (~0–125ms) before running.57 p90 |
03Rank 3 of 5 on the full Turn-taking board — not re-ranked for this page. Turnsense | 86.0% | 88.0% | 16.0% | 127msPads every input to 256 tokens regardless of length — slowest as-shipped.193 p90 |
04Rank 4 of 5 on the full Turn-taking board — not re-ranked for this page. LiveKit EN | 82.0% | 87.0% | 23.0% | 5ms24 p90 |