Speech-to-Speech
Action
higher is better0.000.250.500.751.00
Call completion
higher is betterTask capability
higher is better0.000.250.500.751.00
Latency
lower is better0ms500ms1.00s1.50s2.00s
What we tested
| Scenario | The story it forces | Tool |
|---|---|---|
| default | book a flight to your 3pm-meeting city, then flip the day | book_flight |
| travel_support | infer the wedding city, book there, flip Thu→Sat, spell an email | book_flight |
| restaurant_calendar | infer the anniversary night, make the reservation, move 7→8pm | create_event |
| finance_currency | infer trip budget + currency, convert it, then bump the budget | convert_currency |
| refund_negative | negative-heavy support; infer the one free pickup day, schedule it, spell an order code | create_event |
| multiconstraint_schedule | infer the summit city, book a flight, then change destination Berlin→Munich | book_flight |
Per-scenario scores · 0–1 · higher is better · “—” = dropped every session
| Model | Default | Travel | Restaurant | Finance | Refund | Multi-con. |
|---|---|---|---|---|---|---|
| grok-voice-think-fast-2.0 | 0.99 | 0.87 | 0.84 | 0.85 | 0.60 | 0.62 |
| gemini-3.1-flash-live | 0.92 | 0.99 | 0.74 | 0.84 | 0.87 | 0.32 |
| gpt-realtime | 0.96 | 0.95 | 0.78 | 0.75 | 0.43 | 0.67 |
| gpt-realtime-2.1-mini | 0.97 | 0.88 | 0.60 | 0.83 | 0.76 | 0.47 |
| grok-voice-fast | 0.66 | 0.54 | 0.86 | 0.99 | 0.73 | 0.56 |
| grok-voice-think-fast-1.0 | 0.59 | 0.70 | 0.44 | 0.91 | 0.39 | 0.38 |
| gpt-4.1-nano (cascade) | 0.59 | — | 0.49 | 0.74 | — | — |