S2S
| # | Model | Overall | Action | Call completion | Task capability | Latency |
|---|---|---|---|---|---|---|
| 1 | gemini-3.1-flash-liveGemini | 0.77 | 0.78 | 92% | 0.86 | 1108ms |
| 2 | gpt-realtimeOpenAI | 0.76 | 0.83 | 88% | 0.87 | 494ms |
| 3 | gpt-realtime-2.1-miniOpenAI | 0.75 | 0.83 | 93% | 0.79 | 1011ms |
| 4 | grok-voice-fastxAI⚠ | 0.72 | 0.44 | 95% | 0.91 | 1111ms |
| 5 | gpt-realtime-2OpenAI⚠ | 0.68 | 0.87 | 81% | 0.84 | 1103ms |
| 6 | gpt-realtime-2.1OpenAI | 0.63 | 0.56 | 90% | 0.76 | 1104ms |
| 7 | gpt-realtime-miniOpenAI⚠ | 0.62 | 0.70 | 74% | 0.91 | 614ms |
| 8 | grok-voice-think-fastxAI⚠ | 0.56 | 0.22 | 96% | 0.77 | 1007ms |
| 9 | gpt-4.1-nano (cascade)Inworld⚠ | 0.45 | 1.00 | 51% | 0.86 | — |
What we tested
Six multi-turn concierge calls. Each forces the model to infer a fact, fire a real tool, then handle a change — n=3 runs each · clean audio · latency from us-east4 · 2026-07-19.
| Scenario | The story it forces | Tool |
|---|---|---|
| default | book a flight to your 3pm-meeting city, then flip the day | book_flight |
| travel_support | infer the wedding city, book there, flip Thu→Sat, spell an email | book_flight |
| restaurant_calendar | infer the anniversary night, make the reservation, move 7→8pm | create_event |
| finance_currency | infer trip budget + currency, convert it, then bump the budget | convert_currency |
| refund_negative | negative-heavy support; infer the one free pickup day, schedule it, spell an order code | create_event |
| multiconstraint_schedule | infer the summit city, book a flight, then change destination Berlin→Munich | book_flight |
Per-scenario scores · 0–1 · higher is better · “—” = dropped every session
| Model | Default | Travel | Restaurant | Finance | Refund | Multi-con. |
|---|---|---|---|---|---|---|
| gemini-3.1-flash-live | 0.92 | 0.99 | 0.74 | 0.84 | 0.87 | 0.32 |
| gpt-realtime | 0.96 | 0.95 | 0.78 | 0.75 | 0.43 | 0.67 |
| gpt-realtime-2.1-mini | 0.97 | 0.88 | 0.60 | 0.83 | 0.76 | 0.47 |
| grok-voice-fast | 0.66 | 0.54 | 0.86 | 0.99 | 0.73 | 0.56 |
| grok-voice-think-fast | 0.59 | 0.70 | 0.44 | 0.91 | 0.39 | 0.38 |
| gpt-4.1-nano (cascade) | 0.59 | — | 0.49 | 0.74 | — | — |