Gemini 3.8 Live Tops Our Board. Its Smarter Sibling Comes Fourth.
Google shipped two live audio models on September 15. One of them takes the highest score on our speech-to-speech board. The other reasons better and takes seven and a half seconds to answer a question. We measured every reasoning level to find out whether that could be tuned away.
Google shipped two live audio models on September 15: Gemini 3.8 Live, and Gemini 3.8 Live Extended Thinking. We ran both through our six-scenario concierge harness the same night.
Google positions the two as a choice: 3.8 Live for low-latency voice agents, Extended Thinking for when you need more background reasoning during the call. On the S2S board the plain model takes the highest overall score and the thinking model comes fourth. The gap between those two sentences is the interesting part.
What we measured
Six multi-turn concierge calls, three sessions each, eighteen sessions per model.
| Action | Call completion | Task capability | Answer | |
|---|---|---|---|---|
| Gemini 3.8 Live | 0.83 | 98% | 0.92 | 2.0s |
| GPT-Live-1 | 0.78 | 100% | 0.90 | 4.2s |
| Gemini 3.8 Live Extended Thinking | 0.83 | 96% | 0.88 | 7.5s |
| GPT-Realtime-2.1 | 0.67 | 100% | 0.79 | 2.5s |
Every model here is measured in the configuration our production stack would run it in. One note before anyone quotes the table: “Answer” is not first sound. It is when the caller hears the actual answer to what they asked, which is a different number from when the model starts making noise.
Reasoning is a required parameter, and less of it is better
Extended Thinking requires a reasoning level and ships no default, so we measured all three the API accepts.
| Reasoning level | Overall | Task capability | Answer | Cost per call |
|---|---|---|---|---|
| LOW | 0.83 | 0.88 | 7.5s | $2.17 |
| MEDIUM | 0.79 | 0.82 | 7.3s | $2.74 |
| HIGH | 0.81 | 0.83 | 7.8s | $4.55 |
LOW wins, and costs half of HIGH. More reasoning makes the model a worse agent on this task: at HIGH it loses on brevity and on knowing when to ask rather than act, while gaining on recall and error recovery. The reasoning is real. It is spent where a concierge call does not pay for it.
The number that settles the question is the one that does not move. Answer latency is flat across the entire sweep — 7.5, 7.3, 7.8 seconds. Turning reasoning down does not buy the time back, because the wait is the delegated reasoning loop itself. There is no configuration of this model that answers a phone call quickly.
What we would ship
Gemini 3.8 Live, for a phone agent. Highest capability on the board, an answer in two seconds, and a dollar a call.
Extended Thinking is the stronger reasoner and we would not put it on a call. If your product can wait seven seconds for an answer, it is probably not a phone call.