← Blog

GPT-Live-1 Is Out. We Have Been Benchmarking It Since the Alpha.

OpenAI shipped GPT-Live-1 to the API on September 10. We benchmarked it through the alpha with the OpenAI team. The design splits a voice agent into a talker and a thinker, and that changes how a benchmark of it has to be read. On our S2S board it takes the highest overall score of the ten rows.


OpenAI shipped GPT-Live-1 to the API on September 10. We had been running it before that. Through the alpha we ran it through our harness and sent our findings to the OpenAI team as we went.

How a GPT-Live-1 call splits between a talker and a thinker The caller and the talker exchange speech continuously in both directions at the same time. When something needs real work, the talker hands it to the thinker, keeps the conversation going while the thinker works, and speaks the answer when it comes back. CALLER YOU TALK IT TALKS The talker hears and speaks at once THE HARD PART THE ANSWER The thinker looks it up, works it out
The talker never goes quiet while the thinker works.

One model, two jobs

A speech-to-speech model has had to be two things at once: a fast, natural talker and a careful reasoner. The two jobs pull in opposite directions. Reasoning depth costs latency. A tool call leaves the caller waiting while the result comes back. Audio-native reasoning trailed the text models for most of 2024 and 2025. When it caught up, it caught up inside one model that still had to answer in real time.

GPT-Live-1 splits the jobs. The voice model is full-duplex: it listens while it speaks, so pauses, interruptions, backchannels, and background speech are handled inside the model while it talks. It owns the floor of the conversation. When the caller asks for something that needs real thinking or a tool, the voice model delegates that to a backend of your choice. It keeps the conversation going while the backend works and speaks the answer when it lands. The backend can be an OpenAI reasoning model or your own agent stack.

Voice layer Backend
Owns turn-taking, interruptions, backchannels, tone reasoning, tool calls, lookups
Runs continuously, full-duplex on demand, when the voice layer asks
Billed per minute of audio per token
Swappable no yes
What the S2S board sees Latency Action, Task capability

Call completion sits on the seam between the two columns: a call completes only if the handoff works in both directions.

What the split changes for a builder

Three things, in the order they will show up on the board.

Reasoning becomes a choice. In May we wrote that intelligence was no longer the case against speech-to-speech. Component swap and observability were: one vendor for the whole pipeline, and no text to inspect between the ear and the mouth. GPT-Live-1 answers half of that. The reasoning layer is now a model you pick and can replace, and every tool call and result crosses the seam as text you can log. The voice layer itself is still a black box. The seam is not.

Price becomes two lines. The voice layer is billed at $0.05 per minute of audio. The backend is billed per token, and the total depends on how much thinking you route to it. Our May post showed speech-to-speech cost growing faster than call length, because the audio context re-accumulated every turn. Per-minute billing has no such curve: the voice half costs the same in minute twelve as in minute one. Whether the backend half brings the curve back depends on the backend, so any cost comparison for this model needs two numbers, not one.

Latency becomes two numbers. How fast the model takes its turn is one measurement. How fast an answer arrives once the backend is involved is another, and it is the one a caller feels when they ask for a real lookup. A single latency figure hides one of them.

What the board shows

The S2S board now carries a GPT-Live-1 row.

Measure Result
Overall 0.86, first of ten rows
Call completion 100 percent, 18 of 18 calls finished
Task capability 0.90
Weakest of the six scenarios 0.66, the highest floor on the matrix

The floor is the number worth reading. A booking, a travel support call, a restaurant reschedule, a currency question, a refused refund, a reschedule under several constraints: the call type it handles worst still scores 0.66, and no other row on the scenario matrix has a worst case that high. Peaking on one scenario is easy. Holding across six is what survives a real queue.

One configuration note. The delegated backend is told the current date, which is what our production stack injects on every call. The other rows were measured without it, so the action numbers are not like-for-like.