← Blog

Eleven v3 Can Take Calls Now

Expression costing you latency was never a law of the field, but it was real inside ElevenLabs' own lineup: v3 carried the audio tags, and their docs pointed realtime builders at Flash instead. The conversational build of v3 landed on 19 August and we measured it the same day, on the same rig and vantage as our published cells: 168ms off the median first audio byte, a 39% cut, at the same price per character and the same 74 languages.


Expression and speed pull against each other inside a model. Builds that laugh, sigh and shift register tend to be larger and decode at higher fidelity, and that shows up as a slower first byte. It is a real tension, not a law of the field: several systems on our own boards hold their own on expressive lines while speaking in a fraction of the time.

Eleven v3 was that particular model. ElevenLabs has been candid about it: v3 is where the audio tags live, and their docs pointed realtime builders at Flash instead. A stack that wanted v3’s tag control paid for them in latency, every turn.

ElevenLabs shipped the conversational build of v3 today. We measured it the same day.

It gave back 168ms, which is 39% of the wait

Median time to first audio byte, two sweeps per model100200300400500median ms to first audio bytev3 Conversational266v3434
Two sweeps per model, thirty calls each.

v3 Conversational reports 266ms to first audio against 434ms for v3, measured in the same sweeps on the same connection.

Same price per character, same 74 languages, same audio tags. Credit to them: that speed-up is not paid for with anything.

It is available on Speko now, alongside v3.

How it was measured

Both models went through the same rig our published latency cells go through, from the same place, on the same voice: one keep-alive connection held open for the whole sweep, concurrency of exactly one, a warmup per model thrown away, n=30, and a reading kept only when the pin was honoured with no failover. Every sweep returned 30 of 30.