← Blog

Bland's New Voice Is Among the Most Natural English We Have Measured

Bland shipped a new text-to-speech model this week. We scored its best voice on the same eight sentences and with the same blind-vote-trained scorer we use for every system on our English board. It joins the group at the top — a group our own human panel cannot separate. Naturalness is one axis, though, and it is not the one that decides whether a voice survives a real call.


Bland released a new text-to-speech model this week and called it the most human speech engine available. We scored its best voice against our English board, on the same eight sentences behind every system there, with the scorer we fit on our own blind human votes.

It joins the group at the top.

Hear it before you read the number

Three systems, same sentence, labels hidden. One is Bland.

A

B

C

Reveal which is which

A is ElevenLabs. B is Bland. C is Gemini.

This is the phone-number line, one of two sentences where Bland beats every system that previously led our board. We picked a sentence it wins, and we are telling you so.

Where it lands

-0.5-0.25+0.25+0.5no difference0← the other system sounds betterBland sounds better →difference in naturalness, Bland minus the other systemvs Cartesia[-0.25, +0.04]vs Deepgram[-0.24, +0.26]vs ElevenLabs[-0.40, +0.20]vs Gemini[-0.41, +0.13]
Every interval crosses zero. On eight sentences that means indistinguishable, not identical.

Same eight sentences for every system, so the differences are paired — a far stronger test than comparing two averages. Against all four systems that previously led the board, the difference is within noise: the largest is 0.14 on a five-point scale and its interval runs from −0.41 to +0.13. Not one of the four comes close to significance.

That is the result. Bland’s voice belongs in the group at the top, and our data does not order the systems inside it. Our blind human panel reaches the same verdict about that group on its own: an eight-way statistical tie with confidence intervals that all overlap.

Sentence by sentence, Bland beats every previous leader on two: the phone-number line, and a refund line carrying a dollar amount and a date. Both are entity-heavy, which is where the whole field scores lowest and where a real call lives. It loses the expressive sentences, the disfluency and the emotional close, where the incumbents are still better. Winning the entity sentences and losing the expressive ones is the more useful way round. A launch clip is expressive prose. A collections call is dollar amounts and dates.

One caveat that applies to the number and not to the vendor: it belongs to a specific voice, not to a catalog. The voices inside one model vary more than the systems inside this group do, so pin the voice you tested and re-test if you change it.

Naturalness does not decide production readiness

Naturalness is measured on short clean clips. Every axis that ends a real call is a different one.

Latency. A caller waits through time-to-first-audio, not synthesis time. It streams, and once it starts it renders faster than real time, so the whole cost sits in the first frame — and measured the way we measure every other system, that first frame puts it at the back of the field. Voice choice moves it the opposite way from naturalness: of the voices we timed, the one we scored is the slowest.

Robustness. A thirty-second clip says nothing about minute eight, about what happens when you send your own control tags, or about whether the voice holds its identity across a long call.

Use naturalness as a floor, then choose on the axes that end calls.

Our English board, the human panel behind the scorer, and every clip above are at benchmarks.speko.ai.