← Blog

We Ran the Blind Test. 'Most Natural' Is an Eight-Way Tie.

Every TTS vendor leads with a naturalness score. We ran a blind native-speaker test on fifteen English voices, and the eight best finished in a statistical tie with overlapping confidence intervals. The differences that survive are style, not quality. And the score is measured on 30-second clips, while every axis that decides a production call is one it never sees.


We ran a blind, native-speaker listening test on fifteen English text-to-speech systems: every vote a forced choice between two clips, no brand labels shown. The eight best voices finished in a statistical tie. Gemini and ElevenLabs share the top at an arena Elo of 1591 and 1590; eighth place sits at 1544, and its confidence interval overlaps the leader’s. One of those eight is being sold to you right now, on a landing page, as the most natural voice in the world. Our panel could not pick it out of the other seven.

Naturalness is the number every voice vendor puts first. It is also the number that tells you least about your production call. Two reasons, and this post is about both. Among the models worth considering, naturalness has saturated into a tie. And the score is measured in the one setting production never happens in: a short, clean, single clip of the vendor’s showcase voice.

The tie is the finding, not a footnote

Our eval does not ask people for a 1-to-5 rating. It shows two unlabeled clips of the same sentence and asks which sounds better. That is an arena, so we score it like one: the blind pairwise votes are fit with a preference model and read out as Elo, where the field averages 1500 by construction and 400 points is ten-to-one odds of winning a matchup. We used to project the same fit onto the familiar 1-to-5 MOS scale. It read more easily and it named the measurement after an instrument we never used — MOS is what you get from a panel rating clips on a card, and nobody in this study ever rated anything. Same votes, same ordering, honest units. Here is the English board, with the interval on every score:

leader’s 95% CI1300140015001600Arena Elo — field mean 1500Gemini 3.1 Flash TTS1591ElevenLabs eleven_v3_conversational1590Deepgram aura-21584Cartesia sonic-3.51574Speechify simba-3.21573Soniox tts-rt-v11569Inworld inworld-tts-21561Smallest lightning_v3.11544xAI Grok grok-tts1495Gradium default1448MiniMax speech-2.8-hd1431Rime arcanav31429OpenAI gpt-4o-mini-tts1424Hume octave-21377Qwen qwen3-tts-flash1310
Blind A/B votes from native speakers, brands hidden. The eight best intervals all reach into the leader’s. The ninth does not.

Eight systems from 1544 to 1591, every confidence interval overlapping its neighbors. Then the board does something it never does inside that group: it breaks. xAI lands at 1495, a clean step down, and the floor is Qwen at 1310. The test is not blind noise. It separates good from bad decisively. It just cannot separate good from good, which is the only comparison a “most natural” claim ever makes.

The scores compress because the field is crowded. In a fifteen-way race, even the best voice wins only about 62% of its blind matchups against the pack. There is no runaway. The honest one-line summary of the English naturalness board: eight voices are excellent, six are worse, and asking which of the eight is the most natural asks a question the data refuses to answer.

What survives the tie is taste, not quality

You might guess a tie means the voices converged, all sanded down to one safe timbre. They did not. We measured pitch and expressiveness across the tied leaders. Mean pitch spans 139 to 204 Hz, one voice low and calm, another bright and forward. Expressiveness, the range the pitch travels, spans 4.2 to 6.1 semitones, one animated and one flat. Between-provider variation dwarfs the variation inside any single provider. These are eight distinct voices, and listeners liked them the same.

That kills the axis the marketing leans on. “More expressive, more human” describes a knob, not a rank. The proof is at the bottom of the board: Qwen, dead last at 1310, is as expressive as Cartesia near the top. Expressiveness is a style dial that runs orthogonal to quality. Once a voice clears the bar where it stops sounding like a 2018 car-navigation unit, “most natural” is a preference between roasts of coffee, and there is no objectively best roast.

This only holds where the field is strong. In Spanish, where some providers are plainly bad, a naturalness ranking carries real signal: an acoustic predictor we built tracks the human preference at +0.62 correlation. Run the same predictor on the saturated English top and it hits −0.29, nothing left to learn. A naturalness ranking means something exactly when there are bad models to expose, and nothing among the good ones. Vendors quote it in the second case.

We tried to measure it seven ways. It does not reduce.

None of this is for lack of trying. We spent weeks attempting to compute naturalness automatically, to skip the cost of human panels, and every method failed in an instructive way. A learned MOS predictor (UTMOS) turned out English-biased, rating native Filipino audio lowest. Distance-to-native in a self-supervised embedding space crowned whichever model sat farthest from real recordings, meaning the cleanest synthetic one, a recording-channel artifact rather than a quality signal. Waveform kurtosis measured “real microphone versus polished synthesis” and punished the most polished voice for being polished. Optimal-transport distances, Fréchet audio distance, and a distributional SSL metric each anti-correlated with the human, or collapsed, or would not even install.

They failed for one reason. Every method computes a distance in some feature space, and roboticness is not a distance. A voice a native speaker flags as robotic sits acoustically inside the cloud of natural ones; its distribution overlaps theirs. The listener hears the flaw as a gestalt (vocoder phase, micro-timing, articulation) that no single feature isolates. Naturalness is the one axis that needs a human ear.

Put the two findings together and the metric takes a strange shape. Naturalness is the hardest quality to measure honestly, and once you finally spend on the human ear that can measure it, the models you are choosing between tie. It is at once the least automatable number and the least discriminating one. An odd thing to set in 72-point type on a homepage.

And it is blind to everything production does

The naturalness score comes from a clip of roughly thirty seconds: clean audio, one utterance, the vendor’s best-behaved voice, in English. Your call is none of those things. Every variable the demo strips out is a place where a “natural” model breaks, and we have measured each one separately.

  • Length. The arena’s #2 English model logged 353 severe-failure windows in a single ten-minute take, while #1 and #3 logged zero. The 30-second naturalness score said nothing about minute eight. (the ten-minute test)
  • The specific voice. “Cartesia” ranks #1 as a brand, but its default voice drifts into a different speaker inside a minute, and only 5 of 50 sampled voices held their identity across 90 seconds. The arena ranks the brand; you ship one voice out of hundreds. (the drift probe)
  • Your control tags. Speechify sits at the top of our clarity board, and a single prosody pitch="+15%" tag halves its speaker similarity on a one-second line. Author the audio expressively, the way you will, and the top-of-board voice becomes someone else. (the pitch tag)

None of these surface in a naturalness number, because each lives one level below, or one condition outside, the thing being scored.

Use it as a floor, not a tiebreaker

Do not throw the number away. It is good for one job: catching the bad voices, the 1310s and 1495s a blind test separates cleanly from the pack. As a floor, naturalness works.

As a tiebreaker, it is theater. Above the floor, stop ranking on it and rank on the things that do not tie and do end calls: whether it reads your entities correctly, whether it speaks your language natively, whether it holds its voice across your real session length, whether it survives the control tags you will write. Among whatever clears the floor, pick the voice you like. That last step is taste, and taste is a fine way to choose. The only problem is being sold taste as a measurement.

We publish the full board and every failure clip at benchmarks.speko.ai. The one number we will not print at the top is a single naturalness rank, because we ran the test, and the top of that ranking is a tie.