← Blog

Deepgram Flux TTS Is the Fastest Voice in Our Naturalness Band

Flux TTS returns first audio in 106ms, sits inside our naturalness tie band, reads money and decimals perfectly, and keeps the minus sign that several systems drop. Free to build against until 12 September.


Deepgram shipped Flux TTS today. We wired it through our gateway the same evening and measured it on all four axes our board publishes. It is the strongest new text-to-speech row we have added this year.

106 ms to first audio — second of eighteen systems. Inside the naturalness tie band, on day one. Perfect scores on currency, decimals, large numerals and ordinals. And it keeps the minus sign, which is the mistake that actually costs callers money.

naturalness tie band — these systems overlap on 95% intervals1450150015501600100ms250ms500ms1000msFluxtime to first audio, p50 — lower is better →naturalness Elo →

That corner of the chart is hard to reach. Speed and naturalness usually trade against each other here — the two highest naturalness scores on the board take 481 ms and 978 ms to produce a first sample, four to nine times longer than Flux. Everything faster than Flux sits below the band. Flux is the fastest voice inside it.

Deepgram now holds two of the five fastest models on the board and has both of them inside the naturalness band. No other vendor manages that with even one pair.

Inside the band, where the top is a tie

The top systems on our naturalness column overlap on their 95% intervals. Inside that group there is no ranking to win — the intervals cross, so ordering them reads noise. What matters is whether a voice is in the band. Flux is, on day one.

Here is Flux reading a line from the set, unedited:

Money, decimals and ordinals: perfect

Our robustness axis asks whether a voice says the right words for a number, a date, an amount or an operator. An HTTP 200 cannot see any of this — a system reading “dollar sign one comma two three four” returns a perfectly successful response.

currency10/10large numerals10/10decimals10/10ordinals10/10times9/10dates3/10
Sixty fixtures, spokenForm off so the voice sees raw glyphs. Dates is the one category that separates systems on this axis, and it is where Flux loses points.

Currency, large numerals, decimals and ordinals came back correct on every fixture.

Flux keeps the minus sign. Given -$12.50 it says “negative twelve dollars and fifty cents”, and both independent transcription oracles agree:

Several systems on this board drop that sign entirely. It is the single heaviest penalty our scoring applies, because a lost minus turns a refund into a charge. Flux does not make that class of mistake, and that is the one to care about.

Where it does slip

Two categories cost it points. Both are worth hearing rather than taking on trust.

On numeric date forms, Flux reads the digits without saying the month. 07/15/2026:

Written with the month spelled out, it is correct — “We close on December 24th”:

So this is specific to the numeric form, not to dates. And it is provisional: on the confirmation pass both oracles returned the written form 07/15/2026, which cannot show whether the voice said “July” or “oh seven”. Every published claim on this axis is reproduced three times across independent oracles, and this one is not settled.

The other is a hyphenated range. 08:00-20:00 came back as “eight twenty” from both oracles — the range collapsed into a single time:

Both are recoverable. A caller who hears “seven fifteen twenty twenty-six” can still act on it, and both are fixable upstream by verbalising dates and ranges before they reach the voice. Neither inverts a value.

Price

Flux TTS lists at $45 per million characters pay-as-you-go, $40.50 on the Growth tier — mid-field on a board that runs from $10 to $100.

$10$25$50$75$100Flux TTS · $45list price per million characters across the board — each tick is one system

It is free to build against through 12 September 2026 — 45 concurrent streaming connections globally, 5 in the EU and Australia. Billing starts 13 September. If you are evaluating voices, that window is the cheapest time you will ever get to do it.

What it is actually for

Flux is built for voice agents, and the API is honest about that. Interrupt it and it returns the text the caller actually heard, split from the text they did not, so agent context can be reconciled after a barge-in instead of guessed at. Prosody carries across turns on its own.

It is fast, it is in the band, and it reads money correctly. Until 13 September it is also free.