← Blog

The .99 Problem: The Cheapest Streaming TTS Drops Your Cents

Qwen3 TTS Flash is the cheapest streaming voice in the gateway at $0.0014/min, 3x under the next option. Then we asked it to read an invoice total. Across eight takes it never once read '$2,349.99' correctly; Grok got it right every time. Here is what stress-testing the two cheapest TTS providers turned up.


Listen to the cheapest streaming voice in our gateway read an invoice total. The line is “The invoice total came to $2,349.99 plus a €15 shipping fee.”

It says two thousand three hundred forty-nine, then something mushy, then “plus.” The .99 is gone. Whisper hears it as “…came to $2,349 and plus a 15 euro fee.” The cents weren’t mistranscribed. They were never clearly spoken.

That’s Alibaba Qwen3 TTS Flash, the cheapest streaming TTS in the gateway: $0.0014 per minute, about a third of the next-cheapest option and ~30x under ElevenLabs v3. When someone asks for the cheapest voice, this is the answer. This post is the sentence that goes next to it.

The two cheapest streaming voices

We stress-tested the two cheapest streaming-capable providers on English: Alibaba’s qwen3-tts-flash at $0.0014/min and xAI’s grok-tts at $0.0045/min, both over WebSocket. On clean prose both are excellent, round-trip error near zero. The gap opens once the text stops being prose.

The .99 problem, and it reproduces

One bad take is an anecdote. So we ran the same invoice line eight times through Qwen and four through Grok, and transcribed each with faster-whisper large-v3.

Qwen qwen3-tts-flash
0/8
xAI grok-tts
4/4
Filled = pass · red = fail
Reading "$2,349.99", eight takes for Qwen and four for Grok. Qwen missed the cents every time; Grok nailed it every time. Hover a Qwen dot to see what faster-whisper heard back.

Grok read “$2,349.99” right all four times. Qwen never did in eight. The decimal came back as “.999” five times, “.9” once, and a clean “.99” just once (and that take put a euro sign on it). Its duration for the identical line swung from 7.3 to 9.0 seconds. Three Qwen takes, the number landing differently each time:

And Grok, dead stable take after take:

Zoom: where the cents go

Word-level timestamps show the failure. Qwen emits no clean .99 token. In its place sits a ~0.7-second smear the transcriber best-guesses as “and” at 57% confidence, bleeding into a clipped “plus.” Grok emits a crisp .99 at 99%.

Word-level Whisper timestamps for the invoice line. Qwen's row has no ".99" token — instead a low-confidence "and" (probability 0.57) sits where the cents should be, running into "plus". Grok's row has a clean ".99" token at 99% confidence followed by "plus".

Qwen first, Grok second: the isolated cents-into-”plus” slice from each.

It’s audible, not just a transcription quirk

Maybe Whisper is the unreliable narrator here. So we looked at the audio itself. Against Qwen’s own clean sentence, same voice and model, the pitch is identical to within 0.2%. This is not a voice or language switch. What changes is timbre:

Pitch (F0)+0.2% — same voice
clean
261.6 Hz
currency
262.2 Hz
Brightness (spectral centroid)+23% brighter
clean
2196 Hz
currency
2703 Hz
High-freq energy (85% rolloff)+31% brighter
clean
3571 Hz
currency
4664 Hz
Same voice, two clips. Pitch is identical (it never switched voices); what changes is timbre — the currency clip carries 23–31% more high-frequency energy, smeared over the digits.

Spectrograms of Qwen's clean sentence versus its currency clip. The clean clip has crisp harmonic stacks; the currency clip has a diffuse high-frequency haze layered over the number spans.

The clean clip has crisp harmonic stacks. The currency clip has a grainy high-frequency haze over the numbers, and it lines up exactly with where the words fall apart.

The other caveat: it folds under load

Price assumes the call gets through. We fired 20 synthesis requests in parallel, one short prompt:

Qwen qwen3-tts-flash
5/20
xAI grok-tts
20/20
Filled = pass · red = fail
Twenty synthesis calls fired at once, one short prompt. Grok absorbed all twenty; Qwen dropped fifteen to HTTP 429.

Qwen rate-limits hard; Grok absorbed all twenty in under two seconds. For a batch job at low QPS, Qwen’s 3x saving is real and worth taking. In front of a live agent at scale, it will drop calls unless you put your own queue in front of it.

The verdict

Cheapest is a real answer, not a free one.

  • Qwen qwen3-tts-flash is the cheapest streaming voice in the gateway and genuinely good on clean English. It mangles decimals and currency, and folds under concurrency. Reach for it on offline or low-QPS work where the 3x saving beats money-reading and burst reliability.
  • xAI grok-tts costs 3x more and is the boring, safe pick: clean numbers, stable under load. Caveat: English-only.

So the honest answer to “what’s the cheapest TTS” isn’t a name. It’s a name plus a sentence. For Qwen: don’t hand it a price tag or a rush.