← Blog

Our Tamil Column Read 33% Error. The Transcripts Were Fine.

Adding Tamil to our speech-to-text board produced word error rates that overstated the severity we heard in the transcripts. Character error rate better reflects these long-word samples.


We added Hindi, Tamil and Telugu to our speech-to-text board this month. The Tamil systems came back at 24% to 79% word error rate. For the stronger systems, many transcripts were recognisable sentences with local errors even when WER looked severe. The worst system genuinely failed; changing the metric changes the interpretation, not the transcript.

The metric was not broken. It was the wrong instrument.

A wrong vowel costs the whole word

The full 50-sentence source samples contain 1,264 reference words in Hindi and 835 in Tamil. In those samples, Tamil words average 7.86 characters against Hindi’s 4.08. Tamil attaches case, tense and politeness to the stem, so a local error can spoil a long word.

Word error rate counts insertions, deletions and substitutions over reference words. A one-character change inside a long Tamil word costs a full substitution; omitting that word costs one deletion.

The size of the mistake

We rescored the transcripts we already had, by word and by character.

0%10%20%30%40%Google Chirp 37.4%24.3%OpenAI GPT-4o Transcribe10.1%33.2%Smallest AI Pulse Pro12.1%26.4%
scored by characterscored by wordChirp 3 n=48 | GPT-4o Transcribe n=49 | Pulse Pro n=48
Each line re-scores one system's saved outputs. Chirp 3 and Pulse Pro each have 48 clips; GPT-4o has 49 after one invalid FLEURS audio/reference pair was excluded from every system. Chirp 3 ran directly against Google Speech V2 in the EU, not through the Speko gateway.

Each pair describes the same saved outputs for one system. Word scoring charges the best measured Tamil system roughly one edit per four reference words; character scoring counts about seven edits per 100 reference characters. The second reading better describes what we heard.

Spanish, German, French and Norwegian stay on word error rate because it remains the more interpretable external standard for those samples. Script alone does not choose the metric.

What we do now

Tamil and Telugu are scored by character because word scoring materially overstates the practical severity on these samples. Hindi stays on word error rate because its shorter words make WER more interpretable and WER is the unit most commonly reported in published ASR work. The decision follows observed distortion and external conventions, not script or a hard word-length threshold.

One caveat we’re carrying forward: character error rate is the right number to publish on these languages and the wrong one for judging an agent. A single wrong character in a name or a dollar amount can break the booking, while character scoring may charge only a small fraction for it.