Speech-to-Text

Emotion

Whether a listener hears how a thing was said. The words in every clip carry nothing, so only the delivery can score.

Transfer: human speech → synthetic voices

oruk-resonance91% → 48%
emotion2vec+ large88% → 16%
qwen3-asr-flash65% → 14%
MERaLiON-SER-v161% → 33%
gpt-audio-1.550% → 14%
gemini-3.8-flash50% → 30%
oruk-fourier47% → 33%
SenseVoiceSmall10% → 15%
humansynthetic

Human

SystemEmotionNeutrals
Ooruk-resonancededicated SER91%47/48
Femotion2vec+ largeopen model88%46/48
Qwenqwen3-asr-flashaudio LLM65%47/48
MMERaLiON-SER-v1open model61%36/48
gpt-audio-1.5audio LLM50%33/48
Geminigemini-3.8-flashaudio LLM50%10/48
Ooruk-fourierdedicated SER47%3/48
FSenseVoiceSmallopen model10%41/48

Synthetic

SystemEmotionNeutrals
Ooruk-resonancededicated SER48%41/90
MMERaLiON-SER-v1open model33%43/90
Ooruk-fourierdedicated SER33%18/90
Geminigemini-3.8-flashaudio LLM30%31/90
Femotion2vec+ largeopen model16%80/90
FSenseVoiceSmallopen model15%62/90
Qwenqwen3-asr-flashaudio LLM14%89/90
gpt-audio-1.5audio LLM14%67/90

Top-1, one family mapping across systems; hover a row for mechanism and per-class recall. English only, preliminary.