Speech-to-Text
Emotion
Whether a listener hears how a thing was said. The words in every clip carry nothing, so only the delivery can score.
Transfer: human speech → synthetic voices
oruk-resonance91% → 48%
emotion2vec+ large88% → 16%
qwen3-asr-flash65% → 14%
MERaLiON-SER-v161% → 33%
gpt-audio-1.550% → 14%
gemini-3.8-flash50% → 30%
oruk-fourier47% → 33%
SenseVoiceSmall10% → 15%
humansynthetic
Human
| System | Emotion | Neutrals |
|---|---|---|
| oruk-resonancededicated SER | 91% | 47/48 |
| emotion2vec+ largeopen model | 88% | 46/48 |
| qwen3-asr-flashaudio LLM | 65% | 47/48 |
| MERaLiON-SER-v1open model | 61% | 36/48 |
| gpt-audio-1.5audio LLM | 50% | 33/48 |
| gemini-3.8-flashaudio LLM | 50% | 10/48 |
| oruk-fourierdedicated SER | 47% | 3/48 |
| SenseVoiceSmallopen model | 10% | 41/48 |
Synthetic
| System | Emotion | Neutrals |
|---|---|---|
| oruk-resonancededicated SER | 48% | 41/90 |
| MERaLiON-SER-v1open model | 33% | 43/90 |
| oruk-fourierdedicated SER | 33% | 18/90 |
| gemini-3.8-flashaudio LLM | 30% | 31/90 |
| emotion2vec+ largeopen model | 16% | 80/90 |
| SenseVoiceSmallopen model | 15% | 62/90 |
| qwen3-asr-flashaudio LLM | 14% | 89/90 |
| gpt-audio-1.5audio LLM | 14% | 67/90 |
Top-1, one family mapping across systems; hover a row for mechanism and per-class recall. English only, preliminary.