Speech Emotion Recognition on MELD
63.5AccuracySoundwave
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| Soundwave2025.02 | 63.5 | — | — | — | — | — | |
| Qwen2.5-Omni-7B + CoATBackbone=Qwen2.5-Omni-7B, Method=CoAT2026.06 | 60.8 | — | — | — | — | 59.2 | |
| Audio Flamingo 3 + CoATBackbone=Audio Flamingo 3, Method=CoAT2026.06 | 59.8 | — | — | — | — | 57.2 | |
| EmotionThinkerModel Category=Omni Large Language Models2026.01 | 59.71 | — | — | — | — | — | |
| Kimi-AudioModel Category=General Speech Large Language Models2026.01 | 59.13 | — | — | — | — | — | |
| Qwen2-Audio + CoATBackbone=Qwen2-Audio, Method=CoAT2026.06 | 58 | — | — | — | — | 56.1 | |
| FreezeEmpath2026.04 | 57.5 | — | — | — | — | — | |
| BLSP-EmoModel Category=Emotion Focused Speech Large Language Models2026.01 | 57.3 | — | — | — | — | — | |
| BLSP-Emo2026.04 | 57.3 | — | — | — | — | — | |
| Qwen-Audio-ChatModel Category=General Speech Large Language Models2026.01 | 55.7 | — | — | — | — | — | |
| Qwen2-Audio2025.02 | 55.3 | — | — | — | — | — | |
| Qwen2.5-Omni-7BModel Category=Omni Large Language Models2026.01 | 54.64 | — | — | — | — | — | |
| OSUM-EChatModel Category=Emotion Focused Speech Large Language Models2026.01 | 53.38 | — | — | — | — | — | |
| MiniCPM-OModel Category=Omni Large Language Models2026.01 | 52.78 | — | — | — | — | — | |
| C²SER2026.04 | 51.5 | — | — | — | — | — | |
| Qwen2-Audio-InstructModel Category=General Speech Large Language Models2026.01 | 51.23 | — | — | — | — | — | |
| MERaLiON2Model Category=General Speech Large Language Models2026.01 | 51.1 | — | — | — | — | — | |
| BLSPModel Category=General Speech Large Language Models2026.01 | 50.47 | — | — | — | — | — | |
| Qwen2.5-Omni-7BBackbone=Qwen2.5-Omni-7B, Method=Baseline2026.06 | 49.4 | — | — | — | — | 51.1 | |
| Qwen2-Audio2026.04 | 47.6 | — | — | — | — | — | |
| Audio Flamingo 3Backbone=Audio Flamingo 3, Method=Baseline2026.06 | 40.8 | — | — | — | — | 45.9 | |
| Kimi-Audio2026.04 | 40.4 | — | — | — | — | — | |
| Phi-4-MultimodalModel Category=Omni Large Language Models2026.01 | 39.81 | — | — | — | — | — | |
| MERaLiONModel Category=General Speech Large Language Models2026.01 | 37.98 | — | — | — | — | — | |
| DIVAModel Category=General Speech Large Language Models2026.01 | 35.19 | — | — | — | — | — | |
| SECapModel Category=Emotion Focused Speech Large Language Models2026.01 | 34.2 | — | — | — | — | — | |
| SALMONNModel Category=General Speech Large Language Models2026.01 | 31.32 | — | — | — | — | — | |
| Qwen2-AudioBackbone=Qwen2-Audio, Method=Baseline2026.06 | 25.5 | — | — | — | — | 21.8 | |
| Megrez-3B-OmniModel Category=Omni Large Language Models2026.01 | 21.89 | — | — | — | — | — | |
| GLM-4-VoiceModel Category=General Speech Large Language Models2026.01 | 21.43 | — | — | — | — | — | |
| B-GRPOTraining Protocol=RL-based unsupervised (after 100 epoch supervised pre-training)2026.02 | — | — | — | — | 30.7 | — | |
| BaselineTraining Protocol=Supervised (half labeled, 100 epochs)2026.02 | — | — | — | — | 25.3 | — | |
| CAREParams=160M2026.03 | — | — | 28.8 | 48.1 | — | — | |
| Crab2026.03 | — | 62.57 | 55.65 | — | — | — | |
| data2vecDownstream=Linear2023.12 | — | 45.75 | 24.98 | 43.59 | — | — | |
| data2vecParams=94M2026.03 | — | — | 23.1 | 41.9 | — | — | |
| data2vec 2.0Downstream=Linear2023.12 | — | 48.92 | 26.1 | 45.8 | — | — | |
| DINOTraining Protocol=Unsupervised learning2026.02 | — | — | — | — | 26.6 | — | |
| E2E-TTS-Stage2Training Stage=22026.06 | — | 55.6 | — | — | — | — | |
| Emo2Vec-large2026.06 | — | 57.4 | — | — | — | — | |
| emotion2vecDownstream=Linear2023.12 | — | 51.88 | 28.03 | 48.7 | — | — | |
| emotion2vecParams=94M2026.03 | — | — | 27.4 | 47.6 | — | — | |
| FocalSerSSL Encoders=WavLM and RoBERTa, Configuration=single-model2026.03 | — | 37.51 | 40.85 | — | — | — | |
| Full labeledTraining Protocol=Supervised (full dataset)2026.02 | — | — | — | — | 28.3 | — | |
| HuBERTParams=94M2026.03 | — | — | 24 | 45.3 | — | — | |
| HuBERT-largeClassifier=MLP, Feature=Self-supervised embeddings2026.02 | — | — | — | 56.8 | — | — | |
| MedusaSSL Encoders=WavLM and RoBERTa, Configuration=single-model2026.03 | — | 43.03 | 45.48 | — | — | — | |
| MemoCMTSSL Encoders=HuBERT and BERT2026.03 | — | 64.9 | 43.14 | — | — | — | |
| SALMONNParams=7B2026.03 | — | — | 33.4 | 53.3 | — | — | |
| SALMONNParams=13B2026.03 | — | — | 32.8 | 52.6 | — | — | |
| Same epochsTraining Protocol=Supervised (200 epochs)2026.02 | — | — | — | — | 27.3 | — | |
| SenseVoice-small2026.06 | — | 57.8 | — | — | — | — | |
| SONARParams=600M2026.03 | — | — | 23.3 | 43.2 | — | — | |
| VowelPrompt2026.02 | — | — | — | 69.6 | — | — | |
| w/o LLMAblation=Without LLM2026.06 | — | 54.1 | — | — | — | — | |
| w/o LLM and LFMAblation=Without LLM and LFM2026.06 | — | 55.7 | — | — | — | — | |
| wav2vec-largeClassifier=MLP, Feature=Self-supervised embeddings2026.02 | — | — | — | 55.1 | — | — | |
| WavLMParams=94M2026.03 | — | — | 24.3 | 45.6 | — | — | |
| WavLM BaselineBackbone=WavLM-large2026.03 | — | 30.11 | 30.46 | — | — | — | |
| WavLM-baseDownstream=Linear2023.12 | — | 46.95 | 16.34 | 35.16 | — | — | |
| WavLM-base+Downstream=Linear2023.12 | — | 43.78 | 16.75 | 34.6 | — | — |