Speech Emotion Recognition on MSP-Podcast (test1)
42.24P-MacroF1ADEPT
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| ADEPTvariant=+ GRPO, Avg Size=2.21, classification_mode=primary+secondary prediction2026.02 | 42.24 | 78.74 | 61.92 | 47.51 | |
| WavLMtype=Reported SSL baseline, mode=Fine-tuned, classification_mode=primary-only classification2026.02 | 36.4 | — | — | — | |
| ADEPTvariant=w/o GRPO, Avg Size=2.15, classification_mode=primary+secondary prediction2026.02 | 35.71 | 70.32 | 57.48 | 44.5 | |
| WavLMtype=Reported SSL baseline, mode=Frozen, classification_mode=primary-only classification2026.02 | 29.7 | — | — | — | |
| BLSP-Emotype=Reported SLLM/MLLM baseline, Avg Size=2.08, classification_mode=generative, primary+secondary capable2026.02 | 29.15 | 67.4 | 53.2 | 42.8 | |
| HuBERTtype=Reported SSL baseline, classification_mode=primary-only classification2026.02 | 28.5 | — | — | — | |
| wav2vec2type=Reported SSL baseline, classification_mode=primary-only classification2026.02 | 23.8 | — | — | — | |
| Qwen-3-Omnitype=Baseline, Avg Size=2.02, classification_mode=primary+secondary prediction2026.02 | 23.58 | 60.03 | 49.23 | 38.9 | |
| Qwen-2-Audiotype=Reported SLLM/MLLM baseline, Avg Size=1.95, classification_mode=generative, primary+secondary capable2026.02 | 22.9 | 51.2 | 41.5 | 32.6 | |
| SALMONNtype=Reported SLLM/MLLM baseline, Avg Size=1.76, classification_mode=generative, primary+secondary capable2026.02 | 13.89 | 43.2 | 33.2 | 22.8 |