Speech Emotion Recognition on CREMA-D
95.24Weighted AccuracyStudent
Evaluation Results
| Method | Links | |
|---|---|---|
| Studentdepth=3, num head=5, FLOPs=0.32G2024.03 | 95.24 | |
| Teacherdepth=6, num head=5, FLOPs=1.43G2024.03 | 95.07 | |
| Studentdepth=3, num head=5, FLOPs=0.32G2024.03 | 94.07 | |
| Teacherdepth=12, num head=12, FLOPs=3.58G2024.03 | 92.64 | |
| Gong et al.depth=12, num head=12, Pre-train=ImageNet, FLOPs=3.32G2024.03 | 88.79 | |
| Ensemble (Submission)Ensemble Strategy=Embedding Concatenation2026.01 | 81.5 | |
| Ristea et al.depth=6, num head=5, FLOPs=175.81G2024.03 | 79.94 | |
| Dasheng 1.2BParameter Count=1.2B2026.01 | 79 | |
| Challenge BaselineDescription=Best among Dasheng-base, data2vec, and Whisper2026.01 | 77.2 | |
| BEATs (300M) SpeechParameter Count=300M, Pre-training Mixture=Speech-heavy (70:15:15)2026.01 | 67 | |
| BEATs (300M) BalancedParameter Count=300M, Pre-training Mixture=Balanced (40:30:30)2026.01 | 65.9 | |
| BEATs (90M) iter3Parameter Count=90M, Iteration=32026.01 | 64.2 |