Emotion Recognition on RAVDESS (test)
0.9735AccuracyMM student (sup)
Evaluation Results
| Method | Links | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|
| MM student (sup)modality=images and audios (i, a), Dl=false, Du=trained on true labels2021.03 | 0.9735 | — | — | — | — | — | — | — | — | — | |
| MM student (ours)modality=images and audios (i, a), Dl=false, Du=true2021.03 | 0.9138 | — | — | — | — | — | — | — | — | — | |
| MM student (no reg)modality=images and audios (i, a), Dl=false, Du=true2021.03 | 0.8928 | — | — | — | — | — | — | — | — | — | |
| N-PPBackbone=I3D, Approach=No privacy preservation2026.03 | 0.8773 | — | — | — | — | — | — | — | — | — | |
| N-PPBackbone=R(2+1)D, Approach=No privacy preservation2026.03 | 0.8644 | — | — | — | — | — | — | — | — | — | |
| VQ-MAE-S-12Masking Strategy=Frame, Query2Emo=true2023.04 | 0.841 | 84.4 | — | — | — | — | — | — | — | — | |
| NOISY studentmodality=images (i), Dl=true, Du=true2021.03 | 0.8309 | — | — | — | — | — | — | — | — | — | |
| VQ-MAE-S-12Masking Strategy=Frame, Query2Emo=false2023.04 | 0.808 | 80.5 | — | — | — | — | — | — | — | — | |
| Identity-Disentangled Privacy-Preserving Video FERBackbone=I3D2026.03 | 0.8035 | — | — | — | — | — | — | — | — | — | |
| UM teachermodality=images (i), Dl=true, Du=false2021.03 | 0.8033 | — | — | — | — | — | — | — | — | — | |
| Contr-HLBackbone=R(2+1)D2026.03 | 0.7993 | — | — | — | — | — | — | — | — | — | |
| Identity-Disentangled Privacy-Preserving Video FERBackbone=R(2+1)D2026.03 | 0.792 | — | — | — | — | — | — | — | — | — | |
| Multi-Agent Emotion-to-Response SystemClassification Granularity=4-class2026.01 | 0.785 | — | — | — | — | — | — | — | — | — | |
| GBBackbone=I3D, Gaussian blurring hyperparameter (sigma)=0.62026.03 | 0.7821 | — | — | — | — | — | — | — | — | — | |
| VQ-MAE-S-12Masking Strategy=Patch-tf, Query2Emo=true2023.04 | 0.782 | 77.5 | — | — | — | — | — | — | — | — | |
| UM studentmodality=images (i), Dl=false, Du=true2021.03 | 0.7779 | — | — | — | — | — | — | — | — | — | |
| Contr-HLBackbone=I3D2026.03 | 0.7759 | — | — | — | — | — | — | — | — | — | |
| GBBackbone=R(2+1)D, Gaussian blurring hyperparameter (sigma)=0.62026.03 | 0.7754 | — | — | — | — | — | — | — | — | — | |
| Adver.Backbone=I3D2026.03 | 0.7731 | — | — | — | — | — | — | — | — | — | |
| VQ-MAE-S-12Masking Strategy=Patch-tf, Query2Emo=false2023.04 | 0.767 | 75.9 | — | — | — | — | — | — | — | — | |
| Adver.Backbone=R(2+1)D2026.03 | 0.7508 | — | — | — | — | — | — | — | — | — | |
| Multi-Agent Emotion-to-Response SystemClassification Granularity=6-class2026.01 | 0.713 | — | — | — | — | — | — | — | — | — | |
| Face S.Backbone=I3D2026.03 | 0.6985 | — | — | — | — | — | — | — | — | — | |
| Face S.Backbone=R(2+1)D2026.03 | 0.6839 | — | — | — | — | — | — | — | — | — | |
| Multi-Agent Emotion-to-Response SystemClassification Granularity=8-class2026.01 | 0.642 | — | — | — | — | — | — | — | — | — | |
| Self-attention audio2023.04 | 0.583 | — | — | — | — | — | — | — | — | — | |
| SpecMAE-12Masking Strategy=Patch-tf2023.04 | 0.522 | 52 | — | — | — | — | — | — | — | — | |
| Elizalde et al., 2023bzero-shot=true2024.02 | 0.217 | — | — | — | — | — | — | — | — | — | |
| Audio Flamingozero-shot=true2024.02 | 0.209 | — | — | — | — | — | — | — | — | — | |
| Baseline CNN2026.02 | — | 0.53 | 0.66 | 0.49 | 0.52 | 0.48 | 0.56 | 0.52 | 0.5 | 0.53 | |
| Implemented CNN2026.02 | — | 0.94 | 0.97 | 0.92 | 0.96 | 0.92 | 0.94 | 0.93 | 0.92 | 0.95 |