Group Emotion Recognition on VGAF
82.25AccuracyVE-MD + Wav2Vec 2.0
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VE-MD + Wav2Vec 2.0Year=2026, Modalities=A,V2026.04 | 82.25 | — | |
| VE-MD + Wav2Vec 2.0Year=2025, Modalities=A,V, Features=Pred. Body+Face SR2026.05 | 82.25 | — | |
| VE-MD + WavLMYear=2025, Modalities=A,V, Features=Pred. Body+Face SR2026.05 | 82.11 | — | |
| Streams-fusion (TimeSformer, Wav2Vec 2.0, YOLOv8)Year=2024, Modalities=A,V2026.04 | 81.98 | — | |
| Multimodal fusion (TimeSformer, Wav2Vec 2.0, YOLOv8)Year=2024, Modalities=A,V, Ind. Features=true, Features=Body Pose2026.05 | 81.98 | — | |
| VE-MDYear=2026, Modalities=V2026.04 | 80.81 | — | |
| VE-MD HeatmapYear=2025, Modalities=V, Features=Pred. Face SR2026.05 | 80.81 | — | |
| VE-MD HeatmapYear=2025, Modalities=V, Features=Pred. Body SR2026.05 | 80.68 | — | |
| VE-MD HeatmapYear=2025, Modalities=V, Features=Pred. Body+Face SR2026.05 | 80.55 | — | |
| VE-MD DETRYear=2025, Modalities=V, Features=Pred. Body SR2026.05 | 80.42 | — | |
| ViT + Synthetic dataYear=2023, Modalities=V2026.04 | 79.24 | — | |
| ViT + Synthetic dataYear=2023, Modalities=V, Features=Global Image2026.05 | 79.24 | — | |
| Cross AttentionYear=2023, Modalities=A,V, Features=Global Image2026.05 | 78.72 | — | |
| TimeSformer, YOLOv8Year=2024, Modalities=V2026.04 | 73.1 | — | |
| TimeSformer, YOLOv8Year=2024, Modalities=V, Ind. Features=true, Features=Body Pose2026.05 | 73.1 | — | |
| Wav2Vec 2.0 + WhisperYear=2025, Modalities=A, Features=Acoustic, Content2026.05 | 69.45 | — | |
| Wav2Vec 2.0 +1D CNNYear=2024, Modalities=A, Features=Acoustic2026.05 | 64.09 | — | |
| Synthetic augmentation, VGG19Authors=Petrova et al. [176], Methodology=Synthetic augmentation, VGG192026.05 | 59.12 | 11.24 | |
| CNN+TransformerYear=2023, Modalities=A, Features=Acoustic2026.05 | 56.4 | — |