Audio-visual event classification on AudioSet 20K
42.4mAP (Audio-only)EquiAV
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| EquiAVPretrain=SSL2024.03 | 42.4 | 25.7 | 46.6 | — | |
| MAVILPretrain=SSL2024.03 | 41.8 | 24.8 | 44.9 | — | |
| MAVIL2024.09 | 41.8 | 24.8 | 44.9 | — | |
| CAV-MAEPretrain=SSL2024.03 | 37.7 | 19.8 | 42 | — | |
| CAV-MAE2024.09 | 37.7 | 19.8 | 42 | — | |
| MBT*Pretrain=IN21K SL2024.03 | 31.3 | 27.7 | 43.9 | — | |
| MBT2024.09 | 31.3 | 27.7 | 43.9 | — | |
| GBlend2024.03 | 29.1 | 22.1 | 37.8 | — | |
| G-Blend2024.09 | 29.1 | 22.1 | 37.8 | — | |
| VAB-EncodecAudio Tokenizer=Encodec2024.09 | 29 | 29 | 38.7 | — | |
| VAB-DACAudio Tokenizer=DAC2024.09 | 28.8 | 28.3 | 38.9 | — | |
| CAV-MAEPre-training=AS2M, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | — | — | — | 27.3 | |
| CAV-MAE Scale+++Pre-training=AS2M, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | — | — | — | 25.3 | |
| CAV-MAE SyncPre-training=AS2M, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | — | — | — | 30.5 | |
| MaViLPre-training=AS2M, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | — | — | — | 30 | |
| OursPre-training=AS2M, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | — | — | — | 32 | |
| VAB-EncodecPre-training=AS2M+VGG, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | — | — | — | 33.3 |