Audio-visual Classification on VGGSound
69.8Top-1 AccMirasol3B
Evaluation Results
| Method | Links | ||||||||
|---|---|---|---|---|---|---|---|---|---|
| Mirasol3BLearnable Param.=3B2025.03 | 69.8 | — | — | — | — | — | — | — | |
| CA2STLearnable Param.=62M2025.03 | 68.3 | — | — | — | — | — | — | — | |
| CAVALearnable Param.=44M2025.03 | 68.2 | — | — | — | — | — | — | — | |
| MAViLPT=IN-SSL, AS-SSL2022.12 | 67.1 | 60.8 | 50.9 | — | — | — | — | — | |
| MAVIL(A+V)Learnable Param.=86M2025.03 | 67.1 | — | — | — | — | — | — | — | |
| EquiAV(A+V)2025.03 | 67.1 | — | — | — | — | — | — | — | |
| CrossMAE2025.03 | 67 | — | — | — | — | — | — | — | |
| MAViLPT=AS-SSL2022.12 | 66.5 | 60.6 | 50 | — | — | — | — | — | |
| MMT(A+V)Learnable Param.=52M2025.03 | 66.2 | — | — | — | — | — | — | — | |
| CAV-MAEPT=IN-SSL, AS-SSL2022.12 | 65.5 | 59.5 | 47 | — | — | — | — | — | |
| CAV-MAELearnable Param.=164M2025.03 | 65.5 | — | — | — | — | — | — | — | |
| AudiovisualMAELearnable Param.=36M2025.03 | 65 | — | — | — | — | — | — | — | |
| MBTPT=IN21K-SL2022.12 | 64.1 | 52.3 | 51.2 | — | — | — | — | — | |
| MBT2025.03 | 64.1 | — | — | — | — | — | — | — | |
| VAB-EncodecPre-training=AS2M+VGG, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | 57.6 | — | — | — | — | — | — | — | |
| CASTLearnable Param.=45M2025.03 | 54.7 | — | — | — | — | — | — | — | |
| Uni-Modal Teacher2021.06 | 53.46 | — | — | — | — | — | — | — | |
| CAV-MAE SyncPre-training=AS2M, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | 52.7 | — | — | — | — | — | — | — | |
| OursPre-training=AS2M, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | 52.7 | — | — | — | — | — | — | — | |
| CAV-MAE Scale+++Pre-training=AS2M, Evaluation Protocol=attention probing, Encoder State=frozen2026.04 | 51.6 | — | — | — | — | — | — | — | |
| Modality Dropoutdrop_probability=1/32021.06 | 51.37 | — | — | — | — | — | — | — | |
| Pre-train + Fine-tunestrategy=pre-train uni-modal encoders, then fine-tune classifier2021.06 | 50.81 | — | — | — | — | — | — | — | |
| Gradient-Blending2021.06 | 50.39 | — | — | — | — | — | — | — | |
| Self Distillationstrategy=distilling a pre-trained naive fusion model to a new one2021.06 | 49.86 | — | — | — | — | — | — | — | |
| Dropoutdropout_ratio=0.52021.06 | 49.83 | — | — | — | — | — | — | — | |
| Video-only Distillationdistillation_source=video2021.06 | 49.55 | — | — | — | — | — | — | — | |
| Naive Fusionbaseline=true2021.06 | 49.46 | — | — | — | — | — | — | — | |
| Audio-only Distillationdistillation_source=audio2021.06 | 48.84 | — | — | — | — | — | — | — | |
| RL-MBALabelled samples=3,000, Percentage of training set=2.9%2026.03 | 22.23 | — | — | — | — | — | — | — | |
| RandomLabelled samples=3,000, Percentage of training set=2.9%2026.03 | 21.73 | — | — | — | — | — | — | — | |
| BMMALLabelled samples=3,000, Percentage of training set=2.9%2026.03 | 20.53 | — | — | — | — | — | — | — | |
| EntropyLabelled samples=3,000, Percentage of training set=2.9%2026.03 | 20.43 | — | — | — | — | — | — | — | |
| GCNALLabelled samples=3,000, Percentage of training set=2.9%2026.03 | 20.33 | — | — | — | — | — | — | — | |
| BADGELabelled samples=3,000, Percentage of training set=2.9%2026.03 | 20.23 | — | — | — | — | — | — | — | |
| CoreSetLabelled samples=3,000, Percentage of training set=2.9%2026.03 | 20.13 | — | — | — | — | — | — | — | |
| BALDLabelled samples=3,000, Percentage of training set=2.9%2026.03 | 19.93 | — | — | — | — | — | — | — | |
| DeepFoolLabelled samples=3,000, Percentage of training set=2.9%2026.03 | 19.73 | — | — | — | — | — | — | — | |
| Aud-SlowFastPT=-2022.12 | — | 50.1 | — | — | — | — | — | — | |
| AudioMAEEvaluation Protocol=Linear Probing2024.10 | — | — | — | — | — | 42.35 | — | — | |
| AudioMAEEvaluation Protocol=Fine-tuning2024.10 | — | — | — | — | — | 57.76 | — | — | |
| Audiovisual MAEPretraining=SSL VGGSound2022.12 | — | — | — | 57.2 | 50.3 | 65 | — | — | |
| AudiovisualMAE*Pretrain=SSL2024.03 | — | — | — | 57.2 | 50.3 | 65 | — | — | |
| AV-MAEEvaluation Protocol=Linear Probing2024.10 | — | — | — | — | — | 56.15 | — | — | |
| AV-MAEEvaluation Protocol=Fine-tuning2024.10 | — | — | — | — | — | 65.08 | — | — | |
| AVAGENTEvaluation Protocol=Linear Probing2024.10 | — | — | — | — | — | 61.56 | — | — | |
| AVAGENTEvaluation Protocol=Fine-tuning2024.10 | — | — | — | — | — | 69.24 | — | — | |
| CAV-MAEPretrain=SSL2024.03 | — | — | — | 59.5 | 47 | 65.5 | — | — | |
| CAV-MAE2024.09 | — | 59.5 | 47 | — | — | 65.5 | — | — | |
| CAV-MAEEvaluation Protocol=Linear Probing2024.10 | — | — | — | — | — | 55.27 | — | — | |
| CAV-MAEEvaluation Protocol=Fine-tuning2024.10 | — | — | — | — | — | 65.53 | — | — | |
| EquiAVPretrain=SSL2024.03 | — | — | — | 61 | 50.7 | 67.1 | — | — | |
| ImageBindA-Enc Params.=.09B, Data (M)=32025.12 | — | — | — | — | — | — | 40.8 | 28.2 | |
| Kazakos et al.Pretraining=Sup. Im1K2022.12 | — | — | — | 52.5 | — | — | — | — | |
| LangBindA-Enc Params.=0.3B, Data (M)=102025.12 | — | — | — | — | — | — | 44.1 | 26 | |
| MAEEvaluation Protocol=Linear Probing2024.10 | — | — | — | — | — | 15.61 | — | — | |
| MAEEvaluation Protocol=Fine-tuning2024.10 | — | — | — | — | — | 45.73 | — | — | |
| MAVILPretrain=SSL2024.03 | — | — | — | 60.8 | 50.9 | 67.1 | — | — | |
| MAVIL2024.09 | — | 60.8 | 50.9 | — | — | 67.1 | — | — | |
| MAVILEvaluation Protocol=Linear Probing2024.10 | — | — | — | — | — | 57.36 | — | — | |
| MAVILEvaluation Protocol=Fine-tuning2024.10 | — | — | — | — | — | 67.17 | — | — | |
| MBTPretraining=Sup. Im21K2022.12 | — | — | — | 52.3 | 51.2 | 64.1 | — | — | |
| MBT2024.09 | — | 52.3 | 51.2 | — | — | 64.1 | — | — | |
| MBT*Pretrain=IN21K SL2024.03 | — | — | — | 52.3 | 51.2 | 64.1 | — | — | |
| PEAV BA-Enc Params.=.2B, Data (M)=124, Sampling=16 Frames2025.12 | — | — | — | — | — | — | 61.7 | 45.2 | |
| PEAV BA-Enc Params.=0.2B, Data (M)=124, Sampling=30 FPS2025.12 | — | — | — | — | — | — | 62.1 | 44.5 | |
| PEAV LA-Enc Params.=1.1B, Data (M)=124, Sampling=16 Frames2025.12 | — | — | — | — | — | — | 63.7 | 46.7 | |
| PEAV LA-Enc Params.=1.1B, Data (M)=124, Sampling=30 FPS2025.12 | — | — | — | — | — | — | 63.7 | 47.1 | |
| PEAV L (PT)A-Enc Params.=1.1B, Data (M)=92, Sampling=30 FPS, PT=true2025.12 | — | — | — | — | — | — | 55.7 | 42.4 | |
| PEAV L-OODA-Enc Params.=1.1B, Data (M)=114, Sampling=30 FPS, OOD=true2025.12 | — | — | — | — | — | — | 58.5 | 43.9 | |
| PEAV SA-Enc Params.=.09B, Data (M)=124, Sampling=16 Frames2025.12 | — | — | — | — | — | — | 60.9 | 43 | |
| PEAV SA-Enc Params.=.09B, Data (M)=124, Sampling=30 FPS2025.12 | — | — | — | — | — | — | 61.6 | 43 | |
| PlayItBackPretraining=Sup. Im21K2022.12 | — | — | — | 53.7 | — | — | — | — | |
| PolyViTPretraining=Sup. Im21K, AS2022.12 | — | — | — | 55.1 | — | — | — | — | |
| VAB-DACAudio Tokenizer=DAC2024.09 | — | 48.2 | 55.4 | — | — | 63.9 | — | — | |
| VAB-EncodecAudio Tokenizer=Encodec2024.09 | — | 51.3 | 55.1 | — | — | 65.2 | — | — | |
| VGGSoundPT=-2022.12 | — | 48.8 | — | — | — | — | — | — |