Classification on AudioSet (test)
49.6mAPPaSST-S
Evaluation Results
| Method | Links | |
|---|---|---|
| PaSST-SPatch stride=S10-16, Number of models=92021.10 | 49.6 | |
| PaSST-SPatch stride=S10-16, Number of models=52021.10 | 49.5 | |
| PaSST-SPatch stride=S10-16, Number of models=42021.10 | 49.3 | |
| PaSST-SPatch stride=S16,14, Number of models=22021.10 | 48.6 | |
| ASTEnsemble configuration=Ensemble-M2021.10 | 48.5 | |
| PlayItBackX3Backbone=MViTv2-B [32], Train set=AS-500K2022.10 | 47.7 | |
| ASTEnsemble configuration=Ensemble-S2021.10 | 47.5 | |
| PSLAEnsemble configuration=Ensemble-M2021.10 | 47.4 | |
| Audio-MAEBackbone=ViT-B, Train set=AS-2M2022.10 | 47.3 | |
| PaSSTBackbone=DeiT-B [28], Train set=AS-2M2022.10 | 47.1 | |
| HTS-ATBackbone=Swin-T [29], Train set=AS-2M2022.10 | 47.1 | |
| MaskSpecBackbone=ViT-B, Train set=AS-2M2022.10 | 47.1 | |
| PSLAEnsemble configuration=Ensemble-S2021.10 | 46.9 | |
| LAION-CLAPZero-shot=true2025.03 | 45.85 | |
| LAION-CLAPzero-shot=true2024.06 | 45.85 | |
| PSLABackbone=EffNet-B2 [27], Train set=AS-2M2022.10 | 44.4 | |
| MBTBackbone=ViT-B, Train set=AS-500K2022.10 | 44.3 | |
| SupervisedTrain inputs=waveform + log-mel, Eval inputs=waveform + log-mel2021.03 | 43.9 | |
| PANNBackbone=ResNet38 [26], Train set=AS-2M2022.10 | 43.4 | |
| ConformerBackbone=Conformer, Train set=AS-2M2022.10 | 41.1 | |
| GLAPZero-shot=true2025.03 | 40.9 | |
| PerceiverBackbone=Perceiver, Train set=AS-2M2022.10 | 38.4 | |
| multi-format contrastive audio learning frameworkTrain inputs=waveform + log-mel, Eval inputs=waveform + log-mel2021.03 | 37.6 | |
| multi-format contrastive audio learning frameworkTrain inputs=waveform + log-mel, Eval inputs=log-mel2021.03 | 36.8 | |
| multi-format contrastive audio learning frameworkTrain inputs=waveform + log-mel, Eval inputs=waveform2021.03 | 35.5 | |
| multi-format contrastive audio learning frameworkTrain inputs=waveform, Eval inputs=waveform2021.03 | 33.6 | |
| multi-format contrastive audio learning frameworkTrain inputs=log-mel, Eval inputs=log-mel2021.03 | 32.9 | |
| Baseline2021.10 | 31.4 | |
| MMVTrain inputs=log-mel + video + text, Eval inputs=log-mel2021.03 | 30.9 | |
| MAE-ASTBackbone=ViT-B [4], Train set=mini-AS2022.10 | 30.6 | |
| M2D-CLAP2025Zero-shot=true2025.03 | 30.05 | |
| C³Train inputs=log-mel + video, Eval inputs=log-mel2021.03 | 28.5 | |
| M2D-CLAP2025 stage2Zero-shot=true2025.03 | 28.27 | |
| CPCTrain inputs=waveform, Eval inputs=waveform2021.03 | 27.7 | |
| M2D-CLAP2025 stage2.1Zero-shot=true2025.03 | 27.24 | |
| M2D-CLAPmasking_ratio=0.7, zero-shot=true2024.06 | 27.24 | |
| L³Train inputs=log-mel + video, Eval inputs=log-mel2021.03 | 24.9 | |
| TripletTrain inputs=log-mel, Eval inputs=log-mel2021.03 | 24.4 | |
| M2D-CLAP2025 stage1Zero-shot=true2025.03 | 23.15 | |
| MGA-CLAP+Zero-shot=true2025.03 | 23 | |
| M2D-CLAP2024 stage1Zero-shot=true2025.03 | 20.82 | |
| CDOntologyZero-shot=true2025.03 | 19.98 | |
| WavCapsZero-shot=true2025.03 | 19.6 | |
| WavCapszero-shot=true2024.06 | 19.6 | |
| LTUZero-shot=true2025.03 | 18.7 | |
| LTUzero-shot=true2024.06 | 18.7 | |
| PengiZero-shot=true2025.03 | 16.35 | |
| Pengizero-shot=true2024.06 | 16.35 | |
| CLAP2023Zero-shot=true2025.03 | 10.2 | |
| CLAP2023zero-shot=true2024.06 | 10.2 | |
| COLLATZero-shot=true2025.03 | 9 | |
| AudioCLIPzero-shot=true2024.06 | 6.4 | |
| CLAP2022Zero-shot=true2025.03 | 5.8 | |
| CLAP2022zero-shot=true2024.06 | 5.8 | |
| Proto-LCzero-shot=true2024.06 | 5.21 | |
| Wav2CLIPZero-shot=true2025.03 | 3.02 | |
| Wav2CLIPzero-shot=true2024.06 | 3.02 |