Audio Classification on GTZAN
94.66AccuracyNystromformer
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| NystromformerNumber of Transformer layers=2, Input transformation (VGGish)=VGGish network, Attention mechanism=low-rank Nystromformer attention, Input tokens=120, Landmarks=42025.11 | 94.66 | 845.4 | 0.56 | |
| Co. TransformerNumber of Transformer layers=2, Input transformation (VGGish)=VGGish network, Attention mechanism=Continual Transformers, Input tokens=1202025.11 | 94.28 | 230.7 | 1.02 | |
| TransformerNumber of Transformer layers=2, Input transformation (VGGish)=VGGish network, Attention mechanism=regular Transformers, Input tokens=1202025.11 | 94.19 | 11,134.3 | 1 | |
| DeepCoTNumber of Transformer layers=2, Input transformation (VGGish)=VGGish network, Attention mechanism=DeepCoT (Continual Single Output attention block), Input tokens=1202025.11 | 94.19 | 138.7 | 37.24 | |
| Co. NystromformerNumber of Transformer layers=2, Input transformation (VGGish)=VGGish network, Attention mechanism=Continual Nystromformer, Input tokens=120, Landmarks=42025.11 | 93.53 | 114.3 | 0.71 | |
| DeLoRes-MEvaluation protocol=Linear evaluation2022.10 | 90.4 | — | — | |
| M2DASAudio enc. fine-tuning on AudioSet=true2025.03 | 87.6 | — | — | |
| PaSSTTraining data=AS, Learning type=Supervised, Evaluation protocol=Linear evaluation2025.02 | 87.4 | — | — | |
| M2D-AS# Params=86M, Masking ratio=0.7, Evaluation protocol=Linear evaluation2024.04 | 86.9 | — | — | |
| M2D-CLAP stage1 2025Audio enc. fine-tuning on AudioSet=false2025.03 | 86.6 | — | — | |
| M2D-CLAP2025Audio enc. fine-tuning on AudioSet=true2025.03 | 86.3 | — | — | |
| BEATs iter3+Training data=AS, Learning type=Supervised, Evaluation protocol=Linear evaluation2025.02 | 86 | — | — | |
| MATPACAudio enc. fine-tuning on AudioSet=false2025.03 | 85.9 | — | — | |
| MATPACTraining data=AS, nr,epoch=20, Learning type=SSL, Evaluation protocol=Linear evaluation2025.02 | 85.9 | — | — | |
| HTS-ATTraining data=AS, Learning type=Supervised, Evaluation protocol=Linear evaluation2025.02 | 85.9 | — | — | |
| MATPACTraining data=AS, nr,epoch=10, Learning type=SSL, Evaluation protocol=Linear evaluation2025.02 | 85.3 | — | — | |
| ASTAudio enc. fine-tuning on AudioSet=true2025.03 | 85.1 | — | — | |
| HTS-ATAudio enc. fine-tuning on AudioSet=true2025.03 | 85.1 | — | — | |
| HTS-AT# Params=31M, Evaluation protocol=Linear evaluation2024.04 | 85.1 | — | — | |
| Cacophony (stage 2)Audio enc. fine-tuning on AudioSet=false2025.03 | 85 | — | — | |
| F3-TokenizerProbing=Frozen representation2026.06 | 85 | — | — | |
| BEATsiter3+Audio enc. fine-tuning on AudioSet=true2025.03 | 84.6 | — | — | |
| BEATSiter3+# Params=90M, Evaluation protocol=Linear evaluation2024.04 | 84.6 | — | — | |
| LAION-CLAPAudio enc. fine-tuning on AudioSet=true2025.03 | 84.3 | — | — | |
| AST# Params=86M, Evaluation protocol=Linear evaluation2024.04 | 84.3 | — | — | |
| MATPAC w/o ClassificationTraining data=AS, Learning type=SSL, Evaluation protocol=Linear evaluation2025.02 | 84.2 | — | — | |
| M2DAudio enc. fine-tuning on AudioSet=false2025.03 | 84.1 | — | — | |
| M2D-CLAP stage1 2024Audio enc. fine-tuning on AudioSet=false2025.03 | 84.1 | — | — | |
| M2D# Params=86M, Masking ratio=0.7, Evaluation protocol=Linear evaluation2024.04 | 84.1 | — | — | |
| M2DEvaluation protocol=Linear evaluation, Masking ratio=0.72022.10 | 83.9 | — | — | |
| M2Dratio=0.7, Training data=AS, Learning type=SSL, Evaluation protocol=Linear evaluation2025.02 | 83.9 | — | — | |
| Cacophony (stage 1)Audio enc. fine-tuning on AudioSet=false2025.03 | 83.8 | — | — | |
| M2DEvaluation protocol=Linear evaluation, Masking ratio=0.62022.10 | 83.3 | — | — | |
| ATST-FrameAudio enc. fine-tuning on AudioSet=false2025.03 | 82.9 | — | — | |
| AST-Fusion#5#12Evaluation protocol=Linear evaluation, Variant=#5#122022.10 | 82.9 | — | — | |
| CLAP2023Audio enc. fine-tuning on AudioSet=true2025.03 | 82.3 | — | — | |
| ATST-FrameTraining data=AS, Learning type=SSL, Evaluation protocol=Linear evaluation2025.02 | 80.7 | — | — | |
| SLAP_WavcapsZero-shot=true, training_data=WavCaps2026.01 | 80.5 | — | — | |
| w/o LLMProbing=Frozen representation2026.06 | 80.4 | — | — | |
| WavCapsAudio enc. fine-tuning on AudioSet=true2025.03 | 80.2 | — | — | |
| PengiAudio enc. fine-tuning on AudioSet=true2025.03 | 80 | — | — | |
| BEATS iter3Training data=AS, Learning type=SSL, Evaluation protocol=Linear evaluation2025.02 | 80 | — | — | |
| ATST-ClipTraining data=AS, Learning type=SSL, Evaluation protocol=Linear evaluation2025.02 | 79.9 | — | — | |
| CLAP2022Audio enc. fine-tuning on AudioSet=true2025.03 | 79.3 | — | — | |
| M2D2Zero-shot=true2026.01 | 79.3 | — | — | |
| ATST-ClipAudio enc. fine-tuning on AudioSet=false2025.03 | 78.9 | — | — | |
| PANNs CNN14Audio enc. fine-tuning on AudioSet=true2025.03 | 78.7 | — | — | |
| MSM-MAEEvaluation protocol=Linear evaluation2022.10 | 78.4 | — | — | |
| SF NFNet-F0Evaluation protocol=Linear evaluation2022.10 | 78.2 | — | — | |
| ATST BaseEvaluation protocol=Linear evaluation, Variant=Base2022.10 | 76.4 | — | — | |
| BEATsiter3Audio enc. fine-tuning on AudioSet=false2025.03 | 72.6 | — | — | |
| w/o RQProbing=Frozen representation2026.06 | 72.6 | — | — | |
| Ming-UProbing=Frozen representation2026.06 | 71.17 | — | — | |
| AudioSetCapsZero-shot=true2026.01 | 70.5 | — | — | |
| BYOL-AEvaluation protocol=Linear evaluation2022.10 | 70.1 | — | — | |
| GLAPZero-shot=true2026.01 | 69.6 | — | — | |
| Laion-CLAPzero-shot=true, multimodal query=true2024.10 | 68.07 | — | — | |
| Laion-CLAPzero-shot=true2024.10 | 66.26 | — | — | |
| MAE-ASTTraining data=AS+LS, Learning type=SSL, Evaluation protocol=Linear evaluation2025.02 | 64.1 | — | — | |
| WhisperProbing=Frozen representation2026.06 | 62.2 | — | — | |
| MS-CLAPZero-shot=true2026.01 | 58.4 | — | — | |
| Wav2Vec2Evaluation protocol=Linear evaluation2022.10 | 57.8 | — | — | |
| SLAPZero-shot=true2026.01 | 56.8 | — | — | |
| CEDAudio enc. fine-tuning on AudioSet=false2025.03 | 42.3 | — | — | |
| No repr.Probing=Frozen representation2026.06 | 14.2 | — | — |