Audio Classification on Crema-D
78.9AccuracyF3-Tokenizer
Evaluation Results
| Method | Links | |
|---|---|---|
| F3-TokenizerProbing=Frozen representation2026.06 | 78.9 | |
| M2D/0.7#Params=86M, Protocol=Linear Probing2025.12 | 73 | |
| ATST-Frame#Params=86M, Protocol=Linear Probing2025.12 | 72.3 | |
| LeRaCTraining Regime=LeRaC, Model=SepTr2022.05 | 70.95 | |
| conventionalTraining Regime=conventional, Model=SepTr2022.05 | 70.47 | |
| w/o LLMProbing=Frozen representation2026.06 | 70.4 | |
| CBSTraining Regime=CBS, Model=SepTr2022.05 | 69.98 | |
| LeRaCTraining Regime=LeRaC, Model=DenseNet-1212022.05 | 68.99 | |
| AaPE#Params=87M, Protocol=Linear Probing2025.12 | 68.7 | |
| CBSTraining Regime=CBS, Model=DenseNet-1212022.05 | 68.16 | |
| conventionalTraining Regime=conventional, Model=DenseNet-1212022.05 | 67.21 | |
| Ming-UProbing=Frozen representation2026.06 | 66.85 | |
| BEATs_iter3#Params=90M, Protocol=Linear Probing2025.12 | 64.7 | |
| FedMLAC#Clients=18, Active Ratio=20%, Model Homogeneity=homogeneous2025.06 | 63.04 | |
| FedKAD#Clients=18, Active Ratio=20%, Model Homogeneity=homogeneous2025.06 | 62.17 | |
| w/o RQProbing=Frozen representation2026.06 | 61.2 | |
| FedProx#Clients=18, Active Ratio=20%, Model Homogeneity=homogeneous2025.06 | 61.02 | |
| FedAvg#Clients=18, Active Ratio=20%, Model Homogeneity=homogeneous2025.06 | 60.84 | |
| EfficientLEAFFilterbank=Gabor 8G-opt, Compression=L-M-TBN2022.07 | 60.2 | |
| STFT-MelFilterbank=STFT-Mel*, Compression=L-M-TBN*2022.07 | 58.8 | |
| FedOPT#Clients=18, Active Ratio=20%, Model Homogeneity=homogeneous2025.06 | 58.47 | |
| EfficientLEAFFilterbank=Gabor 4G, Compression=L-M-TBN2022.07 | 58 | |
| WhisperProbing=Frozen representation2026.06 | 57.2 | |
| LEAFFilterbank=Gabor, Compression=PCEN2022.07 | 50.2 | |
| EfficientLEAFFilterbank=Gabor 4G, Compression=PCEN2022.07 | 50 | |
| No repr.Probing=Frozen representation2026.06 | 12.5 |