Audio Classification on FSD50k
46.3ScoreEnsemble (Submission)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Ensemble (Submission)Ensemble Strategy=Embedding Concatenation2026.01 | 46.3 | — | — | |
| Dasheng 1.2BParameter Count=1.2B2026.01 | 45.5 | — | — | |
| BEATs (300M) SpeechParameter Count=300M, Pre-training Mixture=Speech-heavy (70:15:15)2026.01 | 43.2 | — | — | |
| Challenge BaselineDescription=Best among Dasheng-base, data2vec, and Whisper2026.01 | 40.8 | — | — | |
| BEATs (300M) BalancedParameter Count=300M, Pre-training Mixture=Balanced (40:30:30)2026.01 | 38 | — | — | |
| WQ-FusionEncoder Type=Fusion2026.06 | 29.5 | — | — | |
| Concat.Encoder Type=Fusion2026.06 | 29.3 | — | — | |
| Gated Trans.Encoder Type=Fusion2026.06 | 27.8 | — | — | |
| Adapt. and Trans.Encoder Type=Fusion2026.06 | 25.8 | — | — | |
| Qwen2-Audio-7BEncoder Type=Single2026.06 | 25.2 | — | — | |
| BEATs (90M) iter3Parameter Count=90M, Iteration=32026.01 | 21.7 | — | — | |
| Whisper-LargeEncoder Type=Single2026.06 | 17.3 | — | — | |
| AudioMAEEncoder Type=Single2026.06 | 14.3 | — | — | |
| Dasheng-BaseEncoder Type=Single2026.06 | 6.3 | — | — | |
| Audio Flamingo2024.02 | — | 69.7 | — | |
| CAV-MAE-Scale++training_mode=fine-tuning2024.03 | — | — | 45.5 | |
| Deshmukh et al., 20232024.02 | — | 65.6 | — | |
| EquiAVtraining_mode=fine-tuning2024.03 | — | — | 62.6 | |
| F3-TokenizerProbing=Frozen representation2026.06 | — | — | 32.1 | |
| Ming-UProbing=Frozen representation2026.06 | — | — | 24.83 | |
| No repr.Probing=Frozen representation2026.06 | — | — | 2.6 | |
| w/o LLMProbing=Frozen representation2026.06 | — | — | 28.7 | |
| w/o RQProbing=Frozen representation2026.06 | — | — | 21.2 | |
| WhisperProbing=Frozen representation2026.06 | — | — | 26.2 | |
| XKDtraining_mode=fine-tuning2024.03 | — | — | 58.5 |