Acoustic Event Detection on AudioSet (test)
0.462mAPAttention AV-fusion
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Attention AV-fusionModality=A+V2021.03 | 0.462 | — | |
| Perceiver (mel spectrogram - tuned)Input=mel spectrogram, Modality=A+V, Tuned=true2021.03 | 0.442 | — | |
| Perceiver (raw audio)Input=raw audio, Modality=A+V2021.03 | 0.435 | — | |
| Perceiver (mel spectrogram)Input=mel spectrogram, Modality=A+V2021.03 | 0.432 | — | |
| CNN-14Modality=Audio2021.03 | 0.431 | — | |
| G-BlendBackbone depth=101, Modality=Audio + Visual2019.05 | 0.418 | 0.975 | |
| G-blendModality=A+V2021.03 | 0.418 | — | |
| Naive A/VBackbone depth=101, Modality=Audio + Visual2019.05 | 0.402 | 0.973 | |
| SUSTAINTeachers=2 Teachers2020.06 | 0.398 | 0.972 | |
| SUSTAINTeachers=Single Teacher2020.06 | 0.394 | 0.972 | |
| Attention AV-fusionModality=Audio2021.03 | 0.384 | — | |
| Perceiver (mel spectrogram)Input=mel spectrogram, Modality=Audio2021.03 | 0.384 | — | |
| Perceiver (raw audio)Input=raw audio, Modality=Audio2021.03 | 0.383 | — | |
| ResNet-101Pooling=Attention2020.06 | 0.38 | 0.97 | |
| ResNet-50Modality=Audio2021.03 | 0.38 | — | |
| CNN-14 (no balancing & no mixup)Modality=Audio, Balancing=false, Mixup=false2021.03 | 0.375 | — | |
| Kong et al.Scale=Large2020.06 | 0.369 | 0.969 | |
| WEANET2020.06 | 0.366 | 0.958 | |
| TAL-Net2019.05 | 0.362 | 0.965 | |
| TALNetPooling=exp. pooling2020.06 | 0.362 | 0.965 | |
| Kong et al.Scale=Small2020.06 | 0.361 | 0.969 | |
| Multi-level Attn.2019.05 | 0.36 | 0.97 | |
| ResNet-34Pooling=Attention2020.06 | 0.36 | 0.966 | |
| Multi-level AttentionModality=Audio2021.03 | 0.36 | — | |
| TALNetPooling=Attention2020.06 | 0.354 | 0.963 | |
| AttentionModality=Audio2021.03 | 0.327 | — | |
| Audio: R2DBackbone depth=101, Modality=Audio only2019.05 | 0.324 | 0.961 | |
| G-blendModality=Audio2021.03 | 0.324 | — | |
| BenchmarkModality=Audio2021.03 | 0.314 | — | |
| Perceiver (raw audio)Input=raw audio, Modality=Video2021.03 | 0.258 | — | |
| Perceiver (mel spectrogram)Input=mel spectrogram, Modality=Video2021.03 | 0.258 | — | |
| Attention AV-fusionModality=Video2021.03 | 0.257 | — | |
| Visual: R(2+1)DBackbone depth=101, Modality=Visual only2019.05 | 0.188 | 0.918 | |
| G-blendModality=Video2021.03 | 0.188 | — |