Visual Sound Source Localization on Flickr-SoundNet extended (test)
86LocAccSLAVC
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SLAVCTraining Dataset=VGG-Sound 144k, OGL=true2022.08 | 86 | 52.15 | 60.1 | |
| SOUPLEOGL (Object-Guided Learning)=Without, Training Epochs=202026.03 | 84.8 | 80.25 | 83.11 | |
| FNACOGL (Object-Guided Learning)=Without2026.03 | 84.73 | 50.4 | 62.3 | |
| MarginNCEOGL (Object-Guided Learning)=Without2026.03 | 83.94 | 57.99 | 61.8 | |
| SLAVCTraining Dataset=VGG-Sound 144k, OGL=false2022.08 | 83.6 | 51.63 | 59.1 | |
| SLAVCOGL (Object-Guided Learning)=Without2026.03 | 83.6 | 51.63 | 59.1 | |
| SOUPLEOGL (Object-Guided Learning)=Without, Training Epochs=502026.03 | 83.6 | 79.89 | 82.96 | |
| ACL-SSLOGL (Object-Guided Learning)=Without2026.03 | 80.8 | 76.07 | 73.2 | |
| AlignmentOGL (Object-Guided Learning)=Without2026.03 | 79.6 | 64.43 | 66.9 | |
| OGLTraining Dataset=VGG-Sound 144k2022.08 | 77.2 | 40.2 | 55.7 | |
| DSOLTraining Dataset=VGG-Sound 144k2022.08 | 72.91 | 38.32 | 49.4 | |
| EZ-VSLTraining Dataset=VGG-Sound 144k, OGL=true2022.08 | 72.8 | 48.75 | 56.8 | |
| Center PriorTraining Dataset=VGG-Sound 144k2022.08 | 67.6 | — | — | |
| EZ-VSLTraining Dataset=VGG-Sound 144k, OGL=false2022.08 | 66.4 | 46.3 | 54.6 | |
| DMCTraining Dataset=VGG-Sound 144k2022.08 | 52.8 | 25.56 | 41.8 | |
| CoarsetoFineTraining Dataset=VGG-Sound 144k2022.08 | 47.2 | 0 | 38.2 | |
| AudioCLIP2026.03 | 45.22 | 34 | 38.8 | |
| Attention10kTraining Dataset=VGG-Sound 144k2022.08 | 34.16 | 15.98 | 24 | |
| WAV2CLIP2026.03 | 29.6 | 20.99 | 24.8 | |
| LVSTraining Dataset=VGG-Sound 144k2022.08 | 19.6 | 9.8 | 17.9 |