Audio Retrieval on AudioCaps
52R@1VAST
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| VAST2023.05 | 52 | 82.9 | 76.8 | |
| VAST2023.05 | 52 | — | — | |
| MAViL2022.12 | 49.3 | 91.5 | 81.8 | |
| OmniBind2026.05 | 46.7 | 79.7 | — | |
| CLAP2023.10 | 44.1 | 87.67 | 76.8 | |
| CLAP2022.12 | 43.9 | 87.6 | 77.7 | |
| ASK*Interaction Strategy=Local, Knowledge Source=Original training set2025.12 | 43.7 | 86.2 | 75.8 | |
| ASK+Interaction Strategy=Local, Knowledge Source=WavCaps2025.12 | 43.1 | 86.9 | 74 | |
| ASK†Interaction Strategy=Local, Knowledge Source=Gemini-annotated training set2025.12 | 42.9 | 86.4 | 75.1 | |
| CLAPSsentence embedding training=unsupervised (Ls)2023.10 | 42.73 | 87.57 | 75.44 | |
| ONE-PIECE2023.05 | 42.5 | 88.4 | 77.5 | |
| HTSAT-BERTAudio encoder=HTSAT, Text encoder=BERT2023.05 | 42.2 | 87.1 | 76.5 | |
| Method [18]2023.05 | 42.2 | — | — | |
| SOTA2024.06 | 42.2 | — | — | |
| Xie et al., 2024Interaction Strategy=Local2025.12 | 41.1 | 85.2 | 73.8 | |
| MiCo2024.06 | 41 | — | — | |
| CyCLAP2023.10 | 40.65 | 86.52 | 74.19 | |
| VALOR-BModel size=Base2023.05 | 40.1 | 83.1 | 73.9 | |
| VALOR_B2023.04 | 40.1 | 83.1 | 73.9 | |
| CyCLAPSsentence embedding training=unsupervised (Ls)2023.10 | 39.81 | 85.79 | 74.4 | |
| MMTtrained without LAION-630K=true2022.12 | 39.6 | 86.7 | 76.8 | |
| ML-ACTtrained without LAION-630K=true2022.12 | 39.4 | 83.9 | 72 | |
| MAViL2022.12 | 37.3 | 84.5 | 72.8 | |
| MMTtrained without LAION-630K=true2022.12 | 36.1 | 84.5 | 72 | |
| LAION2023.05 | 36.1 | 83.9 | 71.8 | |
| Nagrani et al.2023.05 | 35.5 | 84.5 | — | |
| Nagrani et al.2023.04 | 35.5 | 84.5 | — | |
| CNN14-BERTAudio encoder=CNN14, Text encoder=BERT2023.05 | 35.1 | 82.1 | 70 | |
| CLAP2023.10 | 34.82 | 82.93 | 70.62 | |
| CLAPSsentence embedding training=unsupervised (Ls)2023.10 | 34.69 | 82.99 | 69.8 | |
| CyCLAPSsentence embedding training=unsupervised (Ls)2023.10 | 34.23 | 82.74 | 70.24 | |
| CyCLAP2023.10 | 34.13 | 82.3 | 69.24 | |
| ML-ACTtrained without LAION-630K=true2022.12 | 33.9 | 82.6 | 69.7 | |
| CLAP2022.12 | 32.7 | 81.2 | 68 | |
| FreeBind2026.05 | 29.2 | — | — | |
| Oncescu et al.2023.05 | 25.1 | 73.2 | — | |
| Oncescu et al.2023.04 | 25.1 | 73.2 | — | |
| ARNLQEmergent=false, Supervision=Supervised2023.05 | 24.3 | 72.1 | — | |
| LanguageBind2026.05 | 19.7 | 67.6 | — | |
| C-MCRAnchor Modality=–2026.04 | 15.76 | 48.1 | — | |
| CodeBind-VL2026.05 | 15.6 | 55 | — | |
| ViT-LENS_Lanchor=I+T2023.11 | 14.4 | 54.9 | — | |
| CodeBind-IB2026.05 | 13.3 | 53.8 | — | |
| ViT-LENS_Lanchor=I2023.11 | 12.2 | 48.7 | — | |
| LanguageBindAnchor Modality=Text2026.04 | 12.2 | 53.2 | — | |
| EmergentBridgeAnchor Modality=Text2026.04 | 11.8 | 52.1 | — | |
| UNIALIGN2026.05 | 11.7 | 49.3 | — | |
| EmergentBridgeAnchor Modality=Image, Evaluation Protocol=emergent zero-shot2026.04 | 10.1 | 49.2 | — | |
| ImageBind-Hanchor=I2023.11 | 9.3 | 42.3 | — | |
| IMAGEBINDEmergent=true, Supervision=No audio and text supervision2023.05 | 9.3 | 42.3 | — | |
| ImageBindAnchor Modality=Image, Evaluation Protocol=emergent zero-shot2026.04 | 9.3 | 42.3 | — | |
| AVFIC2023.11 | 8.7 | 37.7 | — | |
| AVFICEmergent=false, Supervision=Uses audio and text loss2023.05 | 8.7 | 37.7 | — | |
| AVFICAnchor Modality=–2026.04 | 8.7 | 37.7 | — | |
| AudioCLIPAnchor Modality=–2026.04 | 3.53 | 31.6 | — | |
| WAV2CLIPAnchor Modality=–2026.04 | 0.88 | 15.3 | — |