Audio-to-Text Retrieval on Sound (test)
11.27R@1Speech-CLAP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Speech-CLAPbackbone=WavLM2023.12 | 11.27 | 47.27 | 64.48 | |
| CLAPvariant=with speech2023.12 | 9.7 | 43.15 | 59.03 | |
| CLAPvariant=general audio2023.12 | 9.45 | 44.36 | 61.7 |
| Method | Links | |||
|---|---|---|---|---|
| Speech-CLAPbackbone=WavLM2023.12 | 11.27 | 47.27 | 64.48 | |
| CLAPvariant=with speech2023.12 | 9.7 | 43.15 | 59.03 | |
| CLAPvariant=general audio2023.12 | 9.45 | 44.36 | 61.7 |