Text-to-Audio Retrieval on Sound (test)
13.33Recall@1Speech-CLAP
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| Speech-CLAPbackbone=WavLM2023.12 | 13.33 | 51.88 | 67.64 | |
| CLAPvariant=with speech2023.12 | 11.15 | 42.42 | 60.36 | |
| CLAPvariant=general audio2023.12 | 11.03 | 45.33 | 63.64 |
| Method | Links | |||
|---|---|---|---|---|
| Speech-CLAPbackbone=WavLM2023.12 | 13.33 | 51.88 | 67.64 | |
| CLAPvariant=with speech2023.12 | 11.15 | 42.42 | 60.36 | |
| CLAPvariant=general audio2023.12 | 11.03 | 45.33 | 63.64 |