Cross-modal Retrieval on InstVL img 1K Instance
50.25T2V R@1InstAP
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InstAP2026.04 | 50.25 | 49.26 | |
| UMT-L (InstVL; g+i)training_corpus=all InstVL captions treated as global2026.04 | 45.74 | 44.27 | |
| UMT-L2026.04 | 38.44 | 35.65 | |
| SigLIP2026.04 | 38.17 | 45.17 | |
| OpenCLIP2026.04 | 37.88 | 44.06 | |
| UMT-L (InstVL; g)training_corpus=InstVL global captions2026.04 | 34.44 | 41.24 | |
| ViCLIP2026.04 | 28.38 | 28.91 | |
| VideoPrism2026.04 | 28.21 | 34.52 | |
| CLIP4Clip2026.04 | 25.1 | 33.21 | |
| CLIP-ViP2026.04 | 24.04 | 32.06 | |
| MCQ2026.04 | 19.33 | 22.11 | |
| Coca2026.04 | 11.83 | 21.79 |