Cross-modal Vision-Language Retrieval on Pitts30
50.2R@1La-BLIP
Evaluation Results
| Method | Links | ||||
|---|---|---|---|---|---|
| La-BLIPAugmentation=Language-augmented (La-)2026.02 | 50.2 | 74.2 | 82.8 | 89.6 | |
| La-CLIPAugmentation=Language-augmented (La-)2026.02 | 49.3 | 74.4 | 82.8 | 89.7 | |
| La-SigLIP-V2Augmentation=Language-augmented (La-)2026.02 | 49.3 | 74.6 | 82.8 | 89.2 | |
| La-EVA-V2Augmentation=Language-augmented (La-)2026.02 | 41.5 | 66.8 | 76.6 | 85.1 | |
| EVA-CLIP-V2Augmentation=None2026.02 | 11.4 | 30.9 | 44.2 | 59.9 | |
| CLIPAugmentation=None2026.02 | 10.9 | 29.6 | 43.5 | 59.6 | |
| SigLIP-V2Augmentation=None2026.02 | 10.1 | 30.1 | 44.3 | 61.9 | |
| BLIPAugmentation=None2026.02 | 8.1 | 26.8 | 41.1 | 56.7 |