Caption Matching and Retrieval on NoCaps (val)
99.5Matching AccuracyCosine Similarity
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| Cosine SimilarityVision Model=CLIP2024.01 | 99.5 | 99.6 | |
| Cosine Similarity*Vision Model=CLIP [40]2024.01 | 99.5 | 99.6 | |
| QAPVision Model=CLIP-V2024.01 | 67.3 | — | |
| QAPVision Model=CLIP-V [40]2024.01 | 67.3 | — | |
| Local CKAVision Model=CLIP-V2024.01 | 65.1 | 60.5 | |
| Local CKAVision Model=CLIP-V [40]2024.01 | 65.1 | 65.9 | |
| Linear regressionVision Model=CLIP-V [40]2024.01 | 63.6 | 70.1 | |
| Relative representationsVision Model=CLIP-V2024.01 | 61.3 | 37.6 | |
| Relative representationsVision Model=CLIP-V [40]2024.01 | 61.3 | 3 | |
| Local CKAVision Model=DINOv22024.01 | 58.7 | 61.8 | |
| QAPVision Model=DINOv2 [37]2024.01 | 58.5 | — | |
| QAPVision Model=DINOv22024.01 | 57.7 | — | |
| Local CKAVision Model=DINOv2 [37]2024.01 | 55.7 | 64.2 | |
| Linear regressionVision Model=DINOv2 [37]2024.01 | 46.8 | 59.9 | |
| QAPVision Model=ConvNeXt2024.01 | 46.7 | — | |
| Relative representationsVision Model=DINOv22024.01 | 46 | 46.4 | |
| Relative representationsVision Model=DINOv2 [37]2024.01 | 45.9 | 38.1 | |
| QAPVision Model=ConvNeXt [47]2024.01 | 45.9 | — | |
| Local CKAVision Model=ConvNeXt [47]2024.01 | 44.8 | 33 | |
| Local CKAVision Model=ConvNeXt2024.01 | 43.7 | 44.4 | |
| Linear regressionVision Model=DINOv22024.01 | 38.1 | 50.3 | |
| Linear regressionVision Model=CLIP-V2024.01 | 29.3 | 44.7 | |
| Relative representationsVision Model=ConvNeXt2024.01 | 25.5 | 17.8 | |
| Relative representationsVision Model=ConvNeXt [47]2024.01 | 25.5 | 2.7 | |
| Linear regressionVision Model=ConvNeXt [47]2024.01 | 22.8 | 38.9 | |
| Linear regressionVision Model=ConvNeXt2024.01 | 19 | 28.5 |