Image-to-Text Retrieval on diagnostic (test-unseen)
85.21Accuracy@50Human
Evaluation Results
| Method | Links | |
|---|---|---|
| Human2022.11 | 85.21 | |
| BLIPModel Category=Both, Backbone=ViT-L, Objective Variant=itm2022.11 | 41.94 | |
| CLIPModel Category=Contrastive, Backbone=ViT-L2022.11 | 39.61 | |
| FLAVAModel Category=Matching, Backbone=ViT-B2022.11 | 38.43 | |
| ALBEFModel Category=Both, Backbone=ViT-B2022.11 | 38.32 | |
| ViLTModel Category=Matching, Backbone=ViT-B2022.11 | 35.34 | |
| OwlVitModel Category=Contrastive, Backbone=ViT-L2022.11 | 32.3 | |
| BLIPModel Category=Both, Backbone=ViT-L, Objective Variant=itc2022.11 | 31 | |
| Random2022.11 | 23.62 |