Text-to-Image Retrieval on diagnostic (test-unseen)
80.23Acc@50Human
Evaluation Results
| Method | Links | |
|---|---|---|
| Human2022.11 | 80.23 | |
| FLAVAModel Category=Matching, Backbone=ViT-B2022.11 | 40.93 | |
| ViLTModel Category=Matching, Backbone=ViT-B2022.11 | 40.77 | |
| BLIPModel Category=Both, Backbone=ViT-L, Objective Variant=itm2022.11 | 40.19 | |
| BLIPModel Category=Both, Backbone=ViT-L, Objective Variant=itc2022.11 | 40.07 | |
| CLIPModel Category=Contrastive, Backbone=ViT-L2022.11 | 39.49 | |
| ALBEFModel Category=Both, Backbone=ViT-B2022.11 | 37.92 | |
| OwlVitModel Category=Contrastive, Backbone=ViT-L2022.11 | 36.44 | |
| Random2022.11 | 23.8 |