Text-to-Image Retrieval on diagnostic (test-seen)
82.02Accuracy@50Human
Evaluation Results
| Method | Links | |
|---|---|---|
| Human2022.11 | 82.02 | |
| FLAVAModel Category=Matching, Backbone=ViT-B2022.11 | 41.44 | |
| ViLTModel Category=Matching, Backbone=ViT-B2022.11 | 40.98 | |
| BLIPModel Category=Both, Backbone=ViT-L, Objective Variant=itm2022.11 | 40.3 | |
| CLIPModel Category=Contrastive, Backbone=ViT-L2022.11 | 39.51 | |
| ALBEFModel Category=Both, Backbone=ViT-B2022.11 | 39.01 | |
| BLIPModel Category=Both, Backbone=ViT-L, Objective Variant=itc2022.11 | 38.35 | |
| OwlVitModel Category=Contrastive, Backbone=ViT-L2022.11 | 36.73 | |
| Random2022.11 | 23.81 |