Image-to-Text Retrieval on diagnostic seen (test)
84.97Acc@50Human
Evaluation Results
| Method | Links | |
|---|---|---|
| Human2022.11 | 84.97 | |
| BLIPModel Category=Both, Backbone=ViT-L, Objective Variant=itm2022.11 | 40.17 | |
| FLAVAModel Category=Matching, Backbone=ViT-B2022.11 | 38.5 | |
| CLIPModel Category=Contrastive, Backbone=ViT-L2022.11 | 38.17 | |
| ALBEFModel Category=Both, Backbone=ViT-B2022.11 | 37.49 | |
| OwlVitModel Category=Contrastive, Backbone=ViT-L2022.11 | 33.25 | |
| ViLTModel Category=Matching, Backbone=ViT-B2022.11 | 32.17 | |
| BLIPModel Category=Both, Backbone=ViT-L, Objective Variant=itc2022.11 | 31.67 | |
| Random2022.11 | 24 |