Text-to-Image Retrieval on MS-COCO fine-tuned
66.7R@1BLIP-2 + MAFA
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| BLIP-2 + MAFAModel=BLIP-2 + MAFA, Evaluation protocol=fine-tuned2023.12 | 66.7 | 86.8 | 91.9 | |
| BLIP-2Backbone=ViT-L/16, Training data=400M, Storage per image=2,359 kB, Joint encoding parameters=167M, Inference time (ms)=98.642025.10 | 66.3 | — | — | |
| BLIP-2Model=BLIP-2, Evaluation protocol=fine-tuned2023.12 | 66.1 | 86.8 | 92 | |
| BLIPBackbone=ViT-L/16, Training data=129M, Storage per image=2,359 kB, Joint encoding parameters=139M, Inference time (ms)=101.612025.10 | 65.1 | — | — | |
| Local (EDJE)Backbone=ViT-L/16, Training data=12M, Storage per image=442kB, Joint encoding parameters=33M, Inference time (ms)=4.142025.10 | 64.9 | — | — | |
| Compressed-128 (EDJE)Backbone=ViT-L/16, Training data=12M, Storage per image=98kB, Joint encoding parameters=33M, Inference time (ms)=2.042025.10 | 64.6 | — | — | |
| Compressed-64 (EDJE)Backbone=ViT-L/16, Training data=12M, Storage per image=49kB, Joint encoding parameters=33M, Inference time (ms)=1.912025.10 | 64.6 | — | — | |
| BLIPBackbone=ViT-B/16, Training data=12M, Storage per image=1,769 kB, Joint encoding parameters=139M, Inference time (ms)=83.272025.10 | 63.1 | — | — | |
| Local (EDJE)Backbone=ViT-B/16, Training data=12M, Storage per image=442kB, Joint encoding parameters=33M, Inference time (ms)=2.862025.10 | 60.9 | — | — | |
| ALBEFBackbone=ViT-B/16, Training data=12M, Storage per image=1,769 kB, Joint encoding parameters=147M, Inference time (ms)=45.922025.10 | 60.7 | — | — | |
| BLIP-2 + GRITModel=BLIP-2 + GRIT, Evaluation protocol=fine-tuned2023.12 | 52.5 | 79.1 | 87.3 |