Text-to-Image Retrieval on Cola (test)
83.88Multi-Obj AccHuman
Evaluation Results
| Method | Links | |
|---|---|---|
| Humantype=human performance2023.05 | 83.88 | |
| MM-PredBase Model=CLIP, Adaptation Strategy=MM-Pred (multimodal head)2023.05 | 41.42 | |
| MM-AdapterBase Model=CLIP, Adaptation Strategy=MM-Adapter (multimodal feature adapter)2023.05 | 40.95 | |
| MM-AdapterBase Model=FLAVA, Adaptation Strategy=MM-Adapter (multimodal feature adapter)2023.05 | 40.47 | |
| MM-PredBase Model=FLAVA, Adaptation Strategy=MM-Pred (multimodal head)2023.05 | 39.04 | |
| CLIP + FT lateBase Model=CLIP, Adaptation Strategy=Late Fine-tuning2023.05 | 36.19 | |
| CLIP + FT allBase Model=CLIP, Adaptation Strategy=Full Fine-tuning2023.05 | 34.76 | |
| CLIP + LinearBase Model=CLIP, Adaptation Strategy=Linear2023.05 | 30.47 | |
| CLIP + Prompt-tuneBase Model=CLIP, Adaptation Strategy=Prompt-tuning2023.05 | 27.14 | |
| Randomtype=baseline2023.05 | 25 | |
| FLAVABase Model=FLAVA2023.05 | 24.76 | |
| FLAVA + LinearBase Model=FLAVA, Adaptation Strategy=Linear2023.05 | 22.38 | |
| FLAVA + FT lateBase Model=FLAVA, Adaptation Strategy=Late Fine-tuning2023.05 | 22.38 | |
| CLIPBase Model=CLIP2023.05 | 21.42 |