Vision-Language Compositional Reasoning on ARO
0.804AccuracyCE-CLIP+
Evaluation Results
| Method | Links | |
|---|---|---|
| CE-CLIP+Training Source=COCO+CC3M, Num Real Images=3M, Num Real Captions=3M, Num Synthetic Images=0, Num Synthetic Captions=15M2025.03 | 0.804 | |
| MosaiCLIPTraining Source=COCO, Num Real Images=109K, Num Real Captions=109K, Num Synthetic Images=0, Num Synthetic Captions=981K2025.03 | 0.803 | |
| CE-CLIPTraining Source=COCO, Num Real Images=82K, Num Real Captions=410K, Num Synthetic Images=0, Num Synthetic Captions=2M2025.03 | 0.797 | |
| AMR-NegCLIPTraining Source=COCO, Num Real Images=100K, Num Real Captions=100K, Num Synthetic Images=0, Num Synthetic Captions=500K2025.03 | 0.794 | |
| SPARCLTraining Source=COCO, Num Real Images=82K, Num Real Captions=410K, Num Synthetic Images=820K, Num Synthetic Captions=820K2025.03 | 0.772 | |
| NegCLIPTraining Source=COCO, Num Real Images=100K, Num Real Captions=100K, Num Synthetic Images=0, Num Synthetic Captions=500K2025.03 | 0.76 | |
| CLOVETraining Source=LAION-COCO, Num Real Images=>1B, Num Real Captions=>1B, Num Synthetic Images=0, Num Synthetic Captions=>1B2025.03 | 0.732 | |
| SPECTraining Source=LAION, Num Real Images=20K, Num Real Captions=20K, Num Synthetic Images=20K, Num Synthetic Captions=20K2025.03 | 0.701 | |
| syn-CLIPTraining Source=SyViC, Num Real Images=0, Num Real Captions=0, Num Synthetic Images=>1M, Num Synthetic Captions=>1M2025.03 | 0.692 | |
| FIGCLIPTraining Source=VidSitu, Num Real Images=20K videos, Num Real Captions=0, Num Synthetic Images=0, Num Synthetic Captions=02025.03 | 0.67 | |
| [79]Training Source=COCO, Num Real Images=0, Num Real Captions=0, Num Synthetic Images=82K, Num Synthetic Captions=82K2025.03 | 0.65 | |
| CLIPTraining Source=COCO, Num Real Images=82K, Num Real Captions=410K, Num Synthetic Images=0, Num Synthetic Captions=0, Training Protocol=Finetune2025.03 | 0.641 | |
| CLIPNum Real Images=0, Num Real Captions=0, Num Synthetic Images=0, Num Synthetic Captions=0, Training Protocol=Zero-Shot2025.03 | 0.611 | |
| SDS-CLIPTraining Source=COCO, Num Real Images=82K, Num Real Captions=410K, Num Synthetic Images=0, Num Synthetic Captions=02025.03 | 0.575 |