Text-to-Video Retrieval on Something-Something CiA-Retrieval v2
85.1mAP (Chiral)TARA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TARAMethod category=MLLMs with fine-tuning, Training dataset size (M)=0.01, Video used during retrieval adaptation=false2025.12 | 85.1 | 47.8 | |
| Qwen2.5VL-7BMethod category=MLLMs zero-shot2025.12 | 67.6 | 20.6 | |
| ArrowRLMethod category=MLLMs with fine-tuning, Training dataset size (M)=0.02, Video used during retrieval adaptation=true2025.12 | 67.5 | 22.5 | |
| CaReMethod category=MLLMs with fine-tuning, Training dataset size (M)=0.3, Video used during retrieval adaptation=false2025.12 | 66.4 | 23.7 | |
| Qwen2VL-7BMethod category=MLLMs zero-shot2025.12 | 60.2 | 17.3 | |
| VLM2Vec-V2Method category=MLLMs with fine-tuning, Training dataset size (M)=1.7, Video used during retrieval adaptation=true2025.12 | 58.8 | 15.9 | |
| LAMRAMethod category=MLLMs with fine-tuning, Training dataset size (M)=1.4, Video used during retrieval adaptation=false2025.12 | 55.3 | 7.8 | |
| XCLIPMethod category=Dual encoder models2025.12 | 54.7 | 16.8 | |
| GVE-7BMethod category=MLLMs with fine-tuning, Training dataset size (M)=13.0, Video used during retrieval adaptation=true2025.12 | 53.4 | 4 | |
| E5-VMethod category=MLLMs with fine-tuning, Training dataset size (M)=0.3, Video used during retrieval adaptation=false2025.12 | 52.6 | 14.7 | |
| InternVideo 2Method category=Dual encoder models2025.12 | 52.5 | 20.6 | |
| DINO.txtMethod category=Dual encoder models2025.12 | 52.1 | 13.1 | |
| CLIP (avg.)Method category=Dual encoder models2025.12 | 52 | 12.7 | |
| ViCLIPMethod category=Dual encoder models2025.12 | 50.8 | 16.2 | |
| Perception Enc.Method category=Dual encoder models2025.12 | 50.1 | 17.2 | |
| Chance2025.12 | 50 | 3.1 |