Video-to-Text Retrieval on Something-Something CiA-Retrieval v2
84R@1 (Chiral)TARA
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| TARAMethod category=MLLMs with fine-tuning, Training dataset size (M)=0.01, Video used during retrieval adaptation=false2025.12 | 84 | 32.8 | |
| ArrowRLMethod category=MLLMs with fine-tuning, Training dataset size (M)=0.02, Video used during retrieval adaptation=true2025.12 | 66.4 | 14.3 | |
| CaReMethod category=MLLMs with fine-tuning, Training dataset size (M)=0.3, Video used during retrieval adaptation=false2025.12 | 63.9 | 23.8 | |
| Qwen2.5VL-7BMethod category=MLLMs zero-shot2025.12 | 63.7 | 12.6 | |
| Qwen2VL-7BMethod category=MLLMs zero-shot2025.12 | 58.4 | 9 | |
| VLM2Vec-V2Method category=MLLMs with fine-tuning, Training dataset size (M)=1.7, Video used during retrieval adaptation=true2025.12 | 55.8 | 8.9 | |
| LAMRAMethod category=MLLMs with fine-tuning, Training dataset size (M)=1.4, Video used during retrieval adaptation=false2025.12 | 54.2 | 10.3 | |
| GVE-7BMethod category=MLLMs with fine-tuning, Training dataset size (M)=13.0, Video used during retrieval adaptation=true2025.12 | 52.5 | 7.3 | |
| XCLIPMethod category=Dual encoder models2025.12 | 52.4 | 5.3 | |
| DINO.txtMethod category=Dual encoder models2025.12 | 52.3 | 5.5 | |
| CLIP (avg.)Method category=Dual encoder models2025.12 | 52.1 | 5.9 | |
| Perception Enc.Method category=Dual encoder models2025.12 | 51.8 | 7.4 | |
| InternVideo 2Method category=Dual encoder models2025.12 | 51.6 | 10.9 | |
| ViCLIPMethod category=Dual encoder models2025.12 | 51.4 | 6.2 | |
| E5-VMethod category=MLLMs with fine-tuning, Training dataset size (M)=0.3, Video used during retrieval adaptation=false2025.12 | 51.2 | 5.3 | |
| Chance2025.12 | 50 | 3.1 |