Consistent Video Retrieval on CrossTask (test)
0.6436AccuracyCAST
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CASTBackbone=InternVideo2-1B, Setting=+ CAST2026.03 | 0.6436 | 0.0171 | |
| CASTBackbone=VideoPrism-B, Setting=+ CAST2026.03 | 0.6211 | 0.0174 | |
| CASTBackbone=GME-Qwen2-VL-2B, Setting=+ CAST2026.03 | 0.5243 | 0.0204 | |
| CASTBackbone=Qwen3-VL-Embedding-2B, Setting=+ CAST2026.03 | 0.4896 | 0.0209 | |
| CASTContext Modeling=State Transition, Backbone=CLIP-B/322026.03 | 0.4739 | 0.0214 | |
| Early FusionContext Modeling=Feature Concat., Backbone=CLIP-B/322026.03 | 0.3529 | 0.0236 | |
| Late Fusion (Learned)Context Modeling=Learned Weighting, Backbone=CLIP-B/322026.03 | 0.2552 | 0.0286 | |
| Late Fusion (Heuristic)Context Modeling=Fixed Weighting, Backbone=CLIP-B/322026.03 | 0.2205 | 0.0286 | |
| InternVideo2-1BBackbone=InternVideo2-1B, Setting=Zero-Shot2026.03 | 0.2061 | 0.0331 | |
| VideoPrism-BBackbone=VideoPrism-B, Setting=Zero-Shot2026.03 | 0.2025 | 0.0324 | |
| Qwen3-VL-Embedding-2BBackbone=Qwen3-VL-Embedding-2B, Setting=Zero-Shot2026.03 | 0.1944 | 0.0356 | |
| GME-Qwen2-VL-2BBackbone=GME-Qwen2-VL-2B, Setting=Zero-Shot2026.03 | 0.194 | 0.0361 | |
| CLIP BaselineContext Modeling=Context-Free, Backbone=CLIP-B/322026.03 | 0.1683 | 0.0415 |