Consistent Video Retrieval on COIN (test)
51.64AccuracyCAST
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CASTBackbone=VideoPrism-B, Setting=+ CAST2026.03 | 51.64 | 1.9 | |
| CASTBackbone=InternVideo2-1B, Setting=+ CAST2026.03 | 51.03 | 1.9 | |
| CASTBackbone=GME-Qwen2-VL-2B, Setting=+ CAST2026.03 | 45.68 | 2.05 | |
| CASTBackbone=Qwen3-VL-Embedding-2B, Setting=+ CAST2026.03 | 44.87 | 2.09 | |
| Late Fusion (Learned)Context Modeling=Learned Weighting, Backbone=CLIP-B/322026.03 | 44.66 | 2.11 | |
| CASTContext Modeling=State Transition, Backbone=CLIP-B/322026.03 | 40.47 | 2.16 | |
| InternVideo2-1BBackbone=InternVideo2-1B, Setting=Zero-Shot2026.03 | 17.99 | 3.36 | |
| Late Fusion (Heuristic)Context Modeling=Fixed Weighting, Backbone=CLIP-B/322026.03 | 17.85 | 3.28 | |
| Qwen3-VL-Embedding-2BBackbone=Qwen3-VL-Embedding-2B, Setting=Zero-Shot2026.03 | 17.73 | 3.5 | |
| VideoPrism-BBackbone=VideoPrism-B, Setting=Zero-Shot2026.03 | 17.6 | 3.32 | |
| GME-Qwen2-VL-2BBackbone=GME-Qwen2-VL-2B, Setting=Zero-Shot2026.03 | 17.17 | 3.44 | |
| Early FusionContext Modeling=Feature Concat., Backbone=CLIP-B/322026.03 | 15.12 | 2.6 | |
| CLIP BaselineContext Modeling=Context-Free, Backbone=CLIP-B/322026.03 | 14.1 | 3.91 |