Consistent Video Retrieval on YouCook2 official (val)
44.77AccuracyCAST
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CASTContext Modeling=State Transition, Backbone=CLIP-B/322026.03 | 44.77 | 2.15 | |
| Late Fusion (Learned)Context Modeling=Learned Weighting, Backbone=CLIP-B/322026.03 | 36.6 | 2.53 | |
| Early FusionContext Modeling=Feature Concat., Backbone=CLIP-B/322026.03 | 35.99 | 2.28 | |
| Late Fusion (Heuristic)Context Modeling=Fixed Weighting, Backbone=CLIP-B/322026.03 | 31.1 | 2.56 | |
| CLIP BaselineContext Modeling=Context-Free, Backbone=CLIP-B/322026.03 | 25.03 | 3.6 |