Consistent Video Retrieval on Diagnostic Average of YouCook2, COIN, CrossTask
53.81State AccuracyCAST
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| CASTContext Modeling=State Transition, Backbone=CLIP-B/322026.03 | 53.81 | 74.67 | |
| CLIP BaselineContext Modeling=Context-Free, Backbone=CLIP-B/322026.03 | 45.52 | 28.9 | |
| Late Fusion (Learned)Context Modeling=Learned Weighting, Backbone=CLIP-B/322026.03 | 40.06 | 76.06 | |
| Early FusionContext Modeling=Feature Concat., Backbone=CLIP-B/322026.03 | 31.14 | 83.59 | |
| Late Fusion (Heuristic)Context Modeling=Fixed Weighting, Backbone=CLIP-B/322026.03 | 28.69 | 68.29 |