Long Video Understanding on LVU 53 (test)
68.2Place AccuracyMovies2Scenes
Evaluation Results
| Method | Links | |||||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Movies2Scenespre-train data=2.5M movie scene-pairs, modalities=visual, frames/scene=9, params=~0.7M2022.02 | 68.2 | 70.9 | 71.2 | 42.2 | 53.7 | 57.8 | 55.9 | 2.79 | 0.192 | |
| ViS4merpre-train data=30K LVU movie clips, modalities=visual, frames/scene=60, params=~3.6M2022.02 | 67.4 | 62.6 | 57.1 | 40.7 | 48.8 | 44.7 | 54.7 | 3.63 | 0.26 | |
| Bridge Formerpre-train data=3.3M image+2.5M video, modalities=vis.+text, frames/scene=4, params=~0.4M2022.02 | 62.8 | 55.7 | 60.4 | 40.9 | 49.7 | 41.4 | 52.9 | 3.97 | 0.312 | |
| Merlot Reservepre-train data=20M Youtube videos, modalities=vis.+text+aud., frames/scene=8, params=~0.7M2022.02 | 59.2 | 54.4 | 60 | 41.1 | 38.5 | 49.7 | 54.6 | 3.03 | 0.217 | |
| OTpre-train data=30K LVU movie clips, modalities=visual, frames/scene=60, params=~27M2022.02 | 56.9 | 51.2 | 53.1 | 39.4 | 34.5 | 39.1 | 54.6 | 3.55 | 0.23 | |
| Video Bertpre-train data=30K LVU movie clips, modalities=visual, frames/scene=60, params=~8.77M2022.02 | 54.9 | 47.3 | 52.8 | 37.9 | 38.5 | 36.1 | 51.9 | 4.46 | 0.32 | |
| SlowFast R101pre-train data=30K LVU movie clips, modalities=visual, frames/scene=60, params=~44M2022.02 | 54.7 | 44.9 | 52.4 | 35.8 | 36.3 | 52.5 | 53 | 3.77 | 0.386 | |
| CLIPpre-train data=400M image-text pairs, modalities=vis.+text, frames/scene=9, params=~0.7M2022.02 | 52.9 | 56.2 | 56.1 | 36.7 | 37.8 | 46.4 | 50.9 | 3.85 | 0.411 | |
| ShotCoLpre-train data=2.5M movie shot pairs, modalities=visual, frames/scene=9, params=~1.3M2022.02 | 45.3 | 49.5 | 46.7 | 31.1 | 29.8 | 36.7 | 43.1 | 4.79 | 0.397 | |
| Hierarchical OTpre-train data=240K Kinetics + 23.5K VidSitu, modalities=visual, frames/scene=16, params=~27M2022.02 | 44.1 | 40.1 | 50.9 | 34.1 | 31.4 | 29.6 | 51.1 | 4.88 | 0.353 |