Video Representation Generalization on SEVERE benchmark
72.8Domain Shift (SSv2)TrackMAE
Evaluation Results
| Method | Links | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| TrackMAEReconstruction Mode=Feature2026.03 | 72.8 | 91.1 | — | — | 59 | 90.1 | 17 | 30.5 | — | 86.7 | 34.4 | |
| SMILEBackbone=ViT-B, Pre-training dataset=K4002025.04 | 72.1 | 90.8 | 86.4 | 35.1 | 55.1 | 88.3 | 0.17 | 32.5 | 67.9 | — | — | |
| SMILEReconstruction Mode=Feature2026.03 | 72.1 | 90.8 | — | — | 55.1 | 88.3 | 17 | 32.5 | — | 86.4 | 35.1 | |
| TrackMAE w/o motionReconstruction Mode=Feature2026.03 | 71.9 | 90 | — | — | 55.1 | 74.6 | 17 | 27.1 | — | 85.5 | 31.4 | |
| SMILE w/o motionBackbone=ViT-B, Pre-training dataset=K4002025.04 | 71.6 | 90 | 85.2 | 32 | 53.4 | 74.3 | 0.175 | 30.5 | 64.9 | — | — | |
| MGMBackbone=ViT-B, Pre-training dataset=K4002025.04 | 71.1 | 89.1 | 78.4 | 26.4 | 38.6 | 86.9 | 0.152 | 22.5 | 62.2 | — | — | |
| MGMReconstruction Mode=Pixel2026.03 | 71.1 | 89.1 | — | — | 38.6 | 86.9 | 15.2 | 22.5 | — | 78.4 | 26.4 | |
| SIGMABackbone=ViT-B, Pre-training dataset=K4002025.04 | 70.9 | 89.7 | 84.1 | 28 | 55.1 | 79.9 | 0.169 | 23.1 | 64.2 | — | — | |
| SIGMAReconstruction Mode=Feature2026.03 | 70.9 | 89.7 | — | — | 55.1 | 79.9 | 16.9 | 23.1 | — | 84.1 | 28 | |
| TrackMAEReconstruction Mode=Pixel2026.03 | 70.3 | 88.7 | — | — | 41.6 | 85.5 | 16.2 | 20.8 | — | 79.8 | 31 | |
| MMEBackbone=ViT-B, Pre-training dataset=K4002025.04 | 70.1 | 89.7 | 79.2 | 29.8 | 55.5 | 87.2 | 0.155 | 23.6 | 65 | — | — | |
| MMEReconstruction Mode=Feature2026.03 | 70.1 | 89.7 | — | — | 55.5 | 87.2 | 15.5 | 23.6 | — | 79.2 | 29.8 | |
| MVDBackbone=ViT-B, Pre-training dataset=K4002025.04 | 70 | 82.5 | 67.1 | 17.5 | 31.3 | 50.5 | 0.184 | 16.1 | 52.1 | — | — | |
| MVDReconstruction Mode=Pixel2026.03 | 70 | 82.5 | — | — | 31.3 | 50.5 | 18.4 | 16.1 | — | 67.1 | 17.5 | |
| MGMAEBackbone=ViT-B, Pre-training dataset=K4002025.04 | 68.9 | 87.2 | 77.2 | 24.1 | 33.7 | 79.5 | 0.181 | 17.9 | 58.8 | — | — | |
| MGMAEReconstruction Mode=Pixel2026.03 | 68.9 | 87.2 | — | — | 33.7 | 79.5 | 18.1 | 17.9 | — | 77.2 | 24.1 | |
| VideoMAEBackbone=ViT-B, Pre-training dataset=K4002025.04 | 68.6 | 86.6 | 74.6 | 25.9 | 42.8 | 65.3 | 0.172 | 14.4 | 57.6 | — | — | |
| VideoMAEReconstruction Mode=Pixel2026.03 | 68.6 | 86.6 | — | — | 42.8 | 65.3 | 17.2 | 17.8 | — | 74.6 | 25.9 |