Video Classification on Something-Something v2 (val)
73.3Top-1 AccMViTv2-L(↑312)
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| MViTv2-L(↑312)Protocol=full fine-tuning, Pretrain=K400†, Views=40×3×1, GFLOPs=8484, Param(M)=2132023.10 | 73.3 | 94.1 | — | |
| Gen4U w data augmPre-training=Diffusion, Data augmentation=true, Evaluation protocol=Frozen representations with trained one-block attention read-out2026.07 | 72.6 | — | — | |
| ZeroI2V ViT-L/14Pretrain=CLIP, Views=32×3×1, GFLOPs=7783, Extra GFLOPs=0, Param(M)=304, New Param(M)=02023.10 | 72.2 | 93 | — | |
| V-JEPA-HPre-training=Masked feature prediction, Model size (M)=635, Evaluation protocol=Frozen representations with trained one-block attention read-out2026.07 | 72.2 | — | — | |
| ZeroI2V ViT-L/14Pretrain=CLIP, Views=16×3×1, GFLOPs=3892, Extra GFLOPs=0, Param(M)=304, New Param(M)=02023.10 | 71.4 | 93 | — | |
| Gen4UPre-training=Diffusion, Data augmentation=false, Evaluation protocol=Frozen representations with trained one-block attention read-out2026.07 | 71.3 | — | — | |
| UniFormer-BProtocol=full fine-tuning, Pretrain=K600†, Views=32×3×1, GFLOPs=777, Param(M)=502023.10 | 71.2 | 92.8 | — | |
| AIM VIT-L/14Protocol=PETL, Pretrain=CLIP, Views=32×3×1, GFLOPs=11508, Extra GFLOPs=3725, Param(M)=354, New Param(M)=502023.10 | 70.6 | 92.7 | — | |
| ZeroI2V ViT-B/16Pretrain=CLIP, Views=32×3×1, GFLOPs=1688, Extra GFLOPs=0, Param(M)=86, New Param(M)=02023.10 | 70.1 | 92.4 | — | |
| ZeroI2V ViT-L/14Pretrain=CLIP, Views=8×3×1, GFLOPs=1946, Extra GFLOPs=0, Param(M)=304, New Param(M)=02023.10 | 70.1 | 91.8 | — | |
| Swin-B w/ STDHAPretrain=IN21K, Views=32×3×1, GFLOPs=741, Param(M)=89, Tunable Param(M)=892023.10 | 70 | 92.1 | — | |
| Swin-BPretrain=Kinetics-400, Views=1 x 3, FLOPs=321, Param=88.82021.06 | 69.6 | 92.7 | — | |
| VideoSwin-BPretrain=K400†, Views=32×3×1, GFLOPs=963, Param(M)=89, Tunable Param(M)=892023.10 | 69.6 | 92.7 | — | |
| VideoSwin-BProtocol=full fine-tuning, Pretrain=K400†, Views=32×3×1, GFLOPs=963, Param(M)=892023.10 | 69.6 | 92.7 | — | |
| ST-Adapter ViT-B/16Protocol=PETL, Pretrain=CLIP, Views=32×3×1, GFLOPs=1955, Extra GFLOPs=267, Param(M)=100, New Param(M)=142023.10 | 69.5 | 92.6 | — | |
| ZeroI2V ViT-B/16Pretrain=CLIP, Views=16×3×1, GFLOPs=844, Extra GFLOPs=0, Param(M)=86, New Param(M)=02023.10 | 69.4 | 91.7 | — | |
| MViT-B-24, 32×3Pretrain=Kinetics-600, Views=1 x 3, FLOPs=236, Param=53.22021.06 | 68.7 | 91.5 | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=MVIT-B, Pretraining=IN-21K+K400, Clip Size=32 x 224^22022.03 | 68.5 | 91 | — | |
| MTV-B(↑320)Protocol=full fine-tuning, Pretrain=K400†, Views=32×3×4, GFLOPs=11160, Param(M)=3102023.10 | 68.5 | 90.4 | — | |
| 4DS-jPre-training=MAE, Model size (M)=21,495, Evaluation protocol=Frozen representations with trained one-block attention read-out2026.07 | 68.2 | — | — | |
| MformerAttention=Trajectory, ViT Architecture=ViT-B, Pretraining=IN-21K+K400, Clip Size=32 x 224^22022.03 | 68.1 | 91.2 | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=MVIT-B, Pretraining=IN-21K+K400, Clip Size=16 x 224^22022.03 | 68 | 91 | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=ViT-B, Pretraining=IN-21K+K400, Clip Size=16 x 336^22022.03 | 67.9 | 90.8 | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=MVIT-B, Pretraining=IN-21K, Clip Size=16 x 224^22022.03 | 67.8 | 90.6 | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=MVIT-B, Pretraining=IN-1K+K400, Clip Size=16 x 224^22022.03 | 67.8 | 90.6 | — | |
| ZeroI2V Swin-BPretrain=IN21K, Views=32×3×1, GFLOPs=741, Param(M)=89, Tunable Param(M)=142023.10 | 67.8 | 91.4 | — | |
| ILA VIT-L/14Protocol=full fine-tuning, Pretrain=CLIP, Views=8×3×4, GFLOPs=10884, Extra GFLOPs=3100, Param(M)=529, New Param(M)=2252023.10 | 67.8 | 90.5 | — | |
| MViT-B, 64×3Pretrain=Kinetics-400, Views=1 x 3, FLOPs=455, Param=36.62021.06 | 67.7 | 90.9 | — | |
| ZeroI2V ViT-B/16Pretrain=CLIP, Views=8×3×1, GFLOPs=422, Extra GFLOPs=0, Param(M)=86, New Param(M)=02023.10 | 67.7 | 90.8 | — | |
| InternVideo2Pre-training=MAE, Contrastive, Captioning, Model size (M)=6,000, Evaluation protocol=Frozen representations with trained one-block attention read-out2026.07 | 67.7 | — | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=MVIT-B, Pretraining=K400, Clip Size=16 x 224^22022.03 | 67.5 | 90.8 | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=MVIT-B, Pretraining=IN-1K, Clip Size=16 x 224^22022.03 | 67.4 | 90.6 | — | |
| PST-BPretrain=IN21K, Views=32×3×1, GFLOPs=741, Param(M)=89, Tunable Param(M)=892023.10 | 67.4 | 90.9 | — | |
| MformerAttention=Trajectory, ViT Architecture=ViT-B, Pretraining=IN-21K+K400, Clip Size=16 x 336^22022.03 | 67.1 | 90.6 | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=ViT-B, Pretraining=IN-21K+K400, Clip Size=16 x 224^22022.03 | 67 | 90.5 | — | |
| EVL VIT-L/14Protocol=PETL, Pretrain=CLIP, Views=32×3×1, GFLOPs=9641, Extra GFLOPs=1858, Param(M)=479, New Param(M)=1752023.10 | 66.7 | — | — | |
| MformerAttention=Trajectory, ViT Architecture=ViT-B, Pretraining=IN-21K+K400, Clip Size=16 x 224^22022.03 | 66.5 | 90.1 | — | |
| E3D-LBackbone=E3D-L, Pretrain=No pretrain, Resolution=16 x 312^2, GFLOPs=18.3, Testing Strategy=2x32023.03 | 65.7 | 89.8 | — | |
| VideoMAEv2-gPre-training=MAE, Model size (M)=1,013, Evaluation protocol=Frozen representations with trained one-block attention read-out2026.07 | 65.6 | — | — | |
| ViViT-L/16x2FLOPs=903, Param=352.12021.06 | 65.4 | 89.8 | — | |
| ViViTAttention=Joint-ST, ViT Architecture=ViT-L, Pretraining=IN-21K, Clip Size=16 x 320^22022.03 | 65.4 | 89.8 | — | |
| ViViT-LProtocol=full fine-tuning, Pretrain=K400†, Views=16×3×4, GFLOPs=11892, Param(M)=3112023.10 | 65.4 | 89.8 | — | |
| VideoPrism-gPre-training=MAE, Contrastive, Model size (M)=1,113, Evaluation protocol=Frozen representations with trained one-block attention read-out2026.07 | 65.4 | — | — | |
| ZeroI2V ViT-B/16Pretrain=IN21K, Views=8×3×1, GFLOPs=422, Extra GFLOPs=0, Param(M)=86, New Param(M)=02023.10 | 65.3 | — | — | |
| blVNetPretrain=SSv2, Views=1 x 1, FLOPs=129, Param=40.22021.06 | 65.2 | 90.3 | — | |
| bLVNetPretraining=IN-1K, Clip Size=32 x 224^22022.03 | 65.2 | 90.3 | — | |
| TEAPretrain=ImageNet-21K, Views=10 x 3, FLOPs=702021.06 | 65.1 | 89.9 | — | |
| DVTAttention=D-ST+MS-A, ViT Architecture=ViT-B, Pretraining=IN-1K, Clip Size=8 x 224^22022.03 | 64.8 | 89.5 | — | |
| TAdaBackbone=ConvNeXt-T, Pretrain=ImageNet, Resolution=16 x 256^2, GFLOPs=47, Testing Strategy=2x32023.03 | 64.8 | 88.8 | — | |
| MSNetPretrain=ImageNet-21K, Views=1 x 1, FLOPs=67, Param=24.62021.06 | 64.7 | 89.4 | — | |
| MSNetPretraining=IN-1K, Clip Size=16 x 224^22022.03 | 64.7 | 89.4 | — | |
| MVITAttention=MHPA, ViT Architecture=MVIT-B, Pretraining=K400, Clip Size=16 x 224^22022.03 | 64.7 | 89.2 | — | |
| E3D-MBackbone=E3D-M, Pretrain=No pretrain, Resolution=16 x 224^2, GFLOPs=4.7, Testing Strategy=2x32023.03 | 64.7 | 89.6 | — | |
| TANetBackbone=ResNet50, Pretrain=ImageNet, Resolution=16 x 256^2, GFLOPs=66, Testing Strategy=2x32023.03 | 64.6 | 89.5 | — | |
| MoViNet-A1Backbone=MoViNet-A1, Pretrain=No pretrain, Resolution=50 x 172^2, GFLOPs=6, Testing Strategy=2x32023.03 | 64.5 | 89.1 | — | |
| ActionNetBackbone=ResNet50, Pretrain=ImageNet, Resolution=16 x 256^2, GFLOPs=69.5, Testing Strategy=2x32023.03 | 64 | 89.3 | — | |
| TSMBackbone=ResNet50, Pretrain=ImageNet, Resolution=16 x 256^2, GFLOPs=65, Testing Strategy=2x32023.03 | 63.4 | 88.5 | — | |
| TSM-RGBPretrain=Kinetics-400, Views=2 x 3, FLOPs=62, Param=42.92021.06 | 63.3 | 88.2 | — | |
| SlowFast R101, 8×8Pretrain=Kinetics-400, Views=1 x 3, FLOPs=106, Param=53.32021.06 | 63.1 | 87.6 | — | |
| ST-Adapter ViT-B/16Protocol=PETL, Pretrain=IN21K, Views=8×3×1, GFLOPs=455, Extra GFLOPs=33, Param(M)=93, New Param(M)=72023.10 | 62.8 | — | — | |
| SIFAR-BPretrain=IN21K, Views=32×3×1, GFLOPs=789, Param(M)=87, Tunable Param(M)=872023.10 | 62.6 | 88.5 | — | |
| TimeSformer-HRPretrain=ImageNet-21K, Views=1 x 3, FLOPs=1703, Param=121.42021.06 | 62.5 | — | — | |
| TformerAttention=Divided-ST, ViT Architecture=ViT-B, Pretraining=IN-1K, Clip Size=96 x 224^22022.03 | 62.4 | 81 | — | |
| TimeSformer-LProtocol=full fine-tuning, Pretrain=IN21K, Views=64×3×1, GFLOPs=7140, Param(M)=1212023.10 | 62.4 | — | — | |
| TformerAttention=Divided-ST, ViT Architecture=ViT-B, Pretraining=IN-1K, Clip Size=16 x 448^22022.03 | 62.2 | 78 | — | |
| X3D-MBackbone=X3D-M, Pretrain=No pretrain, Resolution=16 x 224^2, GFLOPs=4.7, Testing Strategy=2x32023.03 | 62.2 | 87.2 | — | |
| E3D-SBackbone=E3D-S, Pretrain=No pretrain, Resolution=13 x 160^2, GFLOPs=1.9, Testing Strategy=2x32023.03 | 62.1 | 87.6 | — | |
| AIM ViT-B/16Protocol=PETL, Pretrain=IN21K, Views=8×3×1, GFLOPs=624, Extra GFLOPs=202, Param(M)=100, New Param(M)=142023.10 | 62 | — | — | |
| MoViNet-A0Backbone=MoViNet-A0, Pretrain=No pretrain, Resolution=50 x 172^2, GFLOPs=2.7, Testing Strategy=2x32023.03 | 61.9 | 87.2 | — | |
| X3D-SBackbone=X3D-S, Pretrain=No pretrain, Resolution=13 x 160^2, GFLOPs=2, Testing Strategy=2x32023.03 | 60.1 | 85.9 | — | |
| V-WaltPre-training=Diffusion, Model size (M)=1,900, Evaluation protocol=Frozen representations with trained one-block attention read-out2026.07 | 59.7 | — | — | |
| TformerAttention=Divided-ST, ViT Architecture=ViT-B, Pretraining=IN-1K, Clip Size=8 x 224^22022.03 | 59.5 | 74.9 | — | |
| ViT-L/14Protocol=full fine-tuning, Pretrain=CLIP, Views=8×3×1, GFLOPs=1946, Extra GFLOPs=0, Param(M)=304, New Param(M)=02023.10 | 48.7 | 77.5 | — | |
| ILAPretraining=CLIP-400M, Backbone=ViT-B/16-8f, Setting=Zero-shot2023.04 | 43.9 | 71.8 | — | |
| AIMPretraining=CLIP-400M, Backbone=ViT-B/16-8f, Setting=Zero-shot2023.04 | 39.1 | 68.7 | — | |
| X-CLIPPretraining=CLIP-400M, Backbone=ViT-B/16-8f, Setting=Zero-shot2023.04 | 38.1 | 68.1 | — | |
| EVLPretraining=CLIP-400M, Backbone=ViT-B/16-8f, Setting=Zero-shot2023.04 | 35.2 | 65.4 | — | |
| EVL ViT-B/16Backbone=ViT-B/16, #Frames=8 x 3, GFLOPS=5122022.08 | — | — | 61 | |
| EVL ViT-B/16Backbone=ViT-B/16, #Frames=16 x 3, GFLOPS=1,0232022.08 | — | — | 61.7 | |
| EVL ViT-B/16Backbone=ViT-B/16, #Frames=32 x 3, GFLOPS=2,0472022.08 | — | — | 62.4 | |
| EVL ViT-B/16 EnsBackbone=ViT-B/16, Ensemble=Uniformer-B [32], #Frames=32 x 3 + 32 x 3, GFLOPS=2,8242022.08 | — | — | 72.1 | |
| EVL ViT-L/14Backbone=ViT-L/14, #Frames=8 x 3, GFLOPS=2,4112022.08 | — | — | 65.1 | |
| EVL ViT-L/14Backbone=ViT-L/14, #Frames=32 x 3, GFLOPS=9,6412022.08 | — | — | 66.7 | |
| EVL ViT-L/14Backbone=ViT-L/14, Resolution=336px, #Frames=32 x 3, GFLOPS=24,2592022.08 | — | — | 68 | |
| Hiera-Lpretrain_dataset=Kinetics-400, pretrain_method=MAE, FLOPs (G)=413×3×1, Param=213M2023.06 | — | — | 74.7 | |
| Hiera-Lpretrain_dataset=Kinetics-400, pretrain_method=MAE, FLOPs (G)=413×3×2, Param=213M2023.06 | — | — | 75 | |
| Hiera-Lpretrain_dataset=SSv2, pretrain_method=MAE, FLOPs (G)=413×3×1, Param=213M2023.06 | — | — | 74.9 | |
| Hiera-Lpretrain_dataset=SSv2, pretrain_method=MAE, FLOPs (G)=413×3×2, Param=213M2023.06 | — | — | 75.1 | |
| Hiera-Lpretrain_dataset=SSv2, pretrain_method=MAE, variant=L32, FLOPs (G)=1029×3×1, Param=213M2023.06 | — | — | 76.5 | |
| MViTv2-Lpretrain_dataset=Kinetics-400, pretrain_method=MaskFeat, resolution_config=40, 312, FLOPs (G)=2828×3×1, Param=218M2023.06 | — | — | 74.4 | |
| ViT-Lpretrain_dataset=Kinetics-400, pretrain_method=supervised, FLOPs (G)=598×3×1, Param=304M2023.06 | — | — | 55.7 | |
| ViT-Lpretrain_dataset=Kinetics-400, pretrain_method=MAE, FLOPs (G)=597×3×2, Param=305M2023.06 | — | — | 74 | |
| ViT-Lpretrain_dataset=SSv2, pretrain_method=MAE, FLOPs (G)=597×3×2, Param=305M2023.06 | — | — | 74.3 | |
| ViT-Lpretrain_dataset=SSv2, pretrain_method=MAE, variant=L32, FLOPs (G)=1436×3×1, Param=305M2023.06 | — | — | 75.4 |