Action Recognition on Something-Something v2 (test)
77Top-1 AccVideoMAE V2-g
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| VideoMAE V2-gBackbone=ViT-g2023.03 | 77 | 95.9 | — | — | — | |
| VideoMAE V2-HBackbone=ViT-H2023.03 | 76.8 | 95.8 | — | — | — | |
| MAE-ST2023.03 | 75.5 | 95 | — | — | — | |
| VideoMAE2023.03 | 75.4 | 95.2 | — | — | — | |
| MaskFeat 312Backbone=MVIT-L, Extra data=Kinetics-600, Extra labels=Supervised, Frames=40, GFLOPs=2828x1x3, Param=2182022.11 | 75 | 95 | — | — | — | |
| MaskFeat2023.03 | 75 | 95 | — | — | — | |
| Ours - AM/12, TCN (Dinov2-B)Pretrain=IN-21K, #F=24, Model # Params=86+28+45+54M, Trainable # Params=28+45+54M, FLOPS (T)=5.12024.11 | 74.8 | 95 | — | — | — | |
| ATMBackbone=ViT-L/14, Pre-training=Merged-2B, Frames × Crops × Clips=16×3×2, GFLOPs=842×62023.07 | 74.6 | — | — | — | — | |
| Ours - AM/12, Transformer (Dinov2-B)Pretrain=IN-21K, #F=24, Model # Params=86+28+45+103M, Trainable # Params=28+45+103M, FLOPS (T)=11.92024.11 | 74.6 | 95 | — | — | — | |
| MaskFeat 312Backbone=MVIT-L, Extra data=Kinetics-400, Extra labels=Supervised, Frames=40, GFLOPs=2828x1x3, Param=2182022.11 | 74.4 | 94.6 | — | — | — | |
| VideoMAE-L#F=16, Model # Params=305M, Trainable # Params=305M, FLOPS (T)=3.62024.11 | 74.3 | 94.6 | — | — | — | |
| ATMBackbone=ViT-L/14, Pre-training=WIT-400M, Frames × Crops × Clips=16×3×2, GFLOPs=842×62023.07 | 73.5 | — | — | — | — | |
| Ours - AM/12, TCN (CLIP-B)Pretrain=IN-21K, #F=24, Model # Params=86+28+45+54M, Trainable # Params=28+45+54M, FLOPS (T)=5.12024.11 | 73.5 | 94.7 | — | — | — | |
| MViTv2-LPretrain=IN-21K+K400, #F=40, Model # Params=213M, Trainable # Params=213M, FLOPS (T)=8.52024.11 | 73.3 | 92.7 | — | — | — | |
| ATMYear=2023, Backbone=ViT-L/14, Pre-training=WIT-400M, Frames × Crops × Clips=16×3×1, GFLOPs=842×32023.07 | 73.2 | — | — | — | — | |
| UniFormerV2Year=2023, Backbone=ViT-L/14, Pre-training=WIT-400M, Frames × Crops × Clips=32×3×1, GFLOPs=52002023.07 | 73 | — | — | — | — | |
| UniFormerV2-LPretrain=CLIP-400M, #F=32, Model # Params=574M, Trainable # Params=574M, FLOPS (T)=5.22024.11 | 73 | 94.5 | — | — | — | |
| ST-AdapterYear=2022, Backbone=ViT-L/14, Pre-training=WIT-400M, Frames × Crops × Clips=32×3×1, GFLOPs=82482023.07 | 72.3 | — | — | — | — | |
| V-JEPA-HTraining Data=V-2M, Frames x Resolution=16 x 384, Probing Type=Attentive Probing2024.03 | 72.2 | — | — | — | — | |
| DUALPATH-LPretrain=CLIP-400M, #F=48, Model # Params=336M, Trainable # Params=33M, FLOPS (T)=1.92024.11 | 72.2 | 93.7 | — | — | — | |
| MViTv2-BBackbone=MViTv2-B2023.03 | 72.1 | 93.4 | — | — | — | |
| UniFormerV2-LPretrain=CLIP-400M, #F=16, Model # Params=574M, Trainable # Params=574M, FLOPS (T)=82024.11 | 72.1 | 93.6 | — | — | — | |
| ATMBackbone=ViT-B/16, Pre-training=WIT-400M, Frames × Crops × Clips=32×3×1, GFLOPs=378×32023.07 | 71.9 | — | — | — | — | |
| UniFormer2023.03 | 71.2 | 92.8 | — | — | — | |
| ATM_EnBackbone=ResNet101, Pre-training=ImageNet-1K, Frames × Crops × Clips=(8+16+32)×1×1, GFLOPs=4692023.07 | 70.8 | — | — | — | — | |
| CoVeRYear=2022, Backbone=TSFormer (448), Pre-training=JFT-3B, Frames × Crops × Clips=16×3×1, GFLOPs=176002023.07 | 70.8 | — | — | — | — | |
| VideoMAE-B#F=16, Model # Params=87M, Trainable # Params=87M, FLOPS (T)=1.12024.11 | 70.8 | 92.4 | — | — | — | |
| CoVeRPretrain=JFT-3B+KMI, #F=16, Model # Params=431M, Trainable # Params=431M, FLOPS (T)=17.62024.11 | 70.8 | — | — | — | — | |
| BEVTBackbone=Swin-B, Extra data=IN-1K+Kinetics-400+DALLE, Extra labels=Unlabeled, Frames=32, GFLOPs=321x1x3, Param=882022.11 | 70.6 | — | — | — | — | |
| BEVT2023.03 | 70.6 | — | — | — | — | |
| AIMYear=2023, Backbone=ViT-L/14, Pre-training=WIT-400M, Frames × Crops × Clips=32×3×1, GFLOPs=115082023.07 | 70.6 | — | — | — | — | |
| DUALPATH-BPretrain=CLIP-400M, #F=16, Model # Params=99M, Trainable # Params=13M, FLOPS (T)=0.72024.11 | 70.3 | 92.9 | — | — | — | |
| AdaMAE p=95%Backbone=ViT-B, Extra data=no external data, Extra labels=Unlabeled, Pre. Epochs=800, Frames=16, GFLOPs=180x2x3, Param=872022.11 | 70 | 92.7 | — | — | — | |
| SIFA-TransformerBackbone=Swin-B, GFLOPS x views=270x32022.06 | 69.8 | 93.1 | — | — | — | |
| Video-SwinBackbone=Swin-B, GFLOPS x views=321x32022.06 | 69.6 | 92.7 | — | — | — | |
| TDN EnBackbone=ResNet101x2, Extra data=IN1K, Extra labels=Supervised, Frames=8+16, GFLOPs=198x1x3, Param=882022.11 | 69.6 | 92.2 | — | — | — | |
| Video SwinBackbone=Swin-B, Extra data=IN21K+Kinetics-400, Extra labels=Supervised, Frames=32, GFLOPs=321x1x3, Param=882022.11 | 69.6 | 92.7 | — | — | — | |
| TDN2023.03 | 69.6 | 92.2 | — | — | — | |
| Video Swin-BBackbone=Swin-B2023.03 | 69.6 | 92.7 | — | — | — | |
| VideoSwin-BPretrain=IN-21K+K400, #F=32, Model # Params=89M, Trainable # Params=89M, FLOPS (T)=12024.11 | 69.6 | 92.7 | — | — | — | |
| UniFormerV2-BPretrain=CLIP-400M, #F=32, Model # Params=163M, Trainable # Params=163M, FLOPS (T)=1.12024.11 | 69.5 | 92.3 | — | — | — | |
| UniFormerV2-BPretrain=CLIP-400M, #F=16, Model # Params=163M, Trainable # Params=163M, FLOPS (T)=0.62024.11 | 69.5 | 92.3 | — | — | — | |
| ST-AdapterPretrain=CLIP-400M, #F=32, Model # Params=97M, Trainable # Params=11M, FLOPS (T)=22024.11 | 69.5 | 92.6 | — | — | — | |
| ATM_EnYear=2023, Backbone=ResNet101, Pre-training=ImageNet-1K, Frames × Crops × Clips=(8+16)×1×1, GFLOPs=2012023.07 | 69.4 | — | — | — | — | |
| VideoMAEBackbone=ViT-B, Extra data=no external data, Extra labels=Unlabeled, Pre. Epochs=800, Frames=16, GFLOPs=180x2x3, Param=872022.11 | 69.3 | 92.3 | — | — | — | |
| OmniMAEBackbone=ViT-B, Extra data=IN1K, Extra labels=Unlabeled, Pre. Epochs=800, Frames=16, GFLOPs=180x5x3, Param=872022.11 | 69.3 | — | — | — | — | |
| TVTS (Ours)Backbone=ViT-B, Pre-train Dataset=YT-Temporal, CC3M, WebVid2M, Fine-tuning=true2022.09 | 69.1 | — | — | — | — | |
| AIMYear=2023, Backbone=ViT-B/16, Pre-training=WIT-400M, Frames × Crops × Clips=32×3×1, GFLOPs=24962023.07 | 69.1 | — | — | — | — | |
| ML RGB+Flow+DiffEnsemble=Yes, Base architecture=ResNet-101, Number of input frames=16+16+16, Spatial crops X Temporal clips for prediction=1 x 32020.11 | 69.02 | 92.7 | — | — | — | |
| MViTv1 (B-24)Backbone=MViTv1-B-24, Extra data=Kinetics-600, Extra labels=Supervised, Frames=32, GFLOPs=236x1x3, Param=532022.11 | 68.7 | 91.5 | — | — | — | |
| TVTS (Ours)Backbone=ViT-B, Pre-train Dataset=YT-Temporal, Fine-tuning=true2022.09 | 68.5 | — | — | — | — | |
| MTVYear=2022, Backbone=MTV-B, Pre-training=IN-21K, Frames × Crops × Clips=32×3×4, GFLOPs=112002023.07 | 68.5 | — | — | — | — | |
| VideoPrism-gTraining Data=V-619M, Frames x Resolution=16 x 288, Probing Type=Attentive Probing2024.03 | 68.5 | — | — | — | — | |
| MTV-BPretrain=IN-21K+K400, #F=32, Model # Params=310M, Trainable # Params=310M, FLOPS (T)=11.22024.11 | 68.5 | 90.4 | — | — | — | |
| ATMYear=2023, Backbone=ResNet50, Pre-training=ImageNet-1K, Frames × Crops × Clips=32×1×1, GFLOPs=1482023.07 | 68.4 | — | — | — | — | |
| SpatioTemporalMAEBackbone=ViT-B, Extra data=no external data, Extra labels=Unlabeled, Pre. Epochs=800, Frames=16, GFLOPs=180x2x3, Param=872022.11 | 68.3 | 91.8 | — | — | — | |
| ATM_EnBackbone=ResNet50, Pre-training=ImageNet-1K, Frames × Crops × Clips=(8+16)×1×1, GFLOPs=1112023.07 | 68.3 | — | — | — | — | |
| TDN_EnYear=2021, Backbone=ResNet101, Pre-training=ImageNet-1K, Frames × Crops × Clips=(8+16)×1×1, GFLOPs=1982023.07 | 68.2 | — | — | — | — | |
| ATMBackbone=ResNet101, Pre-training=ImageNet-1K, Frames × Crops × Clips=16×1×1, GFLOPs=1342023.07 | 68.2 | — | — | — | — | |
| SSAEnFrames=16+82021.05 | 68.2 | — | — | — | — | |
| RGB-only ensemble (9702_10347) by AnonymousEnsemble=Yes, Base architecture=?, Number of input frames=?, Spatial crops X Temporal clips for prediction=?x?2020.11 | 68.18 | 91.26 | — | — | — | |
| Mformer-LPretrain=IN-21K+K-400, GFLOPs=1185.1, Views=3x12021.06 | 68.1 | 91.2 | — | — | — | |
| SIFA-NetBackbone=R101, GFLOPS x views=157x3, Input clip length=642022.06 | 68.1 | 92 | — | — | — | |
| MotionformerBackbone=ViT-L, Extra data=IN21K+Kinetics-400, Extra labels=Supervised, Frames=32, GFLOPs=1185x1x3, Param=3822022.11 | 68.1 | 91.2 | — | — | — | |
| VIMPACBackbone=ViT-L, Extra data=HowTo100M+DALLE, Extra labels=Unlabeled, Frames=10, Param=3072022.11 | 68.1 | — | — | — | — | |
| MFormer-HR2023.03 | 68.1 | 91.2 | — | — | — | |
| VIMPAC2023.03 | 68.1 | — | — | — | — | |
| VideoMAEBackbone=ViT-B, Pre-train Dataset=YT-Temporal, Fine-tuning=true2022.09 | 67.9 | — | — | — | — | |
| CT-Net_ENBackbone=2D (R50)×4, #Frame=8+12+16+24, GFLOPS=2802021.06 | 67.8 | 91.1 | — | — | — | |
| MViTv1Backbone=MViTv1-B, Extra data=Kinetics-600, Extra labels=Supervised, Frames=32, GFLOPs=170x1x3, Param=372022.11 | 67.8 | 91.3 | — | — | — | |
| TPNEnsemble=No, Base architecture=ResNet-101, Number of input frames=16, Spatial crops X Temporal clips for prediction=3x22020.11 | 67.72 | 91.28 | — | — | — | |
| TSM ResNet-101, RGB+Flow by AnonymousEnsemble=Yes, Base architecture=ResNet-101, Number of input frames=?, Spatial crops X Temporal clips for prediction=?x?2020.11 | 67.71 | 91.95 | — | — | — | |
| SELFYNet-TSM-R50 EN#frame=8+16, FLOPs x clips=114 G x 22021.02 | 67.7 | 91.1 | — | — | — | |
| MVITBackbone=ViT-B, GFLOPS x views=455x32022.06 | 67.7 | 90.9 | — | — | — | |
| MViTv1Backbone=MViTv1-B, Extra data=Kinetics-400, Extra labels=Supervised, Frames=64, GFLOPs=455x1x3, Param=372022.11 | 67.7 | 90.9 | — | — | — | |
| InternVideo2s2-6BTraining Data=IV-400M, Frames x Resolution=16 x 224, Probing Type=Attentive Probing2024.03 | 67.7 | — | — | — | — | |
| MViTv1-BPretrain=K400, #F=32, Model # Params=37M, Trainable # Params=37M, FLOPS (T)=4.12024.11 | 67.7 | 90.9 | — | — | — | |
| MTV (B/2+S/4+Ti/8)Backbone=ViT-B, Extra data=Kinetics-400+I21K, Extra labels=Supervised, Frames=32, GFLOPs=384x4x3, Param=3102022.11 | 67.6 | 90.1 | — | — | — | |
| MTV-BBackbone=MTV-B2023.03 | 67.6 | 90.1 | — | — | — | |
| SELFYNet-TSM-R50 EN#frame=8+16, FLOPs x clips=114 G x 12021.02 | 67.4 | 91 | — | — | — | |
| SELFY_EnYear=2021, Backbone=ResNet50, Pre-training=ImageNet-1K, Frames × Crops × Clips=(8+16)×1×1, GFLOPs=1142023.07 | 67.4 | — | — | — | — | |
| ATMBackbone=ResNet50, Pre-training=ImageNet-1K, Frames × Crops × Clips=16×1×1, GFLOPs=742023.07 | 67.4 | — | — | — | — | |
| SIFA-NetBackbone=R101, GFLOPS x views=78x3, Input clip length=322022.06 | 67.3 | 91.1 | — | — | — | |
| InternVideo2s2-1BTraining Data=IV-25.5M, Frames x Resolution=16 x 224, Probing Type=Attentive Probing2024.03 | 67.3 | — | — | — | — | |
| TAda2DEnBackbone=ResNet-50, Frames x clips x crops=(8f+16f)×2×3, GFLOPs=1292021.10 | 67.2 | 89.8 | — | — | — | |
| MSNet-R50 En#frame=16+8, FLOPs=101G x 10, #param=49.2M2020.07 | 67.1 | 91 | — | — | — | |
| bLVNet-TAM RGB+FlowEnsemble=Yes, Base architecture=ResNet-101, Number of input frames=32+32, Spatial crops X Temporal clips for prediction=3 x 102020.11 | 67.1 | 91.4 | — | — | — | |
| bLVNet-TAMBackbone=bLResNet-101*, Pretrain=SS-V2, Frames=32×2, Modality=RGB+Flow2019.12 | 67.1 | 91.4 | — | — | — | |
| MVIT-BPretrain=K-400, GFLOPs=170, Views=3x12021.06 | 67.1 | 90.8 | — | — | — | |
| Mformer-HRPretrain=IN-21K+K-400, GFLOPs=958.8, Views=3x12021.06 | 67.1 | 90.6 | — | — | — | |
| MSNet-TSM-R50 EN#frame=8+16, FLOPs x clips=101 G x 102021.02 | 67.1 | 91 | — | — | — | |
| TAdaConvNeXt-TBackbone=ConvNeXt-T, Frames x clips x crops=32f×2×3, GFLOPs=942021.10 | 67.1 | 90.4 | — | — | — | |
| BEVT-VBackbone=Swin-B, Extra data=Kinetics-400+DALLE, Extra labels=Unlabeled, Frames=32, GFLOPs=321x1x3, Param=882022.11 | 67.1 | — | — | — | — | |
| TDNBackbone=ResNet-50, Frames x clips x crops=(8f+32f+16f+64f)×1×1, GFLOPs=1412021.10 | 67 | 90.3 | — | — | — | |
| TDN_EnYear=2021, Backbone=ResNet50, Pre-training=ImageNet-1K, Frames × Crops × Clips=(8+16)×1×1, GFLOPs=1082023.07 | 67 | — | — | — | — | |
| SIFA-NetBackbone=R50, GFLOPS x views=112x3, Input clip length=642022.06 | 66.9 | 90.7 | — | — | — | |
| TDNBackbone=R101, GFLOPS x views=1322022.06 | 66.9 | 90.9 | — | — | — | |
| TDNYear=2021, Backbone=ResNet101, Pre-training=ImageNet-1K, Frames × Crops × Clips=16×1×1, GFLOPs=1322023.07 | 66.9 | — | — | — | — | |
| MMLEnsemble=No, Base architecture=ResNet-101, Number of input frames=16, Spatial crops X Temporal clips for prediction=1 x 32020.11 | 66.83 | 91.3 | — | — | — | |
| EVLYear=2022, Backbone=ViT-L/14, Pre-training=WIT-400M, Frames × Crops × Clips=32×3×1, GFLOPs=96412023.07 | 66.7 | — | — | — | — |