Action Recognition on Something-Something v2 (test val)
74.7Top-1 AccuracyMAR
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| MARPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-L, Input Size=16 x 224^2, FLOPs×Cr.xCl. (G)=276 x 3 x 2, Param (M)=311, mask ratio (p)=50%2022.07 | 74.7 | 94.9 | |
| MaskFeatPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=MViT-L, Input Size=40 x 312^2, FLOPs×Cr.xCl. (G)=2828 x 3 x 1, Param (M)=2182022.07 | 74.4 | 94.6 | |
| MaskFeatpre-train=MaskFeat [77], pre-train data=K400, architecture=MViTv2-L, input size=40 x 312^2, FLOPs=2828 x 3 x 1, param.=2182022.05 | 74.4 | 94.6 | |
| VideoMAEPre-training Dataset=SSv2, Supervised Pre-training=false, Architecture=ViT-L, Input Size=16 x 224^2, FLOPs×Cr.xCl. (G)=597 x 3 x 2, Param (M)=3052022.07 | 74.2 | 94.7 | |
| MAE-vPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-H, Input Size=16 x 224^2, FLOPs×Cr.xCl. (G)=1193 x 3 x 1, Param (M)=6322022.07 | 74.1 | 94.5 | |
| MAEpre-train=MAE, pre-train data=K400, architecture=ViT-H, input size=16 x 224^2, FLOPs=1193 x 3 x 1, param.=6322022.05 | 74.1 | 94.5 | |
| MARPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-L, Input Size=16 x 224^2, FLOPs×Cr.xCl. (G)=131 x 3 x 2, Param (M)=311, mask ratio (p)=75%2022.07 | 73.8 | 94.4 | |
| MViTv2-Lpre-train=supervised, pre-train data=K400 + IN21K, architecture=MViTv2-L, input size=40 x 224^2, FLOPs=2828 x 3 x 1, param.=2132022.05 | 73.3 | 94.1 | |
| MViTv2-BPretrain=IN-21K/K400, Model #Params=213M, Trainable #Params=213M, GFLOPs=8484, Views=32×1×32023.03 | 73.3 | 94.1 | |
| DUALPATHArch=ViT-L/14, Pretrain=CLIP, Model #Params=336M, Trainable #Params=33M, GFLOPs=2151, Views=48×1×32023.03 | 72.2 | 93.7 | |
| MAE-vPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-L, Input Size=16 x 224^2, FLOPs×Cr.xCl. (G)=598 x 3 x 1, Param (M)=3042022.07 | 72.1 | 93.9 | |
| MViTv2-Bpre-train=supervised, pre-train data=K400 + IN21K, architecture=MViTv2-B, input size=32 x 224^2, FLOPs=225 x 3 x 1, param.=512022.05 | 72.1 | 93.4 | |
| MAEpre-train=MAE, pre-train data=K400, architecture=ViT-L, input size=16 x 224^2, FLOPs=598 x 3 x 1, param.=3042022.05 | 72.1 | 93.9 | |
| CAST w/ VideoMAE pretrained on Something-Something-V2Backbone=CAST-B, Frames=16, Views=2x3, TFLOPs=2.35, Learnable Param (M)=452023.11 | 71.6 | — | |
| BEVTpre-train=BEVT [73], pre-train data=K400 + IN1K, architecture=Swin-B, input size=32 x 224^2, FLOPs=321 x 3 x 1, param.=882022.05 | 71.4 | — | |
| OmnivorePretrain=IN-21K/K400, Views=32×1×32023.03 | 71.4 | 93.5 | |
| DUALPATHArch=ViT-L/14, Pretrain=CLIP, Model #Params=336M, Trainable #Params=33M, GFLOPs=1932, Views=32×1×32023.03 | 71.4 | 93.4 | |
| VideoMamba-M 800eArch.=SSM, Isotropic=true, Extra Data=CLIP-400M, Input Size=16x288^2, #Param (M)=74, FLOPs (G)=333x3x2, Training=Self-supervised2024.03 | 71.4 | 92.9 | |
| OMNIVOREBackbone=Swin-B, Frames=322023.11 | 71.4 | — | |
| UniFormer-BPretrain=IN-21K/K600, Model #Params=50M, Trainable #Params=50M, GFLOPs=777, Views=32×1×32023.03 | 71.2 | 92.8 | |
| DUALPATHArch=ViT-B/16, Pretrain=CLIP, Model #Params=99M, Trainable #Params=13M, GFLOPs=791, Views=48×1×32023.03 | 71.2 | 93.2 | |
| MARPre-training Dataset=SSv2, Supervised Pre-training=false, Architecture=ViT-B, Input Size=16 x 224^2, FLOPs×Cr.xCl. (G)=86 x 3 x 2, Param (M)=94, mask ratio (p)=50%2022.07 | 71 | 92.8 | |
| VideoMamba-M 800eArch.=SSM, Isotropic=true, Extra Data=CLIP-400M, Input Size=16x224^2, #Param (M)=74, FLOPs (G)=202x3x2, Training=Self-supervised2024.03 | 71 | 92.7 | |
| VideoMAE-B 2400eArch.=Trans., Isotropic=true, Extra Data=Provided, Input Size=16x224^2, #Param (M)=87, FLOPs (G)=180x3x2, Training=Self-supervised2024.03 | 70.8 | 92.4 | |
| UMT-B 800eArch.=Trans., Isotropic=true, Extra Data=CLIP-400M, Input Size=8x224^2, #Param (M)=87, FLOPs (G)=180x3x2, Training=Self-supervised2024.03 | 70.8 | 92.6 | |
| VideoMAEBackbone=ViT-B, Frames=16, Views=2x3, TFLOPs=1.08, Learnable Param (M)=87, Note=Experiment with own implementation2023.11 | 70.8 | — | |
| BEVTPre-training Dataset=IN-1K+K400+DALLE, Supervised Pre-training=false, Architecture=Swin-B, Input Size=32 x 224^2, FLOPs×Cr.xCl. (G)=321 x 3 x 1, Param (M)=882022.07 | 70.6 | — | |
| BEVT-B 800eArch.=Trans., Isotropic=false, Extra Data=IN-1K+K400, Input Size=32x224^2, #Param (M)=88, FLOPs (G)=321x3x1, Training=Self-supervised2024.03 | 70.6 | — | |
| BEVTBackbone=Swin-B, Frames=32, Views=1x3, TFLOPs=0.962023.11 | 70.6 | — | |
| MViTv2-Bpre-train=supervised, pre-train data=K400, architecture=MViTv2-B, input size=32 x 224^2, FLOPs=225 x 3 x 1, param.=512022.05 | 70.5 | 92.7 | |
| MViTv2-BPretrain=K400, Model #Params=51M, Trainable #Params=51M, GFLOPs=675, Views=40×1×32023.03 | 70.5 | 92.7 | |
| MViTv2-BArch.=Trans., Isotropic=false, Extra Data=K400, Input Size=32x224^2, #Param (M)=51, FLOPs (G)=225x3x1, Training=Supervised2024.03 | 70.5 | 92.7 | |
| UniFormer-BArch.=CNN+Trans., Isotropic=false, Extra Data=IN-1K+K400, Input Size=16x224^2, #Param (M)=50, FLOPs (G)=97x3x1, Training=Supervised2024.03 | 70.4 | 92.8 | |
| VideoMAEPre-training Dataset=SSv2, Supervised Pre-training=false, Architecture=ViT-B, Input Size=16 x 224^2, FLOPs×Cr.xCl. (G)=180 x 3 x 2, Param (M)=872022.07 | 70.3 | 92.7 | |
| DUALPATHArch=ViT-B/16, Pretrain=CLIP, Model #Params=99M, Trainable #Params=13M, GFLOPs=716, Views=32×1×32023.03 | 70.3 | 92.9 | |
| DUALPATHArch=ViT-L/14, Pretrain=CLIP, Model #Params=336M, Trainable #Params=33M, GFLOPs=1713, Views=16×1×32023.03 | 70.2 | 92.7 | |
| VideoMamba-M 800eArch.=SSM, Isotropic=true, Extra Data=CLIP-400M, Input Size=8x224^2, #Param (M)=74, FLOPs (G)=101x3x2, Training=Self-supervised2024.03 | 70.2 | 92.6 | |
| MorphMLP-BPretrain=IN-1K, #Frame=32×3×1, GFLOPs=5912021.11 | 70.1 | 92.8 | |
| Video-Swin-BParams=88.8M, #Frame=16, FLOPs x Clips=321G x 32022.03 | 69.6 | 92.7 | |
| Video SwinPre-training Dataset=IN-21K+K400, Supervised Pre-training=true, Architecture=Swin-B, Input Size=32 x 224^2, FLOPs×Cr.xCl. (G)=321 x 3 x 1, Param (M)=882022.07 | 69.6 | 92.7 | |
| Swin-Bpre-train=supervised, pre-train data=K400 + IN21K, architecture=Swin-B, input size=32 x 224^2, FLOPs=321 x 3 x 1, param.=892022.05 | 69.6 | 92.7 | |
| VideoSwin-BPretrain=IN-21K/K400, Model #Params=89M, Trainable #Params=89M, GFLOPs=963, Views=32×1×12023.03 | 69.6 | 92.7 | |
| DUALPATHArch=ViT-B/16, Pretrain=CLIP, Model #Params=99M, Trainable #Params=13M, GFLOPs=642, Views=16×1×32023.03 | 69.6 | 92.5 | |
| Swin-BArch.=Trans., Isotropic=true, Extra Data=K400, Input Size=32x224^2, #Param (M)=89, FLOPs (G)=88x3x1, Training=Supervised2024.03 | 69.6 | 92.7 | |
| Video SwinBackbone=Swin-B, Frames=32, Views=1x3, TFLOPs=0.96, Learnable Param (M)=892023.11 | 69.6 | — | |
| MARPre-training Dataset=SSv2, Supervised Pre-training=false, Architecture=ViT-B, Input Size=16 x 224^2, FLOPs×Cr.xCl. (G)=41 x 3 x 2, Param (M)=94, mask ratio (p)=75%2022.07 | 69.5 | 91.9 | |
| ST-AdapterPretrain=CLIP, Model #Params=97M, Trainable #Params=11M, GFLOPs=1955, Views=32×3×12023.03 | 69.5 | 92.6 | |
| ST-AdapterBackbone=ViT-B, Frames=32, Views=3x1, TFLOPs=1.962023.11 | 69.5 | — | |
| ORVIT-MF-LBackbone=ViT-L, Frames=32, Views=1x32023.11 | 69.5 | — | |
| AIMBackbone=ViT-B, Frames=32, Views=1x3, TFLOPs=2.5, Learnable Param (M)=142023.11 | 69.1 | — | |
| MVITPre-training Dataset=Kinetics-600, Supervised Pre-training=true, Architecture=MViT-B-24, Input Size=32 x 224^2, FLOPs×Cr.xCl. (G)=236 x 3 x 1, Param (M)=532022.07 | 68.7 | 91.5 | |
| MViT-B-24, 32×3Pretrain=K600, #Frame=32×3×1, GFLOPs=7082021.11 | 68.7 | 91.5 | |
| MTV-HRBackbone=MTV-B, Frames=32, Views=4x3, TFLOPs=11.16, Learnable Param (M)=3102023.11 | 68.5 | — | |
| VideoMamba-MArch.=SSM, Isotropic=true, Extra Data=IN-1K, Input Size=16x288^2, #Param (M)=74, FLOPs (G)=333x3x4, Training=Supervised2024.03 | 68.4 | 91.6 | |
| MorphMLP-SPretrain=IN-1K, #Frame=32×3×1, GFLOPs=4052021.11 | 68.3 | 91.3 | |
| VideoMamba-MArch.=SSM, Isotropic=true, Extra Data=IN-1K, Input Size=16x224^2, #Param (M)=74, FLOPs (G)=202x3x4, Training=Supervised2024.03 | 68.3 | 91.4 | |
| TDNFrame=16+8, FLOPs x clips=198G x 12022.02 | 68.2 | 91.6 | |
| MViTv2-SArch.=CNN+Trans., Isotropic=false, Extra Data=K400, Input Size=16x224^2, #Param (M)=35, FLOPs (G)=65x3x1, Training=Supervised2024.03 | 68.2 | 91.4 | |
| MotionformerPre-training Dataset=IN-21K+K400, Supervised Pre-training=true, Architecture=ViT-L, Input Size=32 x 224^2, FLOPs×Cr.xCl. (G)=1185 x 3 x 1, Param (M)=3822022.07 | 68.1 | 91.2 | |
| VIMPACPre-training Dataset=HowTo100M+DALLE, Supervised Pre-training=false, Architecture=ViT-L, Input Size=10 x 224^2, FLOPs×Cr.xCl. (G)=N/A x 3 x 10, Param (M)=3072022.07 | 68.1 | — | |
| Mformer-LPretrain=IN-21K+K400, #Frame=32×3×1, GFLOPs=35552021.11 | 68.1 | 91.2 | |
| Mformer-HRArch.=Trans., Isotropic=true, Extra Data=IN-21K+K400, Input Size=16x336^2, #Param (M)=311, FLOPs (G)=1185x3x1, Training=Supervised2024.03 | 68.1 | 91.2 | |
| VideoMamba-SArch.=SSM, Isotropic=true, Extra Data=IN-1K, Input Size=16x288^2, #Param (M)=26, FLOPs (G)=112x3x2, Training=Supervised2024.03 | 68.1 | 91.2 | |
| MFormerBackbone=ViT-L, Frames=32, Views=1x3, TFLOPs=3.562023.11 | 68.1 | — | |
| GC-TDNParams=27.4M x 2, #Frame=8+16, FLOPs x Clips=110.1G x 12022.03 | 67.8 | 91.2 | |
| TCM-R50 EnFrame=16+8, FLOPs x clips=105G x 10, Params=49.0M2022.02 | 67.8 | 92.2 | |
| MViT-BParams=36.6M, #Frame=64, FLOPs x Clips=455G x 32022.03 | 67.7 | 90.9 | |
| MViT-B, 32×3Pretrain=K400, #Frame=32×3×1, GFLOPs=13652021.11 | 67.7 | 90.9 | |
| MViTv1-Bpre-train=supervised, pre-train data=K400, architecture=MViTv1-B, input size=64 x 224^2, FLOPs=454 x 3 x 1, param.=372022.05 | 67.7 | 90.9 | |
| UniFormer-SArch.=CNN+Trans., Isotropic=false, Extra Data=IN-1K+K400, Input Size=16x224^2, #Param (M)=21, FLOPs (G)=42x3x1, Training=Supervised2024.03 | 67.7 | 91.4 | |
| MVITBackbone=ViT-B, Frames=64, Views=1x3, TFLOPs=1.37, Learnable Param (M)=372023.11 | 67.7 | — | |
| MorphMLP-BPretrain=IN-1K, #Frame=16×3×1, GFLOPs=2942021.11 | 67.6 | 91.3 | |
| MTV-BPretrain=IN-21K, Model #Params=310M, Trainable #Params=310M, GFLOPs=4790, Views=32×4×32023.03 | 67.6 | 90.4 | |
| VideoMamba-SArch.=SSM, Isotropic=true, Extra Data=IN-1K, Input Size=16x224^2, #Param (M)=26, FLOPs (G)=68x3x2, Training=Supervised2024.03 | 67.6 | 90.9 | |
| GC-TSMParams=25.1M x 2, #Frame=8+16, FLOPs x Clips=99.8G x 62022.03 | 67.5 | 90.9 | |
| VideoMamba-MArch.=SSM, Isotropic=true, Extra Data=IN-1K, Input Size=8x224^2, #Param (M)=74, FLOPs (G)=101x3x4, Training=Supervised2024.03 | 67.3 | 91 | |
| MSNetFrame=16+8, FLOPs x clips=101G x 10, Params=49.2M2022.02 | 67.1 | 91 | |
| TAdaConvNeXt-TPre-training Dataset=ImageNet-1K, Supervised Pre-training=true, Architecture=ConvNeXt-T, Input Size=32 x 224^2, FLOPs×Cr.xCl. (G)=94 x 3 x 2, Param (M)=382022.07 | 67.1 | 90.4 | |
| MViT-B, 16×4Pretrain=K400, #Frame=16×3×1, GFLOPs=5102021.11 | 67.1 | 90.8 | |
| MorphMLP-SPretrain=IN-1K, #Frame=16×3×1, GFLOPs=2012021.11 | 67.1 | 90.9 | |
| MVIT-BPretrain=K400, Model #Params=37M, Trainable #Params=37M, GFLOPs=510, Views=32×1×32023.03 | 67.1 | 90.8 | |
| MViTv1-BArch.=Trans., Isotropic=false, Extra Data=K400, Input Size=32x224^2, #Param (M)=37, FLOPs (G)=170x3x1, Training=Supervised2024.03 | 67.1 | 90.8 | |
| TDNParams=26.1M x 2, #Frame=8+16, FLOPs x Clips=108G x 12022.03 | 67 | 90.3 | |
| VideoMAE-S 2400eArch.=Trans., Isotropic=true, Extra Data=Provided, Input Size=16x224^2, #Param (M)=22, FLOPs (G)=57x3x2, Training=Self-supervised2024.03 | 66.8 | 90.3 | |
| GC-TSMParams=25.1M x 2, #Frame=8+16, FLOPs x Clips=99.8G x 22022.03 | 66.7 | 90.6 | |
| TCM-R50 EnFrame=16+8, FLOPs x clips=105G x 1, Params=49.0M2022.02 | 66.7 | 90.7 | |
| EVLPretrain=CLIP, Model #Params=484M, Trainable #Params=175M, GFLOPs=9641, Views=32×1×32023.03 | 66.7 | — | |
| VideoMamba-SArch.=SSM, Isotropic=true, Extra Data=IN-1K, Input Size=8x224^2, #Param (M)=26, FLOPs (G)=34x3x2, Training=Supervised2024.03 | 66.6 | 90.4 | |
| MformerPretrain=IN-21K+K400, #Frame=16×3×1, GFLOPs=11102021.11 | 66.5 | 90.1 | |
| GC-TSNParams=25.1M, #Frame=8+16, FLOPs x Clips=99.8G x 22022.03 | 66.3 | 90.3 | |
| VideoMamba-TiArch.=SSM, Isotropic=true, Extra Data=IN-1K, Input Size=16x288^2, #Param (M)=7, FLOPs (G)=28x3x2, Training=Supervised2024.03 | 66.2 | 90 | |
| TANetParams=25.1M x 2, #Frame=8+16, FLOPs x Clips=99.0G x 62022.03 | 66 | 90.1 | |
| VideoMamba-TiArch.=SSM, Isotropic=true, Extra Data=IN-1K, Input Size=16x224^2, #Param (M)=7, FLOPs (G)=17x3x2, Training=Supervised2024.03 | 66 | 89.6 | |
| GC-TDNParams=27.4M, #Frame=16, FLOPs x Clips=73.4G x 12022.03 | 65.9 | 90 | |
| CT-NetPretrain=IN-1K, #Frame=16×3×2, GFLOPs=4502021.11 | 65.9 | 90.1 | |
| VIVIT FEBackbone=ViT-L, Frames=32, Views=4x3, TFLOPs=11.89, Learnable Param (M)=3112023.11 | 65.9 | — | |
| TEINetParams=30.4M x 2, #Frame=8+16, FLOPs x Clips=99.0G x 12022.03 | 65.5 | 89.8 | |
| TEINetFrame=16+8, FLOPs x clips=99G x 1, Params=30.4M2022.02 | 65.5 | 89.8 | |
| ViViT-LParams=352.1M, #Frame=32, FLOPs x Clips=903G x 42022.03 | 65.4 | 89.8 | |
| ViViT-LPretrain=IN-21K+K400, #Frame=16×3×4, GFLOPs=118922021.11 | 65.4 | 89.8 |