Action Classification on Kinetics-400 (val)
90Top-1 AccuracyVideoMAE V2-g
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| VideoMAE V2-gResolution=64 x 266^2, Views=2 x 3, TFLOPS=160.302023.03 | 90 | 98.4 | |
| MTV-HPre-training Data=WTS 280^2, Views=4 x 3, TFLOPS=73.572023.03 | 89.9 | 98.3 | |
| VideoMAE V2-HViews=5 x 3, TFLOPS=17.882023.03 | 88.6 | 97.9 | |
| VideoMAE V2-gViews=5 x 3, TFLOPS=38.162023.03 | 88.5 | 98.1 | |
| CoVeRPre-training Data=JFT-3B, Views=1 x 32023.03 | 87.2 | — | |
| MaskFeatViews=4 x 3, TFLOPS=45.482023.03 | 87 | 97.4 | |
| MAE-STViews=4 x 3, TFLOPS=25.052023.03 | 86.8 | 97.2 | |
| VideoMAEViews=5 x 3, TFLOPS=17.882023.03 | 86.6 | 97.1 | |
| MViTv2-LResolution=312^2, Views=40 x 3, TFLOPS=42.422023.03 | 86.1 | 97 | |
| MARmask ratio (p)=50%, Pre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-L, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=276 x 3 x 5, Param (M)=3112022.07 | 85.3 | 96.3 | |
| MaskFeatPre-training Dataset=Kinetics-600, Supervised Pre-training=false, Architecture=MVIT-L, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=377 x 1 x 10, Param (M)=2182022.07 | 85.1 | 96.6 | |
| MAE-vPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-H, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=1193 x 3 x 7, Param (M)=6322022.07 | 85.1 | 96.6 | |
| Video Swin-LResolution=384^2, Views=10 x 5, TFLOPS=105.352023.03 | 84.9 | 96.7 | |
| MAE-vPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-L, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=598 x 3 x 7, Param (M)=3042022.07 | 84.8 | 96.2 | |
| MaskFeatPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=MVIT-L, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=377 x 1 x 10, Param (M)=2182022.07 | 84.3 | 96.3 | |
| VideoMAEPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-L, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=597 x 3 x 5, Param (M)=3052022.07 | 83.9 | 96.3 | |
| MARmask ratio (p)=75%, Pre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-L, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=131 x 3 x 5, Param (M)=3112022.07 | 83.9 | 96 | |
| Video SwinPre-training Dataset=ImageNet-21K, Supervised Pre-training=true, Architecture=Swin-L, Input Size=32 x 224^2, FLOPs×Cr.×Cl. (G)=604 x 3 x 4, Param (M)=1972022.07 | 83.1 | 95.9 | |
| MTV-BResolution=320^2, Views=4 x 3, TFLOPS=11.162023.03 | 82.4 | 95.2 | |
| ViViT FEPre-training Dataset=ImageNet-21K, Supervised Pre-training=true, Architecture=ViT-L, Input Size=128 x 224^2, FLOPs×Cr.×Cl. (G)=3980 x 3 x 12022.07 | 81.7 | 93.8 | |
| ViViT-L FEViews=1 x 3, TFLOPS=11.942023.03 | 81.7 | 93.8 | |
| MAE-vPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-B, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=180 x 3 x 7, Param (M)=872022.07 | 81.3 | 94.9 | |
| MVITSupervised Pre-training=false, Architecture=MViT-B, Input Size=64 x 224^2, FLOPs×Cr.×Cl. (G)=455 x 1 x 5, Param (M)=372022.07 | 81.2 | 95.1 | |
| MARmask ratio (p)=50%, Pre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-B, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=86 x 3 x 5, Param (M)=942022.07 | 81 | 94.4 | |
| TimeSformerPre-training Dataset=ImageNet-21K, Supervised Pre-training=true, Architecture=ViT-L, Input Size=96 x 224^2, FLOPs×Cr.×Cl. (G)=8353 x 3 x 1, Param (M)=4302022.07 | 80.7 | 94.7 | |
| VideoMAEPre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-B, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=180 x 3 x 5, Param (M)=872022.07 | 80.7 | 94.7 | |
| TimeSformer-LViews=1 x 3, TFLOPS=7.142023.03 | 80.7 | 94.7 | |
| Video SwinPre-training Dataset=ImageNet-1K, Supervised Pre-training=true, Architecture=Swin-B, Input Size=32 x 224^2, FLOPs×Cr.×Cl. (G)=282 x 3 x 4, Param (M)=882022.07 | 80.6 | 94.6 | |
| BEVTPre-training Dataset=IN-1K+DALLE, Supervised Pre-training=false, Architecture=Swin-B, Input Size=32 x 224^2, FLOPs×Cr.×Cl. (G)=282 x 3 x 5, Param (M)=882022.07 | 80.6 | — | |
| En-VidTr-LInput=32 x 2, GFLOPs=392, Latency (ms)=147.22021.04 | 80.5 | 94.6 | |
| X3D-XXLInput=16 x 5, GFLOPs=1962021.04 | 80.4 | 94.6 | |
| MotionformerPre-training Dataset=ImageNet-21K, Supervised Pre-training=true, Architecture=ViT-L, Input Size=32 x 224^2, FLOPs×Cr.×Cl. (G)=1185 x 3 x 10, Param (M)=3822022.07 | 80.2 | 94.8 | |
| MVITSupervised Pre-training=false, Architecture=MViT-B, Input Size=32 x 224^2, FLOPs×Cr.×Cl. (G)=170 x 1 x 5, Param (M)=372022.07 | 80.2 | 94.4 | |
| SlowFastSupervised Pre-training=false, Architecture=R101+NL, Input Size=(16+64) x 224^2, FLOPs×Cr.×Cl. (G)=359 x 3 x 10, Param (M)=602022.07 | 79.8 | 93.9 | |
| SlowFast R101-NLViews=10 x 3, TFLOPS=7.022023.03 | 79.8 | 93.9 | |
| En-VidTr-MInput=16 x 4, GFLOPs=220, Latency (ms)=98.12021.04 | 79.7 | 94.2 | |
| En-VidTr-SInput=8 x 8, GFLOPs=130, Latency (ms)=73.22021.04 | 79.4 | 94 | |
| MARmask ratio (p)=75%, Pre-training Dataset=Kinetics-400, Supervised Pre-training=false, Architecture=ViT-B, Input Size=16 x 224^2, FLOPs×Cr.×Cl. (G)=41 x 3 x 5, Param (M)=942022.07 | 79.4 | 93.7 | |
| TDNViews=10 x 3, TFLOPS=5.942023.03 | 79.4 | 94.4 | |
| VidTr-LInput=32 x 2, GFLOPs=351, Latency (ms)=110.22021.04 | 79.1 | 93.9 | |
| En-I3D-TPN-101Input=32 x 2, GFLOPs=541, Latency (ms)=207.82021.04 | 79.1 | 94 | |
| TAdaConvNeXt-TPre-training Dataset=ImageNet-1K, Supervised Pre-training=true, Architecture=ConvNeXt-T, Input Size=32 x 224^2, FLOPs×Cr.×Cl. (G)=94 x 3 x 4, Param (M)=382022.07 | 79.1 | 93.7 | |
| SF101 16x8Input=(64+16)x2, GFLOPs=213, Latency (ms)=124.32021.04 | 78.9 | 93.5 | |
| TPN101Input=32 x 2, GFLOPs=374, Latency (ms)=133.42021.04 | 78.9 | 93.9 | |
| VidTr-MInput=16 x 4, GFLOPs=179, Latency (ms)=61.12021.04 | 78.6 | 93.5 | |
| CorrNet101Input=32 x 2, GFLOPs=1872021.04 | 78.5 | — | |
| ip-CSNSupervised Pre-training=false, Architecture=ResNet152, Input Size=32 x 224^2, FLOPs×Cr.×Cl. (G)=109 x 3 x 10, Param (M)=332022.07 | 77.8 | 92.8 | |
| NL101Input=32 x 2, GFLOPs=544, Latency (ms)=134.12021.04 | 77.7 | 93.3 | |
| TPN50Input=32 x 2, GFLOPs=199, Latency (ms)=89.32021.04 | 77.7 | 93.3 | |
| VidTr-SInput=8 x 8, GFLOPs=89, Latency (ms)=36.22021.04 | 77.7 | 93.3 | |
| En-I3D-50-101Input=32 x 2, GFLOPs=509, Latency (ms)=192.72021.04 | 77.7 | 93.2 | |
| I3D NLViews=10 x 3, TFLOPS=10.772023.03 | 77.7 | 93.3 | |
| SF101 8x8Input=(32+8)x2, GFLOPs=106, Latency (ms)=71.92021.04 | 77.5 | 92.3 | |
| Vanilla-TrInput=8 x 8, GFLOPs=89, Latency (ms)=32.82021.04 | 77.5 | 93.2 | |
| I3D101Input=32 x 2, GFLOPs=342, Latency (ms)=118.32021.04 | 77.4 | 92.7 | |
| NonLocal I3DPre-training Dataset=ImageNet-1K, Supervised Pre-training=true, Architecture=ResNet101, Input Size=128 x 224^2, FLOPs×Cr.×Cl. (G)=234 x 3 x 10, Param (M)=622022.07 | 77.3 | 93.3 | |
| CorrNet50Input=32 x 2, GFLOPs=1152021.04 | 77.2 | — | |
| SF50 8x8Input=(32+8)x2, GFLOPs=66, Latency (ms)=49.32021.04 | 77 | 92.6 | |
| NL50Input=32 x 2, GFLOPs=282, Latency (ms)=53.32021.04 | 76.5 | 92.6 | |
| TEINetInput=16 x 2, GFLOPs=66, Latency (ms)=49.52021.04 | 76.2 | 92.5 | |
| TEA50Input=16 x 2, GFLOPs=702021.04 | 76.1 | 92.5 | |
| CIDCInput=32 x 2, GFLOPs=101, Latency (ms)=82.32021.04 | 75.5 | 92.1 | |
| I3D50Input=32 x 2, GFLOPs=167, Latency (ms)=74.42021.04 | 75 | 92.2 |