Video Classification on Kinetics 400 (val)
85.4Top-1 AccTokenLearner 16at18 (L/10)
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TokenLearner 16at18 (L/10)total GFLOPS=4076 x 12, inference views=12, Pre-training dataset=JFT2021.06 | 85.4 | — | — | — | — | |
| Swin-L (384)total GFLOPS=2107 x 50, inference views=50, Pre-training dataset=ImageNet-21K2021.06 | 84.9 | — | — | — | — | |
| ViViT-L#param=>307M, #data=300M, tuning=fine-tuning2022.06 | 84.9 | — | — | — | — | |
| MAEpre-train set=K600, # pre-train data=387k, Backbone=ViT-L, pre-training epochs=16002022.05 | 84.9 | — | — | — | — | |
| Swin-L-384↑Pretrain=IN-21K, sampling_config=32x5x10, FLOPs_G=1053502022.01 | 84.9 | 96.7 | — | — | — | |
| MAEpre-train set=K400, # pre-train data=240k, Backbone=ViT-L, pre-training epochs=16002022.05 | 84.8 | — | — | — | — | |
| ViViT-HPretrain=JFT-300M, sampling_config=16x3x4, FLOPs_G=997922022.01 | 84.8 | 95.8 | — | — | — | |
| TokenLearner 16at18 (L/14)total GFLOPS=1621 x 12, inference views=12, Pre-training dataset=JFT2021.06 | 84.7 | — | — | — | — | |
| Swin-L (384)total GFLOPS=2107 x 12, inference views=12, Pre-training dataset=ImageNet-21K2021.06 | 84.6 | — | — | — | — | |
| TokenLearner 8at18 (L/16)total GFLOPS=1105 x 12, inference views=12, Pre-training dataset=JFT2021.06 | 84.5 | — | — | — | — | |
| MAEpre-train set=IG-uncurated, # pre-train data=1M, Backbone=ViT-L, pre-training epochs=16002022.05 | 84.4 | — | — | — | — | |
| Uni-Perceiver-L + Conditional MoEs#param=303M, #data=44.1M, tuning=fine-tuning2022.06 | 84.2 | — | — | — | — | |
| TokenLearner 16at12 (L/16)total GFLOPS=766 x 12, inference views=12, Pre-training dataset=JFT2021.06 | 83.5 | — | — | — | — | |
| TokenLearner 8at18 (L/16)total GFLOPS=1105 x 6, inference views=6, Pre-training dataset=JFT2021.06 | 83.2 | — | — | — | — | |
| Swin-Ltotal GFLOPS=604 x 12, inference views=12, Pre-training dataset=ImageNet-21K2021.06 | 83.1 | — | — | — | — | |
| UniFormer-BPretrain=IN-1K, sampling_config=32x3x4, FLOPs_G=31082022.01 | 83 | 95.4 | — | — | — | |
| UniFormer-BPretrain=IN-1K, sampling_config=32x1x4, FLOPs_G=10362022.01 | 82.9 | 95.4 | — | — | — | |
| ViViT-L/16total GFLOPS=1446 x 12, inference views=12, Pre-training dataset=JFT2021.06 | 82.8 | — | — | — | — | |
| ViViT-LPretrain=JFT-300M, sampling_config=16x3x4, FLOPs_G=173522022.01 | 82.8 | 95.3 | — | — | — | |
| Swin-BPretrain=IN-21K, sampling_config=32x3x4, FLOPs_G=33842022.01 | 82.7 | 95.5 | — | — | — | |
| ir-CSN-152*pretrain=IG-65M, GFLOPS × crops=96.7×302019.04 | 82.6 | 95.3 | — | — | — | |
| ip-CSN-152*pretrain=IG-65M, GFLOPS × crops=108.8×302019.04 | 82.5 | 95.3 | — | — | — | |
| MAEpre-train set=ImageNet-1K, # pre-train data=1.28M, Backbone=ViT-L2022.05 | 82.3 | — | — | — | — | |
| TokenLearner 16at12 (L/16)total GFLOPS=766 x 6, inference views=6, Pre-training dataset=JFT2021.06 | 82.1 | — | — | — | — | |
| UniFormer-BPretrain=IN-1K, sampling_config=16x1x4, FLOPs_G=3892022.01 | 82 | 95.1 | — | — | — | |
| ST Swin w/ LSTCLClip Size=16 x 224^2, Additional Data (# Samples)=-, TFLOPS=1.80, Params=88.0M2021.06 | 81.5 | 95.2 | — | — | — | |
| MoViNet-A6sampling_config=120x1x1, FLOPs_G=3862022.01 | 81.5 | 95.3 | — | — | — | |
| Collaborative Memory (SlowFast-101+NL 8×8)Pretrain=none, Only RGB=true, GFLOPS=137, crops=302021.04 | 81.4 | — | — | — | — | |
| R(2+1)D-152*pretrain=IG-65M, GFLOPS × crops=329×302019.04 | 81.3 | 95.1 | — | — | — | |
| LGD-3D-101Pretrain=ImageNet, Only RGB=false, GFLOPS=N/A, crops=N/A2021.04 | 81.2 | — | — | — | — | |
| MVIT-BClip Size=64 x 224^2, Additional Data (# Samples)=-, TFLOPS=4.09, Params=36.6M2021.06 | 81.2 | 95.1 | — | — | — | |
| Mformer-HRPretrain=IN-21K, sampling_config=16x3x10, FLOPs_G=287642022.01 | 81.1 | 95.2 | — | — | — | |
| CorrNet-101 (Sports1M)Pretrain=Sports1M, Only RGB=true, GFLOPS=224, crops=302021.04 | 81 | — | — | — | — | |
| CorrNetPretrain=Sports1M, sampling_config=32x3x10, FLOPs_G=67202022.01 | 81 | — | — | — | — | |
| MoViNet-A5sampling_config=120x1x1, FLOPs_G=2812022.01 | 80.9 | 94.9 | — | — | — | |
| MorphMLP-BPretrain=IN-1K, #Frame=32x1x4, GFLOPS=7882021.11 | 80.8 | 94.9 | — | — | — | |
| UniFormer-SPretrain=IN-1K, sampling_config=16x1x4, FLOPs_G=1672022.01 | 80.8 | 94.7 | — | — | — | |
| TimeSformer-LTFLOPS=7.142021.02 | 80.7 | 94.7 | — | — | — | |
| TimeSformer-LClip Size=96 x 224^2, Additional Data (# Samples)=ImageNet-21K (14M), TFLOPS=7.14, Params=121.4M2021.06 | 80.7 | 94.7 | — | — | — | |
| TimeSformer-Ltotal GFLOPS=2380 x 3, inference views=3, Pre-training dataset=ImageNet-21K2021.06 | 80.7 | — | — | — | — | |
| TimeSformer-B#param=121.4M, #data=14.2M, tuning=fine-tuning2022.06 | 80.7 | — | — | — | — | |
| Timesformer-LPretrain=IN-21K, #Frame=96x3x1, GFLOPS=71402021.11 | 80.7 | 94.7 | — | — | — | |
| TimeSformer-LPretrain=IN-21K, sampling_config=96x3x1, FLOPs_G=71402022.01 | 80.7 | 94.7 | — | — | — | |
| ViViT-LClip Size=16 x 224^2, Additional Data (# Samples)=ImageNet-21K (14M), TFLOPS=47.9, Params=310.0M2021.06 | 80.6 | 94.7 | — | — | — | |
| ViViT-LPretrain=IN-21K, #Frame=16x3x4, GFLOPS=173572021.11 | 80.6 | 94.7 | — | — | — | |
| VideoSwin-BPretrain=IN-1K, #Frame=32x3x4, GFLOPS=33842021.11 | 80.6 | 94.6 | — | — | — | |
| ViViT-LPretrain=IN-21K, sampling_config=16x3x4, FLOPs_G=173522022.01 | 80.6 | 94.7 | — | — | — | |
| Swin-BPretrain=IN-1K, sampling_config=32x3x4, FLOPs_G=33842022.01 | 80.6 | 94.6 | — | — | — | |
| Collaborative Memory (R(2+1)D-101 32×2)Pretrain=none, Only RGB=true, GFLOPS=243, crops=302021.04 | 80.5 | — | — | — | — | |
| X3D-XXLTFLOPS=5.82021.02 | 80.4 | 94.6 | — | — | — | |
| X-ViTPretrain=IN-21K, #Frame=16x3x1, GFLOPS=8502021.11 | 80.2 | 94.7 | — | — | — | |
| MformerPretrain=IN-21K, #Frame=32x3x10, GFLOPS=110852021.11 | 80.2 | 94.8 | — | — | — | |
| Mformer-LPretrain=IN-21K, #Frame=32x3x10, GFLOPS=355502021.11 | 80.2 | 94.8 | — | — | — | |
| MVIT-B, 32x3#Frame=32x1x5, GFLOPS=8502021.11 | 80.2 | 94.4 | — | — | — | |
| X-ViTPretrain=IN-21K, sampling_config=16x3x1, FLOPs_G=8502022.01 | 80.2 | 94.7 | — | — | — | |
| MViT-B, 32x3sampling_config=32x1x5, FLOPs_G=8502022.01 | 80.2 | 94.4 | — | — | — | |
| Collaborative Memory (SlowFast-101 8×8)Pretrain=none, Only RGB=true, GFLOPS=128, crops=302021.04 | 80 | — | — | — | — | |
| SlowFast+NLpretrain=none, GFLOPS × crops=234×302019.04 | 79.8 | 93.9 | — | — | — | |
| SlowFast-101+NL 16×8Pretrain=none, Only RGB=true, GFLOPS=234, crops=302021.04 | 79.8 | — | — | — | — | |
| SlowFastTFLOPS=72021.02 | 79.8 | 93.9 | — | — | — | |
| ST Swin w/ LSTCLClip Size=8 x 224^2, Additional Data (# Samples)=-, TFLOPS=0.60, Params=88.0M2021.06 | 79.8 | 94 | — | — | — | |
| SlowFast-NLClip Size=16 x 224^2, Additional Data (# Samples)=-, TFLOPS=7.0, Params=59.9M2021.06 | 79.8 | 93.9 | — | — | — | |
| SlowFast 16x8, R101+NLtotal GFLOPS=234 x 30, inference views=302021.06 | 79.8 | — | — | — | — | |
| CT-NetPretrain=IN-1K, #Frame=(16+16)x3x4, GFLOPS=26412021.11 | 79.8 | 94.2 | — | — | — | |
| CT-NetENPretrain=IN-1K, sampling_config=(16+16)x3x4, FLOPs_G=26412022.01 | 79.8 | 94.2 | — | — | — | |
| SlowFast+NLsampling_config=16x3x10, FLOPs_G=70202022.01 | 79.8 | 93.9 | — | — | — | |
| TimeSformer-HRTFLOPS=5.112021.02 | 79.7 | 94.4 | — | — | — | |
| MformerClip Size=16 x 224^2, Additional Data (# Samples)=ImageNet-21K (14M), TFLOPS=11.1, Params=109.1M2021.06 | 79.7 | 94.2 | — | — | — | |
| MorphMLP-SPretrain=IN-1K, #Frame=32x1x4, GFLOPS=5322021.11 | 79.7 | 94.2 | — | — | — | |
| TimeSformer-HRPretrain=IN-21K, sampling_config=16x3x1, FLOPs_G=51092022.01 | 79.7 | 94.4 | — | — | — | |
| VATT-BClip Size=32 x 320^2, Additional Data (# Samples)=AudioSet + HowTo100M (3.2M), TFLOPS=9.08, Params=88.0M2021.06 | 79.6 | 94.9 | — | — | — | |
| MorphMLP-BPretrain=IN-1K, #Frame=16x1x4, GFLOPS=3922021.11 | 79.5 | 94.4 | — | — | — | |
| LGD-3D RGBBackbone=ResNet-101, Training Frames=1282020.07 | 79.4 | 94.4 | — | — | — | |
| LGD-3D-1012021.02 | 79.4 | 94.4 | — | — | — | |
| TDNPretrain=IN-1K, #Frame=(8+16)x3x10, GFLOPS=59402021.11 | 79.4 | 94.4 | — | — | — | |
| TDNENPretrain=IN-1K, sampling_config=(8+16)x3x10, FLOPs_G=59402022.01 | 79.4 | 94.4 | — | — | — | |
| LGDPretrain=IN-1K, sampling_config=128xN/A2022.01 | 79.4 | 94.4 | — | — | — | |
| STAMClip Size=16 x 224^2, Additional Data (# Samples)=ImageNet-21K (14M), TFLOPS=0.27, Params=96.0M2021.06 | 79.3 | — | — | — | — | |
| Uni-Perceiver-B + Conditional MoEs#param=86M, #data=44.1M, tuning=fine-tuning2022.06 | 79.3 | — | — | — | — | |
| ip-CSN-152pretrain=Sports1M, GFLOPS × crops=108.8×302019.04 | 79.2 | 93.8 | — | — | — | |
| ip-CSN-152Pretrain=Sports1M, Only RGB=true, GFLOPS=109, crops=302021.04 | 79.2 | — | — | — | — | |
| CorrNet-101Pretrain=none, Only RGB=true, GFLOPS=224, crops=302021.04 | 79.2 | — | — | — | — | |
| CorrNetTFLOPS=6.72021.02 | 79.2 | — | — | — | — | |
| CorrNet-101Clip Size=16 x 224^2, Additional Data (# Samples)=-, TFLOPS=7.0, Params=-2021.06 | 79.2 | — | — | — | — | |
| CorrNet-101#Frame=32x3x10, GFLOPS=67202021.11 | 79.2 | — | — | — | — | |
| ip-CSNPretrain=Sports1M, #Frame=32x3x10, GFLOPS=32642021.11 | 79.2 | 93.8 | — | — | — | |
| ip-CSNPretrain=Sports1M, sampling_config=32x3x10, FLOPs_G=32702022.01 | 79.2 | 93.8 | — | — | — | |
| STAMPretrain=IN-21K, sampling_config=64x1x1, FLOPs_G=10402022.01 | 79.2 | — | — | — | — | |
| X3D-XLClip Size=16 x 312^2, Additional Data (# Samples)=-, TFLOPS=1.45, Params=11.0M2021.06 | 79.1 | 93.9 | — | — | — | |
| X3D-XL#Frame=16x3x10, GFLOPS=14522021.11 | 79.1 | 93.9 | — | — | — | |
| VidTr-LPretrain=IN-21K, #Frame=32x3x10, GFLOPS=117602021.11 | 79.1 | 93.9 | — | — | — | |
| X3D-XLsampling_config=16x3x10, FLOPs_G=14522022.01 | 79.1 | 93.9 | — | — | — | |
| ir-CSN-152pretrain=Sports1M, GFLOPS × crops=96.7×302019.04 | 79 | 93.5 | — | — | — | |
| SlowFastpretrain=none, GFLOPS × crops=213×302019.04 | 78.9 | 93.5 | — | — | — | |
| SlowFastBackbone=ResNet-101, Training Frames=16+642020.07 | 78.9 | 93.5 | — | — | — | |
| SlowFast-101 16×8Pretrain=none, Only RGB=true, GFLOPS=213, crops=302021.04 | 78.9 | — | — | — | — | |
| SlowFast R101#Frame=(16+64)x3x10, GFLOPS=63902021.11 | 78.9 | 93.5 | — | — | — | |
| VideoSwin-TPretrain=IN-1K, #Frame=32x3x4, GFLOPS=10562021.11 | 78.8 | 93.6 | — | — | — | |
| Swin-TPretrain=IN-1K, sampling_config=32x3x4, FLOPs_G=10562022.01 | 78.8 | 93.6 | — | — | — | |
| SmallBigPretrain=IN-1K, #Frame=(8+32)x3x4, GFLOPS=57002021.11 | 78.7 | 93.7 | — | — | — |