Video Classification on Kinetics-600 (val)
94.4AccuracyFTP-UniFormerV2-L/14
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FTP-UniFormerV2-L/14Input=32 x 336^2, Crops=2 x 3, TFLOPs=78.72024.03 | 94.4 | 99.3 | — | |
| FTP-UniFormerV2-L/14Input=32 x 224^2, Crops=2 x 3, TFLOPs=18.42024.03 | 93.8 | 99.4 | — | |
| FTP-UniFormerV2-B/16Input=8 x 280^2, Crops=4 x 3, TFLOPs=8.12024.03 | 93 | 99 | — | |
| FTP-UniFormerV2-B/16Input=8 x 224^2, Crops=4 x 3, TFLOPs=2.22024.03 | 92.2 | 98.9 | — | |
| UniFormerV2-L/14Input=32 x 224^2, Crops=2 x 3, TFLOPs=16.02024.03 | 89.5 | 98.3 | — | |
| Hiera-Hbackbone=Hiera-H, pretrain=MAE, FLOPs (G)=1159×3×5, Param=672M2023.06 | 88.8 | — | — | |
| VideoMAE V2-gInput=64 x 266^2, Crops=2 x 3, TFLOPs=38.162024.03 | 88.8 | 98.2 | — | |
| Hiera-Lbackbone=Hiera-L, pretrain=MAE, FLOPs (G)=413×3×5, Param=213M2023.06 | 88.3 | — | — | |
| MaskFeatInput=40 x 312^2, Crops=4 x 3, TFLOPs=33.942024.03 | 88.3 | 98 | — | |
| VideoMAE V2-HInput=16 x 224^2, Crops=2 x 3, TFLOPs=17.882024.03 | 88.3 | 98.1 | — | |
| MViTv2-LInput=40 x 352^2, Crops=5 x 3, TFLOPs=45.482024.03 | 87.9 | 97.9 | — | |
| UniFormerV2-B/16Input=8 x 224^2, Crops=4 x 3, TFLOPs=1.82024.03 | 87.4 | 97.9 | — | |
| MViTv2-Lbackbone=MViTv2-L, pretrain=MaskFeat, FLOPs (G)=377×1×10, Param=218M2023.06 | 86.4 | — | — | |
| MaskFeatInput=64 x 224^2, Crops=4 x 3, TFLOPs=3.772024.03 | 86.4 | 97.4 | — | |
| Swin-L-384↑Pretrain=IN-21K, sampling_config=32x5x10, FLOPs_G=1053502022.01 | 86.1 | 97.3 | — | |
| MViTv2-Lbackbone=MViTv2-L, pretrain=Sup, IN-21K, FLOPs (G)=377×1×10, Param=218M2023.06 | 85.8 | — | — | |
| ViViT-HPretrain=JFT-300M, sampling_config=16x3x4, FLOPs_G=997922022.01 | 85.8 | 96.5 | — | |
| MViTv2-BInput=32 x 312^2, Crops=5 x 3, TFLOPs=1.032024.03 | 85.5 | 97.2 | — | |
| UniFormer-BPretrain=IN-1K, sampling_config=32x3x4, FLOPs_G=31082022.01 | 84.9 | 96.7 | — | |
| UniFormer-BPretrain=IN-1K, sampling_config=32x1x4, FLOPs_G=10362022.01 | 84.8 | 96.7 | — | |
| X-ViT (16x)Views=1x3, GFLOPs=8502021.11 | 84.5 | 96.3 | — | |
| X-ViTPretrain=IN-21K, sampling_config=16x3x1, FLOPs_G=8502022.01 | 84.5 | 96.3 | — | |
| TokenLearner 16at12(L/16)Views=4x3, GFLOPs=9,1922021.11 | 84.4 | 96 | — | |
| X-ViT+ATS (16x)Views=1x3, GFLOPs=5212021.11 | 84.4 | 96.2 | — | |
| ViViT-LPretrain=JFT-300M, sampling_config=16x3x4, FLOPs_G=173522022.01 | 84.3 | 96.2 | — | |
| MViT-B-24, 32x3Views=1x5, GFLOPs=7,0802021.11 | 84.1 | 96.5 | — | |
| Swin-BViews=4x3, GFLOPs=3,3842021.11 | 84 | 96.5 | — | |
| Swin-BPretrain=IN-21K, sampling_config=32x3x4, FLOPs_G=33842022.01 | 84 | 96.5 | — | |
| UniFormer-BPretrain=IN-1K, sampling_config=16x1x4, FLOPs_G=3892022.01 | 84 | 96.4 | — | |
| MVIT-BInput=32 x 224^2, Crops=3 x 3, TFLOPs=4.102024.03 | 83.8 | 96.3 | — | |
| MTV-BInput=32 x 320^2, Crops=4 x 3, TFLOPs=4.792024.03 | 83.6 | 96.1 | — | |
| MoViNet-A6sampling_config=120x1x1, FLOPs_G=3862022.01 | 83.5 | 96.2 | — | |
| MViT-B, 32x3sampling_config=32x1x5, FLOPs_G=8502022.01 | 83.4 | 96.3 | — | |
| ViViT-L FEInput=32 x 224^2, Crops=1 x 3, TFLOPs=11.942024.03 | 82.9 | 94.6 | — | |
| UniFormer-SPretrain=IN-1K, sampling_config=16x1x4, FLOPs_G=1672022.01 | 82.8 | 95.8 | — | |
| MoViNet-A5sampling_config=120x1x1, FLOPs_G=2812022.01 | 82.7 | 95.7 | — | |
| Mformer-HRPretrain=IN-21K, sampling_config=16x3x10, FLOPs_G=287642022.01 | 82.7 | 96.1 | — | |
| ViViT-L/16x2Views=4x3, GFLOPs=17,3522021.11 | 82.5 | 95.6 | — | |
| X-ViTPretrain=IN-21K, sampling_config=8x3x1, FLOPs_G=4252022.01 | 82.5 | 95.4 | — | |
| ViViT-LPretrain=IN-21K, sampling_config=16x3x4, FLOPs_G=173522022.01 | 82.5 | 95.6 | — | |
| TimeSformer-HRViews=1x3, GFLOPs=5,1102021.11 | 82.4 | 96 | — | |
| TimeSformer-HRPretrain=IN-21K, sampling_config=16x3x1, FLOPs_G=51092022.01 | 82.4 | 96 | — | |
| TimeSformer-HR+ATSViews=1x3, GFLOPs=3,1032021.11 | 82.2 | 96 | — | |
| TimeSformer-LInput=96 x 224^2, Crops=10 x 3, TFLOPs=7.142024.03 | 82.2 | 95.6 | — | |
| TimeSformer-LPretrain=IN-21K, sampling_config=96x3x1, FLOPs_G=71402022.01 | 82.2 | 95.5 | — | |
| X3D-XL+ATFRViews=10x3, GFLOPs=7682021.11 | 82.1 | 95.6 | — | |
| MViT-B, 16x4sampling_config=16x1x5, FLOPs_G=3532022.01 | 82.1 | 95.7 | — | |
| X3D-XLViews=10x3, GFLOPs=1,4522021.11 | 81.9 | 95.5 | — | |
| X3D-XLsampling_config=16x3x10, FLOPs_G=14522022.01 | 81.9 | 95.5 | — | |
| SlowFast R101+NLViews=10x3, GFLOPs=3,4802021.11 | 81.8 | 95.1 | — | |
| SlowFast R101-NLInput=64 x 224^2, Crops=10 x 3, TFLOPs=7.022024.03 | 81.8 | 95.1 | — | |
| SlowFast+NLsampling_config=16x3x10, FLOPs_G=70202022.01 | 81.8 | 95.1 | — | |
| HATNET2021.11 | 81.6 | — | — | |
| LGD-3D R101Views=10x32021.11 | 81.5 | 95.6 | — | |
| LGDPretrain=IN-1K, sampling_config=128xN/A2022.01 | 81.5 | 95.6 | — | |
| SlowFastsampling_config=8x3x10, FLOPs_G=31802022.01 | 80.4 | 94.8 | — | |
| AttentionNASGFLOPs=1,0342021.11 | 79.8 | 94.4 | — | |
| X3D-Msampling_config=16x3x10, FLOPs_G=1862022.01 | 78.8 | 94.5 | — | |
| I3DMFLOPs=111331, Params=12.70M, Batch Size=8, Input Frames=64, Resolution=224x2242019.04 | 71.9 | — | — | |
| ResNeXt-101MFLOPs=9652, Params=48.75M, Speed (Titan XP)=122 cps, Speed (Jetson TX2)=7 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 68.3 | — | — | |
| EVA CLIP-gzero-shot evaluation protocol=True2022.11 | 64.4 | — | — | |
| OpenAI CLIP-Lzero-shot evaluation protocol=True2022.11 | 64.2 | — | — | |
| ResNet-101MFLOPs=13664, Params=83.58M, Speed (Titan XP)=142 cps, Speed (Jetson TX2)=8 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 64.18 | — | — | |
| Open CLIP-Hzero-shot evaluation protocol=True2022.11 | 63.6 | — | — | |
| ResNet-50MFLOPs=9835, Params=44.54M, Speed (Titan XP)=183 cps, Speed (Jetson TX2)=11 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 63 | — | — | |
| Open CLIP-gzero-shot evaluation protocol=True2022.11 | 62.2 | — | — | |
| ResNet-18MFLOPs=8323, Params=33.36M, Speed (Titan XP)=334 cps, Speed (Jetson TX2)=17 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 57.65 | — | — | |
| 3D-ShuffleNetV1 2.0xMFLOPs=538, Params=4.76M, Speed (Titan XP)=161 cps, Speed (Jetson TX2)=24 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 56.84 | — | — | |
| 3D-ShuffleNetV2 2.0xMFLOPs=438, Params=6.64M, Speed (Titan XP)=146 cps, Speed (Jetson TX2)=26 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 55.17 | — | — | |
| 3D-ShuffleNetV1 1.5xMFLOPs=347, Params=2.92M, Speed (Titan XP)=204 cps, Speed (Jetson TX2)=31 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 52.75 | — | — | |
| 3D-ShuffleNetV2 1.5xMFLOPs=291, Params=3.16M, Speed (Titan XP)=186 cps, Speed (Jetson TX2)=34 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 52.05 | — | — | |
| 3D-MobileNetV2 1.0xMFLOPs=561, Params=3.12M, Speed (Titan XP)=93 cps, Speed (Jetson TX2)=9 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 50.65 | — | — | |
| 3D-MobileNetV1 2.0xMFLOPs=662, Params=14.10M, Speed (Titan XP)=88 cps, Speed (Jetson TX2)=15 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 48.53 | — | — | |
| 3D-MobileNetV1 1.5xMFLOPs=429, Params=8.22M, Speed (Titan XP)=116 cps, Speed (Jetson TX2)=19 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 48.24 | — | — | |
| 3D-ShuffleNetV2 1.0xMFLOPs=195, Params=1.91M, Speed (Titan XP)=243 cps, Speed (Jetson TX2)=44 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 46.1 | — | — | |
| 3D-MobileNetV2 0.7xMFLOPs=325, Params=2.05M, Speed (Titan XP)=130 cps, Speed (Jetson TX2)=13 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 45.59 | — | — | |
| 3D-ShuffleNetV1 1.0xMFLOPs=199, Params=1.52M, Speed (Titan XP)=269 cps, Speed (Jetson TX2)=49 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 45.31 | — | — | |
| 3D-SqueezeNetMFLOPs=926, Params=2.15M, Speed (Titan XP)=682 cps, Speed (Jetson TX2)=46 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 40.52 | — | — | |
| 3D-MobileNetV1 1.0xMFLOPs=241, Params=3.91M, Speed (Titan XP)=164 cps, Speed (Jetson TX2)=31 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 40.07 | — | — | |
| 3D-MobileNetV2 0.45xMFLOPs=177, Params=1.40M, Speed (Titan XP)=203 cps, Speed (Jetson TX2)=19 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 36.47 | — | — | |
| 3D-ShuffleNetV1 0.5xMFLOPs=78, Params=0.55M, Speed (Titan XP)=398 cps, Speed (Jetson TX2)=69 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 35.51 | — | — | |
| 3D-MobileNetV1 0.5xMFLOPs=98, Params=1.17M, Speed (Titan XP)=290 cps, Speed (Jetson TX2)=57 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 31.74 | — | — | |
| 3D-ShuffleNetV2 0.25xMFLOPs=116, Params=0.83M, Speed (Titan XP)=442 cps, Speed (Jetson TX2)=82 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 25.73 | — | — | |
| 3D-MobileNetV2 0.2xMFLOPs=63, Params=0.96M, Speed (Titan XP)=357 cps, Speed (Jetson TX2)=42 cps, Batch Size=8, Input Frames=16, Resolution=112x1122019.04 | 24.14 | — | — | |
| ActionCLIPEncoder=ViT-B/16, Frames=32, Zero-shot=true2026.06 | — | — | 67.7 | |
| DiSTEncoder=ViT-L/14, Frames=32, Zero-shot=true2026.06 | — | — | 75 | |
| FROSTEREncoder=ViT-B/16, Frames=8, Zero-shot=true2026.06 | — | — | 74.8 | |
| MAXIEncoder=ViT-B/16, Frames=16/32, Zero-shot=true2026.06 | — | — | 71.5 | |
| MoTEEncoder=ViT-B/16, Frames=8, Zero-shot=true2026.06 | — | — | 70.2 | |
| MoTEEncoder=ViT-L/14, Frames=8, Zero-shot=true2026.06 | — | — | 78.4 | |
| MViT-B-24, 32x3FLOPs=236, Views=5x1, Params=52.9M, Training Protocol=from scratch2021.09 | — | — | 83.8 | |
| MViT-B, 32x3FLOPs=170, Views=5x1, Params=36.8M, Training Protocol=from scratch2021.09 | — | — | 83.4 | |
| Open-MeDeEncoder=ViT-B/16, Frames=8, Zero-shot=true2026.06 | — | — | 73.7 | |
| Open-VCLIPEncoder=ViT-B/16, Frames=8, Zero-shot=true2026.06 | — | — | 73 | |
| Open-VCLIPEncoder=ViT-L/14, Frames=8, Zero-shot=true2026.06 | — | — | 81.1 | |
| OSTEncoder=ViT-B/16, Frames=8, Zero-shot=true2026.06 | — | — | 73.9 | |
| OTIEncoder=ViT-L/14, Frames=8, Zero-shot=true2026.06 | — | — | 70.6 | |
| R3D-RS-200FLOPs=205, Views=10x3, Params=122.0M, Input Frames=32, Training Protocol=from scratch2021.09 | — | — | 83.1 | |
| R3D-RS-200 (48↑)FLOPs=307, Views=10x3, Params=122.0M, Input Frames=48, Training Protocol=from scratch2021.09 | — | — | 83.8 | |
| SlowFastFLOPs=213, Views=10x3, Training Protocol=from scratch, Backbone=R101, Setting=16x82021.09 | — | — | 81.1 |