Action Recognition on Kinetics-400 (test)
89.7Top-1 AccuracyUniFormerV2-L
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| UniFormerV2-LBackbone=ViT-L, Pretraining=CLIP-400M+K710-0.66M, Views=32 x 3 x 4, Total Parameters (M)=354, Trainable Parameters (M)=354, FLOPs (T)=16.02024.11 | 89.7 | 98.3 | — | — | — | |
| AM/12, TransformerBackbone=ViT-B(Dinov2), Pretraining=IN-21K, Views=32 x 3 x 1, Total Parameters (M)=86+32+45+103, Trainable Parameters (M)=32+45+103, FLOPs (T)=13.82024.11 | 89.6 | 98.4 | — | — | — | |
| MTV-HBackbone=ViT-H+B+S+T, Pretraining=IN-21K+WTS-600M, Views=32 x 3 x 4, Total Parameters (M)=1000+, Trainable Parameters (M)=1000+, FLOPs (T)=44.52024.11 | 89.1 | 98.2 | — | — | — | |
| AM/12, TransformerBackbone=ViT-B(Dinov2), Pretraining=IN-21K, Views=8 x 3 x 1, Total Parameters (M)=86+32+45+103, Trainable Parameters (M)=32+45+103, FLOPs (T)=6.92024.11 | 89.1 | 98.3 | — | — | — | |
| CoCaBackbone=ViT-g, Pretraining=JFT-3B+ALIGN-1.8B, Views=N/A, Total Parameters (M)=1000+, Trainable Parameters (M)=1000+, FLOPs (T)=N/A2024.11 | 88.9 | — | — | — | — | |
| InternVideo2s2-6BTraining Data=IV-400M, Frames x Resolution=16 x 224, Probing Type=Attentive Probing2024.03 | 88.8 | — | — | — | — | |
| AM/12, LSTMBackbone=ViT-B(Dinov2), Pretraining=IN-21K, Views=8 x 3 x 1, Total Parameters (M)=86+32+45+360, Trainable Parameters (M)=32+45+360, FLOPs (T)=5.32024.11 | 88.8 | 98.2 | — | — | — | |
| ViT-22BTraining Data=I-4B, Frames x Resolution=128 x 224, Probing Type=Attentive Probing2024.03 | 88 | — | — | — | — | |
| CoCa-gTraining Data=I-3B, Frames x Resolution=16 x 576, Probing Type=Attentive Probing2024.03 | 88 | — | — | — | — | |
| InternVideo2s2-1BTraining Data=IV-25.5M, Frames x Resolution=16 x 224, Probing Type=Attentive Probing2024.03 | 87.9 | — | — | — | — | |
| UniFormerV2-LBackbone=ViT-L, Pretraining=CLIP-400M, Views=8 x 3 x 4, Total Parameters (M)=354, Trainable Parameters (M)=354, FLOPs (T)=8.02024.11 | 87.7 | 97.9 | — | — | — | |
| DUALPATH-LBackbone=ViT-L, Pretraining=CLIP-400M, Views=32 x 3 x 1, Total Parameters (M)=330, Trainable Parameters (M)=27, FLOPs (T)=1.92024.11 | 87.7 | 97.8 | — | — | — | |
| VideoPrism-gTraining Data=V-619M, Frames x Resolution=16 x 288, Probing Type=Attentive Probing2024.03 | 87.2 | — | — | — | — | |
| ViT-eTraining Data=I-4B, Frames x Resolution=128 x 224, Probing Type=Attentive Probing2024.03 | 86.5 | — | — | — | — | |
| MViT-LBackbone=MViTv2-L, Pretraining=IN-21K, Views=40 x 3 x 5, Total Parameters (M)=218, Trainable Parameters (M)=218, FLOPs (T)=42.42024.11 | 86.1 | 97 | — | — | — | |
| VideoMAE-LBackbone=ViT-L, Pretraining=N/A, Views=40 x 3 x 4, Total Parameters (M)=305, Trainable Parameters (M)=305, FLOPs (T)=47.52024.11 | 86.1 | 97.3 | — | — | — | |
| UniFormerV2-BBackbone=ViT-B, Pretraining=CLIP-400M+K710-0.66M, Views=8 x 3 x 4, Total Parameters (M)=115, Trainable Parameters (M)=115, FLOPs (T)=1.62024.11 | 85.6 | 97 | — | — | — | |
| PoseConv3DModality=RGB + Pose, Fusion=LateFusion2021.04 | 85.5 | — | — | — | — | |
| DUALPATH-BBackbone=ViT-B, Pretraining=CLIP-400M, Views=32 x 3 x 1, Total Parameters (M)=96, Trainable Parameters (M)=10, FLOPs (T)=0.72024.11 | 85.4 | 97.1 | — | — | — | |
| VideoMAE1600eBackbone=ViT-L, Extra labels=false, Frames=16, GFLOPS=597×5×3, Param=305M2023.11 | 85.2 | 96.8 | — | — | — | |
| VideoMAE-LBackbone=ViT-L, Pretraining=N/A, Views=16 x 3 x 5, Total Parameters (M)=305, Trainable Parameters (M)=305, FLOPs (T)=9.02024.11 | 85.2 | 96.8 | — | — | — | |
| VideoSwinModality=RGB2021.04 | 84.9 | — | — | — | — | |
| Swin-L, 384Pretrain=ImageNet-21K, FLOPs (B) × Views=2107 × 10 × 52021.09 | 84.9 | — | — | — | — | |
| ViViT-H/16×2Pretrain=JFT, FLOPs (B) × Views=3981 × 4 × 32021.09 | 84.9 | — | — | — | — | |
| X-CLIPFrames=16, Views=4 x 3, Protocol=Fully-supervised2025.11 | 84.7 | 96.8 | — | — | — | |
| UniFormerV2-BBackbone=ViT-B, Pretraining=CLIP-400M, Views=8 x 3 x 4, Total Parameters (M)=115, Trainable Parameters (M)=115, FLOPs (T)=1.62024.11 | 84.4 | 96.3 | — | — | — | |
| ActionCLIPFrames=32, Views=10 x 3, Protocol=Fully-supervised2025.11 | 83.8 | 96.2 | — | — | — | |
| ActionCLIPBackbone=ViT-B/16, Frames=322021.09 | 83.8 | 97.1 | — | — | — | |
| R3D-RS-200 (48↑)Pretrain=WVT, FLOPs (B) × Views=307 × 10 × 32021.09 | 83.5 | — | — | — | — | |
| DINOv2-gTraining Data=I-142M, Frames x Resolution=16 x 224, Probing Type=Attentive Probing2024.03 | 83.4 | — | — | — | — | |
| VideoSwin-LBackbone=Swin-L, Pretraining=IN-21K, Views=32 x 3 x 4, Total Parameters (M)=197, Trainable Parameters (M)=197, FLOPs (T)=7.22024.11 | 83.1 | 95.9 | — | — | — | |
| SMILEBackbone=ViT-B, Epochs=600, Pretrain=K400, Evaluation Protocol=Full finetuning2025.04 | 83.1 | — | — | — | — | |
| Swin-LFrames=32, Views=4 x 3, Protocol=Fully-supervised2025.11 | 83.1 | 95.9 | — | — | — | |
| Swin-LPretrain=ImageNet-21K, FLOPs (B) × Views=604 × 4 × 32021.09 | 83.1 | — | — | — | — | |
| Uniformer-BFrames=32, Views=4 x 3, Protocol=Fully-supervised2025.11 | 83 | 95.4 | — | — | — | |
| UMT-LTraining Data=IV-25M, Probing Type=Attentive Probing2024.03 | 82.8 | — | — | — | — | |
| ViViT-L/16x2Frames=32, Pre-training=JFT2021.09 | 82.8 | 95.3 | — | — | — | |
| ST-AdapterBackbone=ViT-B, Pretraining=CLIP-400M, Views=32 x 3 x 1, Total Parameters (M)=93, Trainable Parameters (M)=7, FLOPs (T)=1.82024.11 | 82.7 | 96.2 | — | — | — | |
| ActionCLIPBackbone=ViT-B/16, Frames=162021.09 | 82.6 | 96.2 | — | — | — | |
| AMD800eBackbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×5×3, Param=87M2023.11 | 82.2 | 95.3 | — | — | — | |
| VideoMAEv2-gTraining Data=V-1.35M, Probing Type=Attentive Probing2024.03 | 82.1 | — | — | — | — | |
| V-JEPA-HTraining Data=V-2M, Frames x Resolution=16 x 384, Probing Type=Attentive Probing2024.03 | 81.9 | — | — | — | — | |
| ViViT FEBackbone=ViT-L, Extra data=ImageNet-21K, Extra labels=true, Frames=128, GFLOPS=3980×1×32023.11 | 81.7 | 93.8 | — | — | — | |
| ViViT-L/16×2 FEPretrain=ImageNet-21K, FLOPs (B) × Views=3980 × 1 × 32021.09 | 81.7 | — | — | — | — | |
| VideoMAE1600eBackbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×5×3, Param=87M2023.11 | 81.5 | 95.1 | — | — | — | |
| VideoMAE-BBackbone=ViT-B, Pretraining=N/A, Views=16 x 3 x 5, Total Parameters (M)=87, Trainable Parameters (M)=87, FLOPs (T)=2.72024.11 | 81.5 | 95.1 | — | — | — | |
| SIGMABackbone=ViT-B, Epochs=800, Pretrain=K400, Evaluation Protocol=Full finetuning2025.04 | 81.5 | — | — | — | — | |
| MMEBackbone=ViT-B, Epochs=800, Pretrain=K400, Evaluation Protocol=Full finetuning, evaluation_note=results obtained by authors' evaluation2025.04 | 81.5 | — | — | — | — | |
| MAE-STBackbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×7×3, Param=87M2023.11 | 81.3 | 94.9 | — | — | — | |
| MViTv1-BBackbone=MViTv1-B, Pretraining=N/A, Views=64 x 3 x 3, Total Parameters (M)=37, Trainable Parameters (M)=37, FLOPs (T)=4.12024.11 | 81.2 | 95.1 | — | — | — | |
| MGMAEBackbone=ViT-B, Epochs=800, Pretrain=K400, Evaluation Protocol=Full finetuning2025.04 | 81.2 | — | — | — | — | |
| MVIT-BFrames=642021.09 | 81.2 | 95.1 | — | — | — | |
| DMAE300eBackbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×5×3, Param=87M, is_our_implementation=true2023.11 | 80.8 | 94.6 | — | — | — | |
| OmniMAEBackbone=ViT-B, Epochs=800, Pretrain=K400, Evaluation Protocol=Full finetuning2025.04 | 80.8 | — | — | — | — | |
| MGMBackbone=ViT-B, Epochs=800, Pretrain=K400, Evaluation Protocol=Full finetuning2025.04 | 80.8 | — | — | — | — | |
| TimeSformer-LViews=1 x 3, FLOPs (x10^9)=71402021.06 | 80.7 | 94.7 | — | — | — | |
| TimeSformerBackbone=ViT-L, Extra data=ImageNet-21K, Extra labels=true, Frames=96, GFLOPS=8353×1×3, Param=430M2023.11 | 80.7 | 94.7 | — | — | — | |
| TimeSformer-LBackbone=ViT-B, Pretraining=IN-21K, Views=96 x 3 x 1, Total Parameters (M)=121, Trainable Parameters (M)=121, FLOPs (T)=7.12024.11 | 80.7 | 94.7 | — | — | — | |
| TimeSformer-LFrames=96, Views=1 x 3, Protocol=Fully-supervised2025.11 | 80.7 | 94.7 | — | — | — | |
| TimeSformer-LPretrain=ImageNet-21K, FLOPs (B) × Views=2380 × 1 × 32021.09 | 80.7 | — | — | — | — | |
| TimeSformer-LFrames=962021.09 | 80.7 | 94.7 | — | — | — | |
| ViViT-L/16x2Views=4 x 3, FLOPs (x10^9)=173522021.06 | 80.6 | 94.7 | — | — | — | |
| BEVTBackbone=Swin-B, Extra data=IN-1K+DALLE, Extra labels=false, Frames=32, GFLOPS=282×4×3, Param=88M2023.11 | 80.6 | — | — | — | — | |
| Video SwinBackbone=Swin-B, Extra data=ImageNet-21K, Extra labels=true, Frames=32, GFLOPS=282×4×3, Param=88M2023.11 | 80.6 | 94.6 | — | — | — | |
| BEVTBackbone=ViT-B, Pretrain=K400, Evaluation Protocol=Full finetuning2025.04 | 80.6 | — | — | — | — | |
| ViViT-L/16x2Frames=322021.09 | 80.6 | 94.7 | — | — | — | |
| STAMFrames=642021.09 | 80.5 | — | — | — | — | |
| X3D-XXLViews=10 x 3, FLOPs (x10^9)=58232021.06 | 80.4 | 94.6 | — | — | — | |
| X3D-XXLFrames=162021.09 | 80.4 | 94.7 | — | — | — | |
| X-ViTFrames=16x, Views=1 x 3, FLOPs (x10^9)=8502021.06 | 80.2 | 94.7 | — | — | — | |
| MotionformerBackbone=ViT-L, Extra data=ImageNet-21K, Extra labels=true, Frames=32, GFLOPS=1185×10×3, Param=382M2023.11 | 80.2 | 94.8 | — | — | — | |
| VINDLUBackbone=TimeSformer2022.12 | 80.1 | — | — | — | — | |
| AMD800eBackbone=ViT-S, Extra labels=false, Frames=16, GFLOPS=57×5×3, Param=22M2023.11 | 80.1 | 94.5 | — | — | — | |
| VideoMAE800eBackbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×5×3, Param=87M2023.11 | 80 | 94.4 | — | — | — | |
| VideoMAEBackbone=ViT-B, Epochs=800, Pretrain=K400, Evaluation Protocol=Full finetuning2025.04 | 80 | — | — | — | — | |
| SlowFast 16x8 R101+NLViews=10 x 3, FLOPs (x10^9)=70202021.06 | 79.8 | 93.9 | — | — | — | |
| SlowFastPretrain=Scratch, Frame=16+64, FLOPS x Views=234G × 302022.02 | 79.8 | 93.9 | — | — | — | |
| SlowFast+NLBackbone=3D-ResNet-101, Pre-train=None, Frames=(16+64)x30, GFLOPS=234x302021.07 | 79.8 | 93.9 | — | — | — | |
| TVTS (Ours)Backbone=ViT-B, Pre-train Dataset=YT-Temporal, CC3M, WebVid2M, Fine-tuning=true2022.09 | 79.8 | — | — | — | — | |
| SlowFastBackbone=R101+NL, Extra labels=false, Frames=16+64, GFLOPS=234×10×3, Param=60M2023.11 | 79.8 | 93.9 | — | — | — | |
| SlowFastFrames=16+642021.09 | 79.8 | 93.9 | — | — | — | |
| ViT-B-VTNFrames=2502021.09 | 79.8 | 94.2 | — | — | — | |
| MotionformerBackbone=ViT-B, Extra data=ImageNet-21K, Extra labels=true, Frames=32, GFLOPS=370×10×3, Param=109M2023.11 | 79.7 | 94.2 | — | — | — | |
| LGD-3D R1012021.06 | 79.4 | 94.4 | — | — | — | |
| TDNPretrain=ImageNet, Frame=16+8, FLOPS x Views=198G × 302022.02 | 79.4 | 94.4 | — | — | — | |
| TDNEn(RGB+SDM)Backbone=ResNet-101, Pre-train=ImageNet, Frames=(8+16)x5x30, GFLOPS=198x302021.07 | 79.4 | 94.4 | — | — | — | |
| TDNEnBackbone=ResNet101, Extra data=ImageNet-1K, Extra labels=true, Frames=8+16, GFLOPS=198×10×3, Param=88M2023.11 | 79.4 | 94.4 | — | — | — | |
| TDNFrames=8+162021.09 | 79.4 | 93.9 | — | — | — | |
| TANet-152Pretrain=ImageNet, Frame=16, FLOPS x Views=242G x 122022.02 | 79.3 | 94.1 | — | — | — | |
| TANetBackbone=ResNet152, Extra data=ImageNet-1K, Extra labels=true, Frames=16, GFLOPS=242×4×3, Param=59M2023.11 | 79.3 | 94.1 | — | — | — | |
| TANetFrames=162021.09 | 79.3 | 94.1 | — | — | — | |
| CorrNet-101Views=10 x 3, FLOPs (x10^9)=67002021.06 | 79.2 | — | — | — | — | |
| ip-CSN-152Views=10 x 3, FLOPs (x10^9)=32702021.06 | 79.2 | 93.8 | — | — | — | |
| CorrNetPretrain=Scratch, Frame=32, FLOPS x Views=224G x 302022.02 | 79.2 | — | — | — | — | |
| CorrNetFrames=322021.09 | 79.2 | — | — | — | — | |
| X3DPretrain=Scratch, Frame=16, FLOPS x Views=48.4G x 302022.02 | 79.1 | 93.9 | — | — | — | |
| OmniVLBackbone=ViT-B, Pre-train Dataset=mixture of eight datasets, Fine-tuning=true2022.09 | 79.1 | — | — | — | — | |
| OmniVLBackbone=TimeSformer2022.12 | 79.1 | — | — | — | — | |
| EANEn(RGB+LMC)Backbone=ResNet-50, Pre-train=ImageNet, Frames=(8+16)x5x30, GFLOPS=111x302021.07 | 79 | 94.1 | — | — | — | |
| VideoMAEBackbone=ViT-S, Extra labels=false, Frames=16, GFLOPS=57×5×3, Param=22M2023.11 | 79 | 93.8 | — | — | — |