Action Recognition on Kinetics-400
93.6Top-1 AccOmniVec2
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| OmniVec22025.07 | 93.6 | — | — | — | — | |
| InternVideo2025.07 | 91.1 | — | — | — | — | |
| OmniVec2025.07 | 91.1 | — | — | — | — | |
| InternVideo-T#Params=1.3B2022.12 | 91.1 | — | — | — | — | |
| Tube ViT2025.07 | 90.9 | — | — | — | — | |
| InternVideo-D#Params=1.0B2022.12 | 90.9 | — | — | — | — | |
| UMT-LExtra Data=K710, #P(M)=431, GFLOPS=5860x32025.03 | 90.6 | — | — | — | — | |
| TubeViTArch.=ViT-L, Pre-training Data=K400 + IN1K, Evaluation Protocol=Supervised2023.03 | 90.2 | — | — | — | — | |
| Uniformerv22025.07 | 90 | — | — | — | — | |
| UniFormerV2-LExtra Data=CLIP-400M+K710, #P(M)=354, GFLOPS=12550x62025.03 | 90 | — | — | — | — | |
| FluxViT-Be100Extra Data=K710+MASH, #P(M)=97, GFLOPS=440x12, Evaluation Protocol=Larger resolution2025.03 | 90 | — | — | — | — | |
| MTV-H#Params=1B+2022.12 | 89.9 | — | — | — | — | |
| MTV-HExtra Data=IN-21K+WTS-60M, #P(M)=1000+, GFLOPS=6130x122025.03 | 89.9 | — | — | — | — | |
| FluxViT-Be100Extra Data=K710+MASH, #P(M)=97, GFLOPS=255x12, Evaluation Protocol=Larger resolution2025.03 | 89.7 | — | — | — | — | |
| FluxViT-Be100Extra Data=K710+MASH, #P(M)=97, GFLOPS=440x12, Evaluation Protocol=Standard resolution2025.03 | 89.6 | — | — | — | — | |
| FluxViT-Be200Extra Data=K710+MASH, #P(M)=97, GFLOPS=440x12, Evaluation Protocol=Larger resolution2025.03 | 89.4 | — | — | — | — | |
| FluxViT-Be100Extra Data=K710+MASH, #P(M)=97, GFLOPS=255x12, Evaluation Protocol=Standard resolution2025.03 | 89.3 | — | — | — | — | |
| CoCa#Params=1B+2022.12 | 88.9 | — | — | — | — | |
| CoCa-GExtra Data=JFT-3B+ALIGN-1.8B, #P(M)=1000+, GFLOPS=N/Ax122025.03 | 88.9 | — | — | — | — | |
| FluxViT-Be100Extra Data=K710+MASH, #P(M)=97, GFLOPS=108x12, Evaluation Protocol=Larger resolution2025.03 | 88.9 | — | — | — | — | |
| FluxViT-Be200Extra Data=K710+MASH, #P(M)=97, GFLOPS=440x12, Evaluation Protocol=Standard resolution2025.03 | 88.7 | — | — | — | — | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbone=ViT-E/14, Pre-train=LAION-2B, Frames × Views=16×3×4, TFLOPS=15.96×122023.11 | 88.6 | 98.2 | — | — | — | |
| TubeViTArch.=ViT-B, Pre-training Data=K400 + IN1K, Evaluation Protocol=Supervised2023.03 | 88.6 | — | — | — | — | |
| VideoMAE2-HExtra Data=UnlabeledHybrid, #P(M)=633, GFLOPS=1192x152025.03 | 88.6 | — | — | — | — | |
| InternVideo2-BExtra Data=K710+MASH, #P(M)=96, GFLOPS=440x122025.03 | 88.4 | — | — | — | — | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbone=ViT-E/14, Pre-train=LAION-2B, Frames × Views=8×3×4, TFLOPS=7.98×122023.11 | 88.3 | 98 | — | — | — | |
| BIKE+Evaluation Protocol=Full Fine-tuning, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=16×3×4, TFLOPS=0.83×122023.11 | 88.1 | 97.9 | — | — | — | |
| ATMEvaluation Protocol=Full Fine-tuning, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×3×4, TFLOPS=1.68×122023.11 | 88 | 97.6 | — | — | — | |
| DiST+Evaluation Protocol=Frozen backbone, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×1×3, TFLOPS=2.83×32023.11 | 88 | 97.9 | — | — | — | |
| FluxViT-Se100Extra Data=K710+MASH, #P(M)=24, GFLOPS=154x12, Evaluation Protocol=Larger resolution2025.03 | 88 | — | — | — | — | |
| ATM ViT-Levaluation protocol=Full finetuning, input (#frames×#crops×#clips)=32×12, TFLOPs=1.68×12, text supervision=false2026.04 | 88 | — | — | — | — | |
| DiST ViT-Levaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=32×3, TFLOPs=2.83×3, text supervision=true2026.04 | 88 | — | — | — | — | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K710 + MiT + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 87.8 | — | — | — | — | |
| FluxViT-Se100Extra Data=K710+MASH, #P(M)=24, GFLOPS=154x12, Evaluation Protocol=Standard resolution2025.03 | 87.7 | — | — | — | — | |
| FluxViT-Se100Extra Data=K710+MASH, #P(M)=24, GFLOPS=83x12, Evaluation Protocol=Larger resolution2025.03 | 87.7 | — | — | — | — | |
| XCLIP-Levaluation protocol=Full finetuning, input (#frames×#crops×#clips)=16×12, TFLOPs=3.09×12, text supervision=true2026.04 | 87.7 | — | — | — | — | |
| MOSS-Levaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=16×12, TFLOPs=1.67×12, text supervision=false2026.04 | 87.7 | — | — | — | — | |
| Text4Vis+Evaluation Protocol=Full Fine-tuning, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×3×4, TFLOPS=1.66×122023.11 | 87.6 | 97.8 | — | — | — | |
| Text4Vis-Levaluation protocol=Full finetuning, input (#frames×#crops×#clips)=32×12, TFLOPs=1.66×12, text supervision=true2026.04 | 87.6 | — | — | — | — | |
| AIMEvaluation Protocol=Frozen backbone, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×1×3, TFLOPS=3.74×32023.11 | 87.5 | 97.7 | — | — | — | |
| AIM ViT-Levaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=32×3, TFLOPs=3.74×3, text supervision=false2026.04 | 87.5 | — | — | — | — | |
| UMT-BExtra Data=K710, #P(M)=87, GFLOPS=180x122025.03 | 87.4 | — | — | — | — | |
| FluxViT-Be100Extra Data=K710+MASH, #P(M)=97, GFLOPS=49x12, Evaluation Protocol=Larger resolution2025.03 | 87.4 | — | — | — | — | |
| CLIP4Vis ViT-Levaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=8×12, TFLOPs=0.42×12, text supervision=true2026.04 | 87.4 | — | — | — | — | |
| EVLEvaluation Protocol=Frozen backbone, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×1×3, TFLOPS=2.69×32023.11 | 87.3 | — | — | — | — | |
| FluxViT-Be100Extra Data=K710+MASH, #P(M)=97, GFLOPS=108x12, Evaluation Protocol=Standard resolution2025.03 | 87.3 | — | — | — | — | |
| FluxViT-Se200Extra Data=K710+MASH, #P(M)=24, GFLOPS=154x12, Evaluation Protocol=Larger resolution2025.03 | 87.3 | — | — | — | — | |
| FluxViT-Se100Extra Data=K710+MASH, #P(M)=24, GFLOPS=83x12, Evaluation Protocol=Standard resolution2025.03 | 87.3 | — | — | — | — | |
| COVERPretrain=JFT-3B, Finetune=K400+SSv2+MiT+ImNet, Views=1x32021.12 | 87.2 | — | — | — | — | |
| ST-AdapterEvaluation Protocol=Frozen backbone, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=32×1×3, TFLOPS=2.75×32023.11 | 87.2 | 97.6 | — | — | — | |
| COVERArch.=TimeSFormer-SR, Pre-training Data=JFT-3B+ K400+ MiT + IN1K, Evaluation Protocol=Supervised2023.03 | 87.2 | — | — | — | — | |
| ST-Adapter ViT-Levaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=32×3, TFLOPs=2.75×3, text supervision=false2026.04 | 87.2 | — | — | — | — | |
| CoVeR-LExtra Data=JFT-3B+SMI, #P(M)=431, GFLOPS=5860x32025.03 | 87.1 | — | — | — | — | |
| MaskFeat-L#Params=218M2022.12 | 87 | — | — | — | — | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=16×3×4, TFLOPS=1.74×122023.11 | 87 | 97.5 | — | — | — | |
| Side4Video-Levaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=16×12, TFLOPs=1.74×12, text supervision=false2026.04 | 87 | — | — | — | — | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K400 + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 86.8 | — | — | — | — | |
| FluxViT-Be200Extra Data=K710+MASH, #P(M)=97, GFLOPS=49x12, Evaluation Protocol=Larger resolution2025.03 | 86.7 | — | — | — | — | |
| TrackMAEBackbone=ViT-L, Targets=CLIP, Epochs=800, Pre-training Dataset=K700, Evaluation Protocol=Full finetuning2026.03 | 86.7 | — | — | — | — | |
| TrackMAE†Backbone=ViT-L, Targets=CLIP, Evaluation protocol=Full finetuning, Pre-training dataset=K7002026.03 | 86.7 | — | — | — | — | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbone=ViT-L/14, Pre-train=CLIP-400M, Frames × Views=8×3×4, TFLOPS=0.87×122023.11 | 86.6 | 97.4 | — | — | — | |
| FluxViT-Se100Extra Data=K710+MASH, #P(M)=24, GFLOPS=32x12, Evaluation Protocol=Larger resolution2025.03 | 86.6 | — | — | — | — | |
| FluxViT-Se200Extra Data=K710+MASH, #P(M)=24, GFLOPS=154x12, Evaluation Protocol=Standard resolution2025.03 | 86.4 | — | — | — | — | |
| COVERPretrain=JFT-300M, Finetune=K400+SSv2+MiT+ImNet, Views=1x32021.12 | 86.3 | — | — | — | — | |
| MViTv2-LPre-/Training Data=+(a), gFLOPs=28282022.09 | 86.1 | 97 | — | — | — | |
| MViT-v2-LPre-training=ImageNet-21K, Resolution=312^2, #frames=40, #param.=217.6M, GFLOPs × views=2828×3×52023.08 | 86.1 | — | — | — | — | |
| VideoMAE-LExtra Data=None, #P(M)=305, GFLOPS=3958x212025.03 | 86.1 | — | — | — | — | |
| InternVideo2-SExtra Data=K710+MASH, #P(M)=23, GFLOPS=154x122025.03 | 85.8 | — | — | — | — | |
| MTV-Hevaluation protocol=Full finetuning, input (#frames×#crops×#clips)=32×12, TFLOPs=3.71×12, text supervision=false2026.04 | 85.8 | — | — | — | — | |
| TokenLearnerPretrain=JFT-300M, Finetune=K400, Views=4x32021.12 | 85.4 | — | — | — | — | |
| CASTGFLOPs/View=3912023.11 | 85.3 | — | — | — | — | |
| CASTGFLOPs/View=3912025.03 | 85.3 | — | — | — | — | |
| V-JEPA 2 ViT-Hevaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=16×6, text supervision=false2026.04 | 85.3 | — | — | — | — | |
| VideoMAEArch.=ViT-L, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 85.2 | — | — | — | — | |
| VideoMAEBackbone=ViT-L, Targets=Pixel, Epochs=1600, Pre-training Dataset=K400, Evaluation Protocol=Full finetuning2026.03 | 85.2 | — | — | — | — | |
| VideoMAEBackbone=ViT-L, Targets=Pixel, Evaluation protocol=Full finetuning, Pre-training dataset=K4002026.03 | 85.2 | — | — | — | — | |
| MOSS-Bevaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=32×12, TFLOPs=0.72×12, text supervision=false2026.04 | 85.2 | — | — | — | — | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 85.1 | — | — | — | — | |
| Video SwinTransPretrain=ImageNet21k, Finetune=K400, Views=10×52021.12 | 84.9 | — | — | — | — | |
| Video SwinPre-/Training Data=+(a), gFLOPs=21072022.09 | 84.9 | 96.7 | — | — | — | |
| ViViTEvaluation Protocol=Full Fine-tuning, Backbone=H/14×2, Pre-train=IN-21K, Frames × Views=32×3×4, TFLOPS=3.98×122023.11 | 84.9 | 95.8 | — | — | — | |
| VIC-MAEArch.=ViT-L, Pre-training Data=MiT, Evaluation Protocol=Self-Supervised2023.03 | 84.9 | — | — | — | — | |
| ViViT-HExtra Data=JFT-300M, #P(M)=654, GFLOPS=3981x122025.03 | 84.9 | — | — | — | — | |
| ViViT-Hevaluation protocol=Full finetuning, input (#frames×#crops×#clips)=32×12, TFLOPs=3.98×12, text supervision=false2026.04 | 84.9 | — | — | — | — | |
| Previous SotAModality=V2021.06 | 84.8 | — | — | — | — | |
| ViViTPretrain=JFT-300M, Finetune=K400, Views=4x32021.12 | 84.8 | — | — | — | — | |
| ST-MAEArch.=ViT-L, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 84.8 | — | — | — | — | |
| X-CLIPFrames=16, Views=4 x 32022.12 | 84.7 | 96.8 | 287 | 58.5 | — | |
| FluxViT-Be100Extra Data=K710+MASH, #P(M)=97, GFLOPS=49x12, Evaluation Protocol=Standard resolution2025.03 | 84.7 | — | — | — | — | |
| FluxViT-Se100Extra Data=K710+MASH, #P(M)=24, GFLOPS=32x12, Evaluation Protocol=Standard resolution2025.03 | 84.7 | — | — | — | — | |
| FluxViT-Se100Extra Data=K710+MASH, #P(M)=24, GFLOPS=13x12, Evaluation Protocol=Larger resolution2025.03 | 84.7 | — | — | — | — | |
| AIMGFLOPs/View=4042023.11 | 84.5 | — | — | — | — | |
| AIMGFLOPs/View=404, Parameter-efficient tuning with adapters=true2025.03 | 84.5 | — | — | — | — | |
| V-JEPA ViT-Hevaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=16×6, text supervision=false2026.04 | 84.5 | — | — | — | — | |
| IMD-LPre-training=IN-21K, GFLOPS=8,232, Model Size=Large2022.03 | 84.2 | 96.3 | — | — | — | |
| Side4VideoEvaluation Protocol=Frozen backbone, Backbone=ViT-B/16, Pre-train=CLIP-400M, Frames × Views=32×3×4, TFLOPS=0.72×122023.11 | 84.2 | 96.5 | — | — | — | |
| InternVideo2-s2-6BSetting=16 x 224, Protocol=Linear probing2024.03 | 84.2 | — | — | — | — | |
| Side4Video-Bevaluation protocol=Frozen backbone, input (#frames×#crops×#clips)=32×12, TFLOPs=0.72×12, text supervision=false2026.04 | 84.2 | — | — | — | — | |
| Omnivore2025.07 | 84.1 | — | — | — | — | |
| OMNIVOREArch.=ViT-L, Pre-training Data=IN1K + K400 + SUN RGB-D, Evaluation Protocol=Supervised2023.03 | 84.1 | — | — | — | — |