Action Recognition on Kinetics400 (val)
92.1AccuracyInternVideo2_s1-6B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVideo2_s1-6BTraining Data=IV-2M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 92.1 | — | |
| InternVideo2_s1-6BTraining Data=IV-2M, Setting=8 x 224, Inference Mode=End-to-end finetuning2024.03 | 91.9 | — | |
| InternVideo2_s1-1BTraining Data=IV-1.1M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 91.6 | — | |
| InternVideo2_s1-1BTraining Data=IV-1.1M, Setting=8 x 224, Inference Mode=End-to-end finetuning2024.03 | 91.3 | — | |
| InternVideoTraining Data=V-12M, Setting=ensemble, Inference Mode=End-to-end finetuning2024.03 | 91.1 | — | |
| VideoMAEv2-gTraining Data=V-1.35M, Setting=64 x 266, Inference Mode=End-to-end finetuning2024.03 | 90 | — | |
| UniFormerV2-LTraining Data=IV-401M, Setting=64 x 336, Inference Mode=End-to-end finetuning2024.03 | 90 | — | |
| MTV-HTraining Data=IV-370M, Setting=32 x 280, Inference Mode=End-to-end finetuning2024.03 | 89.9 | — | |
| CoCa-gTraining Data=I-3B, Setting=16 x 576, Inference Mode=End-to-end finetuning2024.03 | 88.9 | — | |
| Hiera-HTraining Data=V-0.25M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 87.8 | — | |
| CoVeR-LPretrain=JFT-3B, #Views=1x32023.04 | 87.2 | — | |
| ST-AdapterPretrain=CLIP, #Frames=32, #Views=1x3, GFLOPs=27492023.04 | 87.2 | — | |
| Text4VisPretrain=CLIP, #Frames=32, #Views=1x3, GFLOPs=16622023.04 | 87.1 | — | |
| X-CLIPPretrain=CLIP, #Frames=8, #Views=4x3, GFLOPs=6582023.04 | 87.1 | — | |
| CoVeRTraining Data=IV-3B, Setting=16 x 448, Inference Mode=End-to-end finetuning2024.03 | 87.1 | — | |
| VicTR (L/14)Pretrain=CLIP, #Frames=8, #Views=4x3, GFLOPs=6562023.04 | 87 | — | |
| Video-SwinV2-G (384↑)Pretrain=IN-21K+, #Frames=8, #Views=4x52023.04 | 86.8 | — | |
| EVLPretrain=CLIP, #Frames=8, #Views=1x3, GFLOPs=6742023.04 | 86.3 | — | |
| MViTv2-L (312↑)#Frames=40, #Views=5x3, GFLOPs=28282023.04 | 86.1 | — | |
| TokenLearnerPretrain=JFT-300M, #Frames=64, #Views=4x3, GFLOPs=40762023.04 | 85.4 | — | |
| Video-Swin-L (384↑)Pretrain=IN-21K, #Frames=32, #Views=10x5, GFLOPs=21072023.04 | 84.9 | — | |
| MTV-LPretrain=JFT-300M, #Frames=32, #Views=4x3, GFLOPs=15042023.04 | 84.3 | — | |
| ViViT-L FEPretrain=JFT-300M, #Frames=32, #Views=1x3, GFLOPs=39802023.04 | 83.5 | — | |
| SMILEBackbone=ViT-B, Epochs=600, Pretrain=K4002025.04 | 83.1 | — | |
| SIGMABackbone=ViT-B, Epochs=800, Pretrain=K4002025.04 | 81.5 | — | |
| MMEBackbone=ViT-B, Epochs=800, Pretrain=K400, evaluated_by_authors=true2025.04 | 81.5 | — | |
| MGMAEBackbone=ViT-B, Epochs=800, Pretrain=K4002025.04 | 81.2 | — | |
| OmniMAEBackbone=ViT-B, Epochs=800, Pretrain=K4002025.04 | 80.8 | — | |
| MGMBackbone=ViT-B, Epochs=800, Pretrain=K4002025.04 | 80.8 | — | |
| TimeSformer-LPretrain=IN-21K, #Frames=96, #Views=1x3, GFLOPs=23802023.04 | 80.7 | — | |
| BEVTBackbone=ViT-B, Pretrain=K4002025.04 | 80.6 | — | |
| VideoMAEBackbone=ViT-B, Epochs=800, Pretrain=K4002025.04 | 80 | — | |
| SVTBackbone=ViT-B, Pretrain=K4002025.04 | 78.4 | — | |
| Supervisedviews=1, T=8, Evaluation Protocol=Linear Evaluation2022.12 | 74.7 | — | |
| SCEviews=2, T=16, Evaluation Protocol=Linear Evaluation2022.12 | 69.6 | — | |
| pBYOL (p=3)views=3, T=8, Evaluation Protocol=Linear Evaluation2022.12 | 68.3 | — | |
| SCEviews=2, T=8, Evaluation Protocol=Linear Evaluation2022.12 | 67.6 | — | |
| pMoCo (p=3)views=3, T=8, Evaluation Protocol=Linear Evaluation2022.12 | 67.3 | — | |
| pSWAV (p=3)views=3, T=8, Evaluation Protocol=Linear Evaluation2022.12 | 62.7 | — | |
| pSimCLR (p=3)views=3, T=8, Evaluation Protocol=Linear Evaluation2022.12 | 62 | — | |
| MJEPA ViT-g + data scalingEncoder Params=1B, Pre-train data=AS+VM2M, Frozen evaluation=true, Attentive probe=true2026.06 | — | 85 | |
| MJEPA ViT-LEncoder Params=300M, Pre-train data=AS, Frozen evaluation=true, Attentive probe=true2026.06 | — | 75.2 | |
| MJEPA ViT-L (video-only loss)Encoder Params=300M, Pre-train data=VM2M, Loss=video-only intra-modal (Lv->v), Frozen evaluation=true, Attentive probe=true2026.06 | — | 80.6 | |
| MJEPA ViT-L + data scalingEncoder Params=300M, Pre-train data=AS+VM2M, Frozen evaluation=true, Attentive probe=true2026.06 | — | 84.7 | |
| VJEPA ViT-HEncoder Params=600M, Pre-train data=VM2M, Frozen evaluation=true, Attentive probe=true2026.06 | — | 82 | |
| VJEPA ViT-LEncoder Params=300M, Pre-train data=VM2M, Frozen evaluation=true, Attentive probe=true2026.06 | — | 80.8 | |
| VJEPA2 ViT-gEncoder Params=1B, Pre-train data=VM22M, Frozen evaluation=true, Attentive probe=true2026.06 | — | 86.6 | |
| VJEPA2 ViT-LEncoder Params=300M, Pre-train data=VM22M, Frozen evaluation=true, Attentive probe=true2026.06 | — | 85.1 |