Action Recognition on K400
96.7Top-1 AccuracyVideo-STAR-7B
Evaluation Results
| Method | Links | |
|---|---|---|
| Video-STAR-7BProtocol=Base-to-Novel (harmonic mean)2026.03 | 96.7 | |
| Video-STAR-7B2025.10 | 96.7 | |
| Gemini-1.5-Pro2025.10 | 92.8 | |
| InternVideo2Protocol=Fine-Tuned2026.03 | 92.1 | |
| Qwen3-VL-8B2025.10 | 89.5 | |
| InternVideo2s2-1BParam.=1B, Resolution=256x256, Frame clips=16, Temporal crops=8, Spatial crops=32026.03 | 89.4 | |
| PEcore GParam.=1.9B, Resolution=256x256, Frame clips=16, Temporal crops=8, Spatial crops=32026.03 | 88.5 | |
| InternVideo2Type=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 87.9 | |
| InternVideo2-1BParam.=1B2026.03 | 87.9 | |
| DINOv3Param.=7B2026.03 | 87.8 | |
| V-JEPA 2.1 ViT-GParam.=2B, Resolution=384x384, Frame clips=16, Temporal crops=8, Spatial crops=32026.03 | 87.7 | |
| VideoPrismType=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 87.6 | |
| VideoPrismParam.=1B2026.03 | 87.6 | |
| Siglip2Type=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.2B, Attentive probing=true2025.12 | 87.3 | |
| SigLIP2Param.=1.2B, Resolution=256x256, Frame clips=16, Temporal crops=8, Spatial crops=32026.03 | 87.3 | |
| V-JEPA 2 ViT-gParam.=1B, Resolution=384x384, Frame clips=16, Temporal crops=8, Spatial crops=32026.03 | 87.3 | |
| V-JEPA 2.1 ViT-gParam.=1B, Resolution=384x384, Frame clips=16, Temporal crops=8, Spatial crops=32026.03 | 87 | |
| DINOv3 ViT-H+Param.=0.8B2026.03 | 86.7 | |
| VJEPA2Type=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 86.6 | |
| Qwen2.5-VL-7B2025.10 | 86.3 | |
| OmniStream2026.03 | 85.7 | |
| V-JEPA2-LModel Size=Large2026.03 | 85.1 | |
| V-JEPA ViT-HParam.=600M, Resolution=256x256, Frame clips=16, Temporal crops=8, Spatial crops=32026.03 | 84.5 | |
| DINOv3-LModel Size=Large2026.03 | 83.6 | |
| DINOv2Param.=1.1B, Resolution=256x256, Frame clips=16, Temporal crops=8, Spatial crops=32026.03 | 83.6 | |
| DINOv2Type=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 83.4 | |
| NExT-Vid-GType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 83.1 | |
| VideoMAEv2Param.=1B2026.03 | 82.8 | |
| VJEPAType=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 82 | |
| OpenCLIPType=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.8B, Attentive probing=true2025.12 | 81.8 | |
| NExT-Vid-HType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=600M, Attentive probing=true2025.12 | 80.6 | |
| VideoMAEType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 79.8 | |
| IJEPAType=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 79.7 | |
| MVDType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-L, Param.=200M, Attentive probing=true2025.12 | 79.4 | |
| NExT-Vid-LType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-L, Param.=300M, Attentive probing=true2025.12 | 78.5 | |
| TotoType=Generative Pretraining, Pretrained on=Videos, Arch.=LLaMA, Param.=1B, Attentive probing=true2025.12 | 74.4 | |
| InternVideo2Protocol=Zero-Shot2026.03 | 73.1 | |
| OmniMAEType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 71.4 | |
| VideoMAEv2Type=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 71.2 |