Action Recognition on SSV2
77.7Top-1 AccV-JEPA 2.1 ViT-G
Evaluation Results
| Method | Links | |
|---|---|---|
| V-JEPA 2.1 ViT-GParam.=2B, Resolution=384x384, Frame clips=16, Temporal crops=2, Spatial crops=32026.03 | 77.7 | |
| V-JEPA 2 ViT-gParam.=1B, Resolution=384x384, Frame clips=16, Temporal crops=2, Spatial crops=32026.03 | 77.3 | |
| V-JEPA 2.1 ViT-gParam.=1B, Resolution=384x384, Frame clips=16, Temporal crops=2, Spatial crops=32026.03 | 76.9 | |
| VJEPA2Type=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 75.3 | |
| V-JEPA ViT-HParam.=600M, Resolution=256x256, Frame clips=16, Temporal crops=2, Spatial crops=32026.03 | 74.3 | |
| CASTGFLOPs/View=3912025.03 | 71.6 | |
| OMNIVORE2025.03 | 71.4 | |
| VJEPAType=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 71.4 | |
| VideoMAE-BGFLOPs/View=180, Parameter-efficient tuning with adapters=true2025.03 | 70.8 | |
| BEVTGFLOPs/View=2822025.03 | 70.6 | |
| DINOv3Param.=7B2026.03 | 70.1 | |
| DINOv3 ViT-H+Param.=0.8B2026.03 | 69.8 | |
| InternVideo2s2-1BParam.=1B, Resolution=256x256, Frame clips=16, Temporal crops=2, Spatial crops=32026.03 | 69.7 | |
| Video Swin-LGFLOPs/View=2822025.03 | 69.6 | |
| ST-AdapterGFLOPs/View=6072025.03 | 69.5 | |
| NExT-Vid-GType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 69.5 | |
| MTV-HRGFLOPs/View=9302025.03 | 68.5 | |
| VideoPrismType=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 68.5 | |
| VideoPrismParam.=1B2026.03 | 68.5 | |
| AIMGFLOPs/View=404, Parameter-efficient tuning with adapters=true2025.03 | 68.1 | |
| MFormer-HRGFLOPs/View=11852025.03 | 68.1 | |
| ORVIT MF2025.03 | 67.9 | |
| MVIT-BGFLOPs/View=1702025.03 | 67.7 | |
| InternVideo2Type=Discriminative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1B, Attentive probing=true2025.12 | 67.3 | |
| InternVideo2-1BParam.=1B2026.03 | 67.3 | |
| NExT-Vid-HType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=600M, Attentive probing=true2025.12 | 67 | |
| MVDType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-L, Param.=200M, Attentive probing=true2025.12 | 66.5 | |
| VideoMAEType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 66.2 | |
| VIVIT FEGFLOPs/View=9902025.03 | 65.9 | |
| OmniMAEType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 65.4 | |
| NExT-Vid-LType=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-L, Param.=300M, Attentive probing=true2025.12 | 63.9 | |
| EVLGFLOPs/View=5922025.03 | 62.4 | |
| TimeSformer-BGFLOPs/View=23802025.03 | 62.4 | |
| RVM-BSize(M)=1172025.12 | 61.4 | |
| VideoMAEv2Type=Generative Pretraining, Pretrained on=Videos, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 61.2 | |
| 4DS-BSize(M)=91, config=dist from e2025.12 | 60.3 | |
| OV-Encoder (Codec)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 60.1 | |
| RVM-SSize(M)=342025.12 | 59.7 | |
| OV-Encoder (Frame)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 58.7 | |
| OV-Encoder (Codec)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 58.5 | |
| DINOv3Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 58.3 | |
| SigLIP2Backbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 58.2 | |
| OV-Encoder (Frame)Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 57.7 | |
| DINOv3Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 57.4 | |
| AIMv2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 57.2 | |
| VideoMAEv2Param.=1B2026.03 | 56.1 | |
| SiamMAE-SSize(M)=272025.12 | 56 | |
| PEcore GParam.=1.9B, Resolution=256x256, Frame clips=16, Temporal crops=2, Spatial crops=32026.03 | 55.4 | |
| AIMv2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 55.1 | |
| SigLIPBackbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 52.7 | |
| SigLIP2Backbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 52.6 | |
| VideoMAE-BSize(M)=872025.12 | 52.3 | |
| SigLIPBackbone=ViT-L/16, Resolution=256, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 50.7 | |
| DINOv2Param.=1.1B, Resolution=256x256, Frame clips=16, Temporal crops=2, Spatial crops=32026.03 | 50.7 | |
| MetaCLIPBackbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 50.6 | |
| DINOv2Type=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.1B, Attentive probing=true2025.12 | 50.6 | |
| IJEPAType=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-H, Param.=630M, Attentive probing=true2025.12 | 50 | |
| Siglip2Type=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.2B, Attentive probing=true2025.12 | 49.9 | |
| SigLIP2Param.=1.2B, Resolution=256x256, Frame clips=16, Temporal crops=2, Spatial crops=32026.03 | 49.9 | |
| 4DS-BSize(M)=912025.12 | 49.6 | |
| MetaCLIP2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=16 Frames / 4096 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 49.3 | |
| CLIPBackbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 48.2 | |
| CLIPGFLOPs/View=140, Parameter-efficient tuning with adapters=true2025.03 | 47.8 | |
| MetaCLIP2Backbone=ViT-L/14, Resolution=224, Input Configuration (Frames/Patches)=8 Frames / 2048 Patches, Evaluation Protocol (Attentive probe with frozen backbones)=Attentive probe with frozen backbones2026.02 | 47.2 | |
| 4DS-SSize(M)=242025.12 | 39.9 | |
| OpenCLIPType=Discriminative Pretraining, Pretrained on=Images, Arch.=ViT-G, Param.=1.8B, Attentive probing=true2025.12 | 34.8 | |
| SimVAK (Shots)=16, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 15.2 | |
| BDC-CLIPK (Shots)=16, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 14.8 | |
| GA2-CLIPK=16, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 14.7 | |
| TC-CLIPK (Shots)=16, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 14 | |
| ViLT-CLIPK=16, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 13.2 | |
| ViFi-CLIPK=16, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 13 | |
| ViFi-CLIPK=16, Adaptation Protocol=Tuning pre-trained image VL models2025.11 | 12.4 | |
| ViFi-CLIPK (Shots)=16, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 12.4 | |
| OSTK (Shots)=16, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 12.2 | |
| ActionCLIPK=16, Adaptation Protocol=Adapting pre-trained image VL models2025.11 | 11.1 | |
| ActionCLIPK (Shots)=16, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 11.1 | |
| GA2-CLIPK=8, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 10.6 | |
| CLIP image-FTK=16, Adaptation Protocol=Tuning pre-trained image VL models2025.11 | 10.4 | |
| ViLT-CLIPK=8, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 10.3 | |
| SimVAK (Shots)=8, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 10.3 | |
| X-CLIPK=16, Adaptation Protocol=Adapting pre-trained image VL models2025.11 | 10 | |
| X-CLIPK (Shots)=16, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 10 | |
| GA2-CLIPK=4, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 9.9 | |
| ViFi-CLIPK=8, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 9.8 | |
| A5K=16, Adaptation Protocol=Adapting pre-trained image VL models2025.11 | 9.7 | |
| A5K (Shots)=16, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 9.7 | |
| BDC-CLIPK (Shots)=8, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 9.6 | |
| ViLT-CLIPK=4, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 9.4 | |
| TC-CLIPK (Shots)=8, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 9.3 | |
| CLIP text-FTK=16, Adaptation Protocol=Tuning pre-trained image VL models2025.11 | 9.1 | |
| SimVAK (Shots)=4, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 9 | |
| OSTK (Shots)=8, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 8.9 | |
| TC-CLIPK (Shots)=4, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 8.6 | |
| ViFi-CLIPK=8, Adaptation Protocol=Tuning pre-trained image VL models2025.11 | 8.5 | |
| ViFi-CLIPK (Shots)=8, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 8.5 | |
| ActionCLIPK=8, Adaptation Protocol=Adapting pre-trained image VL models2025.11 | 8.4 | |
| ActionCLIPK (Shots)=8, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 8.4 | |
| BDC-CLIPK (Shots)=4, Few-shot setting=true, Fine-tuned from CLIP=true2026.05 | 8.3 | |
| ViFi-CLIPK=4, Adaptation Protocol=Prompt tuning pre-trained image VL models2025.11 | 8.1 |