Video Action Classification on Something-Something v2
86.1Top-1 AccOmniVec-2
Evaluation Results
| Method | Links | |
|---|---|---|
| OmniVec-2Multi-modal alignment=true2024.08 | 86.1 | |
| OmniVecMulti-modal alignment=true2024.08 | 85.4 | |
| InternVideo2Resolution=224x224, Param.=6B, Multi-modal alignment=true2024.08 | 77.4 | |
| V-JEPA 2 ViT-g384Parameters=1B, Evaluation Protocol=Higher resolution protocol, Input Resolution=384x384, Input format=64x2x3 frames2025.06 | 77.3 | |
| MVD-HResolution=224x224, Param.=633M, Multi-modal alignment=false2024.08 | 77.3 | |
| InternVideoResolution=224x224, Param.=1.3B, Multi-modal alignment=true2024.08 | 77.2 | |
| Baseline IIResolution=256x256, Param.=1.9B, Evaluation Protocol=fully fine-tuning2024.08 | 75.9 | |
| GenRecResolution=256x256, Param.=2.1B2024.08 | 75.8 | |
| OmniMAEArch.=ViT-H, Pretrain Data=IN1K + SSv22022.06 | 75.5 | |
| OmniMAE-HResolution=224x224, Param.=650M, Multi-modal alignment=false2024.08 | 75.5 | |
| V-JEPA 2 ViT-gParameters=1B, Evaluation Protocol=Attentive probe on top of frozen encoder, Input Resolution=256x256, Input format=16x2x3 frames2025.06 | 75.3 | |
| Hiera-LResolution=224x224, Param.=214M, Multi-modal alignment=false2024.08 | 75.1 | |
| MaskFeat-LResolution=312x312, Param.=218M, Multi-modal alignment=false2024.08 | 75 | |
| V-JEPA ViT-HParameters=600M, Evaluation Protocol=Attentive probe on top of frozen encoder, Input Resolution=256x2562025.06 | 74.3 | |
| VideoMAE-LResolution=224x224, Param.=305M, Multi-modal alignment=false2024.08 | 74.3 | |
| OmniMAEArch.=ViT-L, Pretrain Data=IN1K + SSv22022.06 | 74.2 | |
| VideoMAE-16fArch.=ViT-L, Pretrain Data=SSv22022.06 | 74.2 | |
| MAE-VideoArch.=ViT-H, Pretrain Data=K4002022.06 | 74.1 | |
| V-JEPA 2 ViT-HParameters=600M, Evaluation Protocol=Attentive probe on top of frozen encoder, Input Resolution=256x256, Input format=16x2x3 frames2025.06 | 74 | |
| V-JEPA 2 ViT-LParameters=300M, Evaluation Protocol=Attentive probe on top of frozen encoder, Input Resolution=256x256, Input format=16x2x3 frames2025.06 | 73.7 | |
| SMILEEvaluation Protocol=Finetuning, Intermediate Dataset=K400 [35]2025.04 | 71.9 | |
| BEVTArch.=Swin-B, Pretrain Data=IN1K + K4002022.06 | 71.4 | |
| UMTEvaluation Protocol=Finetuning, Intermediate Dataset=K700 [8]2025.04 | 70.1 | |
| InternVideo2_s2-1BParameters=1B, Evaluation Protocol=Attentive probe on top of frozen encoder, Input Resolution=256x2562025.06 | 69.7 | |
| Swin TransformerBackbone=Swin-B, Pre-train=K400, #Frames=32x3, GFLOPS=455x32022.01 | 69.6 | |
| DUAL-PathEvaluation Protocol=Finetuning, Intermediate Dataset=None2025.04 | 69.6 | |
| OmniMAEArch.=ViT-B, Pretrain Data=IN1K + SSv22022.06 | 69.5 | |
| OmniMAEArch.=ViT-B, Pretrain Data=IN1K + K4002022.06 | 69 | |
| VideoPrismParameters=1B, Evaluation Protocol=Reported in Literature2025.06 | 68.5 | |
| VIMPACArch.=ViT-L/2+, Pretrain Data=HowTo100M2022.06 | 68.1 | |
| ViCLIPEvaluation Protocol=Finetuning, Intermediate Dataset=Intervid-10M [84]2025.04 | 67.9 | |
| InternVideo2-6BParameters=6B, Evaluation Protocol=Reported in Literature2025.06 | 67.7 | |
| InternVideo2-1BParameters=1B, Evaluation Protocol=Reported in Literature2025.06 | 67.3 | |
| CLIPEvaluation Protocol=Finetuning, Intermediate Dataset=None2025.04 | 66.7 | |
| AIMEvaluation Protocol=Finetuning, Intermediate Dataset=None2025.04 | 66.4 | |
| ILAEvaluation Protocol=Finetuning, Intermediate Dataset=None2025.04 | 65 | |
| OursTDNBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=312022.01 | 64.3 | |
| STMBackbone=ResNet50, Pre-train=IN, #Frames=16x30, GFLOPS=67x302022.01 | 64.2 | |
| TDNBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=362022.01 | 64 | |
| Baseline IResolution=256x256, Param.=2.1B, Evaluation Protocol=attentive-probing2024.08 | 63.7 | |
| SmallBigNetBackbone=ResNet50, Pre-train=IN, #Frames=8+16, GFLOPS=1572022.01 | 63.3 | |
| VidTr-LInput=32 x 22021.04 | 63 | |
| GSTBackbone=ResNet50, Pre-train=IN, #Frames=16, GFLOPS=592022.01 | 62.6 | |
| TANBackbone=ResNet50, Pre-train=IN, #Frames=16, GFLOPS=662022.01 | 62.5 | |
| Timsformer-HRBackbone=ViT, Pre-train=IN, #Frames=16x3, GFLOPS=51102022.01 | 62.5 | |
| Timsformer-LBackbone=ViT, Pre-train=IN, #Frames=96x3, GFLOPS=71402022.01 | 62.3 | |
| TEINetInput=16 (TSN)2021.04 | 62.1 | |
| TEINetBackbone=ResNet50, Pre-train=IN, #Frames=16, GFLOPS=662022.01 | 62.1 | |
| VidTr-MInput=16 x 42021.04 | 61.9 | |
| TAMBackbone=bLResNet50, Pre-train=IN, #Frames=16x2, GFLOPS=47.7x22022.01 | 61.7 | |
| GSTBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=292022.01 | 61.6 | |
| TEINetBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=332022.01 | 61.3 | |
| SF101Input=64 x 22021.04 | 60.9 | |
| AdaFocusBackbone=ResNet50, Pre-train=IN, #Frames=8+12, GFLOPS=33.72022.01 | 60.7 | |
| TANBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=332022.01 | 60.5 | |
| OursTSMBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=282022.01 | 60.2 | |
| AdaFuseBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=31.52022.01 | 59.8 | |
| TimsformerBackbone=ViT, Pre-train=IN, #Frames=8x3, GFLOPS=5902022.01 | 59.5 | |
| TSMInput=8(TSN)2021.04 | 59.3 | |
| TSMBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=332022.01 | 59.1 | |
| STMBackbone=ResNet50, Pre-train=IN, #Frames=8x30, GFLOPS=33x302022.01 | 59 | |
| RubiksNetBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=15.82022.01 | 59 | |
| X-CLIPEvaluation Protocol=Finetuning, Intermediate Dataset=None2025.04 | 57.4 | |
| VideoMAEv2Parameters=1B, Evaluation Protocol=Reported in Literature2025.06 | 56.1 | |
| PEcoreGParameters=1.9B, Evaluation Protocol=Attentive probe on top of frozen encoder, Input Resolution=256x2562025.06 | 55.4 | |
| DINOv2Parameters=1.1B, Evaluation Protocol=Attentive probe on top of frozen encoder, Input Resolution=256x2562025.06 | 50.7 | |
| I3DInput=32x22021.04 | 50 | |
| SigLIP2Parameters=1.2B, Evaluation Protocol=Attentive probe on top of frozen encoder, Input Resolution=256x2562025.06 | 49.9 | |
| TRNBackbone=BN-Inception, Pre-train=IN, #Frames=8, GFLOPS=162022.01 | 48.8 | |
| TrajViT-2Probing Method=attentive probe2026.02 | 48.7 | |
| ViFi-CLIPEvaluation Protocol=Finetuning, Intermediate Dataset=None2025.04 | 48.6 | |
| ViT3DProbing Method=attentive probe2026.02 | 46.3 | |
| TrajViTProbing Method=attentive probe2026.02 | 45.7 | |
| RLTProbing Method=attentive probe2026.02 | 43.6 | |
| ViViTProbing Method=attentive probe2026.02 | 43.1 | |
| TokenLearnerProbing Method=attentive probe2026.02 | 42.4 | |
| TSNBackbone=ResNet50, Pre-train=IN, #Frames=8, GFLOPS=33.22022.01 | 27.8 | |
| SMILEEvaluation Protocol=Linear Probing, Intermediate Dataset=K400 [35]2025.04 | 23.7 | |
| ViCLIPEvaluation Protocol=Linear Probing, Intermediate Dataset=Intervid-10M [84]2025.04 | 18.9 | |
| UMTEvaluation Protocol=Linear Probing, Intermediate Dataset=K700 [8]2025.04 | 18.8 | |
| OSTTuning protocol=Fine-tuned on K400, K=162023.11 | 12.6 | |
| ViFi-CLIPK=16, Adaptation strategy=Tuning pre-trained image VL models2022.12 | 12.4 | |
| ViFi-CLIPTuning protocol=Directly tuning on CLIP, K=162023.11 | 12.4 | |
| MAXITuning protocol=Fine-tuned on K400, K=162023.11 | 12.4 | |
| OSTTuning protocol=Directly tuning on CLIP, K=162023.11 | 12.2 | |
| CLIPEvaluation Protocol=Linear Probing, Intermediate Dataset=None2025.04 | 11.3 | |
| ActionCLIPK=16, Adaptation strategy=Adapting pre-trained image VL models2022.12 | 11.1 | |
| ActionCLIPTuning protocol=Directly tuning on CLIP, K=162023.11 | 11.1 | |
| ViFi-CLIPTuning protocol=Fine-tuned on K400, K=162023.11 | 11 | |
| OSTTuning protocol=Fine-tuned on K400, K=82023.11 | 10.5 | |
| CLIP image-FTK=16, Adaptation strategy=Tuning pre-trained image VL models2022.12 | 10.4 | |
| XCLIPK=16, Adaptation strategy=Adapting pre-trained image VL models2022.12 | 10 | |
| XCLIPTuning protocol=Directly tuning on CLIP, K=162023.11 | 10 | |
| A5K=16, Adaptation strategy=Adapting pre-trained image VL models2022.12 | 9.7 | |
| A5Tuning protocol=Directly tuning on CLIP, K=162023.11 | 9.7 | |
| MAXITuning protocol=Fine-tuned on K400, K=82023.11 | 9.3 | |
| CLIP text-FTK=16, Adaptation strategy=Tuning pre-trained image VL models2022.12 | 9.1 | |
| OSTTuning protocol=Directly tuning on CLIP, K=82023.11 | 8.9 | |
| OSTTuning protocol=Fine-tuned on K400, K=42023.11 | 8.9 | |
| ViFi-CLIPTuning protocol=Fine-tuned on K400, K=82023.11 | 8.6 |