Video Action Classification on Kinetics-400
0.894Top-1 AccuracyInternVideo2_s2-1B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVideo2_s2-1BParameters=1B, Evaluation Protocol=Attentive probe2025.06 | 0.894 | — | |
| InternVideo2-6BParameters=6B2025.06 | 0.888 | — | |
| PEcoreGParameters=1.9B, Evaluation Protocol=Attentive probe2025.06 | 0.885 | — | |
| InternVideo2-1BParameters=1B2025.06 | 0.879 | — | |
| VideoPrismParameters=1B2025.06 | 0.876 | — | |
| SigLIP2Parameters=1.2B, Evaluation Protocol=Attentive probe2025.06 | 0.873 | — | |
| V-JEPA 2 ViT-g384Parameters=1B, Input Resolution=384x384, Input format=16x8x3 frames2025.06 | 0.873 | — | |
| Video-SwinV2-Gtrain I(W) size=320(20)^2 x 8(8), test I(W) size=384(24)^2 x 8(8), views=4x52021.11 | 0.868 | — | |
| V-JEPA 2 ViT-gParameters=1B, Input format=16x8x3 frames2025.06 | 0.866 | — | |
| TokenLearnertrain I(W) size=256(8)^2 x 64(64), test I(W) size=256(8)^2 x 64(64), views=4x32021.11 | 0.854 | — | |
| V-JEPA 2 ViT-HParameters=600M, Input format=16x8x3 frames2025.06 | 0.853 | — | |
| V-JEPA 2 ViT-LParameters=300M, Input format=16x8x3 frames2025.06 | 0.851 | — | |
| SwinV1-Ltrain I(W) size=480(12)^2 x 16(8), test I(W) size=480(12)^2 x 16(8), views=10x52021.11 | 0.849 | — | |
| ViViTviews=4x32021.11 | 0.848 | — | |
| V-JEPA ViT-HParameters=600M, Evaluation Protocol=Attentive probe2025.06 | 0.845 | — | |
| DINOv2Parameters=1.1B, Evaluation Protocol=Attentive probe2025.06 | 0.836 | — | |
| Video-SwinV2-Gtrain I(W) size=320(20)^2 x 8(8), test I(W) size=384(24)^2 x 8(8), views=1x12021.11 | 0.834 | — | |
| Video-SwinV2-Gtrain I(W) size=320(20)^2 x 8(8), test I(W) size=320(20)^2 x 8(8), views=1x12021.11 | 0.832 | — | |
| VideoMAEv2Parameters=1B2025.06 | 0.828 | — | |
| VATT + BothPre-training Dataset=HT100M + AudioSet + YT8M, Cross-Modality Gradient Realignment (GR)=true, Gradient-based Curriculum Learning (CL)=true2022.11 | 0.8001 | 0.9469 | |
| VATT + CLPre-training Dataset=HT100M + AudioSet, Gradient-based Curriculum Learning (CL)=true2022.11 | 0.7989 | 0.9471 | |
| VATT + GRPre-training Dataset=HT100M + AudioSet + YT8M, Cross-Modality Gradient Realignment (GR)=true2022.11 | 0.7973 | 0.9457 | |
| VATT + CLPre-training Dataset=HT100M + AudioSet + YT8M, Gradient-based Curriculum Learning (CL)=true2022.11 | 0.797 | 0.948 | |
| VATTPre-training Dataset=HT100M + AudioSet + YT8M2022.11 | 0.7939 | 0.9456 | |
| VATT + GRPre-training Dataset=HT100M + AudioSet, Cross-Modality Gradient Realignment (GR)=true2022.11 | 0.7929 | 0.9432 | |
| VATT + BothPre-training Dataset=HT100M + AudioSet, Cross-Modality Gradient Realignment (GR)=true, Gradient-based Curriculum Learning (CL)=true2022.11 | 0.7926 | 0.9448 | |
| VATTPre-training Dataset=HT100M + AudioSet2022.11 | 0.7923 | 0.943 | |
| VATT + RW (VT)Pre-training Dataset=HT100M + AudioSet, Gradient Re-weighting (RW)=Video-Text2022.11 | 0.7859 | 0.9417 | |
| VATT + RW (VT)Pre-training Dataset=HT100M + AudioSet + YT8M, Gradient Re-weighting (RW)=Video-Text2022.11 | 0.7843 | 0.9438 | |
| VATT + RW (VA)Pre-training Dataset=HT100M + AudioSet + YT8M, Gradient Re-weighting (RW)=Video-Audio2022.11 | 0.777 | 0.9383 | |
| VATT + CLPre-training Dataset=HT100M, Gradient-based Curriculum Learning (CL)=true2022.11 | 0.7725 | 0.9338 | |
| VATT + GRPre-training Dataset=HT100M, Cross-Modality Gradient Realignment (GR)=true2022.11 | 0.7672 | 0.9272 | |
| VATT + BothPre-training Dataset=HT100M, Cross-Modality Gradient Realignment (GR)=true, Gradient-based Curriculum Learning (CL)=true2022.11 | 0.7659 | 0.9326 | |
| VATT + RW (VA)Pre-training Dataset=HT100M + AudioSet, Gradient Re-weighting (RW)=Video-Audio2022.11 | 0.7656 | 0.9352 | |
| Supervised (K400)Backbone=R3D-50 (31.7M), Pre-train data=N/A, Modality=V, Evaluation Protocol=N/A2020.08 | 0.76 | — | |
| VATT + RW (VT)Pre-training Dataset=HT100M, Gradient Re-weighting (RW)=Video-Text2022.11 | 0.7566 | 0.928 | |
| VATTPre-training Dataset=HT100M2022.11 | 0.7471 | 0.9269 | |
| CVRLBackbone=R3D-152 (2x) (328.0M), Pre-train data=K600 (11d), Modality=V, Evaluation Protocol=Linear eval on K4002020.08 | 0.716 | — | |
| VATT + RW (VA)Pre-training Dataset=HT100M, Gradient Re-weighting (RW)=Video-Audio2022.11 | 0.7055 | 0.9041 | |
| CVRLBackbone=R3D-101 (59.7M), Pre-train data=K400 (28d), Modality=V, Evaluation Protocol=Linear eval on K4002020.08 | 0.676 | — | |
| CVRLBackbone=R3D-50 (31.7M), Pre-train data=K400 (28d), Modality=V, Evaluation Protocol=Linear eval on K4002020.08 | 0.661 | — | |
| SeCoBackbone=R-50 (23.5M), Pre-train data=K400 (28d), Modality=V, Evaluation Protocol=Linear eval on K4002020.08 | 0.619 | — | |
| TANSettings=finetuned with TAN, Backbone=S3D, Evaluation Protocol=Linear Probing2022.04 | 0.562 | — | |
| MIL-NCESettings=reproduce of [47], Backbone=S3D, Evaluation Protocol=Linear Probing2022.04 | 0.557 | — | |
| ImageNet inflatedBackbone=R3D-50 (31.7M), Pre-train data=ImageNet (N/A), Modality=V, Evaluation Protocol=Linear eval on K4002020.08 | 0.535 | — | |
| VINCEBackbone=R-50 (23.5M), Pre-train data=K400 (28d), Modality=V, Evaluation Protocol=Linear eval on K4002020.08 | 0.491 | — | |
| SimCLR inflatedBackbone=R3D-50 (31.7M), Pre-train data=K400 (28d), Modality=V, Evaluation Protocol=Linear eval on K4002020.08 | 0.468 | — | |
| VTHCLBackbone=R3D-50 (31.7M), Pre-train data=K400 (28d), Modality=V, Evaluation Protocol=Linear eval on K4002020.08 | 0.378 | — |