Action Recognition on Something-Something V2
77.5Top-1 AccuracyInternVideo2_s1-6B
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| InternVideo2_s1-6BTraining Data=IV-2M, Setting=8 x 224, Inference Mode=End-to-end finetuning2024.03 | 77.5 | — | — | |
| InternVideo2_s1-6BTraining Data=IV-2M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 77.4 | — | — | |
| MVD-HTraining Data=IV-1.25M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 77.3 | — | — | |
| InternVideo2022.12 | 77.2 | — | — | |
| InternVideoTraining Data=V-12M, Setting=ensemble, Inference Mode=End-to-end finetuning2024.03 | 77.2 | — | — | |
| InternVideo2_s1-1BTraining Data=IV-1.1M, Setting=8 x 224, Inference Mode=End-to-end finetuning2024.03 | 77.1 | — | — | |
| InternVideo2_s1-1BTraining Data=IV-1.1M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 77.1 | — | — | |
| VideoMAEv2-gTraining Data=V-1.35M, Setting=64 x 266, Inference Mode=End-to-end finetuning2024.03 | 77 | — | — | |
| V-JEPA-HTraining Data=V-2M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 77 | — | — | |
| Hiera-HTraining Data=V-0.25M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 76.5 | — | — | |
| TubeViT-L2022.12 | 76.1 | 95.2 | — | |
| TubeViTArch.=ViT-L, Pre-training Data=K400 + IN1K, Evaluation Protocol=Supervised2023.03 | 76.1 | — | — | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K710 + MiT + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 75.9 | — | — | |
| MaskFeat2022.12 | 75.6 | — | — | |
| VideoMAE2022.12 | 75.4 | 95.2 | — | |
| MaskFeat2022.12 | 75 | 95 | — | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K400 + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 75 | — | — | |
| UMT-L 200eExtra Data=K710*, #F=8, FLOPs (G)=596x6, Param (M)=3042023.03 | 74.7 | — | — | |
| MaskFeat-L 1600e / 312Extra Data=K400, #F=16, FLOPs (G)=2828x3, Param (M)=2182023.03 | 74.4 | — | — | |
| UMT-L 400e#F=8, FLOPs (G)=596x6, Param (M)=3042023.03 | 74.4 | — | — | |
| VideoMAE-L 2400e#F=16, FLOPs (G)=597x6, Param (M)=3052023.03 | 74.3 | — | — | |
| VideoMAE 2400eBackbone=ViT-L, Extra labels=false, Frames=16, GFLOPS=597×2×3, Param=305M, Parameter Group=>100M2023.11 | 74.3 | 94.6 | — | |
| VideoMAEArch.=ViT-L, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 74.3 | — | — | |
| VideoMAE-L 1600eExtra Data=K400*, #F=16, FLOPs (G)=597x6, Param (M)=3052023.03 | 74 | — | — | |
| VIC-MAEArch.=ViT-L, Pre-training Data=MiT, Evaluation Protocol=Self-Supervised2023.03 | 73.8 | — | — | |
| VIC-MAEArch.=ViT-L, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 73.7 | — | — | |
| ST-MAE-L 1600eExtra Data=K700*, #F=16, FLOPs (G)=598x3, Param (M)=3042023.03 | 73.6 | — | — | |
| OmniMAEArch.=ViT-L, Pre-training Data=K400 + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 73.4 | — | — | |
| MViTv2-LPre-/Training Data=+(a), gFLOPs=28282022.09 | 73.3 | 94.1 | — | |
| MViT-v2Training Data=ImageNet-21K, gFLOPS=28282022.09 | 73.3 | 94.1 | — | |
| MViT2022.12 | 73.3 | 94.1 | — | |
| MViTv2-L (312↑)Pretrain=K400+, GFLOPS=8484, Param (M)=213, Tunable Param (M)=213, Views=32×1×32023.02 | 73.3 | 94.1 | — | |
| AMD 800eBackbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×2×3, Param=87M, Parameter Group=55-100M2023.11 | 73.3 | 94 | — | |
| ST-MAEArch.=ViT-L, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 73.2 | — | — | |
| UniFormerV2-LExtra Data=CLIP-400M, #F=32, FLOPs (G)=1718x3, Param (M)=5742023.03 | 73 | — | — | |
| ST-MAE-L 1600eExtra Data=K600*, #F=16, FLOPs (G)=598x3, Param (M)=3042023.03 | 73 | — | — | |
| UniFormerV2-LTraining Data=IV-401M, Setting=64 x 336, Inference Mode=End-to-end finetuning2024.03 | 73 | — | — | |
| ST-MAE-L 1600eExtra Data=K400*, #F=16, FLOPs (G)=598x3, Param (M)=3042023.03 | 72.1 | — | — | |
| OG-ReG-BPretrain=ImageNet-21K+K400, Views=3 × 1, FLOPs=245, Param=96.62026.04 | 71.7 | 93.1 | — | |
| CASTGFLOPs/View=3912023.11 | 71.6 | — | — | |
| OmnivorePretrain=K400+, Views=32×1×32023.02 | 71.4 | 93.5 | — | |
| OMNIVORE2023.11 | 71.4 | — | — | |
| UniFomer-BPretrain=K600+, GFLOPS=777, Param (M)=50, Tunable Param (M)=50, Views=32×1×32023.02 | 71.2 | 92.8 | — | |
| UniFormer-BExtra Data=IN-1K+K400, #F=32, FLOPs (G)=259x3, Param (M)=502023.03 | 71.2 | — | — | |
| UniFormer-B, 32Pretrain=ImageNet-1K + K400, Views=3 × 1, FLOPs=2592026.04 | 71.2 | 92.8 | — | |
| Video-FocalNet-BPre-training=Kinetics 4002023.07 | 71.1 | — | — | |
| COVERPretrain=JFT-3B, Finetune=K600+SSv2+MiT+ImNet, Views=1x32021.12 | 70.9 | — | — | |
| CoVeRPretrain=JFT-3B2023.02 | 70.9 | — | — | |
| COVERArch.=TimeSFormer-SR, Pre-training Data=JFT-3B+ K400+ MiT + IN1K, Evaluation Protocol=Supervised2023.03 | 70.9 | — | — | |
| COVERPretrain=JFT-3B, Finetune=K400+SSv2+MiT+ImNet, Views=1x32021.12 | 70.8 | — | — | |
| VideoMAE-B 2400e#F=16, FLOPs (G)=180x6, Param (M)=872023.03 | 70.8 | — | — | |
| UMT-B 800e#F=8, FLOPs (G)=180x6, Param (M)=872023.03 | 70.8 | — | — | |
| UMT-B 200eExtra Data=K710*, #F=8, FLOPs (G)=180x6, Param (M)=872023.03 | 70.8 | — | — | |
| VideoMAE 2400eBackbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×2×3, Param=87M, Parameter Group=55-100M2023.11 | 70.8 | 92.4 | — | |
| VideoMAEGFLOPs/View=1802023.11 | 70.8 | — | — | |
| CoVeRTraining Data=IV-3B, Setting=16 x 448, Inference Mode=End-to-end finetuning2024.03 | 70.8 | — | — | |
| UniFormerV2-BExtra Data=CLIP-400M, #F=32, FLOPs (G)=375x3, Param (M)=1632023.03 | 70.7 | — | — | |
| COVERPretrain=JFT-3B, Finetune=K700+SSv2+MiT+ImNet, Views=1x32021.12 | 70.6 | — | — | |
| AIM ViT-L/14Pretrain=CLIP, GFLOPS=11508, Param (M)=354, Tunable Param (M)=50, Views=32×1×32023.02 | 70.6 | 92.7 | — | |
| BEVT 800eExtra Data=IN-1K+K400, #F=32, FLOPs (G)=321x3, Param (M)=882023.03 | 70.6 | — | — | |
| BEVTBackbone=Swin-B, Extra data=IN-1K+K400+DALLE, Extra labels=false, Frames=32, GFLOPS=321×1×3, Param=88M, Parameter Group=55-100M2023.11 | 70.6 | — | — | |
| BEVTGFLOPs/View=2822023.11 | 70.6 | — | — | |
| MViTv2-BPretrain=K400, GFLOPS=675, Param (M)=51, Tunable Param (M)=51, Views=40×1×32023.02 | 70.5 | 92.7 | — | |
| MViTv2-BPre-training=Kinetics 4002023.07 | 70.5 | — | — | |
| MViTv2-BExtra Data=K400, #F=64, FLOPs (G)=225x3, Param (M)=512023.03 | 70.5 | — | — | |
| MViTv2-B, 32Pretrain=K400, Views=3 × 1, FLOPs=225, Param=51.12026.04 | 70.5 | 92.7 | — | |
| MultiTrainTraining Data=K700, SSv2, MiT, ActivityNet, gFLOPS=614, Input Resolution=312p2022.09 | 70.4 | 93.1 | — | |
| Uniformer-BPre-training=Kinetics 4002023.07 | 70.4 | — | — | |
| OG-ReG-SPretrain=ImageNet-1K+K400, Views=3 × 1, FLOPs=139, Param=61.92026.04 | 70.4 | 92.3 | — | |
| AMD 800eBackbone=ViT-S, Extra labels=false, Frames=16, GFLOPS=57×2×3, Param=22M, Parameter Group=20-55M2023.11 | 70.2 | 92.5 | — | |
| DMAE 100Backbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×2×3, Param=87M, Parameter Group=55-100M, Implementation=Current Paper Implementation2023.11 | 70 | 92.5 | — | |
| CoVeRExtra Data=JFT-3B+KMI, #F=16, FLOPs (G)=5860x3, Param (M)=4312023.03 | 69.9 | — | — | |
| COVERPretrain=JFT-300M, Finetune=K600+SSv2+MiT+ImNet, Views=1x32021.12 | 69.8 | — | — | |
| VIC-MAEArch.=ViT-B, Pre-training Data=K710 + MiT + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 69.8 | — | — | |
| COVERPretrain=JFT-300M, Finetune=K700+SSv2+MiT+ImNet, Views=1x32021.12 | 69.7 | — | — | |
| VideoMAE-B 1600eExtra Data=K400*, #F=16, FLOPs (G)=180x6, Param (M)=872023.03 | 69.7 | — | — | |
| Video SwinTransPretrain=ImageNet21k+K400, Finetune=SSv2, Views=1x32021.12 | 69.6 | — | — | |
| Video SwinPre-/Training Data=+(a), gFLOPs=21072022.09 | 69.6 | — | — | |
| VideoSwin-BPretrain=K400+, GFLOPS=963, Param (M)=89, Tunable Param (M)=89, Views=32×1×12023.02 | 69.6 | 92.7 | — | |
| Video-Swin-BPre-training=Kinetics 4002023.07 | 69.6 | — | — | |
| TDN_ENExtra Data=IN-1K, #F=87, FLOPs (G)=198x3, Param (M)=882023.03 | 69.6 | — | — | |
| VideoSwin-BExtra Data=IN-21K+K400, #F=32, FLOPs (G)=321x3, Param (M)=882023.03 | 69.6 | — | — | |
| VideoMAE 800eBackbone=ViT-B, Extra labels=false, Frames=16, GFLOPS=180×2×3, Param=87M, Parameter Group=55-100M2023.11 | 69.6 | 92 | — | |
| Video SwinBackbone=Swin-B, Extra data=IN-21K+K400, Extra labels=true, Frames=32, GFLOPS=321×1×3, Param=88M, Parameter Group=55-100M2023.11 | 69.6 | 92.7 | — | |
| TDN EnBackbone=ResNet101x2, Extra data=ImageNet-1K, Extra labels=true, Frames=8+16, GFLOPS=198×1×3, Param=88M, Parameter Group=55-100M2023.11 | 69.6 | 92.2 | — | |
| Video SwinGFLOPs/View=2822023.11 | 69.6 | — | — | |
| VideoMAEArch.=ViT-B, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 69.6 | — | — | |
| Video-Swin-BPretrain=ImageNet-21K+K400, Views=3 × 1, FLOPs=321, Param=88.12026.04 | 69.6 | 92.7 | — | |
| ST-AdapterGFLOPs/View=6072023.11 | 69.5 | — | — | |
| VIC-MAEArch.=ViT-B, Pre-training Data=K400 + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 69.5 | — | — | |
| AIM ViT-L/14Pretrain=CLIP, GFLOPS=5754, Param (M)=354, Tunable Param (M)=50, Views=16×1×32023.02 | 69.4 | 92.3 | — | |
| COVERPretrain=JFT-300M, Finetune=K400+SSv2+MiT+ImNet, Views=1x32021.12 | 69.3 | — | — | |
| MultiTrain (312p)Pre-/Training Data=(e),(f),(g),(h), gFLOPs=6142022.09 | 69.3 | 92.1 | — | |
| ST-MAEArch.=ViT-B, Pre-training Data=K400, Evaluation Protocol=Self-Supervised2023.03 | 69.3 | — | — | |
| PST-BPretrain=ImageNet-21K+K400, Views=3 × 1, FLOPs=247, Param=88.82026.04 | 69.2 | 91.9 | — | |
| MultiTrainTraining Data=K700, SSv2, MiT, ActivityNet, gFLOPS=2242022.09 | 69.1 | 92.2 | — | |
| AIM ViT-B/16Pretrain=CLIP, GFLOPS=2496, Param (M)=100, Tunable Param (M)=14, Views=32×1×32023.02 | 69.1 | 92.2 | — | |
| OmniMAEArch.=ViT-B, Pre-training Data=K400 + IN1K, Evaluation Protocol=Self-Supervised2023.03 | 69 | — | — | |
| MultiTrainPre-/Training Data=(e),(f),(g),(h), gFLOPs=2242022.09 | 68.9 | 91.6 | — | |
| OG-ReG-TPretrain=ImageNet-1K+K400, Views=3 × 1, FLOPs=73, Param=32.42026.04 | 68.9 | 91.2 | — |