Action Recognition on Kinetics-600 (test)
89.1Top-1 AccuracyInternVideo2s2-6B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVideo2s2-6BTraining Data=IV-400M, Frames x Resolution=16 x 224, Probing Type=Attentive Probing2024.03 | 89.1 | — | |
| CoCa-gTraining Data=I-3B, Frames x Resolution=16 x 576, Probing Type=Attentive Probing2024.03 | 88.5 | — | |
| InternVideo2s2-1BTraining Data=IV-25.5M, Frames x Resolution=16 x 224, Probing Type=Attentive Probing2024.03 | 88 | — | |
| Swin-L (384↑)Pretrain=ImageNet-21K, Views=10 x 5, FLOPs=2107, Param=200.02021.06 | 86.1 | 97.3 | |
| Swin-L, 384Pretrain=ImageNet-21K, FLOPs (B) × Views=2107 × 10 × 52021.09 | 86.1 | — | |
| Swin-L (384↑)Pretrain=ImageNet-21K, Views=4 x 3, FLOPs=2107, Param=200.02021.06 | 85.9 | 97.1 | |
| ViViT-H/16x2Pretrain=JFT-300M, Views=4 x 3, FLOPs=8316, Param=647.52021.06 | 85.8 | 96.5 | |
| ViViT-H/16×2Pretrain=JFT, FLOPs (B) × Views=3891 × 4 × 32021.09 | 85.8 | — | |
| X-ViTFrames=16, Views=1 x 3, FLOPs (x 10^9)=8502021.06 | 84.5 | 96.3 | |
| SIFA-TransformerBackbone=Swin-B, GFLOPs=270, Views=122022.06 | 84.5 | 96.9 | |
| R3D-RS-200 (48↑)Pretrain=WVT, FLOPs (B) × Views=307 × 10 × 32021.09 | 84.3 | — | |
| Swin-BPretrain=ImageNet-21K, Views=4 x 3, FLOPs=282, Param=88.12021.06 | 84 | 96.5 | |
| Video-SwinBackbone=Swin-B, GFLOPs=282, Views=122022.06 | 84 | 96.5 | |
| Swin-BPretrain=ImageNet-21K, FLOPs (B) × Views=282 × 4 × 32021.09 | 84 | — | |
| MViT-B-24, 32x3Views=5 x 1, FLOPs=236, Param=52.92021.06 | 83.8 | 96.3 | |
| MVITBackbone=ViT-B, GFLOPs=236, Views=52022.06 | 83.8 | 96.3 | |
| SIFA-NetBackbone=R101, GFLOPs=157, Views=30, Input clip length=642022.06 | 83.2 | 95.9 | |
| ViViT-L/16x2 320Pretrain=ImageNet-21K, Views=4 x 3, FLOPs=3992, Param=310.82021.06 | 83 | 95.7 | |
| ViViTBackbone=ViT-L, GFLOPs=3,992, Views=122022.06 | 83 | 95.7 | |
| LGD-3D Two-streamBackbone=ResNet-101, Modality=Two-stream, Split=Test set2019.06 | 82.7 | 96 | |
| ViViT-L/16x2Views=4 x 3, FLOPs (x 10^9)=17,3522021.06 | 82.5 | 95.6 | |
| X-ViTFrames=8, Views=1 x 3, FLOPs (x 10^9)=4252021.06 | 82.5 | 95.4 | |
| TimeSformer-HRViews=1 x 3, FLOPs (x 10^9)=5,1102021.06 | 82.4 | 96 | |
| TimeSformer-HRPretrain=ImageNet-21K, Views=1 x 3, FLOPs=1703, Param=121.42021.06 | 82.4 | 96 | |
| TimeSformer-LBackbone=TimeSformer-L (121.4M), Training Dataset=IN+K600, Training Years=0.1, Protocol=Linear2021.03 | 82.4 | — | |
| TimeSformerBackbone=ViT-B, GFLOPs=1,703, Views=32022.06 | 82.4 | 96 | |
| TimeSformer-LPretrain=ImageNet-21K, FLOPs (B) × Views=2380 × 1 × 32021.09 | 82.2 | — | |
| SIFA-NetBackbone=R50, GFLOPs=112, Views=30, Input clip length=642022.06 | 82.1 | 95.8 | |
| X3D-XLViews=10 x 3, FLOPs (x 10^9)=1,4522021.06 | 81.9 | 95.5 | |
| X3D-XLViews=10 x 3, FLOPs=48, Param=11.02021.06 | 81.9 | 95.5 | |
| X3D-XLBackbone=custom, GFLOPs=48, Views=302022.06 | 81.9 | 95.5 | |
| SlowFast R101+NLViews=10 x 3, FLOPs (x 10^9)=3,4802021.06 | 81.8 | 95.1 | |
| SlowFast R101+NLViews=10 x 3, FLOPs=234, Param=59.92021.06 | 81.8 | 95.1 | |
| SlowFastBackbone=R101+R101, GFLOPs=234, Views=302022.06 | 81.8 | 95.1 | |
| SIFA-NetBackbone=R101, GFLOPs=78, Views=30, Input clip length=322022.06 | 81.6 | 95.5 | |
| LGD-3D R101Views=10 x 32021.06 | 81.5 | 95.6 | |
| SIFA-NetBackbone=R101, GFLOPs=39, Views=30, Input clip length=162022.06 | 80.8 | 95.2 | |
| SIFA-NetBackbone=R50, GFLOPs=51, Views=30, Input clip length=322022.06 | 80.5 | 95.2 | |
| AttentionNASFLOPs (x 10^9)=1,0342021.06 | 79.8 | 94.4 | |
| SIFA-NetBackbone=R50, GFLOPs=25, Views=30, Input clip length=162022.06 | 79.6 | 94.5 | |
| SlowFastBackbone=R50+R50, GFLOPs=36, Views=302022.06 | 78.8 | 94 | |
| TC-CLIPWeight-space ensemble=true, LLM-based text augmentation=true, Category names=LLM-rephrased, Backbone=ViT-B/162024.04 | 78.1 | 95.7 | |
| Oct-I3DImageNet Pretrain=false, Backbone=ResNet-50, Spatial Size=224x224, #FLOPs (G)=25.6, alpha=0.12019.04 | 76 | — | |
| TC-CLIPWeight-space ensemble=true, LLM-based text augmentation=false, Category names=Original, Backbone=ViT-B/162024.04 | 75.8 | 94.4 | |
| OSTWeight-space ensemble=true, LLM-based text augmentation=true, Backbone=ViT-B/162024.04 | 75.1 | 94.6 | |
| FROSTERWeight-space ensemble=true, LLM-based text augmentation=true, Backbone=ViT-B/162024.04 | 74.8 | — | |
| I3DImageNet Pretrain=false, Backbone=ResNet-50, Spatial Size=224x224, #FLOPs (G)=28.12019.04 | 74.3 | — | |
| ViFi-CLIPWeight-space ensemble=true, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 73.9 | 93.3 | |
| Open-VCLIPWeight-space ensemble=true, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 73 | 93.2 | |
| TC-CLIPWeight-space ensemble=false, LLM-based text augmentation=false, Category names=Original, Backbone=ViT-B/162024.04 | 72.7 | 93.2 | |
| I3DBackbone=Inception, GFLOPs=108, Views=N/A2022.06 | 71.9 | 90.1 | |
| MAXIWeight-space ensemble=true, LLM-based text augmentation=true, Backbone=ViT-B/162024.04 | 71.5 | 92.5 | |
| ViFi-CLIPSetting=Zero-shot, Train set=Kinetics-400, Category=Tuning pre-trained image VL models2022.12 | 71.2 | 92.2 | |
| BraVeConfig=V+V×3, Backbone=TSM-50 (23.5M), Training Dataset=K600, Training Years=0.1, Training Modalities=V, Protocol=Linear2021.03 | 70.8 | — | |
| ViFi-CLIPWeight-space ensemble=false, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 70.7 | 92.1 | |
| BraVeConfig=V+A, Backbone=TSM-50 (23.5M), Training Dataset=AS, Training Years=1, Training Modalities=VA, Protocol=Linear2021.03 | 70.6 | — | |
| CVRLBackbone=R3D50 (31.8M), Training Dataset=K600, Training Years=0.1, Training Modalities=V, Protocol=Linear2021.03 | 70.4 | — | |
| BraVeConfig=V+FA, Backbone=TSM-50x2 (93.9M), Training Dataset=AS, Training Years=1, Training Modalities=VFA, Protocol=Linear2021.03 | 70.3 | — | |
| BraVeConfig=V+V×3, Backbone=R3D50 (31.8M), Training Dataset=K600, Training Years=0.1, Training Modalities=V, Protocol=Linear2021.03 | 70 | — | |
| BraVeConfig=V+FA, Backbone=TSM-50 (23.5M), Training Dataset=AS, Training Years=1, Training Modalities=VFA, Protocol=Linear2021.03 | 69.3 | — | |
| BraVeConfig=V+F×3, Backbone=R3D50 (31.8M), Training Dataset=K600, Training Years=0.1, Training Modalities=VF, Protocol=Linear2021.03 | 69.2 | — | |
| BraVeConfig=V+FA, Backbone=R(2+1)D-50 (46.9M), Training Dataset=AS, Training Years=1, Training Modalities=VFA, Protocol=Linear2021.03 | 68.7 | — | |
| CLIP text-FTSetting=Zero-shot, Train set=Kinetics-400, Category=Tuning pre-trained image VL models2022.12 | 68.5 | 89.6 | |
| BraVeConfig=V+F×3, Backbone=TSM-50 (23.5M), Training Dataset=K600, Training Years=0.1, Training Modalities=VF, Protocol=Linear2021.03 | 68.1 | — | |
| ActionCLIPWeight-space ensemble=true, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 67.5 | 90.7 | |
| Vita-CLIPWeight-space ensemble=false, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 67.4 | — | |
| BraVeConfig=V+F×3, Backbone=R3D50 (31.8M), Training Dataset=K400, Training Years=0.07, Training Modalities=VF, Protocol=Linear2021.03 | 66.9 | — | |
| BraVeConfig=V+V×3, Backbone=R3D50 (31.8M), Training Dataset=K400, Training Years=0.07, Training Modalities=V, Protocol=Linear2021.03 | 66.7 | — | |
| ActionCLIPSetting=Zero-shot, Train set=Kinetics-400, Category=Adapting pre-trained image VL models2022.12 | 66.7 | 91.6 | |
| XCLIPSetting=Zero-shot, Train set=Kinetics-400, Category=Adapting pre-trained image VL models2022.12 | 65.2 | 86.1 | |
| X-CLIPWeight-space ensemble=false, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 65.2 | 86.1 | |
| BraVeConfig=V+A, Backbone=R(2+1)D-18 (33.3M), Training Dataset=AS, Training Years=1, Training Modalities=VA, Protocol=Linear2021.03 | 64.1 | — | |
| CLIP image-FTSetting=Zero-shot, Train set=Kinetics-400, Category=Tuning pre-trained image VL models2022.12 | 62.4 | 85.8 | |
| Vanilla CLIPSetting=Zero-shot, Train set=Kinetics-400, Category=Adapting pre-trained image VL models2022.12 | 59.8 | 83.5 | |
| Vanilla CLIPWeight-space ensemble=false, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 59.8 | 83.5 | |
| ActionCLIPWeight-space ensemble=false, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 56.1 | 83.2 | |
| A5Setting=Zero-shot, Train set=Kinetics-400, Category=Adapting pre-trained image VL models2022.12 | 55.8 | 81.4 | |
| A5Weight-space ensemble=false, LLM-based text augmentation=false, Backbone=ViT-B/162024.04 | 55.8 | 81.4 | |
| MMVBackbone=R(2+1)D-18 (33.3M), Training Dataset=AS, Training Years=1, Training Modalities=VA, Protocol=Linear2021.03 | 55.5 | — | |
| ERZSARSetting=Zero-shot, Train set=Kinetics-400, Category=Uni-modal2022.12 | 42.1 | 73.1 | |
| DEMSetting=Zero-shot, Train set=Kinetics-400, Category=Uni-modal2022.12 | 23.6 | 49.5 | |
| ESZSLSetting=Zero-shot, Train set=Kinetics-400, Category=Uni-modal2022.12 | 22.9 | 48.3 | |
| SJESetting=Zero-shot, Train set=Kinetics-400, Category=Uni-modal2022.12 | 22.3 | 48.2 | |
| GCNSetting=Zero-shot, Train set=Kinetics-400, Category=Uni-modal2022.12 | 22.3 | 49.7 |