Action Recognition on Kinetics-600 (val)
91.9Top-1 AccInternVideo2_s1-6B
Evaluation Results
| Method | Links | ||
|---|---|---|---|
| InternVideo2_s1-6BTraining Data=IV-2M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 91.9 | — | |
| InternVideo2_s1-6BTraining Data=IV-2M, Setting=8 x 224, Inference Mode=End-to-end finetuning2024.03 | 91.7 | — | |
| InternVideo2_s1-1BTraining Data=IV-1.1M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 91.6 | — | |
| InternVideo2_s1-1BTraining Data=IV-1.1M, Setting=8 x 224, Inference Mode=End-to-end finetuning2024.03 | 91.4 | — | |
| InternVideoTraining Data=V-12M, Setting=ensemble, Inference Mode=End-to-end finetuning2024.03 | 91.3 | — | |
| MTV-HTraining Data=IV-370M, Setting=32 x 280, Inference Mode=End-to-end finetuning2024.03 | 90.3 | — | |
| UniFormerV2-LTraining Data=IV-401M, Setting=64 x 336, Inference Mode=End-to-end finetuning2024.03 | 90.1 | — | |
| VideoMAEv2-gTraining Data=V-1.35M, Setting=64 x 266, Inference Mode=End-to-end finetuning2024.03 | 89.9 | — | |
| CoCa-gTraining Data=I-3B, Setting=16 x 576, Inference Mode=End-to-end finetuning2024.03 | 89.4 | — | |
| Hiera-HTraining Data=V-0.25M, Setting=16 x 224, Inference Mode=End-to-end finetuning2024.03 | 88.8 | — | |
| CoVeRTraining Data=IV-3B, Setting=16 x 448, Inference Mode=End-to-end finetuning2024.03 | 87.9 | — | |
| MoViNet-A6GFLOPS=386, Param=15.7M, Augmentation=AutoAugment2021.03 | 84.8 | 96.5 | |
| MoViNet-A5GFLOPS=281, Param=15.7M, Augmentation=AutoAugment2021.03 | 84.3 | 96.4 | |
| MoViNet-A6GFLOPS=386, Param=15.7M2021.03 | 83.5 | 96.2 | |
| LGD-3D Two-streamBackbone=ResNet-101, Modality=Two-stream2019.06 | 83.1 | 96.2 | |
| MoViNet-A4GFLOPS=105, Param=4.9M, Augmentation=AutoAugment2021.03 | 83 | 96 | |
| MoViNet-A5GFLOPS=281, Param=15.7M2021.03 | 82.7 | 95.7 | |
| Three-stream iTXNBackbone=mixed, Modality=Three-stream2019.06 | 82.4 | 95.8 | |
| Three-stream AttentionBackbone=mixed, Modality=Three-stream2019.06 | 82.3 | 96 | |
| X3D-XLGFLOPS=1452, Param=11.0M2021.03 | 81.9 | 95.5 | |
| SlowFast 16×8, R101+NLinput sampling=16×8, backbone=ResNet-101, Nonlocal=true, GFLOPS × views=234 × 302018.12 | 81.8 | 95.1 | |
| SlowFastsampling=16x8, backbone=ResNet-101, NL=true, GFLOPS x views=234 x 302018.12 | 81.8 | 95.1 | |
| SlowFast-R101GFLOPS=7020, Param=59.9M2021.03 | 81.8 | 95.1 | |
| LGD-3D RGBBackbone=ResNet-101, Modality=RGB2019.06 | 81.5 | 95.6 | |
| MoViNet-A3GFLOPS=56.9, Param=5.3M, Augmentation=AutoAugment2021.03 | 81.3 | 95.3 | |
| MoViNet-A4GFLOPS=105, Param=4.9M2021.03 | 81.2 | 94.9 | |
| SlowFast 16×8, R101input sampling=16×8, backbone=ResNet-101, GFLOPS × views=213 × 302018.12 | 81.1 | 95.1 | |
| SlowFastsampling=16x8, backbone=ResNet-101, GFLOPS x views=213 x 302018.12 | 81.1 | 95.1 | |
| P3D Two-streamBackbone=ResNet-152, Modality=Two-stream2019.06 | 80.9 | 94.9 | |
| MoViNet-A3GFLOPS=56.9, Param=5.3M2021.03 | 80.8 | 94.5 | |
| SlowFast 8×8, R101input sampling=8×8, backbone=ResNet-101, GFLOPS × views=106 × 302018.12 | 80.4 | 94.8 | |
| SlowFastsampling=8x8, backbone=ResNet-101, GFLOPS x views=106 x 302018.12 | 80.4 | 94.8 | |
| SlowFast 8×8, R50input sampling=8×8, backbone=ResNet-50, GFLOPS × views=65.7 × 302018.12 | 79.9 | 94.5 | |
| SlowFastsampling=8x8, backbone=ResNet-50, GFLOPS x views=65.7 x 302018.12 | 79.9 | 94.5 | |
| StNet-IRv2 RGBpretrain=ImgNet+Kin400, GFLOPS × views=N/A2018.12 | 79 | — | |
| StNet-IRv2 RGBpretrain=ImgNet+Kin400, GFLOPS x views=N/A2018.12 | 79 | — | |
| StNet RGBBackbone=Inception-ResNet-v2, Modality=RGB2019.06 | 78.9 | — | |
| SlowFast 4×16, R50input sampling=4×16, backbone=ResNet-50, GFLOPS × views=36.1 × 302018.12 | 78.8 | 94 | |
| SlowFastsampling=4x16, backbone=ResNet-50, GFLOPS x views=36.1 x 302018.12 | 78.8 | 94 | |
| X3D-MGFLOPS=186, Param=3.8M2021.03 | 78.8 | 94.5 | |
| SlowFast-R50GFLOPS=1080, Param=34.4M2021.03 | 78.8 | 94 | |
| NL I3D RGBBackbone=ResNet-101, Modality=RGB2019.06 | 78.6 | — | |
| P3D RGBBackbone=ResNet-152, Modality=RGB2019.06 | 78.4 | 93.9 | |
| MoViNet-A2GFLOPS=10.3, Param=4.8M, Streaming=false2021.03 | 77.5 | 93.4 | |
| MoViNet-A2GFLOPS=10.4, Param=4.8M, Streaming=true2021.03 | 76.5 | 93.3 | |
| TSN RGBBackbone=SENet-152, Modality=RGB2019.06 | 76.2 | — | |
| MoViNet-A1GFLOPS=6.02, Param=4.6M, Streaming=false2021.03 | 76 | 92.6 | |
| MoViNet-A1GFLOPS=6.06, Param=4.6M, Streaming=true2021.03 | 75.6 | 92.8 | |
| LGD-3D FlowBackbone=ResNet-101, Modality=Flow2019.06 | 75 | 92.4 | |
| I3DGFLOPS × views=108 × N/A2018.12 | 71.9 | 90.1 | |
| I3DGFLOPS x views=108 x N/A2018.12 | 71.9 | 90.1 | |
| MAXIgt=no, language=K400 dict, GPT3 verbs, vis.encoder=ViT-B/16, frames=162023.03 | 71.6 | 92.3 | |
| MAXIgt=no, language=K400 dict, GPT3 verbs, vis.encoder=ViT-B/16, frames=16/322023.03 | 71.6 | 92.4 | |
| MoViNet-A0GFLOPS=2.71, Param=3.1M, Streaming=false2021.03 | 71.5 | 90.4 | |
| MAXIgt=no, language=K400 dict, GPT3 verbs, BLIP verbs, vis.encoder=ViT-B/16, frames=16/322023.03 | 71.5 | 92.5 | |
| MAXIgt=no, language=K400 dict, GPT3 verbs, BLIP verbs, vis.encoder=ViT-B/16, frames=162023.03 | 71.4 | 92.5 | |
| TSN FlowBackbone=SENet-152, Modality=Flow2019.06 | 71.3 | — | |
| ViFi-CLIPgt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=322023.03 | 71.2 | 92.2 | |
| P3D FlowBackbone=ResNet-152, Modality=Flow2019.06 | 71 | 90 | |
| MAXIgt=no, language=K400 dict., vis.encoder=ViT-B/16, frames=162023.03 | 70.4 | 91.5 | |
| MoViNet-A0GFLOPS=2.73, Param=3.1M, Streaming=true2021.03 | 70.3 | 90.1 | |
| Text4Visgt=yes, language=K400 dict., vis.encoder=ViT-L/14, frames=162023.03 | 68.9 | — | |
| ViFi-CLIP (re-eval)gt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=162023.03 | 67.7 | 90.8 | |
| ActionCLIPgt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=322023.03 | 66.7 | 91.6 | |
| XCLIPgt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=322023.03 | 65.2 | 86.1 | |
| CLIPgt=no, vis.encoder=ViT-B/16, frames=162023.03 | 63.5 | 86.8 | |
| A5gt=yes, language=K400 dict., vis.encoder=ViT-B/16, frames=322023.03 | 55.8 | 81.4 | |
| ER-ZSARgt=yes, language=Manual description, vis.encoder=TSM, frames=162023.03 | 42.1 | 73.1 |