Video Action Recognition on Kinetics-400 (test)
85.3Top-1 AccuracyCAST
Evaluation Results
| Method | Links | |
|---|---|---|
| CASTBackbone=CAST-B, Frames=16, Views=5x3, TFLOPs=5.87, Learnable Param (M)=45, Pre-training=VideoMAE pretrained on Kinetics-4002023.11 | 85.3 | |
| AIMBackbone=ViT-B, Frames=32, Views=3x1, TFLOPs=2.43, Learnable Param (M)=112023.11 | 84.7 | |
| X-CLIPBackbone=ViT-B, Frames=16, Views=4x3, TFLOPs=3.442023.11 | 84.7 | |
| EVLBackbone=ViT-B, Frames=32, Views=3x1, TFLOPs=1.78, Learnable Param (M)=292023.11 | 84.2 | |
| OMNIVOREBackbone=Swin-B, Frames=322023.11 | 84 | |
| Video-FocalNet-BPre-training=ImageNet-1K, Views=4 x 3, FLOPs (G/view)=1492023.07 | 83.6 | |
| Text4VisBackbone=ViT-B, Frames=16, Views=4x32023.11 | 83.6 | |
| Video SwinBackbone=Swin-L, Frames=32, Views=4x3, TFLOPs=7.25, Learnable Param (M)=1972023.11 | 83.1 | |
| Uniformer-BPre-training=ImageNet-1K, Views=4 x 3, FLOPs (G/view)=2592023.07 | 83 | |
| UniFormerBackbone=Hybrid-B, Frames=32, Views=4x3, TFLOPs=3.12, Learnable Param (M)=502023.11 | 83 | |
| MViTv2-BViews=5 x 1, FLOPs (G/view)=2262023.07 | 82.9 | |
| Video-Swin-BPre-training=ImageNet-21K, Views=4 x 3, FLOPs (G/view)=2822023.07 | 82.7 | |
| ST-AdapterBackbone=ViT-B, Frames=32, Views=3x1, TFLOPs=1.822023.11 | 82.7 | |
| MTV-B (320p)Pre-training=ImageNet-21K, Views=4 x 3, FLOPs (G/view)=9672023.07 | 82.4 | |
| MTV-HRBackbone=MTV-B, Frames=32, Views=4x3, TFLOPs=11.16, Learnable Param (M)=3102023.11 | 82.4 | |
| MTV-BPre-training=ImageNet-21K, Views=4 x 3, FLOPs (G/view)=3992023.07 | 81.8 | |
| ViViT-L FEPre-training=ImageNet-21K, Views=1 x 3, FLOPs (G/view)=39802023.07 | 81.7 | |
| VIVIT FEBackbone=ViT-L, Frames=128, Views=1x3, TFLOPs=11.94, Learnable Param (M)=3112023.11 | 81.7 | |
| MoViNet-A6Pre-training=ImageNet-21K, Views=1 x 1, FLOPs (G/view)=3902023.07 | 81.5 | |
| MoViNetBackbone=MoViNet-A6, Frames=120, Views=1x1, TFLOPs=0.392023.11 | 81.5 | |
| VideoMAEBackbone=ViT-B, Frames=16, Views=5x3, TFLOPs=2.7, Learnable Param (M)=872023.11 | 81.5 | |
| Video-FocalNet-SPre-training=ImageNet-1K, Views=4 x 3, FLOPs (G/view)=1242023.07 | 81.4 | |
| MViTv1-BViews=3 x 3, FLOPs (G/view)=4552023.07 | 81.2 | |
| MFormer-HRPre-training=ImageNet-21K, Views=10 x 3, FLOPs (G/view)=9592023.07 | 81.1 | |
| TimeSformer-LPre-training=ImageNet-21K, Views=1 x 3, FLOPs (G/view)=23802023.07 | 80.7 | |
| TimeSformerBackbone=ViT-L, Frames=96, Views=1x3, TFLOPs=25.06, Learnable Param (M)=4302023.11 | 80.7 | |
| Video-Swin-SPre-training=ImageNet-1K, Views=4 x 3, FLOPs (G/view)=1662023.07 | 80.6 | |
| Video-Swin-BPre-training=ImageNet-1K, Views=4 x 3, FLOPs (G/view)=2822023.07 | 80.6 | |
| BEVTBackbone=Swin-B, Frames=32, Views=4x3, TFLOPs=3.38, Learnable Param (M)=882023.11 | 80.6 | |
| OmniSourcePre-training=ImageNet-21K2023.07 | 80.5 | |
| X3D-XXLPre-training=ImageNet-21K, Views=10 x 3, FLOPs (G/view)=1942023.07 | 80.4 | |
| MVITBackbone=ViT-B, Frames=32, Views=5x1, TFLOPs=0.85, Learnable Param (M)=372023.11 | 80.2 | |
| MFormerBackbone=ViT-L, Frames=32, Views=10x3, TFLOPs=35.552023.11 | 80.2 | |
| SlowFast R101-NLPre-training=ImageNet-21K, Views=10 x 3, FLOPs (G/view)=2342023.07 | 79.8 | |
| Video-FocalNet-TPre-training=ImageNet-1K, Views=4 x 3, FLOPs (G/view)=632023.07 | 79.8 | |
| SlowFastBackbone=ResNet101, Frames=80, Views=10x3, TFLOPs=7.022023.11 | 79.8 | |
| LGD-3D R101Pre-training=ImageNet-21K2023.07 | 79.4 | |
| VidTR-LPre-training=ImageNet-21K, Views=10 x 3, FLOPs (G/view)=3512023.07 | 79.1 | |
| X3DBackbone=X3D-XL, Frames=16, Views=10x3, TFLOPs=1.452023.11 | 79.1 | |
| Video-Swin-TPre-training=ImageNet-1K, Views=4 x 3, FLOPs (G/view)=882023.07 | 78.8 | |
| I3D NLPre-training=ImageNet-21K, Views=10 x 3, FLOPs (G/view)=3592023.07 | 77.7 | |
| CLIP*Backbone=ViT-B, Frames=8, Views=5x3, TFLOPs=2.2, Learnable Param (M)=862023.11 | 77.3 | |
| TSM-ResNeXt-101Pre-training=ImageNet-21K2023.07 | 76.3 | |
| TEAPre-training=ImageNet-21K, Views=10 x 3, FLOPs (G/view)=702023.07 | 76.1 |