Action Recognition on Epic-Kitchens 100 (test)
73.8Top-1 Verb AccTIM
Evaluation Results
| Method | Links | |||||
|---|---|---|---|---|---|---|
| TIMEnsemble=true2024.04 | 73.8 | 54.5 | 65.6 | — | — | |
| TIMEnsemble=false2024.04 | 73.1 | 53 | 64.1 | — | — | |
| CASTBackbone=CAST-B, Frames=16, Views=2x3, TFLOPs=2.35, Learnable Param (M)=45, pretrained on=EPIC-KITCHENS-1002023.11 | 72.5 | 49.3 | 60.9 | — | — | |
| MoViNetSegment Model=MoViNet [30]2022.01 | 72.2 | 47.7 | 57.3 | — | — | |
| MoViNet-A62023.08 | 72.2 | 47.7 | 57.3 | — | — | |
| MoViNetBackbone=MoViNet-A6, Learnable Param (M)=312023.11 | 72.2 | 47.7 | 57.3 | — | — | |
| TAdaFormer-L/14Pre-training=K7102023.08 | 71.7 | 51.8 | 64.1 | — | — | |
| yzhaoEnsemble=true2024.04 | 71.7 | 54.3 | 65.8 | — | — | |
| MeMViT2023.08 | 71.4 | 48.4 | 60.3 | — | — | |
| AV-MAEPretraining=SSL VGGSound2022.12 | 71.4 | 46 | 56.4 | — | — | |
| TAdaConvNeXtV2-SPre-training=K7102023.08 | 71 | 48.9 | 60.2 | — | — | |
| TAdaFormer-B/16Pre-training=K7102023.08 | 71 | 49.1 | 60.5 | — | — | |
| hrgdscsEnsemble=true2024.04 | 71 | 50.4 | 61.3 | — | — | |
| M&M (Ensemble)Model indices for Action=0,1,2,3,5,6,7,8,9,10, Model indices for Noun and Verb=4,5,6,7,8,9,102022.06 | 70.9 | 52.8 | 66.2 | — | — | |
| xxiongEnsemble=true2024.04 | 70.9 | 52.8 | 66.2 | — | — | |
| AV-MAEPretraining=SSL VGGSound2022.12 | 70.8 | 45.8 | 55.9 | — | — | |
| JaesungEnsemble=true2024.04 | 70.6 | 52.3 | 63.9 | — | — | |
| MeMVITBackbone=ViT-B, Frames=16, Views=1x1, TFLOPs=0.062023.11 | 70.6 | 46.2 | 58.5 | — | — | |
| VideoMAEBackbone=ViT-B, Frames=16, Views=2x3, TFLOPs=1.08, Learnable Param (M)=87, own implementation=true2023.11 | 70.5 | 41.7 | 51.4 | — | — | |
| TAdaConvNeXtV2-TPre-training=K7102023.08 | 70.4 | 47.4 | 58.6 | — | — | |
| MTV-B(WTS)Resolution=280x2802023.08 | 69.9 | 50.5 | 63.9 | — | — | |
| OMNIVOREBackbone=Swin-B, Frames=322023.11 | 69.5 | 49.9 | 61.7 | — | — | |
| ctaiEnsemble=true2024.04 | 69.4 | 50 | 63.3 | — | — | |
| GSFBackbone=ResNet50, Frames=16, Views=3x2, TFLOPs=0.42023.11 | 68.8 | 44 | 52.7 | — | — | |
| X-ViT2023.08 | 68.7 | 44.3 | 56.4 | — | — | |
| ORViT MF-HRPre-train=IN+K4002021.10 | 68.4 | 45.7 | 58.7 | — | — | |
| ORViT MF-HRPretrain=IN-21K + K400, Uses Bounding Boxes=true2021.10 | 68.4 | 45.7 | 58.7 | — | — | |
| ir-CSN-1522023.08 | 68.4 | 44.5 | 55.9 | — | — | |
| ORVIT-MF-HRBackbone=ViT-B, Frames=16, Views=10x32023.11 | 68.4 | 45.7 | 58.7 | — | — | |
| MTV-BResolution=320x3202023.08 | 68 | 48.6 | 63.1 | — | — | |
| MTV-HRBackbone=MTV-B, Frames=32, Views=4x1, TFLOPs=3.72, Learnable Param (M)=3102023.11 | 68 | 48.6 | 63.1 | — | — | |
| TSMSegment Model=TSM [36], Pretraining Supervision=Supervised: action labels, Pretraining Dataset=Kinetics2022.01 | 67.9 | 38.3 | 49 | — | — | |
| TSMModalities=Visual, Optical flow2021.06 | 67.9 | 38.3 | 49 | — | — | |
| TSM2023.08 | 67.9 | 38.3 | 49 | — | — | |
| TSMPretraining=Sup. Im1K + K4002022.12 | 67.9 | 38.3 | 49 | — | — | |
| MTVPretraining=Sup. Im21K + K4002022.12 | 67.8 | 46.7 | 60.5 | — | — | |
| VideoSwinBackbone=Swin-B, Learnable Param (M)=892023.11 | 67.8 | 46.1 | 57 | — | — | |
| ST-Adapter-B/162023.08 | 67.6 | — | 55 | — | — | |
| ST-AdapterBackbone=ViT-B, Frames=8, Views=3x12023.11 | 67.6 | — | 55 | — | — | |
| ViViT-B/16x2 FEResolution=384x3842023.08 | 67.2 | 47 | 59 | — | — | |
| MF-LPre-train=IN+K4002021.10 | 67.1 | 44.1 | 57.6 | — | — | |
| MF-LPretrain=IN-21K + K400, Uses Bounding Boxes=false2021.10 | 67.1 | 44.1 | 57.6 | — | — | |
| TimeSformerSegment Model=TimeSformer, Pretraining Supervision=Unsupervised: distant supervision (ours), Pretraining Dataset=HT100M2022.01 | 67.1 | 44.4 | 58.1 | — | — | |
| TAdaConvNeXtV2-TPre-training=IN1K2023.08 | 67.1 | 42.4 | 53.7 | — | — | |
| MFormerBackbone=ViT-L, Frames=32, Views=3x1, TFLOPs=3.562023.11 | 67.1 | 44.1 | 57.6 | — | — | |
| MF-HRPre-train=IN+K4002021.10 | 67 | 44.5 | 58.5 | — | — | |
| MF-HR + STINPre-train=IN+K4002021.10 | 67 | 44.2 | 57.9 | — | — | |
| MF-HRPretrain=IN-21K + K400, Uses Bounding Boxes=false2021.10 | 67 | 44.5 | 58.5 | — | — | |
| MF-HR + STINPretrain=IN-21K + K400, Uses Bounding Boxes=true2021.10 | 67 | 44.2 | 57.9 | — | — | |
| MotionFormerPretraining=Sup. Im21K + K4002022.12 | 67 | 44.5 | 58.5 | — | — | |
| MF-HR + STRG + STINPre-train=IN+K4002021.10 | 66.9 | 44.1 | 57.8 | — | — | |
| MF-HR + STRG + STINPretrain=IN-21K + K400, Uses Bounding Boxes=true2021.10 | 66.9 | 44.1 | 57.8 | — | — | |
| MFPre-train=IN+K4002021.10 | 66.7 | 43.1 | 56.5 | — | — | |
| MFPretrain=IN-21K + K400, Uses Bounding Boxes=false2021.10 | 66.7 | 43.1 | 56.5 | — | — | |
| TimeSformerSegment Model=TimeSformer [8], Pretraining Supervision=Supervised: action labels, Pretraining Dataset=Kinetics2022.01 | 66.6 | 42.3 | 54.4 | — | — | |
| ViViT-LPre-train=IN+K4002021.10 | 66.4 | 44 | 56.8 | — | — | |
| ViViT-LPretrain=IN-21K + K400, Uses Bounding Boxes=false2021.10 | 66.4 | 44 | 56.8 | — | — | |
| ViViT-LSegment Model=ViViT-L [6], Pretraining Supervision=Supervised: action labels, Pretraining Dataset=Kinetics2022.01 | 66.4 | 44 | 56.8 | — | — | |
| ViViT-L/16x2 FE2023.08 | 66.4 | 44 | 56.8 | — | — | |
| ViViT-L Fact. EncoderPretraining=Sup. Im21K + K4002022.12 | 66.4 | 44 | 56.8 | — | — | |
| VIVIT FEBackbone=ViT-L, Frames=32, Views=4x1, TFLOPs=15.92, Learnable Param (M)=3112023.11 | 66.4 | 44 | 56.8 | — | — | |
| TBNSegment Model=TBN [29]2022.01 | 66 | 36.7 | 47.2 | — | — | |
| TBNModalities=Audio, Visual, Optical flow2021.06 | 66 | 36.7 | 47.2 | — | — | |
| TRNSegment Model=TRN [68]2022.01 | 65.9 | 35.3 | 45.4 | — | — | |
| TRNModalities=Visual, Optical flow2021.06 | 65.9 | 35.3 | 45.4 | — | — | |
| TRN2023.08 | 65.9 | 35.3 | 45.4 | — | — | |
| MF-HR + STRGPre-train=IN+K4002021.10 | 65.8 | 42.5 | 55.4 | — | — | |
| MF-HR + STRGPretrain=IN-21K + K400, Uses Bounding Boxes=true2021.10 | 65.8 | 42.5 | 55.4 | — | — | |
| SlowFast, R50Pre-train=K4002021.10 | 65.6 | 38.5 | 50 | — | — | |
| SlowFast, R50Pretrain=K400, Uses Bounding Boxes=false2021.10 | 65.6 | 38.5 | 50 | — | — | |
| SlowFastSegment Model=SlowFast [17], Pretraining Supervision=Supervised: action labels, Pretraining Dataset=Kinetics2022.01 | 65.6 | 38.5 | 50 | — | — | |
| SlowFastModalities=Visual2021.06 | 65.6 | 38.5 | 50 | — | — | |
| SlowFast2023.08 | 65.6 | 38.5 | 50 | — | — | |
| TAda2D2023.08 | 65.1 | 41.6 | 52.4 | — | — | |
| MBTModalities=Audio, Visual2021.06 | 64.8 | 43.4 | 58 | — | — | |
| MBTPretraining=Sup. Im21K2022.12 | 64.8 | 43.4 | 58 | — | — | |
| AIMBackbone=ViT-B, Frames=16, Views=2x3, TFLOPs=2.42, Learnable Param (M)=14, own implementation=true2023.11 | 64.8 | 41.3 | 55.5 | — | — | |
| MBTBackbone=ViT-B, Frames=322023.11 | 64.8 | 43.4 | 58 | — | — | |
| MBTModalities=Visual2021.06 | 62 | 40.7 | 56.4 | — | — | |
| MBTPretraining=Sup. Im21K2022.12 | 62 | 40.7 | 56.4 | — | — | |
| TSNSegment Model=TSN [61]2022.01 | 60.2 | 33.2 | 46 | — | — | |
| TSNModalities=Visual, Optical flow2021.06 | 60.2 | 33.2 | 46 | — | — | |
| TSN2023.08 | 60.2 | 33.2 | 46 | — | — | |
| CLIPBackbone=ViT-B, Frames=8, Views=2x3, TFLOPs=0.84, Learnable Param (M)=86, own implementation=true2023.11 | 55.5 | 33.9 | 52.3 | — | — | |
| SlowFastBackbone=ResNet502023.11 | 54.9 | 38.5 | 50 | — | — | |
| AV-MAEPretraining=SSL VGGSound2022.12 | 52.7 | 19.7 | 27.2 | — | — | |
| PlayItBackPretraining=Sup. Im21K2022.12 | 47 | 15.9 | 23.1 | — | — | |
| AudioSlowFastModalities=Audio, Pre-training=VGGSound2021.06 | 46.5 | 15.4 | 22.78 | — | — | |
| Kazakos et al.Pretraining=Sup. VGGSound2022.12 | 46.1 | 15.2 | 23 | — | — | |
| MBTModalities=Audio2021.06 | 44.3 | 13 | 22.4 | — | — | |
| MBTPretraining=Sup. Im21K2022.12 | 44.3 | 13 | 22.4 | — | — | |
| Damen et al.Pretraining=Sup. Im1K2022.12 | 42.6 | 14.5 | 22.4 | — | — | |
| Damen et al.Modalities=Audio2021.06 | 42.1 | 14.8 | 21.5 | — | — | |
| RepLAIL_AVC=true, L_AStC=true, MoI Sampling=true, AVC Pretraining=true2022.09 | 31.71 | — | 11.25 | 73.54 | 30.54 | |
| RepLAI w/o AVCL_AVC=false, L_AStC=true, MoI Sampling=true, AVC Pretraining=true2022.09 | 29.92 | — | 10.46 | 70.58 | 29 | |
| RepLAI w/o AStCL_AVC=true, L_AStC=false, MoI Sampling=true, AVC Pretraining=true2022.09 | 29.29 | — | 9.67 | 73.33 | 29.54 | |
| RepLAI w/o MoIL_AVC=true, L_AStC=true, MoI Sampling=false, AVC Pretraining=true2022.09 | 28.71 | — | 8.33 | 73.17 | 27.29 | |
| AVIDL_AVC=false, L_AStC=false, MoI Sampling=false, AVC Pretraining=true2022.09 | 26.62 | — | 9 | 69.79 | 25.5 | |
| RepLAI (scratch)L_AVC=true, L_AStC=true, MoI Sampling=true, AVC Pretraining=false2022.09 | 25.75 | — | 8.12 | 71.25 | 27.29 | |
| XDCL_AVC=false, L_AStC=false, MoI Sampling=false, AVC Pretraining=false2022.09 | 24.46 | — | 6.75 | 68.04 | 22.71 |