Action Recognition on Charades v1 (test)
45.2mAPSlowFast + NL
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| SlowFast + NLBackbone=ResNet-101, Pretraining=Kinetics-600, Input Sampling (T × τ)=16 × 8, Inference Cost (GFLOPS × views)=234 × 302018.12 | 45.2 | — | — | |
| SlowFastNL=true, pretrain=Kinetics-600, GFLOPS x views=234 x 302018.12 | 45.2 | — | — | |
| SlowFast + NLBackbone=ResNet-101, Pretraining=Kinetics-400, Input Sampling (T × τ)=16 × 8, Inference Cost (GFLOPS × views)=234 × 302018.12 | 42.5 | — | — | |
| SlowFastNL=true, pretrain=Kinetics-400, GFLOPS x views=234 x 302018.12 | 42.5 | — | — | |
| SlowFastBackbone=ResNet-101, Pretraining=Kinetics-400, Input Sampling (T × τ)=16 × 8, Inference Cost (GFLOPS × views)=213 × 302018.12 | 42.1 | — | — | |
| SlowFastpretrain=Kinetics-400, GFLOPS x views=213 x 302018.12 | 42.1 | — | — | |
| TRGBackbone=ResNet-101, # Frames=322019.08 | 40.2 | — | — | |
| NL I3D + GCNsBackbone=ResNet-101, # Frames=322019.08 | 39.7 | — | — | |
| STRGBackbone=ResNet-101 + NL, Pretraining=ImageNet + Kinetics-400, Inference Cost (GFLOPS × views)=630 × 302018.12 | 39.7 | — | — | |
| STRGbackbone=ResNet-101, NL=true, pretrain=ImageNet+Kinetics400, GFLOPS x views=630 x 302018.12 | 39.7 | — | — | |
| I3D + GCNsBackbone=ResNet-101, # Frames=322019.08 | 39.1 | — | — | |
| Slow-onlyBackbone=ResNet-101, Pretraining=Kinetics-400, Input Sampling (T × τ)=16 × 8, Inference Cost (GFLOPS × views)=187 × 302018.12 | 39 | — | — | |
| Slow-onlypretrain=Kinetics-400, GFLOPS x views=187 x 302018.12 | 39 | — | — | |
| TRGBackbone=ResNet-50, # Frames=322019.08 | 38.4 | — | — | |
| TRGBackbone=Inception-V3, # Frames=322019.08 | 37.9 | — | — | |
| TRGBackbone=ResNet-101, # Frames=162019.08 | 37.9 | — | — | |
| NL I3DBackbone=ResNet-101, # Frames=322019.08 | 37.5 | — | — | |
| NL I3D + GCNsBackbone=ResNet-50, # Frames=322019.08 | 37.5 | — | — | |
| NonlocalBackbone=ResNet-101, Pretraining=ImageNet + Kinetics-400, Inference Cost (GFLOPS × views)=544 × 302018.12 | 37.5 | — | — | |
| Nonlocalbackbone=ResNet-101, pretrain=ImageNet+Kinetics400, GFLOPS x views=544 x 302018.12 | 37.5 | — | — | |
| TRGBackbone=ResNet-50, # Frames=162019.08 | 35.8 | — | — | |
| TRGBackbone=ResNet-101, # Frames=82019.08 | 35.7 | — | — | |
| I3DBackbone=ResNet-101, # Frames=322019.08 | 35.5 | — | — | |
| TRGBackbone=Inception-V3, # Frames=162019.08 | 35.1 | — | — | |
| pMoCopre-train=IG-Curated-1M, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 34.9 | — | — | |
| supervisedpre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 34.7 | — | — | |
| TRGBackbone=Inception, # Frames=322019.08 | 33.5 | — | — | |
| pMoCopre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 33.5 | — | — | |
| I3DBackbone=Inception, # Frames=322019.08 | 32.9 | — | — | |
| TRGBackbone=ResNet-50, # Frames=82019.08 | 32.9 | — | — | |
| TRGBackbone=Inception-V3, # Frames=82019.08 | 32.1 | — | — | |
| pMoCopre-train=IG-Uncurated-1M, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 31.3 | — | — | |
| TRGBackbone=Inception, # Frames=162019.08 | 30.7 | — | — | |
| TRGBackbone=Inception, # Frames=82019.08 | 28.4 | — | — | |
| MultiScale TRNBackbone=Inception, # Frames=82019.08 | 25.2 | — | — | |
| MultiScale TRNPretraining=ImageNet, Inference Cost (GFLOPS × views)=N/A2018.12 | 25.2 | — | — | |
| MultiScale TRNpretrain=ImageNet2018.12 | 25.2 | — | — | |
| Asyn-TFBackbone=VGG16, Pretraining=ImageNet, Inference Cost (GFLOPS × views)=N/A2018.12 | 22.4 | — | — | |
| Asyn-TFbackbone=VGG16, pretrain=ImageNet2018.12 | 22.4 | — | — | |
| CoViARBackbone=ResNet-50, Pretraining=ImageNet, Inference Cost (GFLOPS × views)=N/A2018.12 | 21.9 | — | — | |
| CoViARbackbone=ResNet-50, pretrain=ImageNet2018.12 | 21.9 | — | — | |
| ActionVLADInput modality=RGB only, Backbone=BN-inception, Additional components=iDT2017.04 | 21 | 29.9 | — | |
| pBYOLpre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 21 | — | — | |
| Two-streamAdditional components=iDT, Note=best reported2017.04 | 18.6 | — | — | |
| 2-StreamBackbone=VGG16, # Frames=1/202019.08 | 18.6 | — | — | |
| Asyn-TFBackbone=VGG16, # Frames=12019.08 | 18.3 | — | — | |
| 2-Stream + LSTMBackbone=VGG16, # Frames=1/202019.08 | 17.8 | — | — | |
| ActionVLADInput modality=RGB only, Backbone=BN-inception2017.04 | 17.6 | 25.1 | — | |
| RGB streamBackbone=BN-inception, Training strategy=TSN [56] style training, Input modality=RGB2017.04 | 16.8 | 23.1 | — | |
| pSimCLRpre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 11.4 | — | — | |
| pSwAVpre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 10.7 | — | — | |
| supervisedpre-train=scratch, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 7.4 | — | — | |
| AVSlowFast-101+NLPretrain=K400, GFLOPS x crops=278 x 302021.04 | — | — | 43.7 | |
| Collaborative Memory (SlowFast-101+NL 16x8)Pretrain=K400, GFLOPS x crops=277 x 302021.04 | — | — | 44.6 | |
| Collaborative Memory (SlowFast-50 16x8)Pretrain=K400, GFLOPS x crops=135 x 302021.04 | — | — | 42.9 | |
| I3D-101+NLPretrain=ImageNet+K400, GFLOPS x crops=544 x 302021.04 | — | — | 37.5 | |
| LFB (I3D-101+NL)Pretrain=K400, GFLOPS x crops=N/A2021.04 | — | — | 42.5 | |
| SlowFast-101+NLPretrain=K400, GFLOPS x crops=234 x 302021.04 | — | — | 42.5 | |
| SlowFast-101+NL 16x8 (reproduced)Pretrain=K400, GFLOPS x crops=273 x 302021.04 | — | — | 41.3 | |
| SlowFast-50 16x8 (reproduced)Pretrain=K400, GFLOPS x crops=131 x 302021.04 | — | — | 39.4 | |
| STRGPretrain=ImageNet+K400, GFLOPS x crops=630 x 302021.04 | — | — | 39.7 | |
| TimeceptionPretrain=K400, GFLOPS x crops=N/A2021.04 | — | — | 41.1 | |
| TRNPretrain=ImageNet, GFLOPS x crops=N/A2021.04 | — | — | 25.2 |