Action Detection on AVA v2.1 (val)
32mAPTubeR
Evaluation Results
| Method | Links | |
|---|---|---|
| TubeRDetector=X, Input=32 x 2, Backbone=CSN-152, Pre-train=IG + K400, Inference=2 view, GFLOPS=2402021.04 | 32 | |
| TubeRDetector=X, Input=32 x 2, Backbone=CSN-152, Pre-train=IG + K400, Inference=1 view, GFLOPS=1202021.04 | 31.7 | |
| TubeRDetector=X, Input=32 x 2, Backbone=SF-101, Pre-train=K400 + K700, Inference=1 view, GFLOPS=2402021.04 | 31.6 | |
| AIADetector=F-RCNN, Input=32 x 2, Backbone=SF-101, Pre-train=K400 + K700, Inference=18 views, GFLOPS=NA2021.04 | 31.2 | |
| ACAR-NetBackbone=R-101, Sampling=8 x 8, inputs=V, Pre-training=Kinetics-4002020.06 | 30 | |
| ACAR-NETDetector=F-RCNN, Input=32 x 2, Backbone=SF-101, Pre-train=K400 + K600, Inference=6 views, GFLOPS=NA2021.04 | 30 | |
| TubeRDetector=X, Input=32 x 2, Backbone=CSN-50, Pre-train=K400, Inference=1 view, GFLOPS=782021.04 | 28.8 | |
| TubeRDetector=X, Input=16 x 4, Backbone=I3D-Res101, Pre-train=K400, Inference=1 view, GFLOPS=2462021.04 | 28.6 | |
| ACAR-NetBackbone=R-50, Sampling=8 x 8, inputs=V, Pre-training=Kinetics-4002020.06 | 28.3 | |
| ACAR-NETDetector=F-RCNN, Input=32 x 2, Backbone=SF-50, Pre-train=K400, Inference=6 views, GFLOPS=NA2021.04 | 28.3 | |
| SlowFast*, +NLflow=false, video pretrain=Kinetics-600, Backbone=R101+NL, TxTau=8x8, Region Proposals=Uses authors' region proposals for training2018.12 | 28.2 | |
| SF-101-NLDetector=F-RCNN, Input=32 x 2, Backbone=SF-101 + NL, Pre-train=K400 + K600, Inference=6 views, GFLOPS=9622021.04 | 28.2 | |
| WOODetector=X, Input=8 x 8, Backbone=SF-101, Pre-train=K400 + K600, Inference=1 view, GFLOPS=2462021.04 | 28 | |
| AVSF-50 4x16Pretrain=K4002021.04 | 27.8 | |
| LFBDetector=F-RCNN, Input=32 x 2, Backbone=I3D-101-NL, Pre-train=K400, Inference=18 views, GFLOPS=NA2021.04 | 27.7 | |
| LFBBackbone=R-101+NL, inputs=V, Pre-training=Kinetics-4002020.06 | 27.4 | |
| SlowFast, +NLflow=false, video pretrain=Kinetics-600, Backbone=R101+NL, TxTau=8x82018.12 | 27.3 | |
| CSN-152Detector=F-RCNN, Input=32 x 2, Backbone=CSN-152, Pre-train=IG + K400, Inference=1 views, GFLOPS=3422021.04 | 27.3 | |
| SlowFastflow=false, video pretrain=Kinetics-600, Backbone=R101, TxTau=8x82018.12 | 26.8 | |
| SlowFastflow=false, video pretrain=Kinetics-400, Backbone=R101, TxTau=8x82018.12 | 26.3 | |
| Collaborative Memory (R50+NL)Pretrain=K400, backbone=R50+NL2021.04 | 26.3 | |
| SlowFastBackbone=R-101, Sampling=8 x 8, inputs=V, Pre-training=Kinetics-4002020.06 | 26.3 | |
| TubeRDetector=X, Input=16 x 4, Backbone=I3D-Res50, Pre-train=K400, Inference=1 view, GFLOPS=1322021.04 | 26.1 | |
| X3D-XLDetector=F-RCNN, Input=16 x 5, Backbone=X3D-XL, Pre-train=K400, Inference=1 view, GFLOPS=2902021.04 | 26.1 | |
| LFB (R50+NL)Pretrain=K4002021.04 | 25.8 | |
| Collaborative Memory (SF-50 4x16)Pretrain=K400, backbone=SF-50 4x162021.04 | 25.8 | |
| LFBBackbone=R-50+NL, inputs=V, Pre-training=Kinetics-4002020.06 | 25.8 | |
| 9 model ensembleflow=true, video pretrain=Kinetics-4002018.12 | 25.6 | |
| WOODetector=X, Input=8 x 8, Backbone=SF-50, Pre-train=K400, Inference=1 view, GFLOPS=1422021.04 | 25.2 | |
| AT (I3D)Pretrain=K4002021.04 | 25 | |
| Action TXBackbone=I3D, inputs=V, Pre-training=Kinetics-4002020.06 | 25 | |
| VTrDetector=X, Input=64 x 1, Backbone=I3D-VGG, Pre-train=K400, Inference=1 view, GFLOPS=NA2021.04 | 24.9 | |
| SlowFastBackbone=R-50, Sampling=8 x 8, inputs=V, Pre-training=Kinetics-4002020.06 | 24.8 | |
| Slowfast-50Detector=F-RCNN, Input=16 x 4, Backbone=SF-50, Pre-train=K400, Inference=1 view, GFLOPS=3082021.04 | 24.2 | |
| R50+NLPretrain=K400, reproduced=true2021.04 | 23.6 | |
| SF-50 4x16Pretrain=K400, reproduced=true2021.04 | 23.6 | |
| Zhang et al.Backbone=I3D, inputs=V, Pre-training=Kinetics-4002020.06 | 22.2 | |
| I3Dflow=false, video pretrain=Kinetics-6002018.12 | 21.9 | |
| ATR, R50+NLflow=true, video pretrain=Kinetics-4002018.12 | 21.7 | |
| ATR, R50+NLflow=false, video pretrain=Kinetics-4002018.12 | 20 | |
| STEPDetector=X, Input=32 x 2, Backbone=I3D-VGG, Pre-train=K400, Inference=1 view, GFLOPS=NA2021.04 | 18.6 | |
| ACRN, S3Dflow=true, video pretrain=Kinetics-4002018.12 | 17.4 | |
| ACRNPretrain=K4002021.04 | 17.4 | |
| ACRNBackbone=S3D, inputs=V+F, Pre-training=Kinetics-4002020.06 | 17.4 | |
| ACRNDetector=X, Input=32 x 2, Backbone=S3D-G, Pre-train=K400, Inference=1 view, GFLOPS=NA2021.04 | 17.4 | |
| I3Dflow=true, video pretrain=Kinetics-4002018.12 | 15.6 | |
| I3Dflow=false, video pretrain=Kinetics-4002018.12 | 14.5 | |
| I3DDetector=X, Input=32 x 2, Backbone=I3D-VGG, Pre-train=K400, Inference=1 view, GFLOPS=NA2021.04 | 14.5 |