Action Detection on AVA v2.2 (val)
46.2mAPFTP-UniFormerV2-L/14
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| FTP-UniFormerV2-L/14FLOPS=970G, Param=402M2024.03 | 46.2 | — | — | |
| FTP-UniFormerV2-B/16FLOPS=482G, Param=136M2024.03 | 43.9 | — | — | |
| Hiera-HFLOPS=1158G, Param=672M2024.03 | 43.3 | — | — | |
| VideoMAE V2FLOPS=4220G, Param=1050M2024.03 | 42.6 | — | — | |
| LART-MVITFLOPS=3780G, Param=1640M2024.03 | 42.6 | — | — | |
| MVD-H (Teacher-H)extra data=IN-1K+K400, extra labels=true, GFLOPs=1192, Param=6332022.12 | 41.1 | — | — | |
| MVD-HFLOPS=1192G, Param=633M2024.03 | 41.1 | — | — | |
| InternVideoFLOPS=8733G, Param=1300M2024.03 | 41 | — | — | |
| UniFormerV2-L/14FLOPS=833G, Param=354M2024.03 | 40.3 | — | — | |
| MVD-H (Teacher-H)extra data=IN-1K+K400, extra labels=false, GFLOPs=1192, Param=6332022.12 | 40.1 | — | — | |
| VideoMAEBackbone=ViT-H, Pre-train Dataset=Kinetics-400, Extra Labels=true, T x tau=16x4, GFLOPs=1192, Param=633, Image size=224x2242022.03 | 39.5 | — | — | |
| VideoMAE ViT-Hextra data=K400, extra labels=true, GFLOPs=1192, Param=6332022.12 | 39.5 | — | — | |
| VideoMAE-HFLOPS=1192G, Param=633M2024.03 | 39.5 | — | — | |
| VideoMAEBackbone=ViT-L, Pre-train Dataset=Kinetics-700, Extra Labels=true, T x tau=16x4, GFLOPs=597, Param=305, Image size=224x2242022.03 | 39.3 | — | — | |
| MaskFeatBackbone=MViT-L, Pre-train Dataset=Kinetics-600, Extra Labels=true, T x tau=40x3, GFLOPs=2828, Param=218, Image size=224x2242022.03 | 38.8 | — | — | |
| MVD-L (Teacher-L)extra data=IN-1K+K400, extra labels=true, GFLOPs=597, Param=3052022.12 | 38.7 | — | — | |
| UniFormerV2-B/16FLOPS=406G, Param=115M2024.03 | 38.4 | — | — | |
| MVD-L (Teacher-L)extra data=IN-1K+K400, extra labels=false, GFLOPs=597, Param=3052022.12 | 37.7 | — | — | |
| MaskFeatBackbone=MViT-L, Pre-train Dataset=Kinetics-400, Extra Labels=true, T x tau=40x3, GFLOPs=2828, Param=218, Image size=224x2242022.03 | 37.5 | — | — | |
| MaskFeat MViT-Lextra data=K400, extra labels=true, GFLOPs=2828, Param=2182022.12 | 37.5 | — | — | |
| MaskFeatFLOPS=2828G, Param=218M2024.03 | 37.5 | — | — | |
| VideoMAEBackbone=ViT-L, Pre-train Dataset=Kinetics-400, Extra Labels=true, T x tau=16x4, GFLOPs=597, Param=305, Image size=224x2242022.03 | 37 | — | — | |
| VideoMAE ViT-Lextra data=K400, extra labels=true, GFLOPs=597, Param=3052022.12 | 37 | — | — | |
| VideoMAE-LFLOPS=597G, Param=305M2024.03 | 37 | — | — | |
| VideoMAEBackbone=ViT-H, Pre-train Dataset=Kinetics-400, Extra Labels=false, T x tau=16x4, GFLOPs=1192, Param=633, Image size=224x2242022.03 | 36.5 | — | — | |
| VideoMAE ViT-Hextra data=K400, extra labels=false, GFLOPs=1192, Param=6332022.12 | 36.5 | — | — | |
| ST-MAE ViT-Hextra data=K400, extra labels=true, GFLOPs=1193, Param=6322022.12 | 36.2 | — | — | |
| ST-MAE-HFLOPS=1193G, Param=632M2024.03 | 36.2 | — | — | |
| VideoMAEBackbone=ViT-L, Pre-train Dataset=Kinetics-700, Extra Labels=false, T x tau=16x4, GFLOPs=597, Param=305, Image size=224x2242022.03 | 36.1 | — | — | |
| ST-MAE ViT-Lextra data=K400, extra labels=true, GFLOPs=598, Param=3042022.12 | 35.7 | — | — | |
| MViTv2-Lextra data=IN-21K+K700, extra labels=true, GFLOPs=2828, Param=2132022.12 | 34.4 | — | — | |
| MViTv2-LFLOPS=2828G, Param=213M2024.03 | 34.4 | — | — | |
| VideoMAEBackbone=ViT-L, Pre-train Dataset=Kinetics-400, Extra Labels=false, T x tau=16x4, GFLOPs=597, Param=305, Image size=224x2242022.03 | 34.3 | — | — | |
| VideoMAE ViT-Lextra data=K400, extra labels=false, GFLOPs=597, Param=3052022.12 | 34.3 | — | — | |
| MVD-B (Teacher-L)extra data=IN-1K+K400, extra labels=true, GFLOPs=180, Param=872022.12 | 34.2 | — | — | |
| TubeRbackbone=CSN-152, pre-train=IG + K400, inference=2 views2021.04 | 33.6 | — | — | |
| MVD-B (Teacher-B)extra data=IN-1K+K400, extra labels=true, GFLOPs=180, Param=872022.12 | 33.6 | — | — | |
| TubeRbackbone=CSN-152, pre-train=IG + K400, inference=1 view2021.04 | 33.4 | — | — | |
| ACAR-Netbackbone=ResNet-101, pre-train=Kinetics-700, T x τ=8 x 82020.06 | 33.3 | — | — | |
| ACAR-Netbackbone=SF-101, pre-train=K700+ K400, inference=6 views2021.04 | 33.3 | — | — | |
| HITFLOPS=622G, Param=198M2024.03 | 32.6 | — | — | |
| AIA(SF-101 8x8)Pre-train=K700, Backbone=SlowFast-101 8x8, extra_information=optical flow, audio, and/or detection2021.04 | 32.3 | — | — | |
| AIAbackbone=ResNet-101 + Non-Local, pre-train=Kinetics-700, T x τ=8 x 82020.06 | 32.3 | — | — | |
| AIA (obj)backbone=SF-101, pre-train=K700+ K400, inference=18 views2021.04 | 32.2 | — | — | |
| VideoMAEBackbone=ViT-B, Pre-train Dataset=Kinetics-400, Extra Labels=true, T x tau=16x4, GFLOPs=180, Param=87, Image size=224x2242022.03 | 31.8 | — | — | |
| VideoMAE ViT-Bextra data=K400, extra labels=true, GFLOPs=180, Param=872022.12 | 31.8 | — | — | |
| CM (SF-101 8x8)Pre-train=K700, Backbone=SlowFast-101 8x82021.04 | 31.6 | — | — | |
| ACAR-Netbackbone=ResNet-101 + Non-Local, pre-train=Kinetics-600, T x τ=8 x 82020.06 | 31.4 | — | — | |
| MVD-B (Teacher-L)extra data=IN-1K+K400, extra labels=false, GFLOPs=180, Param=872022.12 | 31.1 | — | — | |
| SlowFast++Backbone=R101+NL, TxTau=16x8, video pretrain=Kinetics-600, Augmentation=multi-scale and horizontal flipping2018.12 | 30.7 | — | — | |
| SlowFasttemporal resolution=16x8, video pretrain=Kinetics-600, Backbone=R-101+NL, flow=false, augmentation=multi-scale and horizontal flipping, region proposals=used for training2018.12 | 30.7 | — | — | |
| SlowFastBackbone=R101+NL, TxTau=16x8, video pretrain=Kinetics-6002018.12 | 29.8 | — | — | |
| SlowFasttemporal resolution=16x8, video pretrain=Kinetics-600, Backbone=R-101+NL, flow=false, region proposals=used for training2018.12 | 29.8 | — | — | |
| AIA(SF-50 4x16)Pre-train=K700, Backbone=SlowFast-50 4x162021.04 | 29.8 | — | — | |
| CM (SF-50 4x16)Pre-train=K700, Backbone=SlowFast-50 4x162021.04 | 29.8 | — | — | |
| SlowFast-101backbone=SF-101, pre-train=K600+ K400, inference=6 views2021.04 | 29.8 | — | — | |
| MVD-B (Teacher-B)extra data=IN-1K+K400, extra labels=false, GFLOPs=180, Param=872022.12 | 29.3 | — | — | |
| TubeRbackbone=CSN-50, pre-train=IG + K400, inference=1 view2021.04 | 29.2 | — | — | |
| SlowFastBackbone=R101+NL, TxTau=8x8, video pretrain=Kinetics-6002018.12 | 29 | — | — | |
| SlowFasttemporal resolution=8x8, video pretrain=Kinetics-600, Backbone=R-101+NL, flow=false, region proposals=used for training2018.12 | 29 | — | — | |
| SF-101+NL 8x8Pre-train=K600, Backbone=SlowFast-101 + Non-Local2021.04 | 29 | — | — | |
| SF-101 8x8Pre-train=K700, reproduced=true2021.04 | 29 | — | — | |
| SlowFastbackbone=ResNet-101 + Non-Local, pre-train=Kinetics-600, T x τ=8 x 82020.06 | 29 | — | — | |
| MViTv2-Bextra data=K400, extra labels=true, GFLOPs=225, Param=512022.12 | 29 | — | — | |
| MViT-B-24, 32×3pretrain=K600, FLOPs=236, Param=52.92021.04 | 28.7 | — | — | |
| M-ViT-B-24backbone=MViT-B-24, pre-train=K600+ K400, inference=1 view2021.04 | 28.7 | — | — | |
| AVSF-101 8x8Pre-train=K400, extra_information=optical flow, audio, and/or detection2021.04 | 28.6 | — | — | |
| VideoMAEBackbone=ViT-S, Pre-train Dataset=Kinetics-400, Extra Labels=true, T x tau=16x4, GFLOPs=57, Param=22, Image size=224x2242022.03 | 28.4 | — | — | |
| WOObackbone=SF-101, pre-train=K600+ K400, inference=1 view2021.04 | 28.3 | — | — | |
| CSN-152backbone=CSN-152, pre-train=IG + K400, inference=1 view2021.04 | 27.9 | — | — | |
| SlowFast, 16×8 R101+NLpretrain=K600, FLOPs=296, Param=59.22021.04 | 27.5 | — | — | |
| MViT-B, 32×3pretrain=K600, FLOPs=170, Param=36.42021.04 | 27.5 | — | — | |
| X3D-XLpretrain=K600, FLOPs=48.4, Param=11.02021.04 | 27.4 | — | — | |
| X3D-XLbackbone=X3D-XL, pre-train=K600+ K400, inference=1 view2021.04 | 27.4 | — | — | |
| MViT-B, 64×3pretrain=K400, FLOPs=455, Param=36.42021.04 | 27.3 | — | — | |
| SlowFast, 8×8 R101+NLpretrain=K600, FLOPs=147, Param=59.22021.04 | 27.1 | — | — | |
| SF-50 4x16Pre-train=K700, reproduced=true2021.04 | 26.9 | — | — | |
| MViT-B, 32×3pretrain=K400, FLOPs=170, Param=36.42021.04 | 26.8 | — | — | |
| VideoMAEBackbone=ViT-B, Pre-train Dataset=Kinetics-400, Extra Labels=false, T x tau=16x4, GFLOPs=180, Param=87, Image size=224x2242022.03 | 26.7 | — | — | |
| VideoMAE ViT-Bextra data=K400, extra labels=false, GFLOPs=180, Param=872022.12 | 26.7 | — | — | |
| MViT-B, 16×4pretrain=K600, FLOPs=70.5, Param=36.32021.04 | 26.1 | — | — | |
| MViT-B, 16×4pretrain=K400, FLOPs=70.5, Param=36.42021.04 | 24.5 | — | — | |
| SlowFast, 8×8, R101pretrain=K400, FLOPs=138, Param=53.02021.04 | 23.8 | — | — | |
| supervisedBackbone=SlowFast-R101, Pre-train Dataset=Kinetics-400, Extra Labels=true, T x tau=8x8, GFLOPs=138, Param=53, Image size=224x2242022.03 | 23.8 | — | — | |
| SlowFast R101extra data=K400, extra labels=true, GFLOPs=138, Param=532022.12 | 23.8 | — | — | |
| SlowFast R101FLOPS=138G, Param=53M2024.03 | 23.8 | — | — | |
| pBYOLpre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 23.4 | — | — | |
| ρBYOLρ=3Backbone=SlowOnly-R50, Pre-train Dataset=Kinetics-400, Extra Labels=false, T x tau=8x8, GFLOPs=42, Param=32, Image size=224x2242022.03 | 23.4 | — | — | |
| VideoMAEBackbone=ViT-S, Pre-train Dataset=Kinetics-400, Extra Labels=false, T x tau=16x4, GFLOPs=57, Param=22, Image size=224x2242022.03 | 22.5 | — | — | |
| supervisedpre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 22.2 | — | — | |
| SlowFast, 4×16, R50pretrain=K400, FLOPs=52.6, Param=33.72021.04 | 21.9 | — | — | |
| pMoCopre-train=IG-Uncurated-1M, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 20.5 | — | — | |
| pMoCopre-train=IG-Curated-1M, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 20.4 | — | — | |
| pMoCopre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 20.3 | — | — | |
| ρMoCoρ=3Backbone=SlowOnly-R50, Pre-train Dataset=Kinetics-400, Extra Labels=false, T x tau=8x8, GFLOPs=42, Param=32, Image size=224x2242022.03 | 20.3 | — | — | |
| pSwAVpre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 18.2 | — | — | |
| pSimCLRpre-train=K400-240K, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 17.6 | — | — | |
| CVRLBackbone=SlowOnly-R50, Pre-train Dataset=Kinetics-400, Extra Labels=false, T x tau=32x2, GFLOPs=42, Param=32, Image size=224x2242022.03 | 16.3 | — | — | |
| supervisedpre-train=scratch, evaluation protocol=finetuning accuracy, epochs=200, rho=32021.04 | 11.7 | — | — | |
| ACARpre-train=supervised, pre-train data=K600, architecture=ACAR, input size=64 x 224^22022.05 | — | — | 31.4 |