Action Detection on AVA v2.2 (test)
39.8mAPUMT-L
Evaluation Results
| Method | Links | |
|---|---|---|
| UMT-LPT Data=K710, Input Size=8×224², FLOPs (G)=596, Param (M)=3042023.03 | 39.8 | |
| VideoMAE-LPT Data=K700, Input Size=16×224², FLOPs (G)=597, Param (M)=3052023.03 | 39.3 | |
| UMT-LPT Data=K400, Input Size=8×224², FLOPs (G)=596, Param (M)=3042023.03 | 39 | |
| MaskFeat-LPT Data=K600, Input Size=40×312², FLOPs (G)=2828, Param (M)=2182023.03 | 37.8 | |
| ST-MAE-LPT Data=K700, Input Size=16×224², FLOPs (G)=598, Param (M)=3042023.03 | 37.3 | |
| VideoMAE-LPT Data=K400, Input Size=16×224², FLOPs (G)=597, Param (M)=3052023.03 | 37 | |
| MaskFeat-LPT Data=K400, Input Size=40×312², FLOPs (G)=2828, Param (M)=2182023.03 | 36.3 | |
| ST-MAE-LPT Data=K400, Input Size=16×224², FLOPs (G)=598, Param (M)=3042023.03 | 34.8 | |
| SlowFast++ ensembleBackbone=R101+NL, video pretrain=Kinetics-600, mode=ensemble2018.12 | 34.3 | |
| SlowFastensemble=true, video pretrain=Kinetics-600, Backbone=R-101+NL, augmentation=multi-scale and horizontal flipping, region proposals=used for training2018.12 | 34.3 | |
| MViTv2-LPT Data=IN-21K+K700, Input Size=40×312², FLOPs (G)=2828, Param (M)=2132023.03 | 33.5 | |
| UMT-BPT Data=K710, Input Size=8×224², FLOPs (G)=180, Param (M)=872023.03 | 33.5 | |
| UMT-BPT Data=K400, Input Size=8×224², FLOPs (G)=180, Param (M)=872023.03 | 32.7 | |
| VideoMAE-BPT Data=K400, Input Size=16×224², FLOPs (G)=180, Param (M)=872023.03 | 31.8 | |
| MViTv2-BPT Data=K700, Input Size=32×224², FLOPs (G)=225, Param (M)=512023.03 | 31.3 | |
| MViTv1-BPT Data=K600, Input Size=32×224², FLOPs (G)=236, Param (M)=532023.03 | 28.7 | |
| MViTv2-BPT Data=K400, Input Size=32×224², FLOPs (G)=225, Param (M)=512023.03 | 28.1 | |
| SlowFastPT Data=K600, Input Size=64×224², FLOPs (G)=296, Param (M)=592023.03 | 27.5 | |
| MViTv1-BPT Data=K400, Input Size=64×224², FLOPs (G)=455, Param (M)=362023.03 | 27.3 | |
| SlowFastPT Data=K400, Input Size=32×224², FLOPs (G)=138, Param (M)=532023.03 | 23.8 |