3D Human Pose Estimation on Human3.6M GT 2D pose sequences (test)
15.9MPJPE (Dire.)MotionBERT (DSTformer)
Evaluation Results
| Method | Links | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MotionBERT (DSTformer)T (Clip length)=243, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=true, Training Strategy (scratch vs. finetune)=finetune2022.10 | 15.9 | 17.3 | 16.9 | 14.6 | 16.8 | 18.6 | 18.6 | 18.4 | 22 | 21.8 | 17.3 | 16.9 | 16.1 | 10.5 | 11.4 | 16.9 | |
| MotionBERT (DSTformer)T (Clip length)=243, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=true, Training Strategy (scratch vs. finetune)=scratch2022.10 | 16.7 | 19.9 | 17.1 | 16.5 | 17.4 | 18.8 | 19.3 | 20.5 | 24 | 22.1 | 18.6 | 16.8 | 16.7 | 10.8 | 11.5 | 17.8 | |
| DiffPoseinput_frames=243, temporal_info=true, diffusion_based=true2024.03 | 18.6 | 19.3 | 18 | 18.4 | 18.3 | 21.5 | 21.5 | 19.1 | 23.6 | — | 18.6 | 18.8 | — | 12.8 | — | 18.9 | |
| D3DPinput_frames=243, temporal_info=true, diffusion_based=true2024.03 | 18.7 | 18.2 | 18.4 | 17.8 | 18.6 | 20.9 | 20.2 | 17.7 | 23.8 | — | 18.5 | 17.4 | — | 13.1 | — | 18.4 | |
| KTPFormerinput_frames=243, temporal_info=true, diffusion_based=true2024.03 | 18.8 | 17.4 | 18.1 | 17.7 | 18.3 | 20.6 | 20.8 | 18.3 | 23.3 | — | 19.6 | 17.7 | — | 13.5 | — | 18.1 | |
| KTPFormerinput_frames=243, temporal_info=true2024.03 | 19.6 | 18.6 | 18.5 | 18.1 | 18.7 | 22.1 | 20.8 | 18.3 | 22.8 | — | 18.8 | 18.1 | — | 12.4 | — | 19 | |
| GLA-GCNinput_frames=243, temporal_info=true2024.03 | 20.1 | 21.2 | 20 | 19.6 | 21.5 | 26.7 | 23.3 | 19.8 | 27 | — | 20.8 | 20.1 | — | 12.8 | — | 21 | |
| STCFormerinput_frames=243, temporal_info=true2024.03 | 21.4 | 22.6 | 21 | 21.3 | 23.8 | 26 | 24.2 | 20 | 28.9 | — | 22.3 | 21.4 | — | 14.2 | — | 22 | |
| MixSTET (Clip length)=243, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=true2022.10 | 21.6 | 22 | 20.4 | 21 | 20.8 | 24.3 | 24.7 | 21.9 | 26.9 | 24.9 | 21.2 | 21.5 | 20.8 | 14.7 | 15.6 | 21.6 | |
| MixSTEinput_frames=243, temporal_info=true2024.03 | 21.6 | 22 | 20.4 | 21 | 20.8 | 24.3 | 24.7 | 21.9 | 26.9 | — | 21.2 | 21.5 | — | 14.7 | — | 21.6 | |
| DUEinput_frames=300, temporal_info=true2024.03 | 22.1 | 23.1 | 20.1 | 22.7 | 21.3 | 24.1 | 23.6 | 21.6 | 26.3 | — | 21.7 | 21.4 | — | 16.7 | — | 22 | |
| KTPFormerinput_frames=81, temporal_info=true2024.03 | 22.5 | 22.4 | 21.3 | 21.4 | 21.2 | 25.5 | 24.2 | 22.4 | 24.4 | — | 22.7 | 21.4 | — | 16.3 | — | 22.2 | |
| UGCNT (Clip length)=96, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=false2022.10 | 23 | 25.7 | 22.8 | 22.6 | 24.1 | 30.6 | 24.9 | 24.5 | 31.1 | 35 | 25.6 | 24.3 | 25.1 | 19.8 | 18.4 | 25.6 | |
| UGCNinput_frames=96, temporal_info=true2024.03 | 23 | 25.7 | 22.8 | 22.6 | 24.1 | 30.6 | 24.9 | 24.5 | 31.1 | — | 25.6 | 24.3 | — | 19.8 | — | 25.6 | |
| MixSTEinput_frames=81, temporal_info=true2024.03 | 25.6 | 27.8 | 24.5 | 25.7 | 24.9 | 29.9 | 28.6 | 27.4 | 29.9 | — | 26.1 | 25 | — | 18.7 | — | 25.9 | |
| STCFormerinput_frames=81, temporal_info=true2024.03 | 26.2 | 26.5 | 23.4 | 24.6 | 25 | 28.6 | 28.3 | 24.6 | 30.9 | — | 25.7 | 25.3 | — | 18.6 | — | 25.7 | |
| StridedFormerinput_frames=351, temporal_info=true2024.03 | 27.1 | 29.4 | 26.5 | 27.1 | 28.6 | 33 | 30.7 | 26.8 | 38.2 | — | 29.1 | 29.8 | — | 19.1 | — | 28.5 | |
| MHFormerT (Clip length)=351, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=true2022.10 | 27.7 | 32.1 | 29.1 | 28.9 | 30 | 33.9 | 33 | 31.2 | 37 | 39.3 | 30 | 31 | 29.4 | 22.2 | 23 | 30.5 | |
| MHFormerinput_frames=351, temporal_info=true2024.03 | 27.7 | 32.1 | 29.1 | 28.9 | 30 | 33.9 | 33 | 31.2 | 37 | — | 30 | 31 | — | 22.2 | — | 30.5 | |
| P-STMOT (Clip length)=243, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=true2022.10 | 28.5 | 30.1 | 28.6 | 27.9 | 29.8 | 33.2 | 31.3 | 27.8 | 36 | 37.4 | 29.7 | 29.5 | 28.1 | 21 | 21 | 29.3 | |
| P-STMOinput_frames=243, temporal_info=true2024.03 | 28.5 | 30.1 | 28.6 | 27.9 | 29.8 | 33.2 | 31.3 | 27.8 | 36 | — | 29.7 | 29.5 | — | 21 | — | 29.3 | |
| PoseFormerT (Clip length)=81, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=true2022.10 | 30 | 33.6 | 29.9 | 31 | 30.2 | 33.3 | 34.8 | 31.4 | 37.8 | 38.6 | 31.7 | 31.5 | 29 | 23.3 | 23.1 | 31.3 | |
| PoseFormerinput_frames=81, temporal_info=true2024.03 | 30 | 33.6 | 29.9 | 31 | 30.2 | 33.3 | 34.8 | 31.4 | 37.8 | — | 31.7 | 31.5 | — | 23.3 | — | 31.3 | |
| GraFormertemporal_info=false2024.03 | 32 | 38 | 30.4 | 34.4 | 34.7 | 43.3 | 35.2 | 31.4 | 38 | — | 34.2 | 35.7 | — | 27.4 | — | 35.2 | |
| POT2024.03 | 32.9 | 38.3 | 28.3 | 33.8 | 34.9 | 38.7 | 37.2 | 30.7 | 34.5 | — | 33.9 | 34.7 | — | 26.1 | — | 33.8 | |
| Xu et al.T (Clip length)=1, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=false2022.10 | 35.8 | 38.1 | 31 | 35.3 | 35.8 | 43.2 | 37.3 | 31.7 | 38.4 | 45.5 | 35.4 | 36.7 | 36.8 | 27.9 | 30.7 | 35.8 | |
| LCNT (Clip length)=1, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=false2022.10 | 36.3 | 38.8 | 29.7 | 37.8 | 34.6 | 42.5 | 39.8 | 32.5 | 36.2 | 39.5 | 34.4 | 38.4 | 38.2 | 31.3 | 34.2 | 36.3 | |
| PoseGTACtemporal_info=false2024.03 | 37.2 | 42.2 | 32.6 | 38.6 | 38 | 44 | 40.7 | 35.2 | 41 | — | 38.2 | 39.5 | — | 29.8 | — | 38.2 | |
| Martinez et al.T (Clip length)=1, Input type=GT 2D pose sequences, Spatio-temporal Transformer design=false2022.10 | 37.7 | 44.4 | 40.3 | 42.1 | 48.2 | 54.9 | 44.4 | 42.1 | 54.6 | 58 | 45.1 | 46.4 | 47.6 | 36.4 | 40.4 | 45.5 |