3D Human Pose and Shape Estimation on Human3.6M (test)
29.5PA-MPJPEHybrIK-Transformer
Evaluation Results
| Method | Links | |||
|---|---|---|---|---|
| HybrIK-TransformerBackbone=HrNet-48, Training Data=w. 3DPW2023.02 | 29.5 | 47.5 | — | |
| HybrIK-TransformerBackbone=HrNet-482023.02 | 29.8 | 48.8 | — | |
| TARNumber of input frames=9, Backbone=HRNet2023.11 | 33.3 | 45.6 | 5.6 | |
| HybrIK (Adaptive)variant=Adaptive, trained with 3DPW=true2020.11 | 33.6 | 55.4 | — | |
| HybrIKTraining Data=w. 3DPW2023.02 | 33.6 | 55.4 | — | |
| HybrIK (Adaptive)variant=Adaptive2020.11 | 34.5 | 54.4 | — | |
| Mesh GraphormerInput Mode=Image-based, Backbone=HRNet-W64, Supervision=w/ 3DPW training dataset2023.03 | 34.5 | 51.2 | — | |
| HybrIK2023.02 | 34.5 | 54.4 | — | |
| HybrIK-TransformerBackbone=ResNet-34, Training Data=w. 3DPW2023.02 | 34.6 | 50.2 | — | |
| HybrIK-TransformerBackbone=ResNet-342023.02 | 34.9 | 52.5 | — | |
| HybrIK (Naive)variant=Naive2020.11 | 35.3 | 55.8 | — | |
| METROInput Mode=Image-based, Backbone=HRNet-W64, Supervision=w/ 3DPW training dataset2023.03 | 36.7 | 54 | — | |
| TARNumber of input frames=9, Backbone=ResNet502023.11 | 37.8 | 53.1 | 5.6 | |
| Lee et al.Input Modality=video, 3DPW Training Involvement=w/o 3DPW2022.10 | 38.4 | 58.4 | 6.1 | |
| INT-2Input Mode=Video-based, Backbone=ResNet-50, Supervision=w/ 3DPW training dataset, Phases=Three phases2023.03 | 38.4 | 54.9 | — | |
| INTNumber of input frames=642023.11 | 38.4 | 54.9 | — | |
| MAEDInput Mode=Video-based, Backbone=ResNet-50, Supervision=w/ 3DPW training dataset2023.03 | 38.7 | 56.4 | — | |
| MAEDNumber of input frames=642023.11 | 38.7 | 56.4 | — | |
| MAEDInput=video, Training Data=w/o 3DPW2021.09 | 38.7 | 56.3 | — | |
| MAEDInput=video, Training Data=w/ 3DPW2021.09 | 38.7 | 56.4 | — | |
| SPS-NetApproach Type=Video based, Number of parameters=51.43M2021.03 | 38.7 | 58.9 | — | |
| INT-1Input Mode=Video-based, Backbone=ResNet-50, Supervision=w/o 3DPW training dataset, Phases=First two phases2023.03 | 39.1 | 57.1 | — | |
| DSRTraining data=standard2021.10 | 40.3 | 60.9 | — | |
| PyMAFInput Modality=single image2022.10 | 40.5 | 57.7 | — | |
| PyMAF+Input Mode=Image-based, Backbone=ResNet-50, Supervision=w/o 3DPW training dataset2023.03 | 40.5 | 57.7 | — | |
| METROInput Mode=Image-based, Backbone=ResNet-50, Supervision=w/ 3DPW training dataset2023.03 | 40.6 | 56.5 | — | |
| SPINInput type=single image2020.11 | 41.1 | — | 18.3 | |
| I2L-MeshNetInput type=single image2020.11 | 41.1 | 55.7 | 13.4 | |
| TCMRInput type=video2020.11 | 41.1 | 62.3 | 5.3 | |
| SPINTraining data=standard2021.10 | 41.1 | 62.5 | — | |
| SPIN2020.11 | 41.1 | — | — | |
| SPINInput Modality=single image2022.10 | 41.1 | — | 18.3 | |
| I2L-MeshNetInput Modality=single image2022.10 | 41.1 | 55.7 | 13.4 | |
| TCMRInput Modality=video, 3DPW Training Involvement=w/o 3DPW2022.10 | 41.1 | 62.3 | 5.3 | |
| SPINInput Mode=Image-based, Backbone=ResNet-50, Supervision=w/o 3DPW training dataset2023.03 | 41.1 | — | — | |
| I2L-MeshNet+Input Mode=Image-based, Backbone=ResNet-50, Supervision=w/o 3DPW training dataset2023.03 | 41.1 | 55.7 | — | |
| TCMRInput Mode=Video-based, Backbone=ResNet (from SPIN), Supervision=w/o 3DPW training dataset2023.03 | 41.1 | 62.3 | — | |
| SPIN2023.02 | 41.1 | — | — | |
| TCMRNumber of input frames=162023.11 | 41.1 | 62.3 | 5.3 | |
| SPINInput=image, Training Data=w/o 3DPW2021.09 | 41.1 | — | — | |
| I2LMeshNetInput=image, Training Data=w/o 3DPW2021.09 | 41.1 | 55.7 | — | |
| SPINApproach Type=Frame based, Number of parameters=26.98M2021.03 | 41.1 | — | — | |
| Mesh Graphormertrained with 3DPW=true2020.11 | 41.2 | 34.5 | — | |
| Mesh GraphormerTraining Data=w. 3DPW2023.02 | 41.2 | 34.5 | — | |
| DSRTraining data=w/ 3DPW train2021.10 | 41.4 | 62 | — | |
| VIBEInput Mode=Video-based, Backbone=ResNet (from SPIN), Supervision=w/ 3DPW training dataset2023.03 | 41.4 | 65.6 | — | |
| VIBEInput=video, Training Data=w/ 3DPW2021.09 | 41.4 | 65.6 | — | |
| VIBEInput type=video2020.11 | 41.5 | 65.9 | 18.3 | |
| VIBEInput Modality=video, 3DPW Training Involvement=w/o 3DPW2022.10 | 41.5 | 65.9 | 18.3 | |
| VIBENumber of input frames=162023.11 | 41.5 | 65.9 | 18.3 | |
| VIBEInput=video, Training Data=w/o 3DPW2021.09 | 41.5 | 65.9 | — | |
| VIBEApproach Type=Video based, Number of parameters=48.30M2021.03 | 41.5 | 65.9 | — | |
| I2Ltrained on different datasets=true2020.11 | 41.7 | 55.7 | — | |
| I2L2023.02 | 41.7 | 55.7 | — | |
| Sun et al.Input type=video2020.11 | 42.4 | 59.1 | — | |
| Sun et al.Input Modality=video2022.10 | 42.4 | 59.1 | — | |
| Skeleton-disentangledInput Mode=Video-based, Backbone=ResNet-50, Supervision=w/o 3DPW training dataset2023.03 | 42.4 | 59.1 | — | |
| DSD-SATNInput=video, Training Data=w/o 3DPW2021.09 | 42.4 | 59.1 | — | |
| DSD-SATNApproach Type=Video based2021.03 | 42.4 | 59.1 | — | |
| EFTTraining data=standard2021.10 | 43.7 | — | — | |
| EFTTraining data=w/ 3DPW train2021.10 | 43.8 | — | — | |
| Pose2MeshInput type=single image2020.11 | 46.3 | 64.9 | 23.9 | |
| GLoTNumber of input frames=162023.11 | 46.3 | 67 | 3.6 | |
| Pose2MeshInput=2D Pose, Training Data=w/o 3DPW2021.09 | 46.3 | 64.9 | — | |
| OursInput Modality=video, 3DPW Training Involvement=w/o 3DPW2022.10 | 46.6 | 67.8 | 3.6 | |
| TePoseReal-time capability=true, Training dataset=3DPW train set, Number of input frames (T+1)=62022.07 | 47.1 | 68.6 | 12.1 | |
| MPS-NetInput Mode=Video-based, Backbone=ResNet (from SPIN), Supervision=w/ 3DPW training dataset2023.03 | 47.4 | 69.4 | — | |
| MPS-NetNumber of input frames=162023.11 | 47.4 | 69.4 | 3.6 | |
| GraphCMRInput type=single image2020.11 | 50.1 | — | — | |
| CMRTraining data=standard2021.10 | 50.1 | — | — | |
| Kolotouros et al.2020.11 | 50.1 | — | — | |
| GraphCMRInput Modality=single image2022.10 | 50.1 | — | — | |
| Mesh RegressionInput Mode=Image-based, Backbone=ResNet-50, Supervision=w/o 3DPW training dataset2023.03 | 50.1 | — | — | |
| CMR2023.02 | 50.1 | — | — | |
| GraphCMRInput=image, Training Data=w/o 3DPW2021.09 | 50.1 | — | — | |
| CMRApproach Type=Frame based, Number of parameters=46.31M2021.03 | 50.1 | — | — | |
| OursInput Modality=video, 3DPW Training Involvement=w/ 3DPW2022.10 | 51.9 | 73.3 | 3.6 | |
| TCMRReal-time capability=false, Training dataset=3DPW train set, Number of input frames (T+1)=162022.07 | 52 | 73.6 | 3.9 | |
| TCMRInput Modality=video, 3DPW Training Involvement=w/ 3DPW2022.10 | 52 | 73.6 | 3.9 | |
| HUNDApproach Type=Frame based2021.03 | 53 | 72 | — | |
| MEVAReal-time capability=false, Training dataset=3DPW train set, Number of input frames (T+1)=902022.07 | 53.2 | 76 | 15.3 | |
| MEVANumber of input frames=902023.11 | 53.2 | 76 | 15.3 | |
| MEVAInput=video, Training Data=w/ 3DPW2021.09 | 53.2 | 76 | — | |
| VIBEReal-time capability=true, Training dataset=3DPW train set, Number of input frames (T+1)=162022.07 | 53.3 | 78 | 27.3 | |
| Arnab et al.2020.11 | 54.3 | 77.8 | — | |
| Temporal ContextInput Mode=Video-based, Backbone=ResNet (from HMR), Supervision=w/o 3DPW training dataset2023.03 | 54.3 | 77.8 | — | |
| Arnab et al.2023.02 | 54.3 | 77.8 | — | |
| TemporalContextInput=video, Training Data=w/o 3DPW2021.09 | 54.3 | 77.8 | — | |
| HybrikInput Mode=Image-based, Backbone=ResNet-34, Supervision=w/o 3DPW training dataset2023.03 | 54.4 | — | — | |
| STRAPSInput=image, Training Data=w/ 3DPW2021.09 | 55.4 | — | — | |
| STRAPSApproach Type=Frame based, Number of parameters=12.48M2021.03 | 55.4 | — | — | |
| HMRInput type=single image2020.11 | 56.8 | 88 | — | |
| HMRTraining data=standard2021.10 | 56.8 | 88 | — | |
| HMR2020.11 | 56.8 | 88 | — | |
| HMRInput Modality=single image2022.10 | 56.8 | 88 | — | |
| HMR+Input Mode=Image-based, Backbone=ResNet-502023.03 | 56.8 | 88 | — | |
| HMR2023.02 | 56.8 | 88 | — | |
| HMRInput=image, Training Data=w/o 3DPW2021.09 | 56.8 | 88 | — | |
| HMRApproach Type=Frame based, Number of parameters=26.98M2021.03 | 56.8 | 88 | — | |
| HMMRInput type=video2020.11 | 56.9 | — | — |