3D Human Pose Estimation on MPI-INF-3DHP
16.3MPJPEPoseMoE
Evaluation Results
| Method | Links | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| PoseMoET=81, Seq2Seq=true2025.12 | 16.3 | 99.1 | 86.9 | — | — | — | — | — | — | — | — | — | |
| KTPFormerT=81, Seq2Seq=true2025.12 | 16.7 | 98.9 | 85.9 | — | — | — | — | — | — | — | — | — | |
| PoseMoET=27, Seq2Seq=true2025.12 | 18.5 | 98.6 | 85.8 | — | — | — | — | — | — | — | — | — | |
| PoseMoET=9, Seq2Seq=true2025.12 | 21.6 | 98.1 | 82.2 | — | — | — | — | — | — | — | — | — | |
| PoseRetNetT=81, Seq2Seq=true2025.12 | 22.2 | 99.1 | 84.4 | — | — | — | — | — | — | — | — | — | |
| STCFormerT=81, Seq2Seq=true2025.12 | 23.1 | 98.7 | 83.9 | — | — | — | — | — | — | — | — | — | |
| GLA-GCNProtocol=Protocol#1, T (Number of frames)=81, Temporal information=true, Intermediate 3D pose sequence reconstruction=false2023.07 | 27.76 | 98.53 | 79.12 | — | — | — | — | — | — | — | — | — | |
| PoseFormerV2T=81, Seq2Seq=false2025.12 | 27.8 | 97.9 | 78.8 | — | — | — | — | — | — | — | — | — | |
| GLA-GCNT=81, Seq2Seq=false2025.12 | 27.8 | 98.5 | 79.1 | — | — | — | — | — | — | — | — | — | |
| Fusionformerf=92022.10 | 28.2 | 97.9 | 70 | — | — | — | — | — | — | — | — | — | |
| DDHPoseInput Source=Ground truth 2D pose, Model Scale=L, H=1, W=12025.12 | 29.2 | 98.5 | 78.1 | — | — | — | — | — | — | — | — | — | |
| FastDDHPoseInput Source=Ground truth 2D pose, Model Scale=L, H=1, W=12025.12 | 29.2 | 98.3 | 78.2 | — | — | — | — | — | — | — | — | — | |
| D3DPInput Source=Ground truth 2D pose, Model Scale=L, H=1, W=12025.12 | 30.2 | 97.7 | 77.8 | — | — | — | — | — | — | — | — | — | |
| GLA-GCNProtocol=Protocol#1, T (Number of frames)=27, Temporal information=true, Intermediate 3D pose sequence reconstruction=false2023.07 | 31.36 | 98.19 | 76.53 | — | — | — | — | — | — | — | — | — | |
| P-STMONumber of input frames (N)=812022.03 | 32.2 | 97.9 | 75.8 | — | — | — | — | — | — | — | — | — | |
| Shan et al.Protocol=Protocol#1, T (Number of frames)=81, Temporal information=true, Intermediate 3D pose sequence reconstruction=false2023.07 | 32.2 | 97.9 | 75.8 | — | — | — | — | — | — | — | — | — | |
| P-STMOT=81, Seq2Seq=false2025.12 | 32.2 | 97.9 | 75.8 | — | — | — | — | — | — | — | — | — | |
| P-STMOInput Source=Ground truth 2D pose, Model Scale=M2025.12 | 32.2 | 97.9 | 75.8 | — | — | — | — | — | — | — | — | — | |
| GazeDNumber of hypotheses (H)=20, Aggregation Strategy (A)=ORC2026.01 | 33.8 | — | — | — | — | — | — | — | — | — | — | — | |
| MixSTEInput Source=Ground truth 2D pose, Model Scale=L2025.12 | 35.4 | 96.9 | 75.8 | — | — | — | — | — | — | — | — | — | |
| multi-view TTA frameworkBackbone=CameraHMR2026.03 | 39 | 99.9 | 83.8 | — | — | — | — | — | — | — | — | — | |
| U-HMR2026.03 | 39.7 | 74 | 99.4 | — | — | — | — | — | — | — | — | — | |
| HeatFormer2026.03 | 39.8 | 99.5 | 72.8 | — | — | — | — | — | — | — | — | — | |
| multi-view TTA frameworkBackbone=HMR2.02026.03 | 40.3 | 99.6 | 80.7 | — | — | — | — | — | — | — | — | — | |
| Hu et al.Number of input frames (N)=962022.03 | 42.5 | 97.9 | 69.5 | — | — | — | — | — | — | — | — | — | |
| Hu et al.f=962022.10 | 42.5 | 97.9 | 69.5 | — | — | — | — | — | — | — | — | — | |
| Hu et al.Protocol=Protocol#1, T (Number of frames)=96, Temporal information=true, Intermediate 3D pose sequence reconstruction=true2023.07 | 42.5 | 97.9 | 69.5 | — | — | — | — | — | — | — | — | — | |
| Zhao2026.01 | 44.7 | — | — | — | — | — | — | — | — | — | — | — | |
| multi-view TTA frameworkBackbone=TokenHMR2026.03 | 45.2 | 99.9 | 79.6 | — | — | — | — | — | — | — | — | — | |
| GazeDNumber of hypotheses (H)=20, Aggregation Strategy (A)=AVG2026.01 | 46.6 | — | — | — | — | — | — | — | — | — | — | — | |
| PaFF2026.03 | 48.4 | 98.6 | 67.3 | — | — | — | — | — | — | — | — | — | |
| ConvFormer2023.04 | 53.6 | 96.4 | 69.8 | — | — | — | — | — | — | — | — | — | |
| Zhang et al.f=272022.10 | 54.9 | 94.4 | 66.5 | — | — | — | — | — | — | — | — | — | |
| NLF-Lvariant=L2024.07 | 54.9 | 97.5 | 61 | — | — | — | — | — | — | — | — | — | |
| Zhang2026.01 | 54.9 | — | — | — | — | — | — | — | — | — | — | — | |
| MixSTET=27, Seq2Seq=true2025.12 | 54.9 | 94.4 | 66.5 | — | — | — | — | — | — | — | — | — | |
| MixSTET=272026.04 | 54.9 | 94.4 | 66.5 | — | — | — | — | — | — | — | — | — | |
| MeTRAbs-ACAE-Lvariant=L2024.07 | 55.4 | 97.1 | 60.1 | — | — | — | — | — | — | — | — | — | |
| Zheng2026.01 | 57.7 | — | — | — | — | — | — | — | — | — | — | — | |
| MeTRAbs-ACAE-Svariant=S2024.07 | 57.9 | 96.3 | 58.7 | — | — | — | — | — | — | — | — | — | |
| Li et al.2023.04 | 58 | 93.8 | 63.3 | — | — | — | — | — | — | — | — | — | |
| Li2026.01 | 58 | — | — | — | — | — | — | — | — | — | — | — | |
| MHFormerT=9, Seq2Seq=false2025.12 | 58 | 93.8 | 63.3 | — | — | — | — | — | — | — | — | — | |
| MHFormerT=92026.04 | 58 | 93.8 | 63.3 | — | — | — | — | — | — | — | — | — | |
| NLF-Svariant=S2024.07 | 59.9 | 96.6 | 57.9 | — | — | — | — | — | — | — | — | — | |
| Mehta2020.04 | 64.7 | 72.5 | — | — | — | — | — | — | — | — | — | — | |
| DenseWarperInput Method=sparse Interleaved, 2D detector=SimpleBaseline2026.05 | 65.89 | — | — | — | — | — | — | — | — | — | — | — | |
| KTP-FormerInput Method=Single, T=243, 2D detector=SimpleBaseline2026.05 | 67.59 | — | — | — | — | — | — | — | — | — | — | — | |
| Wang et al.Number of input frames (N)=962022.03 | 68.1 | 86.9 | 62.1 | — | — | — | — | — | — | — | — | — | |
| Wang et al.f=962022.10 | 68.1 | 86.9 | 62.1 | — | — | — | — | — | — | — | — | — | |
| Wang et al.Protocol=Protocol#1, T (Number of frames)=96, Temporal information=true, Intermediate 3D pose sequence reconstruction=true2023.07 | 68.1 | 86.9 | 62.1 | — | — | — | — | — | — | — | — | — | |
| Wang2026.01 | 68.1 | — | — | — | — | — | — | — | — | — | — | — | |
| GLA-GCNInput Method=Single, T=243, 2D detector=SimpleBaseline2026.05 | 75 | — | — | — | — | — | — | — | — | — | — | — | |
| DG-NetInput sequence length (T)=42021.09 | 76 | 87.5 | 53.8 | — | — | — | — | — | — | — | — | — | |
| Zheng et al.Number of input frames (N)=92022.03 | 77.1 | 88.6 | 56.4 | — | — | — | — | — | — | — | — | — | |
| Zheng et al.f=92022.10 | 77.1 | 88.6 | 56.4 | — | — | — | — | — | — | — | — | — | |
| Zheng et al.2023.04 | 77.1 | 88.6 | 56.4 | — | — | — | — | — | — | — | — | — | |
| Zheng et al.Protocol=Protocol#1, T (Number of frames)=9, Temporal information=true, Intermediate 3D pose sequence reconstruction=false2023.07 | 77.1 | 88.6 | 56.4 | — | — | — | — | — | — | — | — | — | |
| Zheng2026.01 | 77.1 | — | — | — | — | — | — | — | — | — | — | — | |
| PoseFormerV1Input Source=Ground truth 2D pose, Model Scale=S2025.12 | 77.1 | 88.6 | 56.4 | — | — | — | — | — | — | — | — | — | |
| PoseFormerInput frames=92021.03 | 77.1 | 88.6 | 56.4 | — | — | — | — | — | — | — | — | — | |
| AdafuseInput Method=Full, 2D detector=SimpleBaseline2026.05 | 78.57 | — | — | — | — | — | — | — | — | — | — | — | |
| Chen et al.Number of input frames (N)=812022.03 | 78.8 | 87.9 | 54 | — | — | — | — | — | — | — | — | — | |
| Chen et al.2023.04 | 78.8 | 87.6 | 54 | — | — | — | — | — | — | — | — | — | |
| Chen et al.Protocol=Protocol#1, T (Number of frames)=81, Temporal information=true, Intermediate 3D pose sequence reconstruction=false2023.07 | 78.8 | 87.9 | 54 | — | — | — | — | — | — | — | — | — | |
| Chen et al.Input frames=81, Venue=TCSVT 212021.03 | 78.8 | 87.9 | 54 | — | — | — | — | — | — | — | — | — | |
| Anatomy3DInput Source=Ground truth 2D pose, Model Scale=L2025.12 | 79.1 | 87.8 | 53.8 | — | — | — | — | — | — | — | — | — | |
| Trajectory Space FactorizationInput=GT 2D, F=25, temporal=true2019.08 | 79.8 | 83.6 | 51.4 | — | — | — | — | — | — | — | — | — | |
| Lin et al.Number of input frames (N)=252022.03 | 79.8 | 83.6 | 51.4 | — | — | — | — | — | — | — | — | — | |
| Lin et al.f=252022.10 | 79.8 | 83.6 | 51.4 | — | — | — | — | — | — | — | — | — | |
| Lin et al.2023.04 | 79.8 | 83.6 | 51.4 | — | — | — | — | — | — | — | — | — | |
| Lin et al.Protocol=Protocol#1, T (Number of frames)=25, Temporal information=true, Intermediate 3D pose sequence reconstruction=false2023.07 | 79.8 | 83.6 | 51.4 | — | — | — | — | — | — | — | — | — | |
| Lin et al.Input frames=25, Venue=BMVC'192021.03 | 79.8 | 83.6 | 51.4 | — | — | — | — | — | — | — | — | — | |
| Lin et al.Input sequence length (T)=252021.09 | 79.8 | 83.6 | 51.4 | — | — | — | — | — | — | — | — | — | |
| Trajectory Space FactorizationInput=GT 2D, F=50, temporal=true2019.08 | 81.9 | 82.4 | 49.6 | — | — | — | — | — | — | — | — | — | |
| Lin et al.Input sequence length (T)=502021.09 | 81.9 | 82.4 | 49.6 | — | — | — | — | — | — | — | — | — | |
| Adafuse + SLERPInput Method=Interpolation, 2D detector=SimpleBaseline2026.05 | 83.37 | — | — | — | — | — | — | — | — | — | — | — | |
| Pavllo et al.Number of input frames (N)=812022.03 | 84 | 86 | 51.9 | — | — | — | — | — | — | — | — | — | |
| Pavllo et al.f=812022.10 | 84 | 86 | 51.9 | — | — | — | — | — | — | — | — | — | |
| Pavllo et al.2023.04 | 84 | 86 | 51.9 | — | — | — | — | — | — | — | — | — | |
| Pavllo et al.Protocol=Protocol#1, T (Number of frames)=81, Temporal information=true, Intermediate 3D pose sequence reconstruction=false2023.07 | 84 | 86 | 51.9 | — | — | — | — | — | — | — | — | — | |
| Pavvlo2026.01 | 84 | — | — | — | — | — | — | — | — | — | — | — | |
| Pavllo et al.Input frames=81, Venue=CVPR'192021.03 | 84 | 86 | 51.9 | — | — | — | — | — | — | — | — | — | |
| Pavllo et al.Input frames=243, Venue=CVPR'192021.03 | 84.8 | 85.5 | 51.5 | — | — | — | — | — | — | — | — | — | |
| IKOLBackbone=ResNet-34, Fine-tuning=false2023.02 | 88.8 | 87.9 | 48.1 | — | — | — | — | — | — | — | — | — | |
| IKOLBackbone=ResNet-34, Fine-tuning=fine tuned with the HybrIK model2023.02 | 89.7 | 87 | 47.6 | — | — | — | — | — | — | — | — | — | |
| HMRAlignment Protocol=After Rigid Alignment, Supervision Type=Paired, Output Format=More than 3D joints2017.12 | 89.8 | 86.3 | 47.8 | — | — | — | — | — | — | — | — | — | |
| Viewpoint Prediction Branch2020.04 | 90.3 | 84.3 | — | — | — | — | — | — | — | — | — | — | |
| Habibie2020.04 | 91 | 82 | — | — | — | — | — | — | — | — | — | — | |
| HybrIKBackbone=ResNet-342023.02 | 91 | 86.2 | 42.2 | — | — | — | — | — | — | — | — | — | |
| ROMPBackbone=HRNet-w322023.02 | 95.11 | — | — | — | — | — | — | — | — | — | — | — | |
| VNectAlignment Protocol=After Rigid Alignment, Output Format=3D joints2017.12 | 98 | 83.9 | 47.3 | — | — | — | — | — | — | — | — | — | |
| Li et al.2023.04 | 99.7 | 81.2 | 46.1 | — | — | — | — | — | — | — | — | — | |
| Li et al.Venue=CVPR 202021.03 | 99.7 | 81.2 | 46.1 | — | — | — | — | — | — | — | — | — | |
| TUCHEXtraining_data=3DPW2021.04 | 101.5 | — | — | 66.4 | — | — | — | — | — | — | — | — | |
| TP-NetTraining Data Augmentation=No2017.11 | 103.8 | 76.7 | 39.1 | — | — | — | — | — | — | — | — | — | |
| SPINBackbone=ResNet-502023.02 | 105.2 | 76.4 | 37.1 | — | — | — | — | — | — | — | — | — | |
| Kolotouros et al.Training Data=MPI-INF-3DHP, Human3.6M, and MPII2019.11 | 105.2 | — | 37.1 | — | — | — | — | — | — | — | — | 76.4 | |
| HMR-EFTtraining_data=3DPW2021.04 | 105.3 | — | — | 68.4 | — | — | — | — | — | — | — | — | |
| PPTInput Method=Full, 2D detector=SimpleBaseline2026.05 | 106.3 | — | — | — | — | — | — | — | — | — | — | — |