3D Human Pose Estimation on CMU Panoptic (test)
10.6MPJPETriangulation with Transformer Regressor (TTR) + Gröbner basis Corrector (GC) + Temporal Equivariant Rectifier (TER)
Evaluation Results
| Method | Links | ||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Triangulation with Transformer Regressor (TTR) + Gröbner basis Corrector (GC) + Temporal Equivariant Rectifier (TER)Intrinsic parameters=Unavailable, Extrinsic parameters=Unavailable, Temporal information=true, Number of cameras=42026.04 | 10.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Zhang et al.Intrinsic parameters=Available, Extrinsic parameters=Available, Temporal information=false, Number of cameras=42026.04 | 11.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Triangulation with Transformer Regressor (TTR) + Gröbner basis Corrector (GC)Intrinsic parameters=Unavailable, Extrinsic parameters=Unavailable, Temporal information=false, Number of cameras=42026.04 | 12.8 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Iskakov et al. VolumetricIntrinsic parameters=Available, Extrinsic parameters=Available, Temporal information=false, Number of cameras=42026.04 | 13.7 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Voxel.+3DSASupervision Paradigm=Fully-Supervised2026.06 | 13.98 | — | — | — | — | — | — | — | — | — | — | — | — | 94.2 | 98.49 | 99.21 | 99.31 | — | |
| LiCamPoseTraining Manner=supervised2023.12 | 14.44 | — | — | — | — | — | — | — | — | — | — | — | 11.61 | — | — | — | — | — | |
| TEMPOSupervision Paradigm=Fully-Supervised2026.06 | 14.68 | — | — | — | — | — | — | — | — | — | — | — | — | 89.01 | 99.08 | 99.76 | 99.93 | — | |
| VoxelPose(PRN)Training Manner=supervised2023.12 | 14.91 | — | — | — | — | — | — | — | — | — | — | — | 11.88 | — | — | — | — | — | |
| Chharia et al.Intrinsic parameters=Available, Extrinsic parameters=Available, Temporal information=false, Number of cameras=42026.04 | 15.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Liao et al.Intrinsic parameters=Available, Extrinsic parameters=Available, Temporal information=false, Number of cameras=42026.04 | 15.6 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Wu et al.Supervision Paradigm=Fully-Supervised2026.06 | 15.63 | — | — | — | — | — | — | — | — | — | — | — | — | 93.93 | 98.93 | 99.78 | 99.9 | 99.97 | |
| MV-SSMSupervision Paradigm=Fully-Supervised2026.06 | 15.7 | — | — | — | — | — | — | — | — | — | — | — | — | 93.5 | — | — | — | — | |
| MvPSupervision Paradigm=Fully-Supervised2026.06 | 15.76 | — | — | — | — | — | — | — | — | — | — | — | — | 92.28 | 96.6 | 97.45 | 97.69 | — | |
| MVGFormerSupervision Paradigm=Fully-Supervised2026.06 | 15.99 | — | — | — | — | — | — | — | — | — | — | — | — | 92.32 | 97.93 | 99.32 | 99.55 | 99.86 | |
| Plane SweepSupervision Paradigm=Fully-Supervised2026.06 | 16.75 | — | — | — | — | — | — | — | — | — | — | — | — | 92.12 | 98.96 | 99.81 | 99.84 | — | |
| VoxelPoseSupervision Paradigm=Fully-Supervised2026.06 | 17.68 | — | — | — | — | — | — | — | — | — | — | — | — | 83.59 | 98.33 | 99.76 | 99.91 | — | |
| Ye et al.Intrinsic parameters=Available, Extrinsic parameters=Available, Temporal information=false, Number of cameras=42026.04 | 17.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Faster Voxel.Supervision Paradigm=Fully-Supervised2026.06 | 18.26 | — | — | — | — | — | — | — | — | — | — | — | — | 85.22 | 98.08 | 99.32 | 99.48 | — | |
| PlanePoseTraining Manner=supervised2023.12 | 18.54 | — | — | — | — | — | — | — | — | — | — | — | 11.92 | — | — | — | — | — | |
| DisPOSESupervision Paradigm=Self-Supervised2026.06 | 21.2 | — | — | — | — | — | — | — | — | — | — | — | — | 68.59 | 98.59 | 99.6 | 99.8 | 99.91 | |
| Iskakov et al. AlgebraicIntrinsic parameters=Available, Extrinsic parameters=Available, Temporal information=false, Number of cameras=42026.04 | 21.3 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| LiCamPoseTraining Manner=unsupervised, Pretraining=Pretrained on BasketBallSync2023.12 | 22 | — | — | — | — | — | — | — | — | — | — | — | 15.96 | — | — | — | — | — | |
| DSPSupervision Paradigm=Self-Supervised, uses 9 temporal frames=true2026.06 | 23.1 | — | — | — | — | — | — | — | — | — | — | — | — | 57.6 | 86.1 | 94 | — | — | |
| COMPOSESupervision Paradigm=Optimization-Based2026.06 | 23.62 | — | — | — | — | — | — | — | — | — | — | — | — | 54.66 | 97.27 | 98.94 | 99.17 | 99.83 | |
| Jiang et al.Intrinsic parameters=Available, Extrinsic parameters=Unavailable, Temporal information=false, Number of cameras=42026.04 | 24.2 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| SelfPose3dSupervision Paradigm=Self-Supervised2026.06 | 24.47 | — | — | — | — | — | — | — | — | — | — | — | — | 55.13 | 96.44 | 98.46 | 98.98 | 99.6 | |
| Bartol et al.Intrinsic parameters=Available, Extrinsic parameters=Unavailable, Temporal information=false, Number of cameras=42026.04 | 25.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MvPTraining Manner=supervised2023.12 | 25.78 | — | — | — | — | — | — | — | — | — | — | — | 25.02 | — | — | — | — | — | |
| MvPoseSupervision Paradigm=Optimization-Based2026.06 | 26.46 | — | — | — | — | — | — | — | — | — | — | — | — | 37.63 | 95.7 | 97.84 | 98.28 | 99.6 | |
| Gordon et al.Intrinsic parameters=Available, Extrinsic parameters=Unavailable, Temporal information=false, Number of cameras=42026.04 | 28.9 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| DLTTraining Manner=unsupervised2023.12 | 38.26 | — | — | — | — | — | — | — | — | — | — | — | 44.77 | — | — | — | — | — | |
| Distribution-Aware Single-stageMethod category=Bottom-up2022.03 | 53.8 | — | 53.3 | 51.2 | 49.1 | 61.5 | — | — | — | — | — | — | — | — | — | — | — | — | |
| DASMethod category=Bottom-up methods, Backbone=2-stage MSPN, Inference device=single V100 GPU, Batch size=12022.03 | 54.6 | 75 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ours (w/ Synth)training_data=with synthetic2022.07 | 55 | — | 55.2 | 55 | 50.4 | 61.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Wang et al.Method category=Top-down2022.03 | 55.1 | — | 50.9 | 50.5 | 50.7 | 68.2 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Ourstraining_data=original2022.07 | 56.1 | — | 54.7 | 55.2 | 50.1 | 66.4 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Zhen et al.Method category=Bottom-up methods, reproduced=true2022.03 | 61.8 | 108 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Zhen et al.Method category=Bottom-up2022.03 | 61.8 | — | 63.1 | 60.3 | 56.6 | 67.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| SMAP2022.07 | 61.8 | — | 63.1 | 60.3 | 56.6 | 67.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Fabbri et al.Method category=Bottom-up methods2022.03 | 69 | 125 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Fabbri et al.Method category=Bottom-up2022.03 | 69 | — | 45 | 95 | 58 | 79 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Zanfir et al.2022.07 | 72.1 | — | 72.4 | 78.8 | 66.8 | 94.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Zanfir et al. [45]Method category=Bottom-up2022.03 | 78.1 | — | 72.4 | 78.8 | 66.8 | 94.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Moon et al.2022.07 | 87.6 | — | 89.6 | 91.3 | 79.6 | 90.1 | — | — | — | — | — | — | — | — | — | — | — | — | |
| Zanfir et al. [44]Method category=Top-down2022.03 | 153.4 | — | 140 | 165.9 | 150.7 | 156 | — | — | — | — | — | — | — | — | — | — | — | — | |
| ACTORSupervision Paradigm=Optimization-Based2026.06 | 168.4 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Popa et al.Method category=Top-down2022.03 | 203.4 | — | 217.9 | 187.3 | 193.6 | 221.3 | — | — | — | — | — | — | — | — | — | — | — | — | |
| 2D HPE + DepthAnything2# Views=12025.12 | — | — | — | — | — | — | — | — | 248.6 | — | — | — | — | — | — | — | — | — | |
| 2D HPE + ml-depth-pro# Views=12025.12 | — | — | — | — | — | — | — | — | 215.9 | — | — | — | — | — | — | — | — | — | |
| 6DRepNet2024.06 | — | — | — | — | — | — | — | — | — | — | 10.74 | — | — | — | — | — | — | — | |
| AdaFuseNumber of cameras=42024.04 | — | — | — | — | — | — | 69.7 | 595 | — | — | — | — | — | — | — | — | — | — | |
| AdaFuseNumber of cameras=102024.04 | — | — | — | — | — | — | 69.7 | 1,487.6 | — | — | — | — | — | — | — | — | — | — | |
| Epipolar TransformersNumber of cameras=42024.04 | — | — | — | — | — | — | 68.1 | 406.5 | — | — | — | — | — | — | — | — | — | — | |
| Epipolar TransformersNumber of cameras=102024.04 | — | — | — | — | — | — | 68.1 | 1,016.2 | — | — | — | — | — | — | — | — | — | — | |
| Frozen CogVLM + Regressortokens=visual tokens2024.06 | — | — | — | — | — | — | — | — | — | — | 32.11 | — | — | — | — | — | — | — | |
| Frozen CogVLM + Regressortokens=all tokens2024.06 | — | — | — | — | — | — | — | — | — | — | 33.47 | — | — | — | — | — | — | — | |
| FullyConnected + RCS + Rays# Views=22025.12 | — | — | — | — | — | — | — | — | 127.4 | 108.6 | — | — | — | — | — | — | — | — | |
| HopeNet2024.06 | — | — | — | — | — | — | — | — | — | — | 22.16 | — | — | — | — | — | — | — | |
| HPE-CogVLM2024.06 | — | — | — | — | — | — | — | — | — | — | 7.36 | 0.052 | — | — | — | — | — | — | |
| Learn.-triang.# Views=52025.12 | — | — | — | — | — | — | — | — | 130.9 | — | — | — | — | — | — | — | — | — | |
| Learnable TriangulationNumber of cameras=42024.04 | — | — | — | — | — | — | 80.6 | 716.1 | — | — | — | — | — | — | — | — | — | — | |
| Learnable TriangulationNumber of cameras=102024.04 | — | — | — | — | — | — | 80.6 | 1,326.9 | — | — | — | — | — | — | — | — | — | — | |
| Lin et al.Method category=Top-down methods, reproduced=true2022.03 | — | 118 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| Moon et al.Method category=Top-down methods, reproduced=true2022.03 | — | 107 | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | — | |
| MPL + RCS# Views=22025.12 | — | — | — | — | — | — | — | — | 278.8 | 274.9 | — | — | — | — | — | — | — | — | |
| MPL + RCS + Rays# Views=22025.12 | — | — | — | — | — | — | — | — | 55.7 | 45.4 | — | — | — | — | — | — | — | — | |
| MTF-TransformersNumber of cameras=42024.04 | — | — | — | — | — | — | 78.6 | 407 | — | — | — | — | — | — | — | — | — | — | |
| MTF-TransformersNumber of cameras=102024.04 | — | — | — | — | — | — | 78.6 | 1,017.8 | — | — | — | — | — | — | — | — | — | — | |
| Non-merging CogVLM2024.06 | — | — | — | — | — | — | — | — | — | — | 8.18 | 0.13 | — | — | — | — | — | — | |
| PoseBERT2024.06 | — | — | — | — | — | — | — | — | — | — | 22.32 | — | — | — | — | — | — | — | |
| PPT# Views=42025.12 | — | — | — | — | — | — | — | — | 108.3 | 114.2 | — | — | — | — | — | — | — | — | |
| RUMPL# Views=12025.12 | — | — | — | — | — | — | — | — | 95.9 | — | — | — | — | — | — | — | — | — | |
| RUMPL# Views=22025.12 | — | — | — | — | — | — | — | — | 35 | 30.8 | — | — | — | — | — | — | — | — | |
| RUMPL (w/o Conf)# Views=2, 2D pose confidence=false2025.12 | — | — | — | — | — | — | — | — | 41.1 | 36.6 | — | — | — | — | — | — | — | — | |
| TA merging CogVLMlambda=0.52024.06 | — | — | — | — | — | — | — | — | — | — | 7.72 | 68.9 | — | — | — | — | — | — | |
| Triangulation# Views=22025.12 | — | — | — | — | — | — | — | — | 44 | 43 | — | — | — | — | — | — | — | — | |
| UPose3DNumber of cameras=4, Backbone=HRNet-W482024.04 | — | — | — | — | — | — | 65.4 | 208.7 | — | — | — | — | — | — | — | — | — | — | |
| UPose3DNumber of cameras=10, Backbone=HRNet-W482024.04 | — | — | — | — | — | — | 65.4 | 517.7 | — | — | — | — | — | — | — | — | — | — | |
| Vision Tower (ViT) + Regressor2024.06 | — | — | — | — | — | — | — | — | — | — | 10.88 | — | — | — | — | — | — | — | |
| WHENet2024.06 | — | — | — | — | — | — | — | — | — | — | 29.55 | — | — | — | — | — | — | — |