Keypoint Detection on COCO (val)
79.4APViTPose++-H
Evaluation Results
| Method | Links | ||||||
|---|---|---|---|---|---|---|---|
| ViTPose++-HBackbone=ViT-H, Params (M)=632, Speed (fps)=241, FLOPs (G)=122.9, Memory (M)=7293, Input Resolution=256x192, Feature Resolution=1/16, Detector=Faster RCNN2022.12 | 79.4 | — | — | 75.8 | 86.5 | 84.8 | |
| ViTPose-HBackbone=ViT-H, Params (M)=632, Speed (fps)=241, FLOPs (G)=122.9, Memory (M)=7293, Input Resolution=256x192, Feature Resolution=1/16, Detector=Faster RCNN2022.12 | 79.1 | — | — | 75.3 | 86 | 84.1 | |
| ViTPose++-LBackbone=ViT-L, Params (M)=307, Speed (fps)=411, FLOPs (G)=59.8, Memory (M)=5587, Input Resolution=256x192, Feature Resolution=1/16, Detector=Faster RCNN2022.12 | 78.6 | — | — | 75.2 | 85.6 | 84.1 | |
| ViTPose-LBackbone=ViT-L, Params (M)=307, Speed (fps)=411, FLOPs (G)=59.8, Memory (M)=5587, Input Resolution=256x192, Feature Resolution=1/16, Detector=Faster RCNN2022.12 | 78.3 | — | — | 74.5 | 85.4 | 83.5 | |
| ViTPosemethod_category=Heatmap-based, evaluation_protocol=GT bbox, flip_test=false2024.06 | 77.4 | 93.6 | 84.8 | 74.7 | 81.9 | 80.2 | |
| LocLLMmethod_category=Language-based, evaluation_protocol=GT bbox, flip_test=false2024.06 | 77.4 | 94.4 | 85.2 | 74.5 | 81.8 | 80.6 | |
| UDPBackbone=HRNet-W48, Params (M)=64, Speed (fps)=309, FLOPs (G)=32.9, Memory (M)=7339, Input Resolution=384x288, Feature Resolution=1/4, Detector=Faster RCNN2022.12 | 77.2 | — | — | 73.2 | 84.4 | 82 | |
| HRFormer-BBackbone=HRFormer-B, Params (M)=43, Speed (fps)=78, FLOPs (G)=26.8, Memory (M)=7859, Input Resolution=384x288, Feature Resolution=1/4, Detector=Faster RCNN2022.12 | 77.2 | — | — | 73.2 | 84.2 | 82 | |
| ViTPose++-BBackbone=ViT-B, Params (M)=86, Speed (fps)=944, FLOPs (G)=17.1, Memory (M)=4589, Input Resolution=256x192, Feature Resolution=1/16, Detector=Faster RCNN2022.12 | 77 | — | — | 73.4 | 84 | 82.6 | |
| HRNetmethod_category=Heatmap-based, evaluation_protocol=GT bbox, flip_test=false2024.06 | 76.8 | 93.6 | 83.6 | 74 | 81.5 | 79.6 | |
| SimCCmethod_category=Heatmap-based, evaluation_protocol=GT bbox, flip_test=false2024.06 | 76.5 | 93.2 | 83.1 | 73.6 | 81.5 | 79.7 | |
| HRNetBackbone=HRNet-W48, Params (M)=64, Speed (fps)=309, FLOPs (G)=32.9, Memory (M)=7339, Input Resolution=384x288, Feature Resolution=1/4, Detector=Faster RCNN2022.12 | 76.3 | — | — | 72.3 | 83.4 | 81.2 | |
| HRNetBackbone=HRNet-W32, Params (M)=29, Speed (fps)=428, FLOPs (G)=16, Memory (M)=7049, Input Resolution=384x288, Feature Resolution=1/4, Detector=Faster RCNN2022.12 | 75.8 | — | — | 71.9 | 82.8 | 81 | |
| TokenPose-L/D24Backbone=HRNet-W48, Params (M)=28, Speed (fps)=602, FLOPs (G)=11, Memory (M)=3477, Input Resolution=256x192, Feature Resolution=1/4, Detector=Faster RCNN2022.12 | 75.8 | — | — | 72.3 | 82.7 | 80.9 | |
| TransPose-H/A6Backbone=HRNet-W48, Params (M)=18, Speed (fps)=309, FLOPs (G)=21.8, Input Resolution=256x192, Feature Resolution=1/4, Detector=Faster RCNN2022.12 | 75.8 | — | — | 76.4 | 87.2 | 80.8 | |
| ViTPose-BBackbone=ViT-B, Params (M)=86, Speed (fps)=944, FLOPs (G)=17.1, Memory (M)=4589, Input Resolution=256x192, Feature Resolution=1/16, Detector=Faster RCNN2022.12 | 75.8 | — | — | 72.1 | 82.2 | 81.1 | |
| ViTPose++-SBackbone=ViT-S, Params (M)=22, Speed (fps)=1439, FLOPs (G)=5.3, Memory (M)=3438, Input Resolution=256x192, Feature Resolution=1/16, Detector=Faster RCNN2022.12 | 75.8 | — | — | 72.3 | 82.6 | 81 | |
| HRFormer-BBackbone=HRFormer-B, Params (M)=43, Speed (fps)=158, FLOPs (G)=12.2, Memory (M)=4287, Input Resolution=256x192, Feature Resolution=1/4, Detector=Faster RCNN2022.12 | 75.6 | — | — | 71.7 | 82.6 | 80.8 | |
| MixFormer-B4Backbone=MixFormer-B42022.04 | 75.3 | 93.5 | 83.5 | — | — | — | |
| HRFormer-SBackbone=HRFormer-S2022.04 | 74.5 | 92.3 | 82.1 | — | — | — | |
| SimplePosemethod_category=Heatmap-based, evaluation_protocol=GT bbox, flip_test=false2024.06 | 74.4 | 92.6 | 82.5 | 71.5 | 79.2 | 77.6 | |
| Swin-TBackbone=Swin-T2022.04 | 74.2 | 92.5 | 82.5 | — | — | — | |
| X-Pose-TBackbone=Swin-T, Training Datasets=COCO, Human-Art, AP-10K, APT-36K, Dataset Volume=155K, Prompting=Textual2023.10 | 74.2 | — | — | 68.8 | 82.1 | — | |
| X-Pose-VBackbone=Swin-T, Training Datasets=COCO, Human-Art, AP-10K, APT-36K, Dataset Volume=155K, Prompting=Visual2023.10 | 74.1 | — | — | 68.8 | 81.8 | — | |
| RLEmethod_category=Regression-based, evaluation_protocol=GT bbox, flip_test=false2024.06 | 74 | 91.5 | 81.6 | 70.9 | 78.5 | 76.8 | |
| ViTPose-SBackbone=ViT-S, Params (M)=22, Speed (fps)=1439, FLOPs (G)=5.3, Memory (M)=3438, Input Resolution=256x192, Feature Resolution=1/16, Detector=Faster RCNN2022.12 | 73.8 | — | — | 70.5 | 80.4 | 79.2 | |
| SimpleBaselineBackbone=ResNet-152, Params (M)=60, Speed (fps)=829, FLOPs (G)=15.7, Memory (M)=3827, Input Resolution=256x192, Feature Resolution=1/32, Detector=Faster RCNN2022.12 | 73.5 | — | — | 69.9 | 80.2 | 79 | |
| FastPoseBackbone=ResNet-152, Params (M)=60, Speed (fps)=653, FLOPs (G)=16, Memory (M)=6013, Input Resolution=256x192, Feature Resolution=1/32, Detector=YoLov32022.12 | 73.3 | — | — | — | — | — | |
| CLIP baselinemethod_category=Language-based, evaluation_protocol=GT bbox, flip_test=false2024.06 | 73.1 | 92.5 | 81.3 | 70.3 | 77.4 | 76.5 | |
| ResNet50Backbone=ResNet502022.04 | 71.8 | 89.8 | 79.5 | — | — | — | |
| DY-ResNet-18Type=A, Backbone #Param=42.2M, Backbone MAdds=1.81G, Head Operator=dconv, Head #Param=8.4M, Head MAdds=5.4G2019.12 | 68.6 | 88.4 | 76.1 | 65.3 | 75.1 | 74.6 | |
| DY-MobileNetV2 x1.0Type=B, Backbone #Param=9.8M, Backbone MAdds=305.3M, Head Operator=bneck, Head #Param=6.3M, Head MAdds=709.4M2019.12 | 68.2 | 88.4 | 76 | 65 | 74.7 | 74.2 | |
| DY-MobileNetV2 x1.0Type=A, Backbone #Param=9.8M, Backbone MAdds=305.3M, Head Operator=dconv, Head #Param=8.4M, Head MAdds=5.4G2019.12 | 67.6 | 88.1 | 75.5 | 64.4 | 74.1 | 73.8 | |
| ResNet-18Type=A, Backbone #Param=10.6M, Backbone MAdds=1.77G, Head Operator=dconv, Head #Param=8.4M, Head MAdds=5.4G2019.12 | 67 | 87.9 | 74.8 | 63.6 | 73.5 | 73.1 | |
| MADBackbone=ViT-B, Parameter(M)=107M, Infer. Time(s)=0.262024.03 | 66.5 | — | — | — | — | — | |
| Keypoint R-CNNBackbone=R50-FPN, Parameter(M)=59M2024.03 | 65.5 | — | — | — | — | — | |
| OpenPose2021.03 | 65.3 | 85.2 | 71.3 | 62.2 | 70.7 | — | |
| Pix2SeqV2Backbone=ViT-B, Parameter(M)=132M, Infer. Time(s)=40.862024.03 | 64.8 | — | — | — | — | — | |
| MobileNetV2 x1.0Type=A, Backbone #Param=2.2M, Backbone MAdds=292.6M, Head Operator=dconv, Head #Param=8.4M, Head MAdds=5.4G2019.12 | 64.7 | 87.2 | 72.6 | 61.3 | 71 | 71 | |
| MobileNetV2 x1.0Type=B, Backbone #Param=2.2M, Backbone MAdds=292.6M, Head Operator=bneck, Head #Param=1.2M, Head MAdds=701.1M2019.12 | 64.6 | 87 | 72.4 | 61.3 | 71 | 71 | |
| MADBackbone=Swin-B, Parameter(M)=107M, Infer. Time(s)=0.192024.03 | 64.6 | — | — | — | — | — | |
| DY-MobileNetV2 x0.5Type=B, Backbone #Param=2.7M, Backbone MAdds=98.0M, Head Operator=bneck, Head #Param=6.3M, Head MAdds=709.4M2019.12 | 62.8 | 86.1 | 70.4 | 59.9 | 68.6 | 69.1 | |
| DY-MobileNetV2 x0.5Type=A, Backbone #Param=2.7M, Backbone MAdds=98.0M, Head Operator=dconv, Head #Param=8.4M, Head MAdds=5.4G2019.12 | 61.9 | 85.8 | 69.7 | 58.9 | 67.9 | 68.4 | |
| DY-MobileNetV3-SmallType=B, Backbone #Param=2.8M, Backbone MAdds=65.1M, Head Operator=bneck, Head #Param=4.9M, Head MAdds=671.1M2019.12 | 60 | 85 | 67.8 | 57.6 | 65.4 | 66.7 | |
| DY-MobileNetV3-SmallType=A, Backbone #Param=2.8M, Backbone MAdds=65.1M, Head Operator=dconv, Head #Param=8.4M, Head MAdds=5.4G2019.12 | 59.3 | 84.7 | 66.7 | 56.9 | 64.7 | 66.1 | |
| MobileNetV2 x0.5Type=B, Backbone #Param=0.7M, Backbone MAdds=93.7M, Head Operator=bneck, Head #Param=1.2M, Head MAdds=701.1M2019.12 | 59.2 | 84.3 | 66.4 | 56.2 | 65 | 65.6 | |
| MobileNetV3-SmallType=A, Backbone #Param=1.1M, Backbone MAdds=62.7M, Head Operator=dconv, Head #Param=8.4M, Head MAdds=5.4G2019.12 | 57.1 | 83.7 | 63.8 | 54.9 | 62.3 | 64.1 | |
| MobileNetV3-SmallType=B, Backbone #Param=1.1M, Backbone MAdds=62.7M, Head Operator=bneck, Head #Param=1.0M, Head MAdds=664.2M2019.12 | 57.1 | 83.8 | 63.7 | 55 | 62.2 | 64.1 | |
| MobileNetV2 x0.5Type=A, Backbone #Param=0.7M, Backbone MAdds=93.7M, Head Operator=dconv, Head #Param=8.4M, Head MAdds=5.4G2019.12 | 57 | 83.7 | 63.1 | 53.9 | 63.1 | 63.7 | |
| DeepPosemethod_category=Regression-based, evaluation_protocol=GT bbox, flip_test=false2024.06 | 53.8 | 82.6 | 59.2 | 52.2 | 57.3 | 66.8 | |
| PyMAFAuxiliary Supervision (AS)=true2021.03 | 24.6 | 48.9 | 22.7 | 26 | 24.2 | — | |
| SMPLifyimplementation=SPIN [26], optimization iterations=3002021.03 | 22 | 37.7 | 23.1 | 27.7 | 17.6 | — | |
| PyMAFAuxiliary Supervision (AS)=false2021.03 | 20.7 | 43.9 | 17.4 | 22.3 | 19.9 | — | |
| HMR2021.03 | 18.9 | 47.5 | 11.7 | 21.5 | 17 | — | |
| SPIN2021.03 | 17.3 | 39.1 | 13.5 | 19 | 16.6 | — | |
| Baseline2021.03 | 16.8 | 38.2 | 12.8 | 18.5 | 16 | — | |
| GraphCMR2021.03 | 9.3 | 26.9 | 4.2 | 11.3 | 8.1 | — | |
| Grounding DINO-BBackbone=Swin-B, Training Datasets=O365, GoldG, Cap4M, Dataset Volume=1858K, Fine-tuned=false2023.10 | 6.8 | — | — | 6.6 | 7.2 | — | |
| Grounding DINO-TBackbone=Swin-T, Training Datasets=O365, GoldG, Cap4M, Dataset Volume=1858K, Fine-tuned=false2023.10 | 3.1 | — | — | 2.8 | 3.2 | — | |
| Grounding DINO-TBackbone=Swin-T, Training Datasets=COCO, O365, GoldG, Cap4M, OpenImage, ODinW-35, RefCOCO, Dataset Volume=1858K + 155K, Fine-tuned=true2023.10 | 1.8 | — | — | 1.7 | 1.9 | — |