Robust 2D Human Pose Estimation with Parallel Graph–Attention Modeling and Entropy-Aware Feature Decoding
Abstract
1. Introduction
- Parallel modeling framework with adaptive coordination: We propose a dual-branch architecture that explicitly models keypoint correlations via GNNs and implicitly captures long-range dependencies via self-attention, coupled with an adaptive nonlinear fusion mechanism to dynamically coordinate structural and semantic cues for robust 2D heatmap-based pose estimation.
- Adaptive Nonlinear Feature Coordination Module: Instead of simple concatenation or linear projection, we introduce a dual-dimensional (channel-wise and spatial-wise) nonlinear attention-based fusion mechanism that dynamically re-weights structured and contextual feature maps in 2D heatmap space, enabling fine-grained conflict mitigation and feature complementarity.
- Spatial Graph Perception Operator: Instead of applying standard graph convolution on joint embeddings, we design an edge-conditioned spatial graph perception layer that performs attention-weighted feature aggregation directly on keypoint-level feature maps, enabling adaptive structural reasoning under occlusion and spatial ambiguity.
- Information-Theoretic Perspective on Representation Refinement: We provide an information-theoretic perspective to interpret occlusion, background clutter, and feature interference as sources of spatial uncertainty in heatmap representations. Under this view, the proposed attention filtering, parallel modeling, and adaptive fusion mechanisms can be understood as progressively concentrating response distributions and reducing spatial ambiguity in keypoint localization.
- Refinement-Decoding Coupled Bias Mitigation: Rather than proposing a new decoding algorithm, we integrate distribution-aware error-compensation decoding within the parallel refinement framework, demonstrating that structural and contextual refinement substantially reduces systematic coordinate bias prior to decoding and alters the optimal compensation regime.
2. Related Work
2.1. Human Pose Estimation
2.2. Heatmap Decoding Methods
2.3. Graph Neural Network-Based Pose Estimation
2.4. Attention Mechanisms
2.5. Information-Theoretic Interpretation of Uncertainty in Human Pose Estimation
3. Methodology
3.1. Overall Framework
3.2. Feature Filtering Module
3.3. Explicit Modeling Module
3.4. Implicit Modeling Module
3.5. Multi-Head Attention Mechanisms and Feature Fusion Strategies
3.6. Error Compensation Decoding Methods
4. Experimentation
4.1. Datasets and Evaluation Metrics
4.1.1. Datasets
MPII Dataset
MSCOCO Dataset
4.1.2. Evaluation Metrics
MPII Metric
MSCOCO Metric
4.2. Results on the MPII Validation Dataset
4.3. Results on the MSCOCO Dataset
4.4. Ablation Experiment
4.4.1. Feature Fusion Approach
4.4.2. PMNet Layers
4.4.3. Multiple Attention Mechanisms
4.4.4. Ablation Study on Error Compensation Decoding
4.5. Qualitative Analysis
4.5.1. Comparison Results on the MPII Dataset
4.5.2. Output Distribution Visualization
4.5.3. Visualization Results on the MPII Validation Dataset
4.5.4. Visualization Results on the MSCOCO val2017 Dataset
5. Conclusions
5.1. Discussion
5.2. Future Research
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Jiao, Y.; Yao, H.; Xu, C. PEN: Pose-Embedding Network for Pedestrian Detection. IEEE Trans. Circuits Syst. Video Technol. 2021, 31, 1150–1162. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Y.; Yang, J.; Huang, H.; Xie, L. AdaPose: Toward Cross-Site Device-Free Human Pose Estimation with Commodity WiFi. IEEE Internet Things J. 2024, 11, 40255–40267. [Google Scholar] [CrossRef] [Scilit]
- Mohamed, A.; Qian, K.; Elhoseiny, M.; Claudel, C. Social-STGCNN: A Social Spatio-Temporal Graph Convolutional Neural Network for Human Trajectory Prediction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 14412–14420. [Google Scholar]
- Xiao, B.; Wu, H.; Wei, Y. Simple Baselines for Human Pose Estimation and Tracking. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 466–481. [Google Scholar]
- Zhang, F.; Zhu, X.; Dai, H.; Ye, M.; Zhu, C. Distribution-Aware Coordinate Representation for Human Pose Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 7091–7100. [Google Scholar]
- Li, J.; Bian, S.; Zeng, A.; Wang, C.; Pang, B.; Liu, W.; Lu, C. Human Pose Regression with Residual Log-Likelihood Estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 11005–11014. [Google Scholar]
- Newell, A.; Yang, K.; Deng, J. Stacked Hourglass Networks for Human Pose Estimation. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 11–14 October 2016; pp. 483–499. [Google Scholar]
- Yang, F.; Song, Z.; Xiao, Z.; Mo, Y.; Chen, Y.; Pan, Z.; Zhang, M.; Zhang, Y.; Qian, B.; Jin, W. Error Compensation Heatmap Decoding for Human Pose Estimation. IEEE Access 2021, 9, 114514–114522. [Google Scholar] [CrossRef] [Scilit]
- Gao, Z.; Chen, J.; Liu, Y.; Jin, Y.; Tian, D. A Systematic Survey on Human Pose Estimation. Artif. Intell. Rev. 2025, 58, 68. [Google Scholar] [CrossRef] [Scilit]
- Hong, X.; Zhang, L.; Yu, X.; Xie, W.; Xie, Y. MBA-Net: Multi-Branch Attention Network for Occluded Person Re-Identification. Multimed. Tools Appl. 2023, 83, 6393–6412. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Z.; Rahmani, H.; Black, S.; Williams, B.M. A Probabilistic Attention Model with Occlusion-Aware Texture Regression for 3D Hand Reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 758–767. [Google Scholar]
- Efros, A.A.; Berg, A.C.; Mori, G.; Malik, J. Recognizing Action at a Distance. In Proceedings of the IEEE International Conference on Computer Vision, Nice, France, 13–16 October 2003; pp. 726–733. [Google Scholar]
- Felzenszwalb, P.F.; Huttenlocher, D.P. Pictorial Structures for Object Recognition. Int. J. Comput. Vis. 2005, 61, 55–79. [Google Scholar] [CrossRef] [Scilit]
- Toshev, A.; Szegedy, C. DeepPose: Human Pose Estimation via Deep Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 1653–1660. [Google Scholar]
- Rogez, G.; Weinzaepfel, P.; Schmid, C. LCR-Net: Localization-Classification-Regression for Human Pose. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1216–1224. [Google Scholar]
- Rogez, G.; Weinzaepfel, P.; Schmid, C. LCR-Net++: Multi-Person 2D and 3D Pose Detection in Natural Images. IEEE Trans. Pattern Anal. Mach. Intell. 2020, 42, 1146–1161. [Google Scholar] [CrossRef] [Scilit]
- Pavllo, D.; Feichtenhofer, C.; Grangier, D.; Auli, M. 3D Human Pose Estimation in Video with Temporal Convolutions and Semi-Supervised Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 7745–7754. [Google Scholar]
- Zheng, C.; Zhu, S.; Mendieta, M.; Yang, T.; Chen, C.; Ding, Z. 3D Human Pose Estimation with Spatial and Temporal Transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 11636–11645. [Google Scholar]
- Wang, L.; Chen, Y.; Guo, Z.; Qian, K.; Lin, M.; Li, H.; Ren, J.S. Generalizing Monocular 3D Human Pose Estimation in the Wild. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27–28 October 2019; pp. 4024–4033. [Google Scholar]
- Tompson, J.; Jain, A.; LeCun, Y.; Bregler, C. Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS 2014), Montreal, QC, Canada, 8–13 December 2014; pp. 1799–1807. [Google Scholar]
- Cai, Y.; Wang, Z.; Luo, Z.; Yin, B.; Du, A.; Wang, H.; Zhang, X.; Zhou, X.; Zhou, E.; Sun, J. Learning Delicate Local Representations for Multi-Person Pose Estimation. In Proceedings of the European Conference on Computer Vision, Glasgow, UK, 23–28 August 2020; pp. 455–472. [Google Scholar]
- Chu, X.; Yang, W.; Ouyang, W.; Ma, C.; Yuille, A.L.; Wang, X. Multi-Context Attention for Human Pose Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 5669–5678. [Google Scholar]
- Ke, L.; Chang, M.C.; Qi, H.; Lyu, S. Multi-Scale Structure-Aware Network for Human Pose Estimation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 731–746. [Google Scholar]
- Yang, W.; Li, S.; Ouyang, W.; Li, H.; Wang, X. Learning Feature Pyramids for Human Pose Estimation. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 1290–1299. [Google Scholar]
- Chen, Y.; Wang, Z.; Peng, Y.; Zhang, Z.; Yu, G.; Sun, J. Cascaded Pyramid Network for Multi-Person Pose Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 7103–7112. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Wang, J.; Sun, K.; Cheng, T.; Jiang, B.; Deng, C.; Zhao, Y.; Liu, D.; Mu, Y.; Tan, M.; Wang, X.; et al. Deep High-Resolution Representation Learning for Visual Recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 3349–3364. [Google Scholar] [CrossRef] [Scilit]
- Zheng, C.; Wu, W.; Chen, C.; Yang, T.; Zhu, S.; Shen, J.; Kehtarnavaz, N.; Shah, M. Deep Learning-Based Human Pose Estimation: A Survey. ACM Comput. Surv. 2024, 56, 1–37. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Z.; Wan, L.; Xu, W.; Wang, S. Low-Resolution Human Pose Estimation and Action Recognition via Pose-Driven Super-Resolution Reconstruction. Mach. Learn. 2025, 114, 135. [Google Scholar] [CrossRef] [Scilit]
- Bai, X.; Wei, X.; Wang, Z.; Zhang, M. CONet: Crowd and Occlusion-Aware Network for Occluded Human Pose Estimation. Neural Netw. 2024, 172, 106109. [Google Scholar] [CrossRef] [Scilit]
- Li, M.; Wang, Y.; Hu, H.; Zhao, X. InferTrans: Hierarchical Structural Fusion Transformer for Crowded Human Pose Estimation. Inf. Fusion 2025, 117, 102878. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Liu, J.; Tang, J.; Wu, G.; Xu, B.; Chou, Y.; Wang, Y. GTPT: Group-Based Token Pruning Transformer for Efficient Human Pose Estimation. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; pp. 213–230. [Google Scholar]
- Chen, Z.; Dai, J.; Pan, J.; Zhou, F. Diffusion Model with Temporal Constraint for 3D Human Pose Estimation. Vis. Comput. 2025, 41, 5961–5977. [Google Scholar] [CrossRef] [Scilit]
- Bao, W.; Xiang, X. DDBMHT: A Diffusion-Based Double-Branch Multi-Hypothesis Transformer for 3D Human Pose Estimation in Video. In Proceedings of the International Conference on Electronic Technology and Information Science, Hangzhou, China, 17–19 May 2024; pp. 35–39. [Google Scholar]
- Feng, Y.; Dai, S.; Zhang, Q.; Wang, Z.; Zhang, X.; Zhou, Y. M3Pose: Multi-Person 3D Pose Estimation Using Sparse Millimeter-Wave Radar Point Clouds. In Proceedings of the Chinese Conference on Pattern Recognition and Computer Vision, Urumqi, China, 18–20 October 2024; pp. 504–517. [Google Scholar]
- Al, M.A.; Shi, X.; Mondher, B.; Ohtsuki, T. mmGAT: Pose Estimation by Graph Attention with Mutual Features from mmWave Radar Point Cloud. In Proceedings of the IEEE International Conference on Communications, Denver, CO, USA, 9–13 June 2024; pp. 2161–2166. [Google Scholar]
- Zhao, L.; Xu, J.; Zhang, S.; Gong, C.; Yang, J.; Gao, X. Perceiving Heavily Occluded Human Poses by Assigning Unbiased Score. Inf. Sci. 2020, 537, 284–301. [Google Scholar] [CrossRef] [Scilit]
- Ying, J.J.C.; Chen, Y.H.; Zhang, J. Few-Shot Learning-Based Human Pose Estimation Model. Inf. Sci. 2025, 717, 122320. [Google Scholar] [CrossRef] [Scilit]
- Reddy, N.D.; Vo, M.; Narasimhan, S.G. Occlusion-Net: 2D/3D Occluded Keypoint Localization Using Graph Networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 15–20 June 2019; pp. 7318–7327. [Google Scholar]
- Bin, Y.; Chen, Z.M.; Wei, X.S.; Chen, X.; Gao, C.; Sang, N. Structure-Aware Human Pose Estimation with Graph Convolutional Networks. Pattern Recognit. 2020, 106, 107410. [Google Scholar] [CrossRef] [Scilit]
- Tian, L.; Wang, P.; Liang, G.; Shen, C. An Adversarial Human Pose Estimation Network Injected with Graph Structure. Pattern Recognit. 2021, 115, 107863. [Google Scholar] [CrossRef] [Scilit]
- Ke, L.; Tai, Y.W.; Tang, C.K. Deep Occlusion-Aware Instance Segmentation with Overlapping BiLayers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 4018–4027. [Google Scholar]
- Jiang, Y.; Ding, W.; Li, H.; Chi, Z. Multi-Person Pose Tracking with Sparse Key-Point Flow Estimation. IEEE Trans. Image Process. 2024, 33, 3590–3605. [Google Scholar] [CrossRef] [Scilit]
- Chen, A.; Wu, C.; Leng, C. Hourglass-GCN for 3D Human Pose Estimation Using Skeleton Structure and View Correlation. Comput. Mater. Contin. 2025, 82, 173–191. [Google Scholar] [CrossRef] [Scilit]
- Hou, Y.; Wang, C.; Peng, H.; Feng, T.; Li, H.; Oh, Y.P. A Robust Framework for 3D Human Pose Estimation Using Semantic Graph Convolution, Criss-Cross Attention and Transformer Encoder. J. Circuits Syst. Comput. 2025, 34, 2550252. [Google Scholar] [CrossRef] [Scilit]
- Wei, S.E.; Ramakrishna, V.; Kanade, T.; Sheikh, Y. Convolutional Pose Machines. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 4724–4732. [Google Scholar]
- Xu, D.; Wang, T.; Hao, F.; Cheng, J. Global-Local Interplay with Transformer and GCN for 3D Human Pose Estimation. Procedia Comput. Sci. 2025, 271, 169–175. [Google Scholar] [CrossRef] [Scilit]
- Ye, M.; Yang, L.; Zhu, H.; Zheng, Z.; Wang, X.; Lo, Y. Dual-stream Transformer-GCN Model with Contextualized Representations Learning for Monocular 3D Human Pose Estimation. arXiv 2025, arXiv:2504.01764. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision, Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Huang, Z.; Wang, X.; Huang, L.; Huang, C.; Wang, Y.; Liu, W. CCNet: Criss-Cross Attention for Semantic Segmentation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 6896–6908. [Google Scholar] [CrossRef] [Scilit]
- Yang, C.; Tkach, A.; Hampali, S.; Zhang, L.; Crowley, E.J.; Keskin, C. EgoPoseFormer: A Simple Baseline for Stereo Egocentric 3D Human Pose Estimation. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024; pp. 401–417. [Google Scholar]
- Li, Y.; Zhang, S.; Wang, Z.; Yang, S.; Yang, W.; Xia, S.T.; Zhou, E. TokenPose: Learning Keypoint Tokens for Human Pose Estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 11293–11302. [Google Scholar]
- Wang, X.; Tong, J.; Wang, R. Attention Refined Network for Human Pose Estimation. Neural Process. Lett. 2021, 53, 2853–2872. [Google Scholar] [CrossRef] [Scilit]
- Liu, F.H.; Zhang, X.; Wang, H.Y.; Feng, J. Context-Aware Superpixel and Bilateral Entropy-Image Coherence Induces Less Entropy. Entropy 2020, 22, 20. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.Y.; Hu, R.Y.; Xue, C.Q. Enhancing User Perception of Reliability in Computer Vision: Uncertainty Visualization for Probability Distributions. Symmetry 2024, 16, 986. [Google Scholar] [CrossRef] [Scilit]
- Rabiee, S.; Biswas, J. Introspective Perception for Mobile Robots. Artif. Intell. 2023, 324, 103999. [Google Scholar] [CrossRef] [Scilit]
- Ferreira, R.S.; Guérin, J.; Delmas, K.; Guiochet, J.; Waeselynck, H. Safety Monitoring of Machine Learning Perception Functions: A Survey. Comput. Intell. 2025, 41, e70032. [Google Scholar] [CrossRef] [Scilit]
- Gasperini, S.; Haug, J.; Mahani, M.A.N.; Marcos-Ramiro, A.; Navab, N.; Busam, B.; Tombari, F. CertainNet: Sampling-Free Uncertainty Estimation for Object Detection. IEEE Robot. Autom. Lett. 2022, 7, 698–705. [Google Scholar] [CrossRef] [Scilit]
- Su, S.B.; Han, S.Y.; Li, Y.M.; Zhang, Z.L.; Feng, C.; Ding, C.W.; Miao, F. Collaborative Multi-Object Tracking With Conformal Uncertainty Propagation. IEEE Robot. Autom. Lett. 2024, 9, 3323–3330. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Z.B.; Qi, H.Y.; Fan, X.Q.; Xu, G.Z.; Qi, Y.C.; Zhai, Y.J.; Zhang, K. Image Representation Method Based on Relative Layer Entropy for Insulator Recognition. Entropy 2020, 22, 419. [Google Scholar] [CrossRef] [Scilit]
- Liu, C.H.; Chen, H.R.; Deng, L.; Guo, C.T.; Lu, X.T.; Yu, H.; Zhu, L.Q.; Dong, M.L. Modality Specific Infrared and Visible Image Fusion Based on Multi-Scale Rich Feature Representation Under Low-Light Environment. Infrared Phys. Technol. 2024, 140, 105351. [Google Scholar] [CrossRef] [Scilit]
- Zhu, G.L.; Fei, H.X.; Hong, J.K.; Luo, Y.Y.; Long, J. An Information-Reserved and Deviation-Controllable Binary Neural Network for Object Detection. Mathematics 2023, 11, 62. [Google Scholar] [CrossRef] [Scilit]
- Li, H.; Yao, H.; Hou, Y. HPNet: Hybrid Parallel Network for Human Pose Estimation. Sensors 2023, 23, 4425. [Google Scholar] [CrossRef] [Scilit]
- Bruna, J.; Zaremba, W.; Szlam, A.; LeCun, Y. Spectral Networks and Locally Connected Networks on Graphs. arXiv 2014, arXiv:1312.6203. [Google Scholar] [CrossRef] [Scilit]
- Defferrard, M.; Bresson, X.; Vandergheynst, P. Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering. arXiv 2016, arXiv:1606.09375. [Google Scholar]
- Ding, X.; Zhang, X.; Han, J.; Ding, G. Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 11953–11965. [Google Scholar]
- Yang, S.; Quan, Z.; Nie, M.; Yang, W. TransPose: Keypoint Localization via Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 11782–11792. [Google Scholar]
- Andriluka, M.; Pishchulin, L.; Gehler, P.; Schiele, B. 2D Human Pose Estimation: New Benchmark and State-of-the-Art Analysis. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Columbus, OH, USA, 23–28 June 2014; pp. 3686–3693. [Google Scholar]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft COCO: Common Objects in Context. In Proceedings of the European Conference on Computer Vision, Zurich, Switzerland, 6–12 September 2014; pp. 740–755. [Google Scholar]
- Yan, S.; Xiong, Y.; Lin, D. Spatial Temporal Graph Convolutional Networks for Skeleton-Based Action Recognition. In Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018; pp. 7444–7452. [Google Scholar]
- Bulat, A.; Tzimiropoulos, G. Human Pose Estimation via Convolutional Part Heatmap Regression. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 11–14 October 2016; pp. 717–732. [Google Scholar]
- Tang, W.; Yu, P.; Wu, Y. Deeply Learned Compositional Models for Human Pose Estimation. In Proceedings of the European Conference on Computer Vision, Munich, Germany, 8–14 September 2018; pp. 197–214. [Google Scholar]
- Zheng, G.; Wang, S.; Yang, B. Hierarchical Structure Correlation Inference for Pose Estimation. Neurocomputing 2020, 404, 186–197. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Geng, F.; Wang, X. MTPose: Human Pose Estimation with High-Resolution Multi-Scale Transformers. Neural Process. Lett. 2022, 54, 3941–3964. [Google Scholar] [CrossRef] [Scilit]
- Zou, X.; Bi, X.; Yu, C. Improving Human Pose Estimation Based on Stacked Hourglass Network. Neural Process. Lett. 2023, 55, 9521–9544. [Google Scholar] [CrossRef] [Scilit]












| Method | Head | Shoulder | Elbow | Wrist | Hip | Knee | Ankle | PCKh@0.5 |
|---|---|---|---|---|---|---|---|---|
| Newell et al. [7] | 96.5 | 96.0 | 90.3 | 85.4 | 88.8 | 85.0 | 81.9 | 89.2 |
| Yang et al. [24] | 96.8 | 96.0 | 90.4 | 86.0 | 89.5 | 85.2 | 82.3 | 89.6 |
| Xiao et al. [4] | 97.0 | 95.9 | 90.3 | 85.0 | 89.2 | 85.3 | 81.3 | 89.6 |
| Bulat and Tzimiropoulos [71] | 97.9 | 95.1 | 89.9 | 85.3 | 89.4 | 85.7 | 81.7 | 89.7 |
| Tang et al. [72] | 95.6 | 95.9 | 90.7 | 86.5 | 89.9 | 86.6 | 82.5 | 89.8 |
| Li et al. [52] | 97.2 | 95.9 | 90.4 | 86.0 | 89.3 | 87.1 | 82.5 | 90.2 |
| Zheng et al. [73] | 97.1 | 95.8 | 90.1 | 86.2 | 89.2 | 88.1 | 85.2 | 90.6 |
| Chu et al. [22] | 98.5 | 96.3 | 91.9 | 88.1 | 90.6 | 88.0 | 85.0 | 91.5 |
| Li et al. [63] | 97.0 | 96.7 | 92.2 | 88.0 | 91.5 | 88.7 | 85.3 | 91.8 |
| Chen et al. [25] | 98.1 | 96.5 | 92.5 | 88.5 | 90.2 | 89.6 | 86.0 | 91.9 |
| Bin et al. [40] | 98.0 | 96.9 | 92.7 | 89.0 | 91.8 | 89.4 | 86.1 | 92.4 |
| HRNet (baseline) | 96.9 | 96.0 | 90.6 | 85.8 | 88.7 | 86.6 | 83.6 | 90.1 |
| HRNet + PMNet (Ours) | 97.3 | 97.0 | 92.6 | 88.3 | 91.9 | 89.6 | 86.2 | 92.4 |
| ResNet-50 (baseline) | 88.3 | 82.6 | 73.6 | 64.4 | 66.0 | 60.4 | 54.3 | 70.0 |
| ResNet-50 + PMNet (Ours) | 97.2 | 95.8 | 88.1 | 86.4 | 87.8 | 86.5 | 79.9 | 90.2 |
| Method | Backbone | AP | AP50 | AP75 | APM | APL | AR |
|---|---|---|---|---|---|---|---|
| Newell et al. [7] | – | 0.669 | – | – | – | – | – |
| CPN [25] | ResNet-50 | 0.689 | – | – | – | – | – |
| CPN+OHKM [25] | ResNet-50 | 0.694 | – | – | – | – | – |
| CPM [46] | – | 0.669 | 0.823 | 0.763 | 0.653 | 0.759 | 0.744 |
| MTPose [74] | HRNet-W48 | 0.753 | 0.899 | 0.820 | 0.719 | 0.819 | 0.804 |
| Zou et al. [75] | Hourglass-8 | 0.753 | 0.902 | 0.822 | 0.719 | 0.820 | 0.805 |
| HR-ARNet [53] | HRNet-W48 | 0.749 | 0.904 | 0.823 | 0.714 | 0.818 | 0.803 |
| Xiao et al. [4] | ResNet-50 | 0.720 | 0.893 | 0.798 | 0.687 | 0.789 | 0.778 |
| Li et al. [63] | – | 0.768 | 0.913 | 0.835 | 0.733 | 0.827 | 0.809 |
| HRNet | – | 0.755 | 0.925 | 0.833 | 0.729 | 0.815 | 0.801 |
| HRNet + PMNet (Ours) | HRNet-W48 | 0.773 | 0.938 | 0.845 | 0.743 | 0.829 | 0.809 |
| ResNet-50 | ResNet-50 | 0.370 | 0.709 | 0.331 | 0.329 | 0.432 | 0.438 |
| ResNet-50 + PMNet (Ours) | ResNet-50 | 0.743 | 0.925 | 0.815 | 0.711 | 0.794 | 0.772 |
| Fusion Method | Head | Shoulder | Elbow | Wrist | Hip | Knee | Ankle | PCKh@0.5 |
|---|---|---|---|---|---|---|---|---|
| (a) Channel-based nonlinear fusion | 97.58 | 96.72 | 92.47 | 88.61 | 91.64 | 89.38 | 84.69 | 92.16 |
| (b) Spatial-based nonlinear fusion | 97.48 | 96.55 | 92.31 | 88.74 | 91.41 | 89.32 | 84.65 | 92.05 |
| L | Head | Shoulder | Elbow | Wrist | Hip | Knee | Ankle | PCKh@0.5 |
|---|---|---|---|---|---|---|---|---|
| 1 | 97.58 | 96.72 | 92.47 | 88.61 | 91.64 | 89.08 | 84.67 | 92.16 |
| 2 | 97.51 | 96.93 | 92.38 | 88.62 | 91.35 | 89.20 | 85.19 | 92.20 |
| 3 | 97.41 | 96.64 | 92.18 | 88.61 | 91.52 | 89.36 | 84.98 | 92.15 |
| 4 | 97.31 | 96.62 | 92.59 | 88.52 | 91.61 | 89.54 | 85.40 | 92.28 |
| 6 | 97.47 | 96.68 | 92.67 | 88.64 | 91.24 | 89.50 | 85.28 | 92.26 |
| h | Head | Shoulder | Elbow | Wrist | Hip | Knee | Ankle | PCKh@0.5 |
|---|---|---|---|---|---|---|---|---|
| 1 | 97.31 | 96.62 | 92.59 | 88.52 | 91.61 | 89.54 | 85.40 | 92.28 |
| 2 | 97.20 | 96.52 | 92.53 | 88.71 | 91.36 | 89.42 | 84.93 | 92.19 |
| 3 | 97.37 | 96.76 | 92.36 | 88.37 | 91.35 | 89.28 | 84.98 | 92.13 |
| Δ | Head | Shoulder | Elbow | Wrist | Hip | Knee | Ankle | PCKh@0.5 |
|---|---|---|---|---|---|---|---|---|
| 4 | 96.93 | 96.40 | 91.89 | 87.68 | 90.29 | 88.09 | 83.66 | 91.46 |
| 3 | 97.14 | 96.62 | 92.47 | 88.26 | 91.07 | 88.92 | 84.55 | 92.21 |
| 2 | 97.20 | 96.62 | 92.48 | 88.32 | 91.35 | 89.14 | 84.93 | 92.34 |
| 1 | 97.30 | 97.01 | 92.60 | 88.33 | 91.36 | 89.58 | 85.17 | 92.42 |
| 0 | 97.03 | 96.52 | 92.60 | 88.30 | 91.52 | 89.54 | 85.14 | 92.31 |
| −1 | 97.03 | 96.50 | 92.50 | 88.30 | 91.54 | 89.56 | 84.93 | 92.13 |
| −2 | 97.00 | 96.50 | 92.38 | 88.28 | 91.57 | 89.56 | 84.93 | 92.12 |
| −3 | 97.00 | 96.50 | 92.25 | 88.13 | 91.59 | 89.44 | 84.77 | 92.06 |
| −4 | 97.00 | 96.43 | 92.26 | 87.84 | 91.55 | 89.44 | 84.70 | 92.00 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhao, J.; Yu, D.; Han, C.; Xu, Y.; Shi, C. Robust 2D Human Pose Estimation with Parallel Graph–Attention Modeling and Entropy-Aware Feature Decoding. Entropy 2026, 28, 265. https://doi.org/10.3390/e28030265
Zhao J, Yu D, Han C, Xu Y, Shi C. Robust 2D Human Pose Estimation with Parallel Graph–Attention Modeling and Entropy-Aware Feature Decoding. Entropy. 2026; 28(3):265. https://doi.org/10.3390/e28030265
Chicago/Turabian StyleZhao, Jiayuan, Dingyao Yu, Chunjia Han, Yingcheng Xu, and Chunlei Shi. 2026. "Robust 2D Human Pose Estimation with Parallel Graph–Attention Modeling and Entropy-Aware Feature Decoding" Entropy 28, no. 3: 265. https://doi.org/10.3390/e28030265
APA StyleZhao, J., Yu, D., Han, C., Xu, Y., & Shi, C. (2026). Robust 2D Human Pose Estimation with Parallel Graph–Attention Modeling and Entropy-Aware Feature Decoding. Entropy, 28(3), 265. https://doi.org/10.3390/e28030265

