3D Pose Estimation Using Virtual Projection Based on 3D Reconstructed Model
Abstract
1. Introduction
- Ensuring 3D consistency through direct constraints on the 3D model: After PCA-based frontal alignment and multi-face virtual plane projection, we back-project the 2D poses from each view to form intersection regions, which yield geometrically plausible joint candidates inside the reconstructed body volume. This reduces occlusion and depth ambiguity in the evaluated setting, providing a consistent pipeline for stable 3D skeleton estimation from volumetric data.
- Density-based refinement and modularity: Using the bisector plane of the connected bones as a reference, we apply DBSCAN and circle or sphere fitting to remove outlier joints and refine their positions. This refinement module remains applicable even when the 2D pose detector or the joint definition changes, offering strong extensibility and reusability across different sensors and datasets.
2. Related Works
3. 3D Pose Estimation
3.1. Workflow
3.2. 3D Reconstruction
3.3. Multi-View Projection
3.4. 3D Skeleton Generation
3.5. 3D Skeleton Refinement
4. Experimental Result
4.1. Environment
4.2. Result of 3D Reconstruction
4.3. 3D Skeleton Extraction Result
4.4. Joint Error Refinement
4.5. Comparison with Motion Capture
4.6. Comparison with Other Research
5. Discussion & Remarks
5.1. Discussions
5.2. Remarks
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Munea, T.L.; Jembre, Y.Z.; Weldegebriel, H.T.; Chen, L.; Huang, C.; Yang, C. The Progress of Human Pose Estimation: A Survey and Taxonomy of Models Applied in 2D Human Pose Estimation. IEEE Access 2020, 8, 133330–133348. [Google Scholar] [CrossRef] [Scilit]
- Zolfaghari, M.; Jourabloo, A.; Ghareh Gozlou, S.; Pedrood, B.; Manzuri-Shalmani, M.T. 3D human pose estimation from image using couple sparse coding. Mach. Vis. Appl. 2014, 25, 1489–1499. [Google Scholar] [CrossRef] [Scilit]
- Iskakov, K.; Burkov, E.; Lempitsky, V.; Malkov, Y. Learnable triangulation of human pose. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 7718–7727. [Google Scholar]
- Qiu, H.; Wang, C.; Wang, J.; Wang, N.; Zeng, W. Cross view fusion for 3d human pose estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 4342–4351. [Google Scholar]
- Huang, F.; Zeng, A.; Liu, M.; Lai, Q.; Xu, Q. Deepfuse: An imu-aware network for real-time 3d human pose estimation from multi-view image. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision; IEEE: New York, NY, USA, 2020; pp. 429–438. [Google Scholar]
- Chen, T.; Fang, C.; Shen, X.; Zhu, Y.; Chen, Z.; Luo, J. Anatomy-aware 3d human pose estimation with bone-based pose decomposition. IEEE Trans. Circuits Syst. Video Technol. 2021, 32, 198–209. [Google Scholar] [CrossRef] [Scilit]
- Moreno-Noguer, F. 3D Human Pose Estimation from a Single Image via Distance Matrix Regression. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2017; pp. 1561–1570. [Google Scholar] [CrossRef] [Scilit]
- Xu, T.; An, D.; Wang, Z.; Jiang, S.; Meng, C.; Zhang, Y.; Wang, Q.; Pan, Z.; Yue, Y. 3D Joints Estimation of the Human Body in Single-Frame Point Cloud. IEEE Access 2020, 8, 178900–178908. [Google Scholar] [CrossRef] [Scilit]
- Zhang, F.; Zhu, X.; Ye, M. Efficient Human Pose Estimation in Hierarchical Context. IEEE Access 2019, 7, 29365–29379. [Google Scholar] [CrossRef] [Scilit]
- Liao, Z.; Zhu, J.; Wang, C.; Hu, H.; Waslander, S.L. Multiple View Geometry Transformers for 3D Human Pose Estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 708–717. [Google Scholar]
- Zhang, C.; Li, S.; Wang, X.; Sheikh, Y. RePOSE: Refining 3D Human Pose Estimation with Recurrent Pose Reasoning. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Hao, Y.; Zhang, W.; Li, J.; Liu, Z. PersPose: 3D Human Pose Estimation with Perspective Encoding. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2025. [Google Scholar]
- Wang, Z.; Liu, M.; Chen, H.; Li, H. ActionPose: 3D Human Pose Estimation with Action-aware Constraints. arXiv 2024, arXiv:2409.00449. [Google Scholar]
- Kolotouros, N.; Pavlakos, G.; Black, M.J.; Daniilidis, K. Learning to reconstruct 3D human pose and shape via model-fitting in the loop. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 2252–2261. [Google Scholar]
- Wu, H.; Xiao, B. 3D human pose estimation via explicit compositional depth maps. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Palo Alto, CA, USA, 2020; Volume 34, pp. 12378–12385. [Google Scholar]
- Pham, H.H.; Salmane, H.; Khoudour, L.; Crouzil, A.; Velastin, S.A.; Zegers, P. A unified deep framework for joint 3d pose estimation and action recognition from a single rgb camera. Sensors 2020, 20, 1825. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cheng, Y.; Yang, B.; Wang, B.; Yan, W.; Tan, R.T. Occlusion-aware networks for 3D human pose estimation in video. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 723–732. [Google Scholar]
- Sun, X.; Xiao, B.; Wei, F.; Liang, S.; Wei, Y. Integral Human Pose Regression. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 529–545. [Google Scholar]
- Huang, J.; Wang, C.; Li, W. Geometry-aware human pose estimation under occlusion using feature refinement. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2024. [Google Scholar]
- Li, H.; Zhang, W.; Chen, M. Temporal optimization for robust 3D human pose estimation under nonlinear motion. In Proceedings of the European Conference on Computer Vision (ECCV), Milan, Italy, 29 September–4 October 2024. [Google Scholar]
- Mei, J.; Chen, X.; Wang, C.; Yuille, A.L.; Lan, X.; Zeng, W. Learning to Refine 3D Human Pose Sequences. In Proceedings of the International Conference on 3D Vision (3DV); IEEE: New York, NY, USA, 2019; pp. 358–366. [Google Scholar] [CrossRef] [Scilit]
- Garau, N.; Martinelli, G.; Bisagno, N.; Tome, D.; Stoll, C. EPOCH: Jointly Estimating the 3D Pose of Cameras and Humans. In Proceedings of the Computer Vision—ECCV 2024 Workshops, Part XIII; Springer: Berlin/Heidelberg, Germany, 2024; pp. 1–18. [Google Scholar] [CrossRef] [Scilit]
- Arnab, A.; Doersch, C.; Zisserman, A. Exploiting Temporal Context for 3D Human Pose Estimation in the Wild. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2019; pp. 3390–3399. [Google Scholar] [CrossRef] [Scilit]
- Shuai, Q.; Geng, C.; Fang, Q.; Peng, S.; Shen, W.; Zhou, X.; Bao, H. Novel view synthesis of human interactions from sparse multi-view videos. In Proceedings of the ACM SIGGRAPH 2022 Conference Proceedings; Association for Computing Machinery: New York, NY, USA, 2022; pp. 1–10. [Google Scholar]
- Choutas, V.; Pavlakos, G.; Bolkart, T.; Tzionas, D.; Black, M.J. Monocular Expressive Body Regression through Body-Driven Attention. In Proceedings of the European Conference on Computer Vision (ECCV); Springer: Berlin/Heidelberg, Germany, 2020; pp. 20–40. [Google Scholar] [CrossRef] [Scilit]
- Kim, K.J.; Park, B.S.; Kim, J.K.; Kim, D.W.; Seo, Y.H. Holographic augmented reality based on three-dimensional volumetric imaging for a photorealistic scene. Opt. Express 2020, 28, 35972–35985. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kim, K.J.; Park, B.S.; Kim, D.W.; Kwon, S.C.; Seo, Y.H. Real-time 3D Volumetric Model Generation using Multiview RGB-D Camera. J. Broadcast Eng. 2020, 25, 439–448. [Google Scholar]
- Nocedal, J.; Wright, S.J. Numerical Optimization, 2nd ed.; Springer: New York, NY, USA, 2006. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Xu, H. Demixed Sparse Principal Component Analysis Through Hybrid Structural Regularizers. IEEE Access 2021, 9, 103075–103090. [Google Scholar] [CrossRef] [Scilit]
- Cao, Z.; Simon, T.; Wei, S.E.; Sheikh, Y. Realtime multi-person 2d pose estimation using part affinity fields. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2017; pp. 7291–7299. [Google Scholar]
- He, Y.; Tan, H.; Luo, W.; Mao, H.; Ma, D.; Feng, S.; Fan, J. Mr-dbscan: An efficient parallel density-based clustering algorithm using mapreduce. In Proceedings of the 2011 IEEE 17th International Conference on Parallel and Distributed Systems; IEEE: New York, NY, USA, 2011; pp. 473–480. [Google Scholar]
- Free3D. Free3D, Rigged Model. 2021. Available online: https://free3d.com/3d-models/ (accessed on 18 May 2021).
- Luvizon, D. Machine Learning for Human Action Recognition and Pose Estimation Based on 3D Information. Ph.D. Thesis, Cergy Paris Universite, Cergy-Pontoise, France, 2019. [Google Scholar]
- Zhang, J.; Joo, H.; Liu, Y.S.; Ramakrishna, V.; Sheikh, Y. EgoHumans: An Egocentric 3D Multi-Human Benchmark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023. [Google Scholar]
- Mehta, D.; Sridhar, S.; Sotnychenko, O.; Rhodin, H.; Shafiei, M.; Seidel, H.P.; Xu, W.; Casas, D. Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision. In Proceedings of the 3DV (International Conference on 3D Vision); IEEE: New York, NY, USA, 2017. [Google Scholar]
- Mehta, D.; Sotnychenko, O.; Mueller, F.; Xu, W.; Sridhar, S.; Pons-Moll, G.; Theobalt, C. Single-shot multi-person 3D pose estimation from monocular RGB. In Proceedings of the 2018 International Conference on 3D Vision (3DV); IEEE: New York, NY, USA, 2018. [Google Scholar]
- Noitom. Perception Neuron. 2021. Available online: https://neuronmocap.com/products/perception_neuron (accessed on 21 November 2021).
- Microsoft. Azure Kinect Body Tracking Joints. 2021. Available online: https://learn.microsoft.com/en-us/azure/kinect-dk/body-joints (accessed on 21 November 2021).
- Girshick, R.; Donahue, J.; Darrell, T.; Malik, J. Rich feature hierarchies for accurate object detection and semantic segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2014; pp. 580–587. [Google Scholar]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2014; pp. 740–755. [Google Scholar]




















| Model | Method | Value (mm) | Enhanced Ratio (%) | ||
|---|---|---|---|---|---|
| Mean | SD | Mean | SD | ||
| Eric | Openpose | 33.215 | 28.325 | 100.00% | 100.00% |
| Proposed | 1.929 | 2.217 | 5.81% | 7.83% | |
| Sophia | Openpose | 21.015 | 27.626 | 100.00% | 100.00% |
| Proposed | 0.916 | 1.627 | 4.36% | 5.89% | |
| Method | MPJPE (mm) ↓ | Avg. MPJPE (mm) ↓ | |
|---|---|---|---|
| Sophia | Eric | ||
| EasyMocap [24] | 108.13 | 117.61 | 112.87 |
| Ours | 23.23 | 33.56 | 28.40 |
| Dataset | Mean (mm) | SD (mm) |
|---|---|---|
| EgoHumans [34] | 33.403 | 8.226 |
| MPI-INF-3DHP [35] | 69.730 | 18.533 |
| Panoptic [36] | 30.736 | 3.803 |
| Number of View Point | MPJPE (mm) |
|---|---|
| 1 | 15.59 |
| 2 | 12.78 |
| 3 | 10.47 |
| 4 | 10.46 |
| Index | Joint | Sol | Sunjong |
|---|---|---|---|
| 0 | Head | 2.980 | 3.435 |
| 1 | Neck | 2.345 | 3.361 |
| 2 | Right Shoulder | 4.330 | 5.098 |
| 3 | Right Elbow | 3.976 | 0.644 |
| 4 | Right Wrist | 3.318 | 4.694 |
| 5 | Left Shoulder | 2.380 | 5.185 |
| 6 | Left Elbow | 2.825 | 1.125 |
| 7 | Left Wrist | 1.932 | 6.392 |
| 8 | Pelvis | 3.007 | 5.196 |
| 9 | Right Hip | 4.751 | 1.310 |
| 10 | Right Knee | 1.999 | 2.565 |
| 11 | Right Ankle | 2.315 | 1.880 |
| 12 | Left Hip | 4.625 | 3.649 |
| 13 | Left Knee | 3.436 | 2.889 |
| 14 | Left Ankle | 2.441 | 4.195 |
| Average | 3.111 | 3.441 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kim, J.-W.; Lee, S.; Park, B.-S.; Lee, H.-B.; Kang, D.-H.; Seo, Y.-H. 3D Pose Estimation Using Virtual Projection Based on 3D Reconstructed Model. Sensors 2026, 26, 3302. https://doi.org/10.3390/s26113302
Kim J-W, Lee S, Park B-S, Lee H-B, Kang D-H, Seo Y-H. 3D Pose Estimation Using Virtual Projection Based on 3D Reconstructed Model. Sensors. 2026; 26(11):3302. https://doi.org/10.3390/s26113302
Chicago/Turabian StyleKim, Jung-Woo, Sol Lee, Byung-Seo Park, Hak-Bum Lee, Dong-Ho Kang, and Young-Ho Seo. 2026. "3D Pose Estimation Using Virtual Projection Based on 3D Reconstructed Model" Sensors 26, no. 11: 3302. https://doi.org/10.3390/s26113302
APA StyleKim, J.-W., Lee, S., Park, B.-S., Lee, H.-B., Kang, D.-H., & Seo, Y.-H. (2026). 3D Pose Estimation Using Virtual Projection Based on 3D Reconstructed Model. Sensors, 26(11), 3302. https://doi.org/10.3390/s26113302

