Lightweight LiDAR-Based 3D Human Pose Estimation via 2D Depth Images for Autonomous Driving
Abstract
1. Introduction
- We present a lightweight design for LiDAR-based 3D human pose estimation that integrates 2D pose estimation with depth-based 3D lifting to achieve efficient 3D joint reconstruction.
- We propose an algorithm to compensate for self-occlusion issues that arise during projection into 2D representations. This algorithm corrects 3D joint coordinates that are lost during depth image generation. It can be extended to compensate for limitations inherent to the observational characteristics of existing LiDAR sensors.
2. Methodology
2.1. Method Design
2.1.1. Overall Structure
2.1.2. Depth-Based 2D Pose Estimation
2.1.3. Depth-Based 3D Lifting
2.2. Self-Occlusion Correction Algorithm
| Algorithm 1: Self-occlusion correction algorithm |
|
2.2.1. Side Occlusion
2.2.2. Bending Occlusion
3. Experiments
3.1. Experimental Settings
- MPJPE: The average Euclidean distance between predicted joint coordinates and ground truth.
- PA-MPJPE: MPJPE that removed the effects of rotation and scale differences through rigid alignment.
- PCK0.5: The proportion of joints satisfying an error within 50% of the torsal length.
3.2. Performance and Efficiency Evaluation
3.3. Ablation Work
3.4. Failure Cases
4. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Schuetz, E.; Flohr, F.B. A Review of Trajectory Prediction Methods for the Vulnerable Road User. Robotics 2023, 13, 1. [Google Scholar] [CrossRef]
- Jiang, J.; Yan, K.; Xia, X.; Yang, B. A Survey of Deep Learning-Based Pedestrian Trajectory Prediction: Challenges and Solutions. Sensors 2025, 25, 957. [Google Scholar] [CrossRef] [PubMed]
- Kooij, J.F.P.; Schneider, N.; Flohr, F.; Gavrila, D.M. Context-Based Pedestrian Path Prediction. In Proceedings of the Computer Vision—ECCV 2014, Zurich, Switzerland, 6–12 September 2014; Springer International Publishing: Cham, Switzerland, 2014; pp. 618–633. [Google Scholar] [CrossRef]
- Kloukiniotis, A.; Papandreou, A.; Lalos, A.; Kapsalas, P.; Nguyen, D.-V.; Moustakas, K. Countering Adversarial Attacks on Autonomous Vehicles Using Denoising Techniques: A Review. IEEE Open J. Intell. Transp. Syst. 2022, 3, 61–80. [Google Scholar] [CrossRef]
- Fang, Z.; López, A.M. Is the Pedestrian Going to Cross? Answering by 2D Pose Estimation. In Proceedings of the 2018 IEEE Intelligent Vehicles Symposium (IV), Changshu, China, 26–30 June 2018; IEEE: Gothenburg, Sweden, 2018; pp. 1271–1276. [Google Scholar] [CrossRef]
- Zhang, S.; Abdel-Aty, M.; Wu, Y.; Zheng, O. Pedestrian Crossing Intention Prediction at Red-Light Using Pose Estimation. IEEE Trans. Intell. Transp. Syst. 2022, 23, 2331–2339. [Google Scholar] [CrossRef]
- Cadena, P.R.G.; Yang, M.; Qian, Y.; Wang, C. Pedestrian Graph: Pedestrian Crossing Prediction Based on 2D Pose Estimation and Graph Convolutional Networks. In Proceedings of the 2019 IEEE Intelligent Transportation Systems Conference (ITSC), Auckland, New Zealand, 27–30 October 2019; IEEE: Toronto, ON, Canada, 2019; pp. 2000–2005. [Google Scholar] [CrossRef]
- Li, N.; Pan, W.; Xu, B.; Liu, H.; Dai, S.; Xu, C. IHENet: An Illumination Invariant Hierarchical Feature Enhancement Network for Low-Light Object Detection. Multimed. Syst. 2025, 31, 407. [Google Scholar] [CrossRef]
- Zheng, Z.; Cheng, Y.; Xin, Z.; Yu, Z.; Zheng, B. Robust Perception under Adverse Conditions for Autonomous Driving Based on Data Augmentation. IEEE Trans. Intell. Transp. Syst. 2023, 24, 13916–13929. [Google Scholar] [CrossRef]
- Hasan, M.; Hanawa, J.; Goto, R.; Suzuki, R.; Fukuda, H.; Kuno, Y.; Kobayashi, Y. LiDAR-Based Detection, Tracking, and Property Estimation: A Contemporary Review. Neurocomputing 2022, 506, 393–405. [Google Scholar] [CrossRef]
- Ohno, M.; Ukyo, R.; Amano, T.; Rizk, H.; Yamaguchi, H. Privacy-Preserving Pedestrian Tracking Using Distributed 3D LiDARs. In Proceedings of the 2023 IEEE International Conference on Pervasive Computing and Communications (PerCom), Atlanta, GA, USA, 13–17 March 2023; IEEE: New York, NY, USA, 2023. [Google Scholar] [CrossRef]
- Li, J.; Zhang, J.; Wang, Z.; Shen, S.; Wen, C.; Ma, Y.; Xu, L.; Yu, J.; Wang, C. LiDARCap: Long-Range Marker-Less 3D Human Motion Capture with LiDAR Point Clouds. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; pp. 20502–20512. [Google Scholar] [CrossRef]
- Jang, D.-K.; Yang, D.; Jang, D.-Y.; Choi, B.; Jin, T.; Lee, S.-H. MOVIN: Real-time Motion Capture Using a Single LiDAR. Comput. Graph. Forum 2023, 42, e14961. [Google Scholar] [CrossRef]
- Zhang, J.; Mao, Q.; Shen, S.; Wen, C.; Xu, L.; Wang, C. LiDARCapV2: 3D Human Pose Estimation with Human–Object Interaction from LiDAR Point Clouds. Pattern Recognit. 2024, 156, 110848. [Google Scholar] [CrossRef]
- Zhang, J.; Mao, Q.; Hu, G.; Shen, S.; Wang, C. Neighborhood-Enhanced 3D Human Pose Estimation with Monocular LiDAR in Long-Range Outdoor Scenes. Proc. AAAI Conf. Artif. Intell. 2024, 38, 7169–7177. [Google Scholar] [CrossRef]
- Ren, Y.; Han, X.; Zhao, C.; Wang, J.; Xu, L.; Yu, J.; Ma, Y. LiveHPS: LiDAR-Based Scene-Level Human Pose and Shape Estimation in Free Environment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 1281–1291. [Google Scholar] [CrossRef]
- Qi, C.R.; Yi, L.; Su, H.; Guibas, L.J. PointNet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space. Adv. Neural Inf. Process. Syst. 2017, 30, 5099–5108. Available online: https://arxiv.org/abs/1706.02413 (accessed on 20 January 2026).
- Ye, D.; Xie, Y.; Chen, W.; Zhou, Z.; Ge, L.; Foroosh, H. LPFormer: LiDAR Pose Estimation Transformer with Multi-Task Network. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: New York, NY, USA, 2024; pp. 16432–16438. [Google Scholar] [CrossRef]
- Shi, J.; Wonka, P. VoxelKP: A Voxel-Based Network Architecture for Human Keypoint Estimation in LiDAR Data. arXiv 2023, arXiv:2312.08871. Available online: https://arxiv.org/abs/2312.08871 (accessed on 20 January 2026).
- Kovács, L.; Bódis, B.M.; Benedek, C. LidPose: Real-Time 3D Human Pose Estimation in Sparse Lidar Point Clouds with Non-Repetitive Circular Scanning Pattern. Sensors 2024, 24, 3427. [Google Scholar] [CrossRef] [PubMed]
- Milioto, A.; Vizzo, I.; Behley, J.; Stachniss, C. RangeNet ++: Fast and Accurate LiDAR Semantic Segmentation. In Proceedings of the 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Macau, China, 3–8 November 2019; IEEE: New York, NY, USA, 2019. [Google Scholar] [CrossRef]
- Ando, A.; Gidaris, S.; Bursuc, A.; Puy, G.; Boulch, A.; Marlet, R. RangeViT: Towards Vision Transformers for 3D Semantic Segmentation in Autonomous Driving. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 18–22 June 2023; IEEE: New York, NY, USA, 2023. [Google Scholar] [CrossRef]
- Chen, C.; Zhao, L.; Guo, W.; Yuan, X.; Tan, S.; Hu, J.; Yang, Z.; Wang, S.; Ge, W. FARVNet: A Fast and Accurate Range-View-Based Method for Semantic Segmentation of Point Clouds. Sensors 2025, 25, 2697. [Google Scholar] [CrossRef] [PubMed]
- Yang, J.; Lee, C.; Ahn, P.; Lee, H.; Yi, E.; Kim, J. PBP-Net: Point Projection and Back-Projection Network for 3D Point Cloud Segmentation. In Proceedings of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 25–29 October 2020; IEEE: New York, NY, USA, 2020. [Google Scholar] [CrossRef]
- Mirzaei, K.; Arashpour, M.; Asadi, E.; Masoumi, H.; Bai, Y.; Behnood, A. 3D Point Cloud Data Processing with Machine Learning for Construction and Infrastructure Applications: A Comprehensive Review. Adv. Eng. Inform. 2022, 51, 101501. [Google Scholar] [CrossRef]
- Kaushik, P.; Lohani, B.P.; Thakur, A.; Gupta, A.; Khan, A.K.; Kumar, A. Body Posture Detection and Comparison between OpenPose, MoveNet and PoseNet. In Proceedings of the 2023 6th International Conference on Contemporary Computing and Informatics (IC3I), Gautam Buddha Nagar, India, 14–16 September 2023; IEEE: New York, NY, USA, 2023. [Google Scholar] [CrossRef]
- Zhou, Q.-Y.; Park, J.; Koltun, V. Open3D: A Modern Library for 3D Data Processing. arXiv 2018, arXiv:1801.09847. Available online: https://arxiv.org/abs/1801.09847 (accessed on 20 January 2026).
- Ionescu, C.; Papava, D.; Olaru, V.; Sminchisescu, C. Human3.6M: Large Scale Datasets and Predictive Methods for 3D Human Sensing in Natural Environments. IEEE Trans. Pattern Anal. Mach. Intell. 2014, 36, 1325–1339. [Google Scholar] [CrossRef] [PubMed]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An Imperative Style, High-Performance Deep Learning Library. Adv. Neural Inf. Process. Syst. 2019, 32, 8024–8035. Available online: https://arxiv.org/abs/1912.01703 (accessed on 20 January 2026).







| Model | MPJPE (mm) | PA-MPJPE (mm) | PCK0.5 | Parameters | FPS (GPU) |
|---|---|---|---|---|---|
| LiDARCap [12] | 79.31 | 66.72 | 95.00 | 34.93 M | 138.86 |
| LiDARCapV2 [14] | 73.21 | 63.42 | 96.02 | - | - |
| NE-LiDARCap [15] | 72.23 | 61.67 | 95.79 | - | - |
| LPFormer [18] | 95.72 | 79.03 | 94.87 | - | - |
| Ours | 138.1 | 108.5 | 89.6 | 1.9 M | 440 |
| Body Part | Joints | PA-MPJPE (mm) |
|---|---|---|
| Torso | hip, spine, thorax, rshoulder, lshoulder, rhip, lhip | 93.05 |
| Head | head, neck | 88.55 |
| Arms | relbow, rwrist, lelbow, lwrist | 158.28 |
| Legs | rknee, lknee, rfoot, lfoot | 120.05 |
| Distance Range (m) | Samples | MPJPE (mm) | PA-MPJPE (mm) |
|---|---|---|---|
| Near (12–17 m) | 8963 | 138.9 | 109.5 |
| Mid (17–22 m) | 11,872 | 134.7 | 106.4 |
| Far (22–28 m) | 3173 | 149.9 | 115.7 |
| Setting | Side | Bending | MPJPE (mm) | PA-MPJPE (mm) |
|---|---|---|---|---|
| Base | X | X | 152 | 122 |
| +Side | O | X | 140 | 109 |
| +Bending | X | O | 151 | 121 |
| +Side + Bending (Full) | O | O | 138 | 108 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kim, G.-Y.; Park, S.; Lee, S.; Seo, B.; Choi, S.-H.; Park, S.-M. Lightweight LiDAR-Based 3D Human Pose Estimation via 2D Depth Images for Autonomous Driving. Sensors 2026, 26, 1631. https://doi.org/10.3390/s26051631
Kim G-Y, Park S, Lee S, Seo B, Choi S-H, Park S-M. Lightweight LiDAR-Based 3D Human Pose Estimation via 2D Depth Images for Autonomous Driving. Sensors. 2026; 26(5):1631. https://doi.org/10.3390/s26051631
Chicago/Turabian StyleKim, Gyu-Yeon, Somi Park, Sunkyung Lee, Bobin Seo, Seon-Han Choi, and Sung-Min Park. 2026. "Lightweight LiDAR-Based 3D Human Pose Estimation via 2D Depth Images for Autonomous Driving" Sensors 26, no. 5: 1631. https://doi.org/10.3390/s26051631
APA StyleKim, G.-Y., Park, S., Lee, S., Seo, B., Choi, S.-H., & Park, S.-M. (2026). Lightweight LiDAR-Based 3D Human Pose Estimation via 2D Depth Images for Autonomous Driving. Sensors, 26(5), 1631. https://doi.org/10.3390/s26051631

