LiDAR-Aided Human–Machine Shared Control Optimization for Unknown Complex Environments via Model Predictive Control and Deep Reinforcement Learning
Abstract
1. Introduction
- A HMSC framework with adaptive control authority allocation is proposed. The system dynamically adjusts fusion weights based on the agent’s confidence level regarding the current environmental state. The proposed HMSC method was implemented and validated, with its performance and the rationality of the confidence model verified through both local and global control tasks.
- This paper designed and implemented an online safety correction mechanism based on parallel multi-trajectory prediction and constraint projection. MPC is employed to generate multiple candidate trajectories for online safety evaluation of the policy. When potential collision risks are detected, a constrained optimization problem is solved to project the control action onto the nearest feasible region on the safe manifold, thereby preserving the human or DRL intent to the greatest extent possible.
- A quantitative evaluation model for jointly optimized machine decision confidence is proposed. By performing randomized forward passes of the policy network during inference, an action sample distribution is obtained, from which statistical measures are used to construct a decision confidence metric. The confidence is further fused with geometric feasibility information derived from LiDAR measurements to obtain a comprehensive composite confidence measure.
2. Problem Description and Related Work
2.1. Problem Description
2.2. Related Work
2.2.1. Human–Machine Shared Control
2.2.2. Model Predictive Control and Safety Filtering
2.2.3. Bayesian Neural Networks
2.2.4. Twin Delayed Deep Deterministic Policy Gradient
3. Method
| Algorithm 1: Human–Machine Shared Control |
![]() |
4. Experiments
4.1. Experimental Setup
4.2. Local Safety Capability Verification: Obstacle Avoidance Task
4.3. Global Capability Verification: Goal-Driven Navigation Task
4.4. Analysis of the Rationality of the Confidence Model
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Bernardo, R.; Sousa, J.M.; Gonçalves, P.J. Survey on robotic systems for internal logistics. J. Manuf. Syst. 2022, 65, 339–350. [Google Scholar] [CrossRef] [Scilit]
- Koung, D.; Kermorgant, O.; Fantoni, I.; Belouaer, L. Cooperative multi-robot object transportation system based on hierarchical quadratic programming. IEEE Robot. Autom. Lett. 2021, 6, 6466–6472. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Wang, S.; Xie, Y.; Xiong, T.; Wu, M. A review of sensing technologies for indoor autonomous mobile robots. Sensors 2024, 24, 1222. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hart, P.E.; Nilsson, N.J.; Raphael, B. A formal basis for the heuristic determination of minimum cost paths. IEEE Trans. Syst. Sci. Cybern. 1968, 4, 100–107. [Google Scholar] [CrossRef] [Scilit]
- Dijkstra, E.W. A note on two problems in connexion with graphs. In Edsger Wybe Dijkstra: His life, Work, and Legacy; Association for Computing Machinery: New York, NY, USA, 2022; pp. 287–290. [Google Scholar]
- Orthey, A.; Chamzas, C.; Kavraki, L.E. Sampling-based motion planning: A comparative review. Annu. Rev. Control. Robot. Auton. Syst. 2023, 7, 285–310. [Google Scholar] [CrossRef] [Scilit]
- Lee, D.H.; Lee, S.S.; Ahn, C.K.; Shi, P.; Lim, C.C. Finite distribution estimation-based dynamic window approach to reliable obstacle avoidance of mobile robot. IEEE Trans. Ind. Electron. 2020, 68, 9998–10006. [Google Scholar] [CrossRef] [Scilit]
- Yang, W.; Wu, P.; Zhou, X.; Lv, H.; Liu, X.; Zhang, G.; Hou, Z.; Wang, W. Improved artificial potential field and dynamic window method for amphibious robot fish path planning. Appl. Sci. 2021, 11, 2114. [Google Scholar] [CrossRef] [Scilit]
- Nascimento, T.P.; Dórea, C.E.; Gonçalves, L.M.G. Nonholonomic mobile robots’ trajectory tracking model predictive control: A survey. Robotica 2018, 36, 676–696. [Google Scholar] [CrossRef] [Scilit]
- Engelsman, D.; Klein, I. Information-aided inertial navigation: A review. IEEE Trans. Instrum. Meas. 2023, 72, 1–18. [Google Scholar] [CrossRef] [Scilit]
- Luo, Y.; Lu, F.; Guo, C.; Liu, J. Matrix lie group-based extended Kalman filtering for inertial-integrated navigation in the navigation frame. IEEE Trans. Instrum. Meas. 2023, 73, 1–16. [Google Scholar] [CrossRef] [Scilit]
- Wu, Z.; Yue, Y.; Wen, M.; Zhang, J.; Yi, J.; Wang, D. Infrastructure-free hierarchical mobile robot global localization in repetitive environments. IEEE Trans. Instrum. Meas. 2021, 70, 1–12. [Google Scholar] [CrossRef] [Scilit]
- Cimurs, R.; Suh, I.H.; Lee, J.H. Goal-driven autonomous exploration through deep reinforcement learning. IEEE Robot. Autom. Lett. 2021, 7, 730–737. [Google Scholar] [CrossRef] [Scilit]
- Silver, D.; Huang, A.; Maddison, C.J.; Guez, A.; Sifre, L.; Van Den Driessche, G.; Schrittwieser, J.; Antonoglou, I.; Panneershelvam, V.; Lanctot, M.; et al. Mastering the game of Go with deep neural networks and tree search. Nature 2016, 529, 484–489. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kang, Y.; Di, J.; Li, M.; Zhao, Y.; Wang, Y. Autonomous multi-drone racing method based on deep reinforcement learning. Sci. China Inf. Sci. 2024, 67, 180203. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. arXiv 2015, arXiv:1509.02971. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; Proceedings of Machine Learning Research; pp. 1861–1870. [Google Scholar]
- Chiang, H.T.L.; Faust, A.; Fiser, M.; Francis, A. Learning navigation behaviors end-to-end with autorl. IEEE Robot. Autom. Lett. 2019, 4, 2007–2014. [Google Scholar] [CrossRef] [Scilit]
- Linial, O.; Tennenholtz, G.; Shalit, U. Benchmarks for Reinforcement Learning with Biased Offline Data and Imperfect Simulators. arXiv 2024, arXiv:2407.00806. [Google Scholar]
- Wagenmaker, A.; Huang, K.; Ke, L.; Jamieson, K.; Gupta, A. Overcoming the Sim-to-Real Gap: Leveraging Simulation to Learn to Explore for Real-World RL. Adv. Neural Inf. Process. Syst. 2024, 37, 78715–78765. [Google Scholar] [CrossRef] [Scilit]
- Losey, D.P.; McDonald, C.G.; Battaglia, E.; O’Malley, M.K. A review of intent detection, arbitration, and communication aspects of shared control for physical human–robot interaction. Appl. Mech. Rev. 2018, 70, 010804. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Wang, X.; Zheng, X.; Jin, J.; Huang, Y.; Zhang, J.J.; Wang, F.Y. SADRL: Merging human experience with machine intelligence via supervised assisted deep reinforcement learning. Neurocomputing 2022, 467, 300–309. [Google Scholar] [CrossRef] [Scilit]
- Li, X.; Wang, Y.; Su, C.; Gong, X.; Huang, J.; Yang, D. Adaptive authority allocation approach for shared steering control system. IEEE Trans. Intell. Transp. Syst. 2022, 23, 19428–19439. [Google Scholar] [CrossRef] [Scilit]
- Marcano, M.; Díaz, S.; Pérez, J.; Irigoyen, E. A review of shared control for automated vehicles: Theory and applications. IEEE Trans. Hum. Mach. Syst. 2020, 50, 475–491. [Google Scholar] [CrossRef] [Scilit]
- Ames, A.D.; Coogan, S.; Egerstedt, M.; Notomista, G.; Sreenath, K.; Tabuada, P. Control barrier functions: Theory and applications. In Proceedings of the 2019 18th European Control Conference (ECC), Naples, Italy, 25–28 June 2019; IEEE: New York, NY, USA, 2019; pp. 3420–3431. [Google Scholar]
- Loquercio, A.; Segu, M.; Scaramuzza, D. A general framework for uncertainty estimation in deep learning. IEEE Robot. Autom. Lett. 2020, 5, 3153–3160. [Google Scholar] [CrossRef] [Scilit]
- Abbink, D.A.; Mulder, M.; Boer, E.R. Haptic shared control: Smoothly shifting control authority? Cogn. Technol. Work. 2012, 14, 19–28. [Google Scholar] [CrossRef] [Scilit]
- Flad, M.; Fröhlich, L.; Hohmann, S. Cooperative shared control driver assistance systems based on motion primitives and differential games. IEEE Trans.-Hum.-Mach. Syst. 2017, 47, 711–722. [Google Scholar] [CrossRef] [Scilit]
- Warnell, G.; Waytowich, N.; Lawhern, V.; Stone, P. Deep tamer: Interactive agent shaping in high-dimensional state spaces. In Proceedings of the AAAI Conference on Artificial Intelligence, New Orleans, LA, USA, 2–7 February 2018; Volume 32. [Google Scholar]
- Mandel, T.; Liu, Y.E.; Brunskill, E.; Popović, Z. Where to add actions in human-in-the-loop reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA, 4–9 February 2017; Volume 31. [Google Scholar]
- Wu, J.; Huang, Z.; Huang, C.; Hu, Z.; Hang, P.; Xing, Y.; Lv, C. Human-in-the-loop deep reinforcement learning with application to autonomous driving. arXiv 2021, arXiv:2104.07246. [Google Scholar]
- Saeidi, H.; Opfermann, J.D.; Kam, M.; Raghunathan, S.; Léonard, S.; Krieger, A. A confidence-based shared control strategy for the smart tissue autonomous robot (STAR). In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; IEEE: New York, NY, USA, 2018; pp. 1268–1275. [Google Scholar]
- Amirshirzad, N.; Kumru, A.; Oztop, E. Human adaptation to human–robot shared control. IEEE Trans. -Hum.-Mach. Syst. 2019, 49, 126–136. [Google Scholar] [CrossRef] [Scilit]
- Zeng, J.; Zhang, B.; Sreenath, K. Safety-critical model predictive control with discrete-time control barrier function. In Proceedings of the 2021 American Control Conference (ACC), Virtual, 25–28 May 2021; IEEE: New York, NY, USA, 2021; pp. 3882–3889. [Google Scholar]
- Wabersich, K.P.; Zeilinger, M.N. A predictive safety filter for learning-based control of constrained nonlinear dynamical systems. Automatica 2021, 129, 109597. [Google Scholar] [CrossRef] [Scilit]
- Ames, A.D.; Xu, X.; Grizzle, J.W.; Tabuada, P. Control barrier function based quadratic programs for safety critical systems. IEEE Trans. Autom. Control 2016, 62, 3861–3876. [Google Scholar] [CrossRef] [Scilit]
- Fisac, J.F.; Akametalu, A.K.; Zeilinger, M.N.; Kaynama, S.; Gillula, J.; Tomlin, C.J. A general safety framework for learning-based control in uncertain robotic systems. IEEE Trans. Autom. Control 2018, 64, 2737–2752. [Google Scholar] [CrossRef] [Scilit]
- Gal, Y. Uncertainty in Deep Learning. Ph.D. Thesis, University of Oxford, Oxford, UK, 2016. [Google Scholar]
- Gal, Y.; Ghahramani, Z. Dropout as a bayesian approximation: Representing model uncertainty in deep learning. In Proceedings of the International Conference on Machine Learning, New York, NY, USA, 19–24 June 2016; Proceedings of Machine Learning Research. pp. 1050–1059. [Google Scholar]
- Kendall, A.; Gal, Y. What uncertainties do we need in bayesian deep learning for computer vision? Adv. Neural Inf. Process. Syst. 2017, 30, 5580–5590. [Google Scholar]
- Fujimoto, S.; Hoof, H.; Meger, D. Addressing function approximation error in actor-critic methods. In Proceedings of the International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; Proceedings of Machine Learning Research. pp. 1587–1596. [Google Scholar]














| Parameter | Symbol | Value |
|---|---|---|
| Robot Kinematics | ||
| Max. Linear Velocity | 1.0 m/s | |
| Max. Angular Velocity | 1.0 rad/s | |
| Robot Radius | 0.35 m | |
| Safety Buffer | 0.1 m | |
| Goal Reward | 100 | |
| Collision Penalty | ||
| MPC Safety Layer | ||
| Prediction Horizon | T | 1 |
| Time Step | 0.1 s | |
| Horizon Steps | N | 10 |
| Human–Machine Shared Control | ||
| MC-Dropout Samples | M | 20 |
| Curvature Factor | 2 | |
| Action Dimension | d | 2 |
| Confidence Threshold | 0.15 | |
| Mapping Slope | 10 | |
| Mapping Offset | b | 3 |
| Method | Success Rate (%) | Collision Rate (%) | Cumulative Reward |
|---|---|---|---|
| HMSC | |||
| MPC-TD3 | |||
| TD3 |
| Component | Mean (ms) | Std. (ms) | Max. (ms) |
|---|---|---|---|
| TD3 policy inference | 4.697 | 6.307 | 48.029 |
| MPC safety layer | 25.361 | 11.884 | 76.038 |
| MC-Dropout () | 22.993 | 26.917 | 180.166 |
| Confidence and control fusion | 0.345 | 1.819 | 35.242 |
| Complete HMSC cycle | 53.476 | 30.882 | 218.874 |
| Method | Training Episodes | Total Training Time (h) | Average Time/Episode (s) |
|---|---|---|---|
| HMSC | 4000 | 29.21 | 26.29 |
| MPC-TD3 | 4000 | 26.87 | 24.18 |
| TD3 | 4000 | 17.58 | 15.82 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Cheng, Z.; Zhang, Q.; Li, Z. LiDAR-Aided Human–Machine Shared Control Optimization for Unknown Complex Environments via Model Predictive Control and Deep Reinforcement Learning. Machines 2026, 14, 1081. https://doi.org/10.3390/machines14091081
Cheng Z, Zhang Q, Li Z. LiDAR-Aided Human–Machine Shared Control Optimization for Unknown Complex Environments via Model Predictive Control and Deep Reinforcement Learning. Machines. 2026; 14(9):1081. https://doi.org/10.3390/machines14091081
Chicago/Turabian StyleCheng, Zhiao, Qianqian Zhang, and Zerui Li. 2026. "LiDAR-Aided Human–Machine Shared Control Optimization for Unknown Complex Environments via Model Predictive Control and Deep Reinforcement Learning" Machines 14, no. 9: 1081. https://doi.org/10.3390/machines14091081
APA StyleCheng, Z., Zhang, Q., & Li, Z. (2026). LiDAR-Aided Human–Machine Shared Control Optimization for Unknown Complex Environments via Model Predictive Control and Deep Reinforcement Learning. Machines, 14(9), 1081. https://doi.org/10.3390/machines14091081


