Reinforcement Learning-Enhanced Adaptive NMPC for Safe Autonomous Driving
Abstract
1. Introduction
2. Problem Formulation
2.1. Vehicle Dynamic Model
2.2. NMPC Controller for Trajectory Tracking
- s.t.
3. Construction of PPO+NMPC+CBF Framework
3.1. PPO Model Training
| Algorithm 1: Offline PPO Training Framework. |
| , reference path environment, NMPC solver. |
| . Initialize networks and empty rollout buffer ; while training budget not exhausted do ; while episode not finish do ) diagonals; NMPC solves horizon problem; apply first control; ; ) in ; end Estimate GAE and value targets from ; via clipped surrogate with entropy; by regression to value targets; end . |
3.2. PPO+NMPC
3.3. PPO+NMPC+CBF
4. Experiments
4.1. Simulation Setup
4.2. Simulation Results and Analysis
4.3. Robustness Evaluation
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Appendix A. Derivation of Vehicle Model
Appendix B. Training Dataset



References
- Marino, R.; Scalzi, S.; Orlando, G.; Netto, M. A nested PID steering control for lane keeping in vision-based autonomous vehicles. In Proceedings of the American Control Conference (ACC), St. Louis, MO, USA, 10–12 June 2009; pp. 2885–2890. [Google Scholar]
- Du, X.; Tan, K.K.; Htet, K.K.K. Vision approach towards fully self-reverse parking system. In Proceedings of the IEEE International Conference on Mechatronics and Automation, Tianjin, China, 3–6 August 2014; pp. 186–191. [Google Scholar]
- Du, X.; Tan, K.K. Autonomous reverse parking system based on robust path generation and improved sliding mode control. IEEE Trans. Intell. Transp. Syst. 2015, 16, 1225–1237. [Google Scholar] [CrossRef]
- Chen, X.; Bao, Q.; Zhang, B. Research on 4WIS electric vehicle path tracking control based on adaptive fuzzy PID algorithm. In Proceedings of the Chinese Control Conference (CCC), Guangzhou, China, 27–30 July 2019; pp. 6753–6760. [Google Scholar]
- Yeh, Y.-C.; Li, T.-H.S.; Chen, C.-Y. Adaptive fuzzy sliding-mode control of dynamic model-based car-like mobile robot. Int. J. Fuzzy Syst. 2009, 11, 272–281. [Google Scholar]
- Limón, D.; Ferramosca, A.; Alvarado, I.; Alamo, T. Nonlinear MPC for tracking piece-wise constant reference signals. IEEE Trans. Autom. Control 2018, 63, 3735–3750. [Google Scholar] [CrossRef]
- Zhang, C.; Chu, D.; Liu, S.; Deng, Z.; Wu, C.; Su, X. Trajectory planning and tracking for autonomous vehicle based on state lattice and model predictive control. IEEE Intell. Transp. Syst. Mag. 2019, 11, 29–40. [Google Scholar] [CrossRef]
- Wang, H.; Liu, B.; Ping, X.; An, Q. Path tracking control for autonomous vehicles based on an improved MPC. IEEE Access 2019, 7, 161064–161073. [Google Scholar] [CrossRef]
- Stano, P.; Montanaro, U.; Tavernini, D.; Tufo, M.; Fiengo, G.; Novella, L.; Sorniotti, A. Model predictive path tracking control for automated road vehicles: A review. Annu. Rev. Control 2023, 55, 194–236. [Google Scholar] [CrossRef]
- Köhler, J.; Müller, M.A.; Allgöwer, F. A nonlinear tracking model predictive control scheme for dynamic target signals. Automatica 2020, 118, 109030. [Google Scholar] [CrossRef]
- Yuan, H.; Sun, X.; Gordon, T. Unified decision-making and control for highway collision avoidance using active front steer and individual wheel torque control. Veh. Syst. Dyn. 2019, 57, 1188–1205. [Google Scholar] [CrossRef]
- Zhu, G.; Jie, H.; Hong, W. NMPC-based path tracking control strategy for 4WID autonomous vehicle considering handling stability under extreme conditions. In Proceedings of the 7th CAA International Conference on Vehicular Control and Intelligence (CVCI), Changsha, China, 27–29 October 2023; pp. 1–6. [Google Scholar]
- Du, X.; Htet, K.K.K.; Tan, K.K. Development of a genetic-algorithm-based nonlinear model predictive control scheme on velocity and steering of autonomous vehicles. IEEE Trans. Ind. Electron. 2016, 63, 6970–6977. [Google Scholar] [CrossRef]
- Le, V.; Malikopoulos, A. Controller adaptation via learning solutions of contextual Bayesian optimization. IEEE Robot. Autom. Lett. 2025, 10, 8308–8315. [Google Scholar] [CrossRef]
- Alcalá, E.; Puig, V.; Quevedo, J.; Rosolia, U. Autonomous racing using Linear Parameter Varying-Model Predictive Control (LPV-MPC). Control Eng. Pract. 2020, 95, 104270. [Google Scholar] [CrossRef]
- Zarrouki, B.; Spanakakis, M.; Betz, J. A safe reinforcement learning driven weights-varying model predictive control for autonomous vehicle motion control. In Proceedings of the IEEE Intelligent Vehicles Symposium (IV), Jeju Island, Republic of Korea, 2–5 June 2024; pp. 1401–1408. [Google Scholar] [CrossRef]
- Ostafew, C.J.; Schoellig, A.P.; Barfoot, T.D. Robust constrained learning-based NMPC enabling reliable mobile robot path tracking. Int. J. Robot. Res. 2016, 35, 1547–1563. [Google Scholar] [CrossRef]
- Stella, L.; Themelis, A.; Sopasakis, P.; Patrinos, P. A simple and efficient algorithm for nonlinear model predictive control. In Proceedings of the 56th IEEE Conference on Decision and Control (CDC), Melbourne, Australia, 12–15 December 2017; pp. 1939–1944. [Google Scholar]
- Diehl, M.; Ferreau, H.J.; Haverbeke, N. Efficient numerical methods for nonlinear MPC and moving horizon estimation. In Nonlinear Model Predictive Control: Towards New Challenging Applications; Magni, L., Raimondo, D.M., Allgöwer, F., Eds.; Springer: Berlin/Heidelberg, Germany, 2009; pp. 391–417. [Google Scholar]
- Kayacan, E.; Saeys, W.; Ramon, H.; Belta, C.; Peschel, J.M. Experimental validation of linear and nonlinear MPC on an articulated unmanned ground vehicle. IEEE/ASME Trans. Mechatron. 2018, 23, 2023–2030. [Google Scholar] [CrossRef]
- Goodwin, G.C.; Cea, M.G.; Seron, M.M.; Ferris, D.; Middleton, R.H.; Campos, B. Opportunities and challenges in the application of nonlinear MPC to industrial problems. In Proceedings of the FAC Nonlinear Model Predictive Control Conference, Noordwijkerhout, The Netherlands, 23–27 August 2012; pp. 39–49. [Google Scholar]
- Pane, Y.P.; Nageshrao, S.P.; Babuška, R. Actor-critic reinforcement learning for tracking control in robotics. In Proceedings of the IEEE 55th Conference on Decision and Control (CDC), Las Vegas, NV, USA, 12–14 December 2016; pp. 5819–5826. [Google Scholar]
- Kosta, A.K.; Anwar, M.A.; Panda, P.; Raychowdhury, A.; Roy, K. RAPID-RL: A reconfigurable architecture with preemptive-exits for efficient deep-reinforcement learning. In Proceedings of the International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 23–27 May 2022; pp. 7492–7498. [Google Scholar]
- Radac, M.; Chirla, D. Near real-time online reinforcement learning with synchronous or asynchronous updates. Sci. Rep. 2025, 15, 2025. [Google Scholar] [CrossRef] [PubMed]
- Shan, Y.; Zheng, B.; Chen, L.; Chen, L.; Chen, D. A reinforcement learning-based adaptive path tracking approach for autonomous driving. IEEE Trans. Veh. Technol. 2020, 69, 10581–10595. [Google Scholar] [CrossRef]
- Sierra-Garcia, J.E.; Santos, M. Combining reinforcement learning and conventional control to improve automatic guided vehicles tracking of complex trajectories. Expert Syst. 2023, 41, e13076. [Google Scholar] [CrossRef]
- Bellegarda, G.; Byl, K. An online training method for augmenting MPC with deep reinforcement learning. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 24 October 2020–24 January 2021; pp. 5453–5459. [Google Scholar]
- Chen, Z.; Lai, J.; Li, P.; Awad, O.I.; Zhu, Y. Prediction horizon-varying model predictive control (MPC) for autonomous vehicle control. Electronics 2024, 13, 1442. [Google Scholar] [CrossRef]
- Rokonuzzaman, M.; Mohajer, N.; Nahavandi, S.; Mohamed, S. Model predictive control with learned vehicle dynamics for autonomous vehicle path tracking. IEEE Access 2021, 9, 128233–128249. [Google Scholar] [CrossRef]
- Wang, R.; Li, H.; Liang, B.; Shi, Y.; Xu, D. Policy learning for nonlinear model predictive control with application to USVs. IEEE Trans. Ind. Electron. 2022, 71, 4089–4097. [Google Scholar] [CrossRef]
- Reiter, R.; Ghezzi, A.; Baumgärtner, K.; Hoffmann, J.; McAllister, R.; Diehl, M. AC4MPC: Actor-critic reinforcement learning for guiding model predictive control. IEEE Trans. Control Syst. Technol. 2026, 34, 395–410. [Google Scholar] [CrossRef]
- Martinsen, A.B.; Lekkas, A.M.; Gros, S. Reinforcement learning-based NMPC for tracking control of ASVs: Theory and experiments. Control Eng. Pract. 2022, 120, 105024. [Google Scholar] [CrossRef]
- Berg, H.S.; Menges, D.; Tengesdal, T.; Rasheed, A. Digital twin syncing for autonomous surface vessels using reinforcement learning and nonlinear model predictive control. Sci. Rep. 2025, 15, 9344. [Google Scholar] [CrossRef]
- Mehndiratta, M.; Camci, E.; Kayacan, E. Automated tuning of nonlinear model predictive controller by reinforcement learning. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; pp. 3016–3021. [Google Scholar]
- Ceusters, G.; Putratama, M.; Franke, R.; Nowé, A.; Messagie, M. An adaptive safety layer with hard constraints for safe reinforcement learning in multi-energy management systems. Sustain. Energy Grids Netw. 2023, 36, 101202. [Google Scholar] [CrossRef]
- Malu, S.K.; Majumdar, J. Kinematics, localization and control of differential drive mobile robot. Glob. J. Res. Eng. 2014, 14, 1–7. [Google Scholar]
- Yang, H.; Deng, F.; He, Y.; Jiao, D.; Han, Z. Robust nonlinear model predictive control for reference tracking of dynamic positioning ships based on nonlinear disturbance observer. Ocean Eng. 2020, 215, 107885. [Google Scholar] [CrossRef]
- Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; Moritz, P. Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 6–11 July 2015; Volume 37, pp. 1889–1897. [Google Scholar]
- Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.I.; Abbeel, P. High-dimensional continuous control using generalized advantage estimation. In Proceedings of the 4th International Conference on Learning Representations (ICLR), San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar]
- Devlin, S.; Kudenko, D. Dynamic potential-based reward shaping. In Proceedings of the 11th International Conference on Autonomous Agents and Multiagent Systems, Valencia, Spain, 4–8 June 2012; pp. 433–440. [Google Scholar]
- Ames, A.D.; Coogan, S.; Egerstedt, M.; Notomista, G.; Sreenath, K.; Tabuada, P. Control barrier functions: Theory and applications. In Proceedings of the 18th European Control Conference (ECC), Naples, Italy, 25–28 June 2019; pp. 3420–3431. [Google Scholar]
- Ames, A.D.; Xu, X.; Grizzle, J.W.; Tabuada, P. Control barrier function based quadratic programs for safety critical systems. IEEE Trans. Autom. Control 2017, 62, 3861–3876. [Google Scholar] [CrossRef]
- Suwartadi, E.; Kungurtsev, V.; Jäschke, J. Sensitivity-Based Economic NMPC with a Path-Following Approach. Processes 2017, 5, 8. [Google Scholar] [CrossRef]
- Zhang, H.; Li, P.; García, C.E. Robust stability of nonlinear model predictive control based on extended Kalman filter. J. Process Control 2012, 22, 82–89. [Google Scholar]









| Parameter | Value |
|---|---|
| Vehicle mass | |
| Inertia | |
| Wheel Radius | |
| Wheelbase | |
| Rolling Friction coef. | |
| Velocity Damping coef. | |
| Angular Damping coef. | |
| NMPC Horizon | |
| Time Step | |
| Max Velocity | |
| Max Torque |
| Metric | NMPC | PPO+NMPC | NMPC+CBF | PPO+NMPC+CBF |
|---|---|---|---|---|
| Steps | 240 | |||
| Mean position error | 0.086 | |||
| Path length | 15.48 | |||
| Total energy | 49.72 | |||
| Energy efficiency | 0.311 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jin, S.; Loh, J.Y.Y. Reinforcement Learning-Enhanced Adaptive NMPC for Safe Autonomous Driving. Electronics 2026, 15, 2577. https://doi.org/10.3390/electronics15122577
Jin S, Loh JYY. Reinforcement Learning-Enhanced Adaptive NMPC for Safe Autonomous Driving. Electronics. 2026; 15(12):2577. https://doi.org/10.3390/electronics15122577
Chicago/Turabian StyleJin, Sheng, and Joel Yi Yang Loh. 2026. "Reinforcement Learning-Enhanced Adaptive NMPC for Safe Autonomous Driving" Electronics 15, no. 12: 2577. https://doi.org/10.3390/electronics15122577
APA StyleJin, S., & Loh, J. Y. Y. (2026). Reinforcement Learning-Enhanced Adaptive NMPC for Safe Autonomous Driving. Electronics, 15(12), 2577. https://doi.org/10.3390/electronics15122577

