Autonomous Navigation of an Unmanned Underwater Vehicle via Safe Reinforcement Learning and Active Disturbance Rejection Control
Abstract
1. Introduction
2. Dynamic Model of the Unmanned Underwater Vehicle
2.1. Vehicle Kinematics
2.2. Vehicle Dynamics
2.3. Control Objective
3. Hierarchical Control Framework
3.1. Lower Layer: ADRC-Based 6-DOF Motion Control
3.1.1. Design of the Attitude Controller
3.1.2. Design of the Velocity Controller
3.2. Upper Layer: Safe Reinforcement Learning for Goal-Directed Navigation
3.2.1. Reinforcement Learning Controller
- (1)
- State and action definition
- (2)
- Policy learning algorithm
- (3)
- Network architecture
- (4)
- Reward shaping with task and safety terms
3.2.2. QP-Based Safety Filter with Control Barrier Functions
3.3. Stability and Safety Discussion
4. Simulation and Discussion
4.1. Validation of the Lower-Layer ADRC
4.2. Obstacle-Avoidance Navigation Using Safe Reinforcement Learning
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| UUV | Unmanned Underwater Vehicle |
| PID | Proportional–Integral–Derivative |
| SMC | Sliding Mode Control |
| ADRC | Active Disturbance Rejection Control |
| ESO | Extended State Observer |
| RL | Reinforcement Learning |
| TD3 | Twin Delayed Deep Deterministic Policy Gradient |
| CBF | Control Barrier Function |
| QP | Quadratic Programming |
| IAE | Integral of Absolute Error |
References
- Li, J.; Zhang, G.; Jiang, C.; Zhang, W. A Survey of Maritime Unmanned Search System: Theory, Applications and Future Directions. Ocean Eng. 2023, 285, 115359. [Google Scholar] [CrossRef] [Scilit]
- Pan, W.; Wang, Y.; Song, F.; Peng, L.; Zhang, X. UUV-assisted Icebreaking Application in Polar Environments using GA-SPSO. J. Mar. Sci. Eng. 2024, 12, 1845. [Google Scholar] [CrossRef] [Scilit]
- Liu, S.; Ma, C.; Juan, R. AUV Obstacle Avoidance Framework Based on Event-triggered Reinforcement Learning. Electronics 2024, 13, 2030. [Google Scholar] [CrossRef] [Scilit]
- Liu, H.; Qi, Z.; Yuan, J.; Tian, X. Robust Fixed-time H∞ Tracking Control of UUVs with Partial and Full State Constraints and Prescribed Performance under Input Saturation. Ocean Eng. 2023, 283, 115023. [Google Scholar] [CrossRef] [Scilit]
- Er, M.J.; Gong, H.; Liu, Y.; Liu, T. Intelligent Trajectory Tracking and Formation Control of Underactuated Autonomous Underwater Vehicles: A Critical Review. IEEE Trans. Syst. Man Cybern. Syst. 2024, 54, 543–555. [Google Scholar] [CrossRef] [Scilit]
- Antonelli, G. Underwater Robots, 3rd ed.; Springer International Publishing: Cham, Switzerland, 2014. [Google Scholar]
- Wang, Y.; Bao, H.; Guo, C.; Williams, G.; Li, Y.; Zhang, H. An SO(3)-Based Attitude Control With Saturation Constraints for Underactuated Underwater Vehicles. IEEE Robot. Autom. Lett. 2026, 11, 1402–1409. [Google Scholar] [CrossRef] [Scilit]
- Dong, B.; Lu, Y.; Xie, W.; Huang, L.; Chen, W.; Yang, Y. Robust Performance-Prescribed Attitude Control of Foldable Wave-Energy Powered AUV Using Optimized Backstepping Technique. IEEE Trans. Intell. Veh. 2023, 8, 1230–1240. [Google Scholar] [CrossRef] [Scilit]
- Tang, J.; Dang, Z.; Deng, Z.; Li, C. Adaptive Fuzzy Nonlinear Integral Sliding Mode Control for Unmanned Underwater Vehicles Based on ESO. Ocean Eng. 2022, 266, 113154. [Google Scholar] [CrossRef] [Scilit]
- Han, J. From PID to Active Disturbance Rejection Control. IEEE Trans. Ind. Electron. 2009, 56, 900–906. [Google Scholar] [CrossRef] [Scilit]
- She, J.; Miyamoto, K.; Han, Q.L.; Wu, M.; Hashimoto, H.; Wang, Q.G. Generalized-Extended-State-Observer and Equivalent-Input-Disturbance Methods for Active Disturbance Rejection: Deep Observation and Comparison. IEEE/CAA J. Autom. Sin. 2023, 10, 957–968. [Google Scholar] [CrossRef] [Scilit]
- Zhang, W.; Wu, W.; Li, Z.; Du, X.; Yan, Z. Three-Dimensional Trajectory Tracking of AUV Based on Nonsingular Terminal Sliding Mode and Active Disturbance Rejection Decoupling Control. J. Mar. Sci. Eng. 2023, 11, 959. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Liang, S.; Guo, M.; Zhang, H.; Wang, H.; Li, Z. ADRC-Based Underwater Navigation Control and Parameter Tuning of an Amphibious Multirotor Vehicle. IEEE J. Ocean. Eng. 2024, 49, 775–792. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Y.; Zhou, H.; Pan, X.; Jin, Y.; Tian, Z.; Zhao, Y. Optimized Line-of-Sight Active Disturbance Rejection Control for Depth Tracking of Hybrid Underwater Gliders in Disturbed Environments. J. Mar. Sci. Eng. 2025, 13, 1835. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Cai, W.; Lu, J.; Ding, X.; Yang, J. Design, Modeling, Control, and Experiments for Multiple AUVs Formation. IEEE Trans. Autom. Sci. Eng. 2022, 19, 2776–2787. [Google Scholar] [CrossRef] [Scilit]
- Liu, T.; Huang, J.; Zhao, J. Research on obstacle avoidance of underactuated autonomous underwater vehicle based on offline reinforcement learning. Robotica 2025, 43, 194–218. [Google Scholar] [CrossRef] [Scilit]
- Fan, Y.; Dong, H.; Zhao, X.; Denissenko, P. Path-Following Control of Unmanned Underwater Vehicle Based on an Improved TD3 Deep Reinforcement Learning. IEEE Trans. Control Syst. Technol. 2024, 32, 1904–1919. [Google Scholar] [CrossRef] [Scilit]
- Hadi, B.; Khosravi, A.; Sarhadi, P. Deep Reinforcement Learning for Adaptive Path Planning and Control of an Autonomous Underwater Vehicle. Appl. Ocean Res. 2022, 129, 103326. [Google Scholar] [CrossRef] [Scilit]
- Zhang, C.; Cheng, P.; Du, B.; Dong, B.; Zhang, W. AUV Path Tracking with Real-time Obstacle Avoidance via Reinforcement Learning under Adaptive Constraints. Ocean Eng. 2022, 256, 111453. [Google Scholar] [CrossRef] [Scilit]
- Gu, S.; Yang, L.; Du, Y.; Chen, G.; Walter, F.; Wang, J.; Knoll, A. A Review of Safe Reinforcement Learning: Methods, Theories, and Applications. IEEE Trans. Pattern Anal. Mach. Intell. 2024, 46, 11216–11235. [Google Scholar] [CrossRef] [Scilit]
- Emam, Y.; Notomista, G.; Glotfelter, P.; Kira, Z.; Egerstedt, M. Safe Reinforcement Learning Using Robust Control Barrier Functions. IEEE Robot. Autom. Lett. 2025, 10, 2886–2893. [Google Scholar] [CrossRef] [Scilit]
- Cao, F.; Xu, H.; Ru, J.; Li, Z.; Zhang, H.; Liu, H. Collision Avoidance of Multi-UUV Systems Based on Deep Reinforcement Learning in Complex Marine Environments. J. Mar. Sci. Eng. 2025, 13, 1615. [Google Scholar] [CrossRef] [Scilit]
- Brunke, L.; Greeff, M.; Hall, A.W.; Yuan, Z.; Zhou, S.; Panerati, J.; Schoellig, A.P. Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning. Annu. Rev. Control Robot. Auton. Syst. 2022, 5, 411–444. [Google Scholar] [CrossRef] [Scilit]
- Stooke, A.; Achiam, J.; Abbeel, P. Responsive Safety in Reinforcement Learning by PID Lagrangian Methods. In Proceedings of the 37th International Conference on Machine Learning (ICML 2020), Vienna, Austria, 12–18 July 2020. [Google Scholar]
- Chow, Y.; Nachum, O.; Duenez-Guzman, E.; Ghavamzadeh, M. A Lyapunov-Based Approach to Safe Reinforcement Learning. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, QC, Canada, 3–8 December 2018. [Google Scholar]
- Wabersich, K.P.; Zeilinger, M.N. A Predictive Safety Filter for Learning-Based Control of Constrained Nonlinear Dynamical Systems. Automatica 2021, 129, 109597. [Google Scholar] [CrossRef] [Scilit]
- Cheng, R.; Orosz, G.; Murray, R.M.; Burdick, J.W. End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI 2019), Honolulu, HI, USA, 27 January–1 February 2019. [Google Scholar]
- Zheng, Q.; Gao, L.Q.; Gao, Z. On Validation of Extended State Observer through Analysis and Experimentation. J. Dyn. Syst. Meas. Control 2012, 134, 024505. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Q.; Chen, Z.; Gao, Z. A Practical Approach to Disturbance Decoupling Control. Control Eng. Pract. 2009, 17, 1016–1025. [Google Scholar] [CrossRef] [Scilit]
- Gao, Z. Scaling and Bandwidth-parameterization based Controller Tuning. In Proceedings of the 2003 American Control Conference (ACC 2003), Denver, CO, USA, 4–6 June 2003. [Google Scholar]
- Jin, H.; Song, J.; Lan, W.; Gao, Z. On the characteristics of ADRC: A PID interpretation. Sci. Chin. Inf. Sci. 2020, 63, 209201. [Google Scholar] [CrossRef] [Scilit]









| Methods | Parameters for the Attitude Controller | Parameters for the Velocity Controller |
|---|---|---|
| ADRC | ||
| PID |
| Methods | ||||||
|---|---|---|---|---|---|---|
| ADRC | 0.30614 | 0.31274 | 0.30954 | 0.85030 | 0.23829 | 0.21255 |
| PID | 1.4627 | 1.5134 | 4.8503 | 0.86009 | 0.88165 | 0.38204 |
| MPC | 5.5458 | 1.4810 | 1.7914 | 1.6026 | 0.14381 | 0.14008 |
| Methods | Mean Runtime (s) | Standard Deviation (s) | Min Runtime (s) | Max Runtime (s) |
|---|---|---|---|---|
| ADRC | 5.2367 | 0.075477 | 5.1502 | 5.3374 |
| PID | 2.3978 | 0.19478 | 2.2104 | 2.6717 |
| MPC | 11.974 | 0.19002 | 11.682 | 12.164 |
| Sample Time | Episode Horizon | Experience Buffer Size | Mini-Batch Size | Discount Factor | Target Smooth Factor |
|---|---|---|---|---|---|
| 0.5 s | 40 s | 5 × 105 | 256 | 0.98 | 0.005 |
| Learning Frequency | Policy Update Frequency | Target Update Frequency | Optimizer | Critic Learning Rate | Actor Learning Rate |
| 2 | 4 | 2 | Adam | 0.0001 | 0.00003 |
| Algorithms | Settling Time (s) | Average IAE | Minimum Distance (m) |
|---|---|---|---|
| TD3 + QP + SR | 11.393 | 106.991 | 3.080 |
| TD3 + SR | 12.027 | 123.263 | 5.166 |
| TD3 + QP | 12.710 | 125.763 | 3.209 |
| TD3 | 11.556 | 104.102 | 3.903 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Chen, Q.; Cheng, Y.; Yuan, Y.; Hua, L. Autonomous Navigation of an Unmanned Underwater Vehicle via Safe Reinforcement Learning and Active Disturbance Rejection Control. J. Mar. Sci. Eng. 2026, 14, 425. https://doi.org/10.3390/jmse14050425
Chen Q, Cheng Y, Yuan Y, Hua L. Autonomous Navigation of an Unmanned Underwater Vehicle via Safe Reinforcement Learning and Active Disturbance Rejection Control. Journal of Marine Science and Engineering. 2026; 14(5):425. https://doi.org/10.3390/jmse14050425
Chicago/Turabian StyleChen, Qinze, Yun Cheng, Yinlong Yuan, and Liang Hua. 2026. "Autonomous Navigation of an Unmanned Underwater Vehicle via Safe Reinforcement Learning and Active Disturbance Rejection Control" Journal of Marine Science and Engineering 14, no. 5: 425. https://doi.org/10.3390/jmse14050425
APA StyleChen, Q., Cheng, Y., Yuan, Y., & Hua, L. (2026). Autonomous Navigation of an Unmanned Underwater Vehicle via Safe Reinforcement Learning and Active Disturbance Rejection Control. Journal of Marine Science and Engineering, 14(5), 425. https://doi.org/10.3390/jmse14050425

