ODARRL: Obstacle- and Disturbance-Aware End-to-End Residual Reinforcement Learning for Underwater Robot Trajectory Tracking with Obstacle Avoidance
Abstract
1. Introduction
- This paper develops an end-to-end trajectory tracking controller under full six-degree-of-freedom dynamics, actuator dynamics, and practical control constraints. By reducing the gap between theoretical validation and practical execution caused by the simplified numerical models widely used in existing studies, the proposed setting improves the physical fidelity of the learned policy and its potential for real-world application.
- This paper proposes a three-stage curriculum learning framework with policy initialization guided by non-human expert data. In this framework, MPC-guided imitation learning provides a feasible initial policy, while reinforcement learning further improves the initialized policy beyond expert demonstrations. This design greatly reduces the difficulty of obtaining expert datasets, improves the feasibility of data collection, and enhances training efficiency and stability.
- This paper designs a Dual-Horizon Attention Disturbance Encoder (DADE) to learn control-relevant disturbance representations from long-horizon and short-horizon observation histories. This module improves disturbance-aware decision making without explicitly estimating flow-field variables.
2. Background
2.1. Mathematical Model of ROV
2.2. Problem Formulation in the Framework of RL
- Observation O: Denotes the observation space available to the agent at each time step. For the ROV trajectory execution task considered in this work, the observation consists of the current motion state of the robot, reference-trajectory-related information, local environmental observations, and temporal context. This definition allows the policy to make decisions based not only on the current state but also on the complex interaction of underwater disturbances and obstacle constraints.
- Action A: Denotes the action space available to the agent. Unlike methods that output only high-level control commands, such as desired velocity or desired attitude, the action space in this work is directly defined as thruster-level continuous control commands, thereby forming an ST end-to-end control scheme.
- Transition probability P: Denotes the state transition probability, which represents the probability distribution over transitions from the current state to the next state after the execution of an action. In the task considered here, state transitions are jointly affected by robot dynamics, current disturbances, and obstacle constraints, and therefore exhibit significant uncertainty and environment dependence.
- Reward function r: Denotes the reward function, which evaluates the immediate return of the agent after taking an action under a given observation. The reward design in this work jointly considers trajectory tracking accuracy, motion safety, control smoothness, and effective progress along the reference trajectory, so as to guide the policy toward safe and robust trajectory execution.
- Discount factor : Denotes the discount factor, which determines the relative importance of future rewards with respect to immediate rewards and controls the contribution of long-term return in policy optimization.
3. Methodology
3.1. MPC-Guided Policy Initialization via Imitation Learning
3.2. Dual-Horizon Attention Disturbance Encoder
3.3. Disturbance-Aware Residual Reinforcement Learning
| Algorithm 1 Disturbance-Aware Residual Reinforcement Learning (DARRL) |
|
3.4. Training Strategy and Network Optimization
4. Simulation, Results and Analysis
4.1. Simulation Platform and Experimental Setup
4.2. Stage-I MPC-Guided Policy Initialization
4.3. Stage-II Trajectory Execution Under Current Disturbances
4.3.1. Comparison with Baseline Methods
4.3.2. Generalization Under Time-Varying Current Disturbances
4.3.3. Component Ablation of DARRL
4.3.4. Sensitivity Analysis of the DADE Architecture
4.4. Stage-III Trajectory Execution Under Current Disturbances and Obstacles
4.4.1. Evaluation Protocol and Path-Progress Metric
4.4.2. Ablation of Stage-Wise Curriculum Initialization
4.5. Discussion
4.5.1. Mechanistic Interpretation of ODARRL
4.5.2. Sim-to-Real Considerations and Limitations
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Aguirre-Castro, O.A.; Inzunza-González, E.; García-Guerrero, E.E.; Tlelo-Cuautle, E.; López-Bonilla, O.R.; Olguín-Tiznado, J.E.; Cárdenas-Valdez, J.R. Design and Construction of an ROV for Underwater Exploration. Sensors 2019, 19, 5387. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khalid, O.; Hao, G.; Desmond, C.; Macdonald, H.; McAuliffe, F.D.; Dooly, G.; Hu, W. Applications of robotics in floating offshore wind farm operations and maintenance: Literature review and trends. Wind Energy 2022, 25, 1880–1899. [Google Scholar] [CrossRef] [Scilit]
- Teague, J.; Allen, M.J.; Scott, T.B. The potential of low-cost ROV for use in deep-sea mineral, ore prospecting and monitoring. Ocean Eng. 2018, 147, 333–339. [Google Scholar] [CrossRef] [Scilit]
- Bingul, Z.; Gul, K. Intelligent-PID with PD feedforward trajectory tracking control of an autonomous underwater vehicle. Machines 2023, 11, 300. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Liu, X.; Luo, M.; Yang, C. MPC-based 3-D trajectory tracking for an autonomous underwater vehicle with constraints in complex ocean environments. Ocean Eng. 2019, 189, 106309. [Google Scholar] [CrossRef] [Scilit]
- Yan, Z.; Wang, M.; Xu, J. Robust adaptive sliding mode control of underactuated autonomous underwater vehicles with uncertain dynamics. Ocean Eng. 2019, 173, 802–809. [Google Scholar] [CrossRef] [Scilit]
- Tijjani, A.S.; Chemori, A.; Creuze, V. A survey on tracking control of unmanned underwater vehicles: Experiments-based approach. Annu. Rev. Control 2022, 54, 125–147. [Google Scholar] [CrossRef] [Scilit]
- Hernández-Alvarado, R.; García-Valdovinos, L.G.; Salgado-Jiménez, T.; Gómez-Espinosa, A.; Fonseca-Navarro, F. Neural network-based self-tuning PID control for underwater vehicles. Sensors 2016, 16, 1429. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liang, J.; Huang, W.; Zhou, F.; Liang, J.; Lin, G.; Xiao, E.; Li, H.; Zhang, X. Double-loop PID-type neural network sliding mode control of an uncertain autonomous underwater vehicle model based on a nonlinear high-order observer with unknown disturbance. Mathematics 2022, 10, 3332. [Google Scholar] [CrossRef] [Scilit]
- Liu, T.; Zhao, J.; Huang, J.; Li, Z.; Xu, L.; Zhao, B. Research on model predictive control of autonomous underwater vehicle based on physics informed neural network modeling. Ocean Eng. 2024, 304, 117844. [Google Scholar] [CrossRef] [Scilit]
- Eski, İ.; Yildirim, S. Design of neural network control system for controlling trajectory of autonomous underwater vehicles. Int. J. Adv. Robot. Syst. 2014, 11, 7. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yu, T.; Zhang, Q.; Liu, T. Reinforcement learning approaches in the motion systems of autonomous underwater vehicles. Appl. Ocean Res. 2025, 161, 104682. [Google Scholar] [CrossRef] [Scilit]
- Mirza, K.Z.; Singh, S. Imitation learning for legged robot locomotion: A survey. Front. Robot. AI 2025, 12, 1678567. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vítor, G.d.A.; Melo, D.C.; Maximo, M.R.; Afonso, R.J. Imitation learning of a model predictive controller for real-time humanoid robot walking. Eng. Appl. Artif. Intell. 2025, 143, 109919. [Google Scholar] [CrossRef] [Scilit]
- Zhou, J.; Mei, J.; Zhao, F.; Chen, J.; Li, S. Online Motion Planning for Quadrotor Multi-Point Navigation Using Efficient Imitation Learning-Based Strategy. In Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2025; pp. 5386–5393. [Google Scholar]
- Chrysomallis, I.; Chalkiadakis, G. Imitation Learning in the Deep Learning Era: A Novel Taxonomy and Recent Advances. arXiv 2025, arXiv:2511.03565. [Google Scholar]
- Tang, C.; Abbatematteo, B.; Hu, J.; Chandra, R.; Martín-Martín, R.; Stone, P. Deep reinforcement learning for robotics: A survey of real-world successes. Annu. Rev. Control Robot. Auton. Syst. 2025, 8, 153–188. [Google Scholar] [CrossRef] [Scilit]
- Kaufmann, E.; Bauersfeld, L.; Loquercio, A.; Müller, M.; Koltun, V.; Scaramuzza, D. Champion-level drone racing using deep reinforcement learning. Nature 2023, 620, 982–987. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hadi, B.; Khosravi, A.; Sarhadi, P. Deep reinforcement learning for adaptive path planning and control of an autonomous underwater vehicle. Appl. Ocean Res. 2022, 129, 103326. [Google Scholar] [CrossRef] [Scilit]
- Deowan, M.E.; Yousha, M.S.Y.; Hossain, T.M.; Hassan, S.; Marxer, R. Optimizing Underwater Robot Navigation: A Study of DRL Algorithms and Multi-Modal Sensor Fusion. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA); IEEE: New York, NY, USA, 2025; pp. 11270–11277. [Google Scholar]
- Wu, D.; Feng, Z.; Hou, D.; Liu, R.; Yin, Y. DRL-based path planning and obstacle avoidance of autonomous underwater vehicle. In Proceedings of the 2023 IEEE International Conference on Mechatronics and Automation (ICMA); IEEE: New York, NY, USA, 2023; pp. 948–953. [Google Scholar]
- Xing, J.; Romero, A.; Bauersfeld, L.; Scaramuzza, D. Bootstrapping reinforcement learning with imitation for vision-based agile flight. arXiv 2024, arXiv:2403.12203. [Google Scholar]
- Lu, Y.; Fu, J.; Tucker, G.; Pan, X.; Bronstein, E.; Roelofs, R.; Sapp, B.; White, B.; Faust, A.; Whiteson, S.; et al. Imitation is not enough: Robustifying imitation with reinforcement learning for challenging driving scenarios. In Proceedings of the 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2023; pp. 7553–7560. [Google Scholar]
- Zhang, Y.; Zhang, T.; Li, Y.; Zhuang, Y.; Wang, D. A novel reward-shaping-based soft actor–critic for random trajectory tracking of AUVs. Ocean Eng. 2025, 322, 120505. [Google Scholar] [CrossRef] [Scilit]
- Niu, S.; Pan, X.; Wang, J.; Li, G. Deep reinforcement learning from human preferences for ROV path tracking. Ocean Eng. 2025, 317, 120036. [Google Scholar] [CrossRef] [Scilit]
- Lyu, X.; Sun, Y.; Wang, L.; Tan, J.; Zhang, L. End-to-end AUV local motion planning method based on deep reinforcement learning. J. Mar. Sci. Eng. 2023, 11, 1796. [Google Scholar] [CrossRef] [Scilit]
- Tong, R.; Feng, Y.; Wang, J.; Wu, Z.; Tan, M.; Yu, J. A survey on reinforcement learning methods in bionic underwater robots. Biomimetics 2023, 8, 168. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Gao, W.; Zhang, T.; Li, Y.; Zhuang, Y.; Zhang, Y. FRRS-SAC: A deep reinforcement learning framework for three-dimensional AUV docking under ocean disturbances. Ocean Eng. 2026, 357, 125409. [Google Scholar] [CrossRef] [Scilit]
- Gunnarson, P.; Mandralis, I.; Novati, G.; Koumoutsakos, P.; Dabiri, J.O. Learning efficient navigation in vortical flow fields. Nat. Commun. 2021, 12, 7143. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yao, Q.; Meng, L.; Zhang, Q.; Zhao, J.; Pajarinen, J.; Wang, X.; Li, Z.; Wang, C. Learning-based propulsion control for amphibious quadruped robots with dynamic adaptation to changing environment. IEEE Robot. Autom. Lett. 2023, 8, 7889–7896. [Google Scholar] [CrossRef] [Scilit]
- Chu, S.; Feng, H.; Ou, Y.; Lin, M.; Li, D. FINDER: Flow-aware intelligent navigation through distilled experience and reinforcement learning for UUVs in complex flow fields. Ocean Eng. 2026, 343, 123188. [Google Scholar] [CrossRef] [Scilit]
- Chu, S.; Huang, Z.; Li, Y.; Lin, M.; Li, D.; Carlucho, I.; Petillot, Y.R.; Yang, C. MarineGym: A high-performance reinforcement learning platform for underwater robotics. In Proceedings of the 2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: New York, NY, USA, 2025; pp. 17146–17153. [Google Scholar]
- Fossen, T.I. Nonlinear Modelling and Control of Underwater Vehicles. Ph.D. Thesis, Norwegian Institute of Technology, Trondheim, Norway, 1991. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Chang, S.R.; Huh, U.Y. Curvature-continuous 3d path-planning using qpmi method. Int. J. Adv. Robot. Syst. 2015, 12, 76. [Google Scholar] [CrossRef] [Scilit] [PubMed]











| Parameter | Configuration |
|---|---|
| Short-horizon window | 10 steps |
| Long-horizon window | 50 steps |
| Input embedding dimension | 64 |
| Attention heads | 4 |
| Feed-forward dimension | 128 |
| Dropout | 0.1 |
| Pooling | Mean pooling |
| Fusion MLP | 128-64-64 |
| Disturbance feature dimension | 64 |
| Symbol | Description | Value |
|---|---|---|
| Reference speed | 0.85 m/s | |
| Path length | 35.0 m | |
| Reference window length | 20 | |
| Path look-ahead distance | 1.5 m | |
| Ocean current speed range | m/s | |
| Ocean current direction | Random | |
| Planar residual coefficient | 0.1 | |
| Vertical residual coefficient | 0.05 | |
| Number of parallel environments | 64 | |
| T | Rollout steps | 1024 |
| K | PPO epochs | 4 |
| Number of minibatches | 64 | |
| Actor learning rate | ||
| Critic learning rate | ||
| Discount factor | 0.99 | |
| GAE parameter | 0.95 | |
| PPO clip parameter | 0.15 | |
| Entropy coefficient | ||
| Progress reward weight | 15.0 | |
| XY error penalty weight | 2.0 | |
| Yaw error penalty weight | 0.8 | |
| Depth error penalty weight | 0.1 |
| Symbol | Description | Value |
|---|---|---|
| Number of obstacles | 2 | |
| Relative obstacle positions along path | ||
| Obstacle detection range | 5.0 m | |
| Obstacle radius | 0.45 m | |
| Target obstacle-surface clearance | 0.8 m | |
| Pass-decision clearance | 0.6 m | |
| Additional pass-progress margin | 0.5 m | |
| Pass-progress offset | 1.75 m | |
| Recovery duration | 100 steps | |
| Early recovery-error threshold | 0.35 m | |
| XY-tracking scale in AVOIDING | 0.15 | |
| Yaw-tracking scale in AVOIDING | 0.10 | |
| Additional avoidance-progress weight | 80.0 | |
| Clearance-growth reward weight | 6.0 | |
| Safety-shortfall penalty weight | 6.0 | |
| Safety-zone entry penalty weight | 4.0 | |
| Collision penalty weight | 50.0 |
| Controller | Horizontal Mean Error | Vertical Mean Error | Mean 3D Error | 3D RMSE |
|---|---|---|---|---|
| DARRL | 0.146 | 0.068 | 0.175 | 0.312 |
| MPC-IL policy | 0.514 | 0.175 | 0.570 | 0.894 |
| PPO | 0.228 | 0.074 | 0.257 | 0.436 |
| A2C | 0.574 | 0.202 | 0.649 | 0.903 |
| SAC | 0.277 | 0.113 | 0.323 | 0.582 |
| VNRS-SAC | 0.216 | 0.095 | 0.236 | 0.391 |
| Method | Residual Learning | DADE | Horizontal Mean Error | Vertical Mean Error | Mean 3D Error | 3D RMSE |
|---|---|---|---|---|---|---|
| PPO | No | No | 0.228 | 0.074 | 0.257 | 0.436 |
| Residual PPO w/o DADE | Yes | No | 0.213 | 0.071 | 0.224 | 0.388 |
| DARRL | Yes | Yes | 0.146 | 0.068 | 0.175 | 0.312 |
| Configuration | Short Horizon | Long Horizon | 3D RMSE (m) |
|---|---|---|---|
| Short-horizon only | 10 | – | 0.405 |
| Long-horizon only | – | 50 | 0.347 |
| Dual horizon | 10 | 50 | 0.312 |
| Configuration | Attention Heads | Dimension per Head | 3D RMSE (m) |
|---|---|---|---|
| Dual-H2 | 2 | 32 | |
| Dual-H4 | 4 | 16 | |
| Dual-H8 | 8 | 8 |
| Initialization | Budget | Terminated by Yaw Deviation (%) | Passed at Least One Obstacle (%) | Reached Goal (%) |
|---|---|---|---|---|
| From scratch | 100k | 36 | 0 | 0 |
| From policy | 100k | 14 | 26 | 0 |
| From policy | 100k | 0 | 72 | 0 |
| From scratch | 200k | 28 | 40 | 0 |
| From policy | 200k | 4 | 68 | 2 |
| From policy | 200k | 0 | 92 | 10 |
| From scratch | 300k | 26 | 36 | 0 |
| From policy | 300k | 0 | 74 | 16 |
| From policy | 300k | 0 | 100 | 72 |
| From scratch | 400k | 10 | 68 | 6 |
| From policy | 400k | 0 | 100 | 46 |
| From policy | 400k | 0 | 100 | 94 |
| From scratch | 500k | 10 | 82 | 10 |
| From policy | 500k | 0 | 100 | 74 |
| From policy | 500k | 0 | 100 | 92 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Meng, L.; Huang, Z.; Yao, Q.; Zhang, Y.; Zhang, Q. ODARRL: Obstacle- and Disturbance-Aware End-to-End Residual Reinforcement Learning for Underwater Robot Trajectory Tracking with Obstacle Avoidance. J. Mar. Sci. Eng. 2026, 14, 1501. https://doi.org/10.3390/jmse14161501
Meng L, Huang Z, Yao Q, Zhang Y, Zhang Q. ODARRL: Obstacle- and Disturbance-Aware End-to-End Residual Reinforcement Learning for Underwater Robot Trajectory Tracking with Obstacle Avoidance. Journal of Marine Science and Engineering. 2026; 14(16):1501. https://doi.org/10.3390/jmse14161501
Chicago/Turabian StyleMeng, Linghan, Zebin Huang, Qingfeng Yao, Yunxiu Zhang, and Qifeng Zhang. 2026. "ODARRL: Obstacle- and Disturbance-Aware End-to-End Residual Reinforcement Learning for Underwater Robot Trajectory Tracking with Obstacle Avoidance" Journal of Marine Science and Engineering 14, no. 16: 1501. https://doi.org/10.3390/jmse14161501
APA StyleMeng, L., Huang, Z., Yao, Q., Zhang, Y., & Zhang, Q. (2026). ODARRL: Obstacle- and Disturbance-Aware End-to-End Residual Reinforcement Learning for Underwater Robot Trajectory Tracking with Obstacle Avoidance. Journal of Marine Science and Engineering, 14(16), 1501. https://doi.org/10.3390/jmse14161501

