Deep Reinforcement Learning for Dynamic Obstacle Avoidance of Mobile Robots in Indoor Environments: A Review
Highlights
- Systematically classifies DRL algorithms for indoor mobile robot dynamic obstacle avoidance.
- Analyzes core challenges and improved strategies across five key technical dimensions.
- Provides a unified reference framework for robot navigation and DRL algorithm selection.
- Clarifies promising future directions for indoor dynamic obstacle avoidance research.
Abstract
1. Introduction
2. The Theoretical Basis of Obstacle Avoidance in DRL Algorithms
3. The Classification and Scene Adaptability of the Basic Algorithm System of DRL
3.1. Value-Based DRL Algorithm
3.2. Policy-Based DRL Algorithm
3.2.1. TRPO Algorithm
3.2.2. PPO Algorithm
3.3. DRL Algorithm Based on Actor-Critic Architecture
3.3.1. DDPG Algorithm
3.3.2. A3C/A2C Algorithms
3.3.3. SAC Algorithm
3.3.4. TD3 Algorithm
3.4. Classification of Basic Algorithm Systems and Summary of Scenario Adaptability
4. Improved DRL Algorithm for Indoor Dynamic Obstacle Avoidance
4.1. Value Assessment Optimization
4.2. Enhancement of Spatio-Temporal Multimodal Perception
4.3. Layered Mixing and Expert Experience Control
4.4. Interactive Security Perception
4.5. Challenges and Solutions: Summary of Indoor Dynamic Obstacle Avoidance
5. Research Trends and Future Directions
5.1. Tight Coupling of Depth Perception and SLAM Module
5.2. Explainable Hierarchical Hybrid Intelligent Planning Architecture
5.3. Establish Standardized Interaction Security Testing Standards
5.4. Lifelong Learning and Cloud-Edge Collaboration for Continuous Adaptive Learning
5.5. Bridging the Sim-to-Real Gap for Robust Deployment
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| DRL | Deep Reinforcement Learning |
| AGVs | Automated Guided Vehicles |
| APF | Artificial Potential Field |
| DWA | Dynamic Window Approach |
| MPC | Model Predictive Control |
| TEB | Timed Elastic Band |
| VO | Velocity Obstacle |
| MDP | Markov Decision Process |
| POMDP | Partially Observable Markov Decision Process |
| MDPs | Markov Decision Processes |
| DQN | Deep Q-Network |
| CNNs | Convolutional Neural Networks |
| TD errors | Time Difference errors |
| TRPO | Trust Region Policy Optimization |
| PPO | Proximal Policy Optimization |
| DKL | Kullback–Leibler Divergence |
| DDPG | Deep Deterministic Policy Gradient |
| DPG | Deterministic Policy Gradient |
| OU noise | Ornstein-Uhlenbeck noise |
| A3C | Asynchronous Advantage Actor-Critic |
| A2C | Advantage Actor-Critic |
| SAC | Soft Actor-Critic |
| TD3 | Twin Delayed Deep Deterministic Policy Gradient |
| DDQN | Double DQN |
| PER | Priority Experience Replay |
| ICM | Intrinsic Curiosity Module |
| BiGRU | Bidirectional Gated Recurrent Units |
| SLAM | Simultaneous Localization and Mapping |
Appendix A
References
- Chen, L.; Jiang, Z.; Cheng, L.; Knoll, A.C.; Zhou, M. Deep Reinforcement Learning Based Trajectory Planning Under Uncertain Constraints. Front. Neurorobot. 2022, 16, 883562. [Google Scholar] [CrossRef]
- Sun, H.; Zhang, W.; Yu, R.; Zhang, Y. Motion Planning for Mobile Robots—Focusing on Deep Reinforcement Learning: A Systematic Review. IEEE Access 2021, 9, 69061–69081. [Google Scholar] [CrossRef]
- Farias, G.; Garcia, G.; Zamora, G.M.; Fabregas, E. Position control of a mobile robot using reinforcement learning. In Proceedings of the 21st International Federation of Automatic Control(IFAC) World Congress, Berlin, Germany, 11–17 July 2020; Available online: https://www.researchgate.net/publication/344351874_Position_control_of_a_mobile_robot_using_reinforcement_learning (accessed on 11 June 2026).
- Duan, C.; Junginger, S.; Huang, J.; Jin, K.; Thurow, K. Deep Learning for Visual SLAM in Transportation Robotics: A Review. Transp. Saf. Environ. 2019, 1, 177–184. [Google Scholar] [CrossRef]
- Sangiovanni, B.; Incremona, G.P.; Piastra, M.; Ferrara, A. Self-Configuring Robot Path Planning with Obstacle Avoidance via Deep Reinforcement Learning. IEEE Control Syst. Lett. 2021, 5, 397–402. [Google Scholar] [CrossRef]
- Khatib, O. Real-Time Obstacle Avoidance for Manipulators and Mobile Robots. Auton. Robot Veh. 1986, 1, 396–404. [Google Scholar] [CrossRef]
- Fox, D.; Burgard, W.; Thrun, S. The dynamic window approach to collision avoidance. IEEE Robot. Autom. Mag. 1997, 4, 23–33. [Google Scholar] [CrossRef]
- Falcone, P.; Borrelli, F.; Asgari, J.; Tseng, H.E.; Hrovat, D. Predictive Active Steering Control for Autonomous Vehicle Systems. IEEE Trans. Control Syst. Technol. 2007, 15, 566–580. [Google Scholar] [CrossRef]
- Roesmann, C.; Feiten, W.; Woesch, T.; Hoffmann, F.; Bertram, T. Trajectory modification considering dynamic constraints of autonomous robots. ROBOTIK 2012. In Proceedings of the 7th German Conference on Robotics, Munich, Germany, 21–22 May 2012; pp. 1–6. [Google Scholar]
- Fiorini, P.; Shiller, Z. Motion Planning in Dynamic Environments Using Velocity Obstacles. Int. J. Robot. Res. 1998, 17, 760–772. [Google Scholar] [CrossRef]
- Fan, T.; Long, P.; Liu, W.; Pan, J. Distributed multi-robot collision avoidance via deep reinforcement learning for navigation in complex scenarios. Int. J. Robot. Res. 2020, 39, 856–892. [Google Scholar] [CrossRef]
- Zhu, Y.; Hasan, W.Z.W.; Ramli, H.R.H.; Norsahperi, N.M.H.; Kassim, M.S.M.; Yao, Y. Deep Reinforcement Learning of Mobile Robot Navigation in Dynamic Environment: A Review. Sensors 2025, 25, 3394. [Google Scholar] [CrossRef] [PubMed]
- Bellman, R. A Markovian Decision Process. Ind. Univ. Math. J. 1957, 6, 679–684. [Google Scholar] [CrossRef]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [PubMed]
- Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. arXiv 2015, arXiv:1509.02971v2. [Google Scholar] [CrossRef]
- Parooei, M.; Masouleh, M.T.; Kalhor, A. MAP3F: A decentralized approach to multi-agent pathfinding and collision avoidance with scalable 1D, 2D and 3D feature fusion. Intell. Serv. Robot 2024, 17, 401–418. [Google Scholar] [CrossRef]
- Zhang, H.; Liu, L.; Xie, H.; Jiang, Y.; Zhou, J.; Wang, Y. Deep Learning-Based Robot Vision: High-End Tools for Smart Manufacturing. IEEE Instrum. Meas. Mag. 2022, 25, 27–35. [Google Scholar] [CrossRef]
- Kaelbling, L.P.; Littman, M.L.; Cassandra, A.R. Planning and acting in partially observable stochastic domains. Artif. Intell. 1998, 101, 99–134. [Google Scholar] [CrossRef]
- Sheng, S.; Yu, P.; Parker, D.; Kwiatkowska, M.; Feng, L. Safe POMDP Online Planning Among Dynamic Agents via Adaptive Conformal Prediction. IEEE Robot. Autom. Lett. 2024, 9, 9946–9953. [Google Scholar] [CrossRef]
- Sharma, G.; Jain, S.; Sharma, R.S. Path Planning for Fully Autonomous UAVs-A Taxonomic Review and Future Perspectives. IEEE Access 2025, 13, 13356–13379. [Google Scholar] [CrossRef]
- Nasti, S.M.; Chishti, M.A. A Review of AI-Enhanced Navigation Strategies for Mobile Robots in Dynamic Environments. In Proceedings of the 2024 ASU International Conference in Emerging Technologies for Sustainability and Intelligent Systems (ICETSIS), Manama, Bahrain, 28–29 January 2024; pp. 1239–1244. [Google Scholar]
- Li, C.; Wu, F.; Zhao, J. A Review of Deep Reinforcement Learning Exploration Methods: Prospects and Challenges for Application to Robot Attitude Control Tasks. In Cognitive Systems and Information Processing, Proceedings of the 7th International Conference, ICCSIP 2022, Fuzhou, China, 17–18 December 2022; Sun, F., Cangelosi, A., Zhang, J., Yu, Y., Liu, H., Fang, B., Eds.; Springer: Singapore, 2023; pp. 247–273. [Google Scholar]
- Zhao, Y.; Zhang, Y.; Wang, S. A Review of Mobile Robot Path Planning Based on Deep Reinforcement Learning Algorithm. J. Phys. Conf. Ser. 2021, 2138, 012011. [Google Scholar] [CrossRef]
- Zhu, K.; Zhang, T. Deep Reinforcement Learning Based Mobile Robot Navigation: A Review. Tsinghua Sci. Technol. 2021, 26, 674–691. [Google Scholar] [CrossRef]
- Jiang, H.; Wang, H.; Yau, W.Y.; Wan, K.W. A Brief Survey: Deep Reinforcement Learning in Mobile Robot Navigation. In Proceedings of the 15th IEEE Conference on Industrial Electronics and Applications(ICIEA), Kristiansand, Norway, 9–13 November 2020. [Google Scholar] [CrossRef]
- Fan, F.; Xu, G.; Feng, N.; Li, L.; Jiang, W.; Yu, L.; Xiong, X. Spatiotemporal path tracking via deep reinforcement learning of robot for manufacturing internal logistics. J. Manuf. Syst. 2023, 69, 150–169. [Google Scholar] [CrossRef]
- Cao, Y.; Ni, K.; Kawaguchi, T.; Hashimoto, S. Path Following for Autonomous Mobile Robots with Deep Reinforcement Learning. Sensors 2024, 24, 561. [Google Scholar] [CrossRef] [PubMed]
- Moustafa, E.Y.; Dusparic, I. Context-Aware Model-Based Reinforcement Learning for Autonomous Racing. In Proceedings of the IEEE International Conference on Advanced Robotics (ICAR), San Juan, Argentina, 2–5 December 2025. [Google Scholar] [CrossRef]
- Smith, D.K. Dynamic Programming and Optimal Control. Volume 1: Dynamic Programming and Optimal Control. Volume 2. J. Oper. Res. Soc. 1996, 47, 833–834. [Google Scholar] [CrossRef]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef]
- Paquet, S.; Tobin, L.; Chaib-draa, B. An online POMDP algorithm for complex multiagent environments. In Proceedings of the AAMAS ‘05-Fourth International Joint Conference on Autonomous Agents and Multiagent Systems, Utrecht, The Netherlands, 25–29 July 2005. [Google Scholar] [CrossRef]
- Mnih, V.; Badia, A.P.; Mirza, M.; Graves, A.; Lillicrap, T.P.; Harley, T.; Silver, D.; Kavukcuoglu, K. Asynchronous Methods for Deep Reinforcement Learning. In Proceedings of the 33rd International Conference on Machine Learning(ICML), New York, NY, USA, 20–22 June 2016. [Google Scholar] [CrossRef]
- Chen, J.; Zhang, Y.; Qian, T.; Wang, M. Optimal ancillary service disaggregation for EV charging station aggregators: A hybrid on–off policy reinforcement learning framework. Expert Syst. Appl. 2026, 316, 131763. [Google Scholar] [CrossRef]
- Arce, D.; Solano, J.; Beltrán, C. A Comparison Study between Traditional and Deep-Reinforcement-Learning-Based Algorithms for Indoor Autonomous Navigation in Dynamic Scenarios. Sensors 2023, 23, 9672. [Google Scholar] [CrossRef] [PubMed]
- Zheng, J.; Mao, S.; Wu, Z.; Kong, P.; Qiang, H. Improved Path Planning for Indoor Patrol Robot Based on Deep Reinforcement Learning. Symmetry 2022, 14, 132. [Google Scholar] [CrossRef]
- Schulman, J.; Levine, S.; Moritz, P.; Jordan, M.I.; Abbeel, P. Trust Region Policy Optimization. In Proceedings of the 32nd International Conference on Machine Learning(ICML), Lille, France, 7–9 July 2015. [Google Scholar] [CrossRef]
- Liu, H.; Shen, Y.; Yu, S.; Gao, Z.; Wu, T. Deep Reinforcement Learning for Mobile Robot Path Planning. J. Theory Pract. Eng. Sci. 2024, arXiv:2404.06974. [Google Scholar] [CrossRef]
- Xue, J.; He, M.; Chen, J.; Dong, B.; Zheng, Y. Improved DDPG based on enhancing decision evaluation for path planning in high-density environments. Expert Syst. Appl. 2025, 279, 127378. [Google Scholar] [CrossRef]
- Dobrevski, M.; Skočaj, D. Deep reinforcement learning for map-Less goal-driven robot navigation. Int. J. Adv. Robot. Syst. 2021. [Google Scholar] [CrossRef]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning(ICML), Stockholm, Sweden, 10–15 July 2018. [Google Scholar] [CrossRef]
- Shi, J.; Du, J.; Wang, J.; Wang, J.; Yuan, J. Priority-Aware Task Offloading in Vehicular Fog Computing Based on Deep Reinforcement Learning. IEEE Trans. Veh. Technol. 2020, 69, 16067–16081. [Google Scholar] [CrossRef]
- Fujimoto, S.; Van Hoof, H.; Meger, D. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning(ICML), Stockholm, Sweden, 10–15 July 2018. [Google Scholar] [CrossRef]
- Xiao, L.; Yu, G.; Zhou, B.; Zhang, J. Robot Path Planning Method Based on an Improved TD3 Algorithm. In Proceedings of the 2025 International Conference on Computational Intelligence and Robotics(CIR), Guangzhou, China, 12–14 September 2025. [Google Scholar] [CrossRef]
- Li, P.; Chen, D.; Wang, Y.; Zhang, L.; Zhao, S. Path planning of mobile robot based on improved TD3 algorithm in dynamic environment. Heliyon 2024, 10, e32167. [Google Scholar] [CrossRef] [PubMed]
- Kwon, R.; Kwon, G. Safety Constraint-Guided Reinforcement Learning with Linear Temporal Logic. Systems 2023, 11, 535. [Google Scholar] [CrossRef]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep Reinforcement Learning with Double Q-Learning. In Proceedings of the 30th AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016. [Google Scholar] [CrossRef]
- Wang, Y.; Li, X.; Wan, P.; Chang, L.; Deng, X. Dueling deep Q-networks for social awareness-aided spectrum sharing. Complex Intell. Syst. 2022, 8, 1975–1986. [Google Scholar] [CrossRef]
- Gu, Y.; Zhu, Z.; Lv, J.; Shi, L.; Hou, Z.; Xu, S. DM-DQN: Dueling Munchausen Deep Q Network for Robot Path Planning. Complex Intell. Syst. 2022, 9, 4287–4300. [Google Scholar] [CrossRef]
- Chen, C.; Yu, J.; Qian, S. An Enhanced Deep Q Network Algorithm for Localized Obstacle Avoidance in Indoor Robot Path Planning. Appl. Sci. 2024, 14, 11195. [Google Scholar] [CrossRef]
- Yin, Y.; Chen, Z.; Liu, G.; Guo, J. A Mapless Local Path Planning Approach Using Deep Reinforcement Learning Framework. Sensors 2023, 23, 2036. [Google Scholar] [CrossRef] [PubMed]
- Yang, X.; Wang, Q.; Li, J.; Jiang, X. NM-TD3: A Hybrid Noise-Driven TD3 Algorithm With Long-Term Reward Propagation for Mobile Robot Path Planning. IEEE Access 2025, 13, 149921–149932. [Google Scholar] [CrossRef]
- Tao, B.; Kim, J.H. Deep reinforcement learning-based local path planning in dynamic environments for mobile robot. J. King Saud Univ. Comput. Inf. Sci. 2024, 36, 102254. [Google Scholar] [CrossRef]
- Huang, B.; Xie, J.; Yan, J. Inspection Robot Navigation Based on Improved TD3 Algorithm. Sensors 2024, 24, 2525. [Google Scholar] [CrossRef] [PubMed]
- Quan, H.; Li, Y.; Zhang, Y. A novel mobile robot navigation method based on deep reinforcement learning. Int. J. Adv. Robot. Syst. 2020, 17. [Google Scholar] [CrossRef]
- Zhang, Y.; Chen, P. Path Planning of a Mobile Robot for a Dynamic Indoor Environment Based on an SAC-LSTM Algorithm. Sensors 2023, 23, 9802. [Google Scholar] [CrossRef] [PubMed]
- Zhang, H.; Sun, Y.; Liu, P.; Ding, D.; Sun, R. BMTD3: An Enhanced TD3 for Mapless Autonomous Navigation with BiGRU and Multi-Head Attention. J. Intell. Robot Syst. 2026, 112, 11. [Google Scholar] [CrossRef]
- Zhang, J.; Chen, H.; Sun, H.; Xu, H.; Yan, T. Convolutional neural network-based deep Q-network (CNN-DQN) path planning method for mobile robots. Intel. Serv. Robot. 2025, 18, 929–950. [Google Scholar] [CrossRef]
- Lu, Z.; He, L.; Wang, H.; Yuan, L.; Xiao, W.; Liu, Z.; Chen, Y. CMADRL: Cross-modal attention based deep reinforcement learning for mobile robot obstacle avoidance. Meas. Sci. Technol. 2025, 36, 036306. [Google Scholar] [CrossRef]
- Meng, J.; Zou, J.; Wang, S.; Yang, R.; Kumar, A.; Kim, J. Deep reinforcement learning for robust robot navigation in complex and crowded environments. J. King Saud Univ. Comput. Inf. Sci. 2025, 37, 333. [Google Scholar] [CrossRef]
- Zhang, Z.; Fu, H.; Yang, J.; Lin, Y. Deep reinforcement learning for path planning of autonomous mobile robots in complicated environments. Complex Intell. Syst. 2025, 11, 277. [Google Scholar] [CrossRef]
- Samsani, S.S.; Mutahira, H.; Muhammad, M.S. Memory-based crowd-aware robot navigation using deep reinforcement learning. Complex Intell. Syst. 2023, 9, 2147–2158. [Google Scholar] [CrossRef]
- Zhou, Y.; Zhang, W. End-to-end robot intelligent obstacle avoidance method based on deep reinforcement learning with spatiotemporal transformer architecture. Front. Neurorobot. 2025, 19, 1646336. [Google Scholar] [CrossRef] [PubMed]
- Tong, S.; Liu, Q.; Ma, Q.; Qin, J. Integrating deep reinforcement learning and improved artificial potential field method for safe path planning for mobile robots. Robot. Intell. Autom. 2024, 44, 871–886. [Google Scholar] [CrossRef]
- Bi, Y.; Fang, X. A Hybrid Path Planning Framework Integrating Deep Reinforcement Learning and Variable-Direction Potential Fields. Mathematics 2025, 13, 2312. [Google Scholar] [CrossRef]
- Zhang, Y.; Cui, C.; Zhao, Q. Path Planning of Mobile Robot Based on A Star Algorithm Combining DQN and DWA in Complex Environment. Appl. Sci. 2025, 15, 4367. [Google Scholar] [CrossRef]
- Yu, X.; Fan, Y.; Xu, S.; Ou, L. A self-adaptive SAC-PID control approach based on reinforcement learning for mobile robots. Int. J. Robust Nonlinear Control 2022, 32, 9625–9643. [Google Scholar] [CrossRef]
- Tang, Z.; Fu, F.; Lu, G.; Chen, D. Reinforcement Learning for Autonomous Agents: Scene-Specific Dynamic Obstacle Avoidance and Target Pursuit in Unknown Environments. IEEE Access 2024, 12, 145496–145510. [Google Scholar] [CrossRef]
- Sheng, J.; Xin, W.; Qing, M. Spatial memory-augmented visual navigation based on hierarchical deep reinforcement learning in unknown environments. Knowl.-Based Syst. 2024, 2, 111358. [Google Scholar] [CrossRef]
- Montero, E.E.; Mutahira, H.; Pico, N.; Muhammad, M.S. Dynamic warning zone and a short-distance goal for autonomous robot navigation using deep reinforcement learning. Complex Intell. Syst. 2024, 10, 1149–1166. [Google Scholar] [CrossRef]
- Long, Z.; Zhang, X.; Mi, J.; Wang, J. Human-Risk-Aware Safe Path Planning Based on Reinforcement Learning for Autonomous Mobile Robots. Sensors 2025, 25, 7211. [Google Scholar] [CrossRef] [PubMed]
- Sun, X.; Zhang, Q.; Wei, Y.; Liu, M. Risk-Aware Deep Reinforcement Learning for Robot Crowd Navigation. Electronics 2023, 12, 4744. [Google Scholar] [CrossRef]
- Xiang, K.; Yuan, X.; Zhong, S.; Di-Hua, Z.; Yun, D.; Si, Z. Differential High Order Control Barrier Function-Based Safe Reinforcement Learning. IEEE Robot. Autom. Lett. 2025, 7, 7524–7531. [Google Scholar] [CrossRef]
- Zhu, X.; Liang, Y.; Sun, H.; Wang, X.; Ren, B. Robot obstacle avoidance system using deep reinforcement learning. Ind. Robot. 2022, 49, 301–310. [Google Scholar] [CrossRef]
- Konar, A.; Baghi, B.H.; Dudek, G. Learning goal conditioned socially compliant navigation from demonstration using risk-based features. IEEE Robot. Autom. Lett. 2021, 4, 651–658. [Google Scholar] [CrossRef]
- Mei, L.; Xu, P. Path Planning for Robots Combined with Zero-Shot and Hierarchical Reinforcement Learning in Novel Environments. Actuators 2024, 13, 458. [Google Scholar] [CrossRef]
- Zhao, H.; Guo, Y.; Liu, Y.; Jin, J. Multirobot unknown environment exploration and obstacle avoidance based on a voronoi diagram and reinforcement learning. Expert Syst. Appl. 2025, 264, 125900. [Google Scholar] [CrossRef]
- Zhao, E.; Zhou, N.; Liu, C.; Su, H.; Liu, Y.; Cong, J. Time-aware MADDPG with LSTM for multi-agent obstacle avoidance: A comparative study. Complex Intell. Syst. 2024, 10, 4141–4155. [Google Scholar] [CrossRef]
- Dong, L.; He, Z.; Song, C.; Yuan, X.; Zhang, H. Multi-robot social-aware cooperative planning in pedestrian environments using attention-based actor-critic. Artif. Intell. Rev. 2024, 57, 108. [Google Scholar] [CrossRef]
- Le, H.; Saeedvand, S.; Hsu, C.C. A Comprehensive Review of Mobile Robot Navigation Using Deep Reinforcement Learning Algorithms in Crowded Environments. J. Intell. Robot Syst. 2024, 110, 158. [Google Scholar] [CrossRef]
- Placed, J.A.; Castellanos, J.A. A Deep Reinforcement Learning Approach for Active SLAM. Appl. Sci. 2020, 10, 8386. [Google Scholar] [CrossRef]
- Malczyk, G.; Kulkarni, M.; Alexis, K. Reinforcement Learning for Active Perception in Autonomous Navigation. arXiv 2026, arXiv:2602.01266v1. [Google Scholar] [CrossRef]
- Liu, B.; Xiao, X.; Stone, P. A Lifelong Learning Approach to Mobile Robot Navigation. IEEE Robot. Autom. Lett. 2021, 6, 1090–1096. [Google Scholar] [CrossRef]
- Lv, T.; Zhang, J.; Chen, Y. A SLAM Algorithm Based on Edge-Cloud Collaborative Computing. J. Sens. 2022, 7213044, 17. [Google Scholar] [CrossRef]
- Ju, H.; Juan, R.; Gomez, R.; Nakamura, K.; Li, G. Transferring Policy of Deep Reinforcement Learning from Simulation to Reality for Robotics. Nat. Mach. Intell. 2022, 4, 1077–1087. [Google Scholar] [CrossRef]
- Wu, J.; Zhou, Y.; Yang, H.; Huang, Z.; Lv, C. Human-Guided Reinforcement Learning with Sim-to-Real Transfer for Autonomous Navigation. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 14745–14759. [Google Scholar] [CrossRef] [PubMed]
- Muratore, F.; Ramos, F.; Turk, G.; Yu, W.; Gienger, M.; Peters, J. Robot Learning from Randomized Simulations: A Review. Front. Robot. AI 2022, 9, 799893. [Google Scholar] [CrossRef] [PubMed]
- Jang, Y.; Baek, J.; Jeon, S.; Han, S. Bridging the Simulation-to-Real Gap of Depth Images for Deep Reinforcement Learning. Expert Syst. Appl. 2024, 253, 124310. [Google Scholar] [CrossRef]


| References | Key Findings |
|---|---|
| [20] | This paper categorizes existing dynamic path planning methods into four distinct paradigms: space-based, time-based, environment-based, and learning-based approaches, with a primary focus on outlining the evolutionary trajectory from classical techniques to advanced state-of-the-art methods. |
| [21] | This article focuses on the utilization of artificial intelligence to assist mobile robot navigation, providing a comparative analysis of the strategic differences among various AI-driven methodologies. |
| [22] | This paper emphasizes how DRL algorithms enhance robotic exploration capabilities, further elaborating on specific implementations and the potential development of DRL in robotic attitude control tasks. |
| [23] | Starting from the foundational principles of DRL algorithms, this paper explores their application and integration within mobile robot path planning. |
| [24] | This study summarizes the performance, functional characteristics, and structural differences in DRL algorithms across multiple typical robotic navigation tasks. |
| Tuple Elements | MDP | POMDP |
|---|---|---|
| Core Tuple | ||
| State Observability | Fully observable | Partially observable |
| Applicability | Ideal simulation scenarios only | Real-world physical scenarios |
| Algorithm Category | References | Base Algorithm | Scenario Applicability |
|---|---|---|---|
| Value-based | [32] | DQN | Suitable for simple indoor obstacle avoidance scenarios with low-dimensional discrete action spaces. |
| Policy-based | [36] | TRPO | Suitable for scenarios where stability requirements outweigh real-time performance. |
| Policy-based | [30] | PPO | Suitable for continuous control scenarios requiring a balance between training stability, sample efficiency, and computational cost. |
| Actor-Critic Architecture | [15] | DDPG | Suitable for continuous control of mobile robot linear/angular velocities; well-adapted to low-dimensional observation spaces. |
| [32,39] | A3C/A2C | Suitable for multi-scenario parallel training and multi-scenario generalization training. | |
| [40] | SAC | Strong exploration capability and robustness; suitable for complex dynamic indoor obstacle avoidance scenarios. | |
| [42] | TD3 | Addresses DDPG’s training instability and overfitting issues in dynamic environments; suitable for complex dynamic indoor obstacle avoidance scenarios. |
| Core Challenge | Algorithm Subcategory | Algorithms | Solution |
|---|---|---|---|
| Value Evaluation Optimization | Architecture Decoupling | Double DQN [46], Dueling DQN [47] | Dual networks and advantage functions are adopted to solve overestimation bias and low evaluation efficiency under redundant actions. |
| Replay and Reward Optimization | DM-DQN [48], PER-D2MQN [49], RND3QN [50] | Prioritized experience replay and multi-step bootstrapping are introduced to solve low sample efficiency under sparse rewards. | |
| Hybrid Exploration and Curiosity | NM-TD3 [51], ASAC [52], LP-TD3 [53] | Hybrid noise and intrinsic curiosity modules are combined to solve the robot’s fear of exploring unknown regions. | |
| Spatio-temporal Multimodal Perception | Temporal Prediction | GRU-DDPG [54], SAC-LSTM [55], BMTD3 [56] | Temporal prediction modules are introduced to overcome the inability to predict future trajectories of dynamic obstacles in POMDP environments. |
| Cross-Modal Feature Fusion | CNN-DQN [57], CMADRL [58] | RGB-D and LiDAR data are fused to extract high-dimensional semantic features, eliminating blind spots inherent in single sensors. | |
| Attention-Enhanced Prediction | Attention-Enhanced PPO [59], GAP_SAC [60], CAM-RL [61], Spatio-temporal Transformer DRL [62] | Gated/multi-head and global self-attention mechanisms are adopted to accelerate key threat extraction in dense crowds and realize active obstacle avoidance. | |
| Hierarchical Hybrid and Expert Experience Control | Expert Prior Guidance | IDDPG-IAPF [63], VDPF-TD3 [64] | Traditional obstacle avoidance knowledge is embedded into the network as prior information to solve blind collisions in early training and local optima. |
| Hybrid Control | A*-DQN-DWA [65], SAC-PID [66], CNN-DQN-B [57] | Traditional algorithms are integrated to resolve kinematic constraint shortages and trajectory oscillations in pure end-to-end control. | |
| Task Decoupling | SSRL-PPO [67], HRL [68] | It alleviates reward conflicts among multiple tasks and slows generalization to novel environments. | |
| Interactive Safety Perception | Safety Constraint | DRL with Dynamic Warning Zones [69], DRL with Human Risk Predictor [70], Risk-Aware DRL [71], DSRL [72] | Hard risk thresholds are defined based on a constrained MDP to provide a rigorous mathematical safety lower bound for black box decisions. |
| Interactive Response Enhancement | Self-Supervised DRL [73], IRL [74], Zero-Shot Hierarchical DRL [75] | Control policies are learned from raw data to improve slow responses to sudden human behaviors. | |
| Interactive Collaboration Compliance | Voronoi-DRL [76], MADDPG-LSTM Actor [77], MARL [78] | Temporal multi-agent networks and spatio-temporal social encoders are introduced to ensure compliant human–robot and multi-robot collaboration. |
| Algorithm | Environment Type | Success Rate | Convergence Behavior (Episodes) | Comparison Algorithms |
|---|---|---|---|---|
| Double DQN [46] | Atari 2600 games | The normalized hyperparameter score increased from 233% to 617% on Road Runner | / | Reduces the overestimation of DQN. |
| DM-DQN [48] | Static and dynamic obstacle environments | 67.6% | Approx. 120 | Improves the success rate by 18.3% compared with DQN. |
| PER-D2MQN [49] | Static, dynamic and complex environments | 65.8% (static), 52.7% (dynamic), 44.6% (complex) | Approx. 250 | Outperforms the comparison algorithms in all obstacle scenarios. |
| NM-TD3 [51] | Complex static and dynamic environments | 92% (static), 82% (dynamic) | Approx. 800 | Improves the success rate by 37–28% compared with the original TD3. |
| SAC-LSTM [55] | Dynamic indoor environments | 100% (obstacle-free), 95.5% (static), 89% (dynamic) | Approx. 130 (obstacle-free), Approx. 500 (static), Approx. 600 (dynamic) | Outperforms SAC in all scenarios. |
| GAP_SAC [60] | Dynamic and narrow environments | 96% ((simple), 95% (normal), 93% (complex) | / | Outperforms the comparison algorithms in all obstacle scenarios. |
| CAM-RL [61] | Cross-shaped pedestrian flow environments | 100% (5 people), 99% (10 people), | / | Outperforms CADRL, LSTM-RL and SARL. |
| VDPF-TD3 [64] | Complex dynamic environments | Effective path planning performance in dynamic environments | Path length shortened by a factor of 0.951 | Outperforms the comparison algorithms in all obstacle scenarios. |
| A*-DQN-DWA [65] | Complex dynamic environments | 99.36% | Approx. 100 | Outperforms DWA and DQN. |
| SSRL-PPO [67] | Static and dynamic obstacle environments | Significantly improved obstacle avoidance performance | / | Outperforms hierarchical reinforcement learning. |
| Risk-Aware DRL [71] | High-density crowd environments | 98.0% (10 people), 93.2% (20 people), 90.0% (25 people) | / | Outperforms ORCA, DS-RNN and CrowdNav++. |
| Zero-Shot Hierarchical DRL [75] | Unseen complex environments | 31.88% | Approx. 480 | Outperforms traditional HRL (16.28%). |
| MARL [78] | Dynamic environments with pedestrians | 98.8% (5p3r), 99.8% (10p3r), 92.6% (20p3r) | / | Outperforms the comparison algorithms in all scenarios. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhao, J.; Zhao, H.; Li, B.; Gong, X. Deep Reinforcement Learning for Dynamic Obstacle Avoidance of Mobile Robots in Indoor Environments: A Review. Sensors 2026, 26, 4797. https://doi.org/10.3390/s26154797
Zhao J, Zhao H, Li B, Gong X. Deep Reinforcement Learning for Dynamic Obstacle Avoidance of Mobile Robots in Indoor Environments: A Review. Sensors. 2026; 26(15):4797. https://doi.org/10.3390/s26154797
Chicago/Turabian StyleZhao, Jiandong, Honghua Zhao, Benwang Li, and Xuyin Gong. 2026. "Deep Reinforcement Learning for Dynamic Obstacle Avoidance of Mobile Robots in Indoor Environments: A Review" Sensors 26, no. 15: 4797. https://doi.org/10.3390/s26154797
APA StyleZhao, J., Zhao, H., Li, B., & Gong, X. (2026). Deep Reinforcement Learning for Dynamic Obstacle Avoidance of Mobile Robots in Indoor Environments: A Review. Sensors, 26(15), 4797. https://doi.org/10.3390/s26154797













