A Systematic Review of Deep Reinforcement Learning for Legged Robot Locomotion
Abstract
1. Introduction
2. Previous Studies
2.1. Types of Legged Robots and Limitations of Traditional Model-Based Control Methods
2.2. DRL for Legged Robots
2.2.1. End-to-End Learning
2.2.2. Hierarchical RL
2.2.3. Integration of Imitation Learning and Reinforcement Learning
2.2.4. Model-Based RL
2.2.5. Hybrid Control
2.2.6. Meta and Adaptive RL
2.3. Representative Applications of DRL in Legged Robots
3. Review Methodology
3.1. Research Questions
3.2. Literature Sources and Search Strategies
3.3. Data Collection and Analysis
4. Results
4.1. What Types of DRL Algorithms Have Been Proposed for Legged Robot Locomotion Control (RQ1)?
4.2. What Are the Key Factors That Influence the Performance of DRL-Based Legged Robot Control Systems (RQ2)?
4.3. What Methodological Approaches Have Been Developed to Improve Sample Efficiency, Accelerate Training, and Enhance Generalization in DRL-Based Locomotion Control (RQ3)?
4.4. How Do DRL-Based Methods Compare with Traditional Model-Based or Optimization-Based Locomotion Control Methods in Terms of Stability, Robustness, Adaptability, and Computational Efficiency (RQ4)?
4.5. What Strategies Have Been Proposed to Bridge the Sim-to-Real Gap for DRL Policies Deployed on Real Legged Robots (RQ5)?
4.6. What Are the Commonly Used Simulation Environments, Datasets, Benchmarks, and Evaluation Metrics in DRL-Based Legged Robot Research (RQ6)?
5. Analysis
5.1. Terrain Adaptation Performance Analysis
5.2. Robust Analysis Under External Disturbances
5.3. Performance Under Different Robot Forms
5.4. Performance Under Different Sensory Modalities
- The single-modality vision scheme performs reasonably well in terms of energy efficiency, but its stability and task responsiveness are relatively weak.
- The dual-modality vision–inertial scheme improves stability and responsiveness while maintaining high energy efficiency.
- The multimodal fusion scheme (vision, force, and inertia) achieves the best performance across all metrics, with the highest average reward, best stability, and high energy efficiency, while also minimizing response latency, demonstrating greater environmental adaptability and robustness.
5.5. Performance Under Different Energy Consumption Conditions
5.6. Performance Under Actuator Constraints
- As actuator constraints increase, performance in stability and recovery decreases.
- Under low constraint, DRL strategies can fully utilize the robot’s dynamic potential to perform higher-speed and more complex movements.
- Moderate constraints achieve a compromise between safety and performance, while high constraints significantly degrade performance, especially in dynamic maneuvers and complex terrain.
6. Conclusions
6.1. The Answer to Seven Research Questions
6.2. Future Research Directions
Supplementary Materials
Author Contributions
Funding
Conflicts of Interest
Abbreviations
| DRL | Deep Reinforcement Learning |
| RL | Reinforcement Learning |
| PPO | Proximal Policy Optimization |
| SAC | Soft Actor–Critic |
| TD3 | Twin Delayed Deep Deterministic Policy Gradient |
| CNN | Convolutional Neural Networks |
| Meta-RL | Meta-Reinforcement Learning |
| AdPO | Adaptive Policy Optimization |
| HRL | Hierarchical Reinforcement Learning |
| HIRO | Hierarchical Reinforcement Learning with Off-policy Correction |
| HAC | Hierarchical Actor–Critic |
| FuN | Feudal Networks |
| MoE | Mixture-of-Experts |
| GAIL | Generative Adversarial Imitation Learning |
| GANs | Generative Adversarial Networks |
| BC | Behavior Cloning |
| AIRL | Adversarial Inverse Reinforcement Learning |
| MoCap | Motion Capture |
| MPC | Model Predictive Control |
| PETS | Probabilistic Ensembles with Trajectory Sampling |
| MBPO | Model-Based Policy Optimization |
| CPG | Central Pattern Generator |
| MAML | Model-Agnostic Meta-Learning |
| EMAML | Exploration-Enhanced Model-Agnostic Meta-Learning |
| ALPaCA | Adaptive Learning from Probabilistic Context Adaptation |
| CAVIA | Context-Adaptive Variational Inference Algorithm |
| RL2 | Fast Reinforcement Learning via Slow Reinforcement Learning |
| RNN | Recurrent Neural Network |
| PRISMA 2020 | Preferred Reporting Items for Systematic Reviews and Meta-Analyses |
| RQ | Research Questions |
| MTL | Multi-Task Learning |
| GNN | Graph Neural Network |
| VMC | Virtual Model Control |
References
- Yang, H.; Zhou, J.; Wu, J.; Yao, Y.A. Research on high-smooth walking and adaptive obstacle-crossing of closed-chain multi-legged robot. Mech. Mach. Theory 2025, 214, 106123. [Google Scholar] [CrossRef] [Scilit]
- Menon, U.V.; Kumaravelu, V.B.; Kumar, C.V.; Rammohan, A.; Chinnadurai, S.; Venkatesan, R.; Hai, H.; Selvaprabhu, P. AI-powered IoT: A survey on integrating artificial intelligence with IoT for enhanced security, efficiency, and smart applications. IEEE Access 2025, 13, 50296–50339. [Google Scholar]
- Zhao, Y.; Wang, J.; Cao, G.; Yuan, Y.; Yao, X.; Qi, L. Intelligent control of multilegged robot smooth motion: A review. IEEE Access 2023, 11, 86645–86685. [Google Scholar] [CrossRef] [Scilit]
- Cao, W.; Bukhari, A.A.S.; Aarniovuori, L. Review of electrical motor drives for electric vehicle applications. Mehran Univ. Res. J. Eng. Technol. 2019, 38, 525–540. [Google Scholar] [CrossRef] [Scilit]
- Guo, Z.; Dong, Z.; Lee, K.-H.; Cheung, C.L.; Fu, H.-C.; Ho, J.D.; He, H.; Poon, W.-S.; Chan, D.T.-M.; Kwok, K.-W. Compact design of a hydraulic driving robot for intraoperative MRI-guided bilateral stereotactic neurosurgery. IEEE Robot. Autom. Lett. 2018, 3, 2515–2522. [Google Scholar] [CrossRef] [Scilit]
- Hu, K.; Shi, H.; He, Y.; Wang, W.; Liu, C.K.; Song, S. Robot Trains Robot: Automatic Real-World Policy Adaptation and Learning for Humanoids. arXiv 2025, arXiv:2508.12252. [Google Scholar] [CrossRef] [Scilit]
- Lu, Y.; Yang, R.; Kou, Q.; Chen, M.; Fan, T.; Cui, P.; Dong, Y.; Lu, P. Contrastive Representation Learning for Robust Sim-to-Real Transfer of Adaptive Humanoid Locomotion. arXiv 2025, arXiv:2509.12858. [Google Scholar]
- Sebastián, E.; Duong, T.; Atanasov, N.; Montijano, E.; Sagüés, C. Physics-informed multi-agent reinforcement learning for distributed multi-robot problems. IEEE Trans. Robot. 2025, 41, 4499–4517. [Google Scholar] [CrossRef] [Scilit]
- Mirza, K.Z.; Singh, S. Imitation learning for legged robot locomotion: A survey. Front. Robot. AI 2025, 12, 1678567. [Google Scholar] [CrossRef] [Scilit]
- Araki, T.; Mukuta, Y.; Osa, T.; Harada, T. Few-shot Imitation Learning by Variable-Length Trajectory Retrieval from a Large and Diverse Dataset. In Proceedings of the 2025 IEEE-RAS 24th International Conference on Humanoid Robots (Humanoids), Seoul, Republic of Korea, 30 September–2 October 2025; IEEE: New York, NY, USA, 2025; pp. 829–836. [Google Scholar]
- Watanabe, T.; Kubo, A.; Tsunoda, K.; Matsuba, T.; Akatsuka, S.; Noda, Y.; Kioka, H.; Izawa, J.; Ishii, S.; Nakamura, Y. Hierarchical reinforcement learning with central pattern generator for enabling a quadruped robot simulator to walk on a variety of terrains. Sci. Rep. 2025, 15, 11262. [Google Scholar] [CrossRef] [Scilit]
- Azimi, D.; Hoseinnezhad, R. Hierarchical Reinforcement Learning for Quadrupedal Robots: Efficient Object Manipulation in Constrained Environments. Sensors 2025, 25, 1565. [Google Scholar] [CrossRef] [Scilit]
- Peng, X.B.; Berseth, G.; Yin, K.; Van De Panne, M. Deeploco: Dynamic locomotion skills using hierarchical deep reinforcement learning. Acm Trans. Graph. 2017, 36, 1–13. [Google Scholar] [CrossRef] [Scilit]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Rudin, N.; Hoeller, D.; Reist, P.; Hutter, M. Learning to walk in minutes using massively parallel deep reinforcement learning. In Conference on Robot Learning; PMLR: London, UK, 2022; pp. 91–100. [Google Scholar]
- Fujimoto, S.; Hoof, H.; Meger, D. Addressing function approximation error in actor-critic methods. In International Conference on Machine Learning; PMLR: London, UK, 2018; pp. 1587–1596. [Google Scholar]
- Makoviychuk, V.; Wawrzyniak, L.; Guo, Y.; Lu, M.; Storey, K.; Macklin, M.; Hoeller, D.; Rudin, N.; Allshire, A.; Handa, A.; et al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv 2021, arXiv:2108.10470. [Google Scholar] [CrossRef] [Scilit]
- Kim, Y.; Oh, H.; Lee, J.; Choi, J.; Ji, G.; Jung, M.; Youm, D.; Hwangbo, J. Not only rewards but also constraints: Applications on legged robot locomotion. IEEE Trans. Robot. 2024, 40, 2984–3003. [Google Scholar] [CrossRef] [Scilit]
- Gupta, A.; Mendonca, R.; Liu, Y.; Abbeel, P.; Levine, S. Meta-reinforcement learning of structured exploration strategies. Adv. Neural Inf. Process. Syst. 2018, 31, 5125–5134. [Google Scholar]
- Barto, A.G.; Mahadevan, S. Recent advances in hierarchical reinforcement learning. Discret. Event Dyn. Syst. 2003, 13, 341–379. [Google Scholar] [CrossRef] [Scilit]
- Parr, R.; Russell, S. Reinforcement learning with hierarchies of machines. Adv. Neural Inf. Process. Syst. 1997, 10, 1043–1049. [Google Scholar]
- Nachum, O.; Gu, S.; Lee, H.; Levine, S. Near-optimal representation learning for hierarchical reinforcement learning. arXiv 2018, arXiv:1810.01257. [Google Scholar]
- Bacon, P.-L.; Harb, J.; Precup, D. The option-critic architecture. In Proceedings of the AAAI Conference on Artificial Intelligence, San Francisco, CA, USA, 4–9 February 2017. [Google Scholar]
- Luo, X.; Li, Q.; Su, K. Multi-terrain Motion Control Method for Quadruped Robot Based on Reinforcement Learning. In The World Conference on Intelligent and 3D Technologies; Springer: Berlin/Heidelberg, Germany, 2023; pp. 113–123. [Google Scholar]
- Vezhnevets, A.S.; Osindero, S.; Schaul, T.; Heess, N.; Jaderberg, M.; Silver, D.; Kavukcuoglu, K. Feudal networks for hierarchical reinforcement learning. In International Conference on Machine Learning; PMLR: London, UK, 2017; pp. 3540–3549. [Google Scholar]
- Levy, A.; Konidaris, G.; Platt, R.; Saenko, K. Learning multi-level hierarchies with hindsight. arXiv 2017, arXiv:1712.00948. [Google Scholar]
- Rajeswaran, A.; Kumar, V.; Gupta, A.; Vezzani, G.; Schulman, J.; Todorov, E.; Levine, S. Learning complex dexterous manipulation with deep reinforcement learning and demonstrations. arXiv 2017, arXiv:1709.10087. [Google Scholar]
- Ho, J.; Ermon, S. Generative adversarial imitation learning. Adv. Neural Inf. Process. Syst. 2016, 29, 4565–4573. [Google Scholar]
- Fu, J.; Luo, K.; Levine, S. Learning robust rewards with adversarial inverse reinforcement learning. arXiv 2017, arXiv:1710.11248. [Google Scholar]
- Ross, S.; Gordon, G.; Bagnell, D. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics; JMLR Workshop and Conference Proceedings: Cambridge, MA, USA, 2011; pp. 627–635. [Google Scholar]
- Clegg, A.; Yu, W.; Tan, J.; Liu, C.K.; Turk, G. Learning to dress: Synthesizing human dressing motion via deep reinforcement learning. ACM Trans. Graph. 2018, 37, 1–10. [Google Scholar] [CrossRef] [Scilit]
- Moerland, T.M.; Broekens, J.; Plaat, A.; Jonker, C.M. Model-based reinforcement learning: A survey. Found. Trends Mach. Learn. 2023, 16, 1–118. [Google Scholar] [CrossRef] [Scilit]
- Tassa, Y.; Erez, T.; Todorov, E. Synthesis and stabilization of complex behaviors through online trajectory optimization. In Proceedings of the 2012 IEEE/RSJ International Conference on Intelligent Robots and Systems, Vilamoura-Algarve, Portugal, 7–12 October 2012; IEEE: New York, NY, USA, 2012; pp. 4906–4913. [Google Scholar]
- Dai, H.; Valenzuela, A.; Tedrake, R. Whole-body motion planning with centroidal dynamics and full kinematics. In Proceedings of the 2014 IEEE-RAS International Conference on Humanoid Robots, Madrid, Spain, 18–20 November 2014; IEEE: New York, NY, USA, 2014; pp. 295–302. [Google Scholar]
- Chua, K.; Calandra, R.; McAllister, R.; Levine, S. Deep reinforcement learning in a handful of trials using probabilistic dynamics models. Adv. Neural Inf. Process. Syst. 2018, 31, 4754–4765. [Google Scholar]
- Janner, M.; Fu, J.; Zhang, M.; Levine, S. When to trust your model: Model-based policy optimization. Adv. Neural Inf. Process. Syst. 2019, 32, 12519–12530. [Google Scholar]
- Wensing, P.M.; Posa, M.; Hu, Y.; Escande, A.; Mansard, N.; Del Prete, A. Optimization-based control for dynamic legged robots. IEEE Trans. Robot. 2023, 40, 43–63. [Google Scholar] [CrossRef] [Scilit]
- Hafner, D.; Lillicrap, T.; Ba, J.; Norouzi, M. Dream to control: Learning behaviors by latent imagination. arXiv 2019, arXiv:1912.01603. [Google Scholar]
- Ha, D.; Schmidhuber, J. Recurrent world models facilitate policy evolution. Adv. Neural Inf. Process. Syst. 2018, 31, 2450–2462. [Google Scholar]
- Kurutach, T.; Clavera, I.; Duan, Y.; Tamar, A.; Abbeel, P. Model-ensemble trust-region policy optimization. arXiv 2018, arXiv:1802.10592. [Google Scholar]
- Sekar, R.; Rybkin, O.; Daniilidis, K.; Abbeel, P.; Hafner, D.; Pathak, D. Planning to explore via self-supervised world models. In International Conference on Machine Learning; PMLR: London, UK, 2020; pp. 8583–8592. [Google Scholar]
- Bernabeu, P. Language and Sensorimotor Simulation in Conceptual Processing: Multilevel Analysis and Statistical Power. Doctoral Dissertation, Lancaster University, Lancaster, UK, 2022. [Google Scholar]
- Rudin, N.; Kolvenbach, H.; Tsounis, V.; Hutter, M. Cat-like jumping and landing of legged robots in low gravity using deep reinforcement learning. IEEE Trans. Robot. 2021, 38, 317–328. [Google Scholar] [CrossRef] [Scilit]
- Gangapurwala, S.; Geisert, M.; Orsolino, R.; Fallon, M.; Havoutis, I. Rloc: Terrain-aware legged locomotion using reinforcement learning and optimal control. IEEE Trans. Robot. 2022, 38, 2908–2927. [Google Scholar] [CrossRef] [Scilit]
- Haarnoja, T.; Moran, B.; Lever, G.; Huang, S.H.; Tirumala, D.; Humplik, J.; Wulfmeier, M.; Tunyasuvunakool, S.; Siegel, N.Y.; Hafner, R.; et al. Learning agile soccer skills for a bipedal robot with deep reinforcement learning. Sci. Robot. 2024, 9, eadi8022. [Google Scholar] [CrossRef] [Scilit]
- Johannink, T.; Bahl, S.; Nair, A.; Luo, J.; Kumar, A.; Loskyll, M.; Ojea, J.A.; Solowjow, E.; Levine, S. Residual reinforcement learning for robot control. In Proceedings of the 2019 International Conference on Robotics and Automation (ICRA), Montreal, QC, Canada, 20–24 May 2019; IEEE: New York, NY, USA, 2019; pp. 6023–6029. [Google Scholar]
- Yu, C.; Rosendo, A. Multi-modal legged locomotion framework with automated residual reinforcement learning. IEEE Robot. Autom. Lett. 2022, 7, 10312–10319. [Google Scholar] [CrossRef] [Scilit]
- Yazdi, M.R.H.; Pourghavam, M.; Bidari, R. Hybrid Control of Advanced Quadruped Locomotion: Integrating Model Predictive Control with Deep Reinforcement Learning. In Proceedings of the 2024 12th RSI International Conference on Robotics and Mechatronics (ICRoM), Tehran, Iran, 17–19 December 2024; IEEE: New York, NY, USA, 2024; pp. 201–208. [Google Scholar]
- Ijspeert, A.J. Biorobotics: Using robots to emulate and investigate agile locomotion. Science 2014, 346, 196–203. [Google Scholar] [CrossRef] [Scilit]
- Finn, C.; Abbeel, P.; Levine, S. Model-agnostic meta-learning for fast adaptation of deep networks. In International Conference on Machine Learning; PMLR: London, UK, 2017; pp. 1126–1135. [Google Scholar]
- Duan, Y.; Schulman, J.; Chen, X.; Bartlett, P.L.; Sutskever, I.; Abbeel, P. Rl $^ 2$: Fast reinforcement learning via slow reinforcement learning. arXiv 2016, arXiv:1611.02779. [Google Scholar]
- Nichol, A.; Achiam, J.; Schulman, J. On first-order meta-learning algorithms. arXiv 2018, arXiv:1803.02999. [Google Scholar]
- Rakelly, K.; Zhou, A.; Finn, C.; Levine, S.; Quillen, D. Efficient off-policy meta-reinforcement learning via probabilistic context variables. In International Conference on Machine Learning; PMLR: London, UK, 2019; pp. 5331–5340. [Google Scholar]
- Harrison, J.; Sharma, A.; Pavone, M. Meta-learning priors for efficient online bayesian regression. In International Workshop on the Algorithmic Foundations of Robotics; Springer: Berlin/Heidelberg, Germany, 2018; pp. 318–337. [Google Scholar]
- Zintgraf, L.; Shiarli, K.; Kurin, V.; Hofmann, K.; Whiteson, S. Fast context adaptation via meta-learning. In International Conference on Machine Learning; PMLR: London, UK, 2019; pp. 7693–7702. [Google Scholar]
- Clavera, I.; Rothfuss, J.; Schulman, J.; Fujita, Y.; Asfour, T.; Abbeel, P. Model-based reinforcement learning via meta-policy optimization. In Conference on Robot Learning; PMLR: London, UK, 2018; pp. 617–629. [Google Scholar]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit]
- Yang, Y.; Caluwaerts, K.; Iscen, A.; Zhang, T.; Tan, J.; Sindhwani, V. Data efficient reinforcement learning for legged robots. In Conference on Robot Learning; PMLR: London, UK, 2020; pp. 1–10. [Google Scholar]
- Belmonte-Baeza, A.; Lee, J.; Valsecchi, G.; Hutter, M. Meta reinforcement learning for optimal design of legged robots. IEEE Robot. Autom. Lett. 2022, 7, 12134–12141. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.-M.; Bui, H.; Posa, M. Reinforcement learning for reduced-order models of legged robots. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: New York, NY, USA, 2024; pp. 5801–5807. [Google Scholar]
- Ha, S.; Kim, J.; Yamane, K. Automated deep reinforcement learning environment for hardware of a modular legged robot. In Proceedings of the 2018 15th International Conference on Ubiquitous Robots (UR), Honolulu, HI, USA, 26–30 June 2018; IEEE: New York, NY, USA, 2018; pp. 348–354. [Google Scholar]
- Smith, L.; Kew, J.C.; Bin Peng, X.; Ha, S.; Tan, J.; Levine, S. Legged robots that keep on learning: Fine-tuning locomotion policies in the real world. In Proceedings of the 2022 International Conference on Robotics and Automation (ICRA), Philadelphia, PA, USA, 23–27 May 2022; IEEE: New York, NY, USA, 2022; pp. 1593–1599. [Google Scholar]
- Chen, X.; Ghadirzadeh, A.; Folkesson, J.; Bjorkman, M.; Jensfelt, P. Deep reinforcement learning to acquire navigation skills for wheel-legged robots in complex environments. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; IEEE: New York, NY, USA, 2018; pp. 3110–3116. [Google Scholar]
- Gan, L.; Grizzle, J.W.; Eustice, R.M.; Ghaffari, M. Energy-based legged robots terrain traversability modeling via deep inverse reinforcement learning. IEEE Robot. Autom. Lett. 2022, 7, 8807–8814. [Google Scholar] [CrossRef] [Scilit]
- Chamorro, S.; Klemm, V.; Valls, M.d.L.I.; Pal, C.; Siegwart, R. Reinforcement learning for blind stair climbing with legged and wheeled-legged robots. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: New York, NY, USA, 2024; pp. 8081–8087. [Google Scholar]
- Konen, K.; Korthals, T.; Melnik, A.; Schilling, M. Biologically-inspired deep reinforcement learning of modular control for a six-legged robot. In Proceedings of the 2019 IEEE International Conference on Robotics and Automation Workshop on Learning Legged Locomotion Workshop,(ICRA) 2019, Montreal, CA, USA, 20–25 May 2019. [Google Scholar]
- Li, S.; Pang, Y.; Bai, P.; Hu, S.; Wang, L.; Wang, G. Dynamic fall recovery control for legged robots via reinforcement learning. Biomimetics 2024, 9, 193. [Google Scholar] [CrossRef] [Scilit]
- Qin, B.; Gao, Y.; Bai, Y. Sim-to-real: Six-legged robot control with deep reinforcement learning and curriculum learning. In Proceedings of the 2019 4th International Conference on Robotics and Automation Engineering (ICRAE), Singapore, 22–24 November 2019; IEEE: New York, NY, USA, 2019; pp. 1–5. [Google Scholar]
- Lee, J.; Bjelonic, M.; Hutter, M. Control of wheeled-legged quadrupeds using deep reinforcement learning. In Climbing and Walking Robots Conference; Springer: Berlin/Heidelberg, Germany, 2022; pp. 119–127. [Google Scholar]
- Chen, G.; Lu, Y.; Yang, X.; Hu, H. Reinforcement learning control for the swimming motions of a beaver-like, single-legged robot based on biological inspiration. Robot. Auton. Syst. 2022, 154, 104116. [Google Scholar] [CrossRef] [Scilit]
- Lyu, S.; Lang, X.; Zhao, H.; Zhang, H.; Ding, P.; Wang, D. Rl2ac: Reinforcement learning-based rapid online adaptive control for legged robot robust locomotion. In Proceedings of the Robotics: Science and Systems, Delft, The Netherlands, 16–19 July 2024. [Google Scholar]
- Weerakoon, K.; Sathyamoorthy, A.J.; Elnoor, M.; Manocha, D. Vapor: Legged robot navigation in unstructured outdoor environments using offline reinforcement learning. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: New York, NY, USA, 2024; pp. 10344–10350. [Google Scholar]
- Chen, C.; Xiang, P.; Zhang, J.; Xiong, R.; Wang, Y.; Lu, H. Deep reinforcement learning based co-optimization of morphology and gait for small-scale legged robot. IEEE/ASME Trans. Mechatron. 2023, 29, 2697–2708. [Google Scholar] [CrossRef] [Scilit]
- Lee, J.; Bjelonic, M.; Reske, A.; Wellhausen, L.; Miki, T.; Hutter, M. Learning robust autonomous navigation and locomotion for wheeled-legged robots. Sci. Robot. 2024, 9, eadi9641. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cui, L.; Wang, S.; Zhang, J.; Zhang, D.; Lai, J.; Zheng, Y.; Zhang, Z.; Jiang, Z.-P. Learning-based balance control of wheel-legged robots. IEEE Robot. Autom. Lett. 2021, 6, 7667–7674. [Google Scholar] [CrossRef] [Scilit]
- Morimoto, D.; Iwamoto, Y.; Hiraga, M.; Ohkura, K. Generating collective behavior of a multi-legged robotic swarm using deep reinforcement learning. J. Robot. Mechatron. 2023, 35, 977–987. [Google Scholar] [CrossRef] [Scilit]
- Yang, T.-Y.; Zhang, T.; Luu, L.; Ha, S.; Tan, J.; Yu, W. Safe reinforcement learning for legged locomotion. In Proceedings of the 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan, 23–27 October 2022; IEEE: New York, NY, USA, 2022; pp. 2454–2461. [Google Scholar]
- Margolis, G.B.; Yang, G.; Paigwar, K.; Chen, T.; Agrawal, P. Rapid locomotion via reinforcement learning. Int. J. Robot. Res. 2024, 43, 572–587. [Google Scholar] [CrossRef] [Scilit]
- Bing, Z.; Lemke, C.; Cheng, L.; Huang, K.; Knoll, A. Energy-efficient and damage-recovery slithering gait design for a snake-like robot based on reinforcement learning and inverse reinforcement learning. Neural Netw. 2020, 129, 323–333. [Google Scholar] [CrossRef] [Scilit]
- Xu, Z.; Raj, A.H.; Xiao, X.; Stone, P. Dexterous legged locomotion in confined 3d spaces with reinforcement learning. In Proceedings of the 2024 IEEE International Conference on Robotics and Automation (ICRA), Yokohama, Japan, 13–17 May 2024; IEEE: New York, NY, USA, 2024; pp. 11474–11480. [Google Scholar]


















| Algorithm | Policy Type | Core Innovation | Control Characteristics | Performance Features |
|---|---|---|---|---|
| PPO [14,15] | On-policy, stochastic | Clipped surrogate objective constrains update magnitude | Smooth policy updates, good convergence stability | Reliable gait learning, moderate sample efficiency |
| SAC [8,12] | Off-policy, stochastic | Maximum entropy regularization enhances exploration | Robust to stochastic environments, adaptive to terrain | High stability, strong exploration ability |
| TD3 [6,16] | Off-policy, deterministic | Double critic and delayed target update mitigate overestimation | Continuous, low-variance action control | Efficient energy use, stable high-frequency execution |
| Distributed RL [17] | Parallel off policy | Multi-environment sampling and gradient averaging | Scale to large systems, improves data throughput | Fast training, enhanced generalization |
| DRL [18] | End-to-end perceptual policy | CNN/Transformer-based visual encoder integrated with policy network | Perception-driven decision making | Terrain-aware locomotion, improved adaptability |
| AdPO [19] | Adaptive metapolicy | Learns generalizable priorities for fast adaptation | Rapid online adaptation to new conditions | High flexibility, suitable for dynamic terrain |
| Algorithm | Core Mechanism | Hierarchical Structure | Application |
|---|---|---|---|
| HIRO [15,22] | High-level policy outputs subgoals; low-level executes them with off-policy correction to prevent experience mismatch | Two-layer hierarchy (high-level subgoals, low-level control) | Quadruped robot adaptive gait and velocity control on uneven terrain |
| Option-Critic Architecture [23,24] | Learns when to start, switch, or terminate behavioral options in an end-to-end differentiable framework | Hierarchical policy with learnable “options” (sub-policies) and termination functions | Hexapod obstacle avoidance, climbing, and turning |
| HAC [25] | Decomposes rewards between levels for better credit assignment and temporal abstraction | Multi-level actor–critic with reward decomposition | Multi-terrain locomotion planning and balance control |
| FuN [26] | High-level “manager” sets latent goals guiding low-level “worker” via vector representations | Feudal manager–worker model | Long-term locomotion control with semantic planning |
| Expert-based HRL [18] | Dynamically selects sub-policies or adapts meta-parameters for cross-task generalization | Adaptive hierarchical modularity | Cross-task and multimodal locomotion learning |
| Algorithm | Core Mechanism | Learning Strategy | Application |
|---|---|---|---|
| GAIL [15,28] | Uses a generator–discriminator framework to imitate expert distributions without explicit rewards | Adversarial Imitation + Policy Gradient | Quadruped gait learning, balance on uneven terrain |
| AIRL [29] | Recovers latent reward functions jointly with policy | Imitation + Inverse RL | Transferable locomotion control across terrains |
| DeepMimic [10] | Combines imitation loss from MoCap with RL rewards in simulation | Imitation pretraining + RL fine-tuning | Quadruped motion imitation and optimization |
| Constrained Imitation-RL Framework [18] | Adds physical and safety constraints during imitation and policy optimization | Constraint-guided hybrid learning | Legged locomotion under safety-critical tasks |
| Multimodal and Hierarchical Imitation-RL [31] | Using multi-level control | Hierarchical policy + multimodal perception | All-terrain adaptive locomotion and manipulation |
| Algorithm | Core Mechanism | Key Innovation | Application |
|---|---|---|---|
| PETS [35] | Learn probabilistic dynamics using an ensemble of neural networks to predict future state distributions | Introduces uncertainty estimation via Gaussian process modeling to improve prediction reliability | Enhance stability and risk-aware control under unseen terrains or external disturbances |
| MBPO [36] | Uses short-horizon model rollouts to generate synthetic samples for policy optimization | Balances real and model-generated data for efficient and stable training | Achieves high sample efficiency and stable performance in locomotion and obstacle-crossing tasks |
| Dreamer [38] | Builds a latent-space world model for policy and value learning using visual inputs | Performs long-horizon prediction and decision-making entirely in the latent space | Reduces real-world interaction needs while maintaining high learning performance in complex terrains |
| Algorithm | Core Mechanism | Integration Strategy | Application |
|---|---|---|---|
| Residual RL [46,47] | Learns a residual term on top of a classical controller to compensate for model errors or external disturbances | RL module adds corrective residuals to traditional control outputs | Enhance stability and adaptability on uneven terrains such as mud or sand by minimizing control residuals |
| RL-MPC [48] | Combines high-level RL with low-level MPC for hierarchical decision-making | RL defines long-term objectives; MPC ensures short-term constraint satisfaction | Achieves dynamic walking and terrain-adaptive planning with improved safety and interpretability |
| RL-CPG [49] | Integrates Reinforcement Learning with CPG neural oscillatory models | RL adjusts CPG oscillation parameters (frequency, phase, amplitude) for adaptive gait control | Enables rhythmic, bio-inspired locomotion adaptable to various speeds and terrains |
| Algorithm | Core Principle | Key Mechanism | Application |
|---|---|---|---|
| MAML [52] | Learns optimal initial parameters for fast adaptation across tasks. | Meta-level gradient optimization enables quick fine-tuning with few updates. | Achieves efficient gait switching across terrains such as slopes and gravel. |
| E-MAML [19] | Extends MAML by integrating exploration to improve robustness. | Adds exploration terms in meta-updates to handle uncertain dynamics. | Enhances terrain adaptability under high uncertainty. |
| PEARL [53] | Models task distribution through latent probabilistic embeddings. | Infers implicit environ-mental features from few samples. | Enables rapid adaptation to new or unseen terrains. |
| ALPaCA [54] | Employs Bayesian regression for online task inference. | Learn task-specific context to support continual adaptation. | Improves control stability in non-stationary environments. |
| CAVIA [55] | Learns task context variables to separate shared and task-specific knowledge. | Reduces inter-task interference via con-text-conditioned updates. | Maintains stability and efficiency during task transitions. |
| RL2 [56] | Internalizes learning dynamics via RNN memory. | Uses recurrent policy networks to encode previous experiences. | Achieves fast terrain adaptation with minimal data. |
| Literature Sources | Search Strings and Keywords |
|---|---|
| Web of Science | TI = (“deep reinforcement learning” OR “multi-legged robot locomotion” OR “legged robot control” OR “quadruped locomotion” OR “biped locomotion” OR “gait adaptation” OR “terrain-aware locomotion”) AND AB = (“deep reinforcement learning” OR “multi-legged robot locomotion” OR “legged robot control” OR “quadruped locomotion” OR “biped locomotion” OR “gait adaptation” OR “terrain-aware locomotion”) AND AK = (“deep reinforcement learning” OR “multi-legged robot locomotion” OR “legged robot control” OR “quadruped locomotion” OR “biped locomotion” OR “gait adaptation” OR “terrain-aware locomotion”) |
| IEEE | (“Document Title”: “deep reinforcement learning” OR “Document Title”: “multi-legged robot locomotion” OR “Document Title”: “legged robot control” OR “Document Title”: “quadruped locomotion” OR “Document Title”: “biped locomotion” OR “Document Title”: “gait adaptation”) AND (“Abstract”: “deep reinforcement learning” OR “Abstract”: “multi-legged robot locomotion” OR “Abstract”: “legged robot control” OR “Abstract”: “quadruped locomotion” OR “Abstract”: “biped locomotion” OR “Abstract”: “gait adaptation”) AND (“Index Terms”: “deep reinforcement learning” OR “Index Terms”: “multi-legged robot locomotion” OR “Index Terms”: “legged robot control” OR “Index Terms”: “quadruped locomotion” OR “Index Terms”: “biped locomotion” OR “Index Terms”: “gait adaptation”) |
| Scopus | TITLE-ABS-KEY (“deep reinforcement learning” OR “multi-legged robot locomotion” OR “legged robot control” OR “quadruped locomotion” OR “biped locomotion” OR “gait adaptation” OR “terrain-aware locomotion”) |
| Science Direct | Title, abstract, or keywords (“deep reinforcement learning” OR “multi-legged robot locomotion” OR “legged robot control”) AND (“gait adaptation” OR “terrain-aware locomotion”) |
| Wiley | (“deep reinforcement learning” OR “multi-legged robot locomotion” OR “legged robot control” OR “quadruped locomotion” OR “biped locomotion”) |
| ACM | Title, Abstract, and Keywords (“deep reinforcement learning” OR “multi-legged robot locomotion” OR “legged robot control” OR “quadruped locomotion” OR “biped locomotion”) |
| Springer | (“deep reinforcement learning” OR “multi-legged robot locomotion” OR “legged robot control” OR “quadruped locomotion” OR “biped locomotion”) |
| Inclusion Criteria | Exclusion Criteria | ||
|---|---|---|---|
| IC1 | Articles published between 2018 and 2025 that explicitly address legged robotic systems within the broader fields of robotics and machine learning, with an emphasis on learning-based locomotion or control problems | EX1 | Reviews, books, book chapters, and other publications that have not been peer-reviewed |
| IC2 | Articles written in English | EX2 | Articles not written in English |
| IC3 | Articles not duplicated in other databases | EX3 | Articles that have been selected in other databases |
| IC4 | Studies focusing on legged robot locomotion, including quadruped, hexapod, or biped systems, using deep reinforcement learning | EX4 | Studies unrelated to legged robot motion control or not employing reinforcement learning approaches |
| IC5 | The full text of the article is available | EX5 | The full text of the article is not available |
| IC6 | Research addressing DRL frameworks, policy learning, sim-to-real transfer, terrain adaptation, or energy-efficient locomotion control | EX6 | Articles not involving DRL-based locomotion control or lacking relevance to motion learning and adaptation |
| Authors and Citation | Year | Venue | Citations | Publication Type | Country |
|---|---|---|---|---|---|
| Yang et al. [58] | 2020 | PMLR CoRL | 200 | Conference | USA |
| Belmonte-Baeza et al. [59] | 2022 | IEEE RA-L | 46 | Journal | Spain |
| Chen et al. [60] | 2024 | IEEE ICRA | 12 | Conference | USA |
| Rudin et al. [43] | 2021 | IEEE T-RO | 143 | Journal | Switzerland |
| Ha et al. [61] | 2018 | IEEE UR | 61 | Conference | Japan |
| Smith et al. [62] | 2022 | IEEE ICRA | 153 | Conference | USA |
| Chen et al. [63] | 2018 | IEEE IROS | 68 | Conference | Sweden |
| Gan et al. [64] | 2022 | IEEE RA-L | 43 | Journal | USA |
| Chamorro et al. [65] | 2024 | IEEE ICRA | 14 | Conference | Switzerland |
| Konen et al. [66] | 2019 | IEEE ICRA | 18 | Conference | Germany |
| Li et al. [67] | 2024 | MDPI Biomimetics | 4 | Journal | China |
| Qin et al. [68] | 2019 | IEEE ICRAE | 22 | Conference | China |
| Lee et al. [69] | 2022 | Springer CWR | 14 | Conference | Germany |
| Chen et al. [70] | 2022 | Elsevier RAS | 44 | Journal | China |
| Lyu et al. [71] | 2024 | RSS | 12 | Conference | China |
| Weerakoon et al. [72] | 2024 | IEEE ICRA | 9 | Conference | USA |
| Chen et al. [73] | 2023 | ASME T-MECH | 9 | Journal | China |
| Lee et al. [74] | 2024 | Science Robotics | 93 | Journal | Germany |
| Cui et al. [75] | 2021 | IEEE RA-L | 99 | Journal | China |
| Gangapurwala et al. [44] | 2022 | IEEE T-RO | 171 | Journal | UK |
| Kim et al. [18] | 2024 | IEEE T-RO | 73 | Journal | South Korea |
| Morimoto et al. [76] | 2023 | JRM | 5 | Journal | Japan |
| Yu et al. [47] | 2025 | IEEE RA-L | 22 | Journal | USA |
| Yang et al. [77] | 2022 | IEEE IROS | 55 | Conference | USA |
| Margolis et al. [78] | 2024 | SAGE IJRR | 286 | Journal | USA |
| Bing et al. [79] | 2020 | Elsevier NN | 82 | Journal | China |
| Xu et al. [80] | 2024 | IEEE ICRA | 14 | Conference | USA |
| Algorithm Type | Core Idea | Typical Application Tasks |
|---|---|---|
| Data-Efficient and Sim-to-Real RL [58,61,62,68] | Focuses on improving data efficiency and transferability from simulation to real-world environments through curriculum learning, domain randomization, and policy fine-tuning. | Sample-efficient learning, sim-to-real adaptation, online fine-tuning |
| Meta-Learning and Morphology Optimization [59,73] | Uses meta-reinforcement learning and co-optimization techniques to jointly adapt control policy and robot morphology for optimal performance. | Adaptive morphology design, cross-task generalization |
| Safety-Constrained and Energy-Aware RL [18,44,64,77] | Integrates safety filters, energy models, and constrained policy optimization to ensure stable and efficient locomotion under dynamic environments. | Energy-efficient locomotion, safe policy learning |
| Multi-Modal and Hybrid Control [47,63,69,74,75] | Combines visual, proprioceptive, and contact modalities or integrates RL with traditional control methods to enhance robustness and adaptability. | Vision–contact fusion, terrain adaptation, robust navigation |
| Task-Specific DRL Applications [43,67,78,80] | Designs specialized DRL frameworks for specific tasks such as recovery, jumping, or high-speed running in structured and unstructured terrains. | Fall recovery, dynamic jumping, agile locomotion |
| Performance Dimension | Evaluation Indicators | DRL-Based Methods (No. of Studies) | Traditional Methods (No. of Studies) | Representative References |
|---|---|---|---|---|
| Stability | Disturbance recovery, fall rate, sustained locomotion | 18/27 | 9/27 | [43,58,71,73] |
| Robustness | Terrain variation, parameter uncertainty, load change | 20/27 | 7/27 | [59,60,68,72] |
| Adaptability | Task generalization, morphology change, damage tolerance | 19/27 | 6/27 | [47,62,73,76] |
| Computational Efficiency | Control frequency, inference latency | 14/27 | 17/27 | [44,58,60] |
| Strategy Category | Technical Principle | Representative Methods and Mechanisms | Performance Improvement Dimensions |
|---|---|---|---|
| Domain Randomization and Multi-Environment Training [43,68,72] | Parameter perturbation and random sampling to expand training distribution | Randomizing friction, mass, inertia, latency, etc.; multi-task parallel training | Generalization, robustness, cross-environment adaptability |
| Model Reduction and Dynamics-Consistent Modeling [58,60] | Using reduced-order models (ROMs) and physics constraints to improve model fidelity | Reduced-order state representation, structured priors, dynamics-consistent optimization | Physical consistency, data efficiency |
| Online Fine-Tuning and Continual Learning [62,67,71] | Real-time policy adaptation after deployment | RL2AC, Keep-on-Learning, adaptive gait optimization | Adaptability, real-time responsiveness |
| Safe Reinforcement Learning and Constrained Optimization [18,44,77] | Integrating safety constraints and energy penalties into RL | Constrained policy gradients, safe reward functions, penalty regularization | Safety, physical feasibility |
| Multi-Modal Fusion and Energy/Terrain-Aware Control [47,64,74] | Fusing vision, inertia, energy consumption, and terrain information | Residual RL + classical control, energy-aware path planning, sensor fusion | Environmental adaptability, energy efficiency, perception robustness |
| Co-Optimization and Meta-Reinforcement Learning [59,73] | Meta-adaptation across tasks and structural co-optimization | Meta-RL, morphology-gait co-optimization | Transferability, system generality |
| Disturbance Type | Observed Performance | Main Weakness | Effective Strategies | Representative Studies |
|---|---|---|---|---|
| Impact/Perturbation | Strong short-term recovery | Limited multi-impact tolerance | Curriculum training, safety RL | [43,44,67,77] |
| Terrain Uncertainty | Adaptive gait modulation | Slippage, loss of balance | Domain randomization, model fusion | [44,64,68] |
| Perception Noise | Partial compensation | Visual dependency | Multimodal sensing, residual RL | [47,63,65,74] |
| Environmental Force | Partial adaptation | Domain overfitting | Physics-informed randomization | [43,58,77] |
| Load Variation | Fast meta-adaptation | Slow convergence | Meta-RL, online fine-tuning | [59,71,73] |
| Robot Morphology | Key Strengths | Main Limitations | Effective Strategies | Representative Studies |
|---|---|---|---|---|
| Biped | Dynamic balance, efficient gait learning | Sensitive to perturbation, poor recovery | Visual feedback + policy regularization | [60,66,74] |
| Quadruped | Strong stability, terrain adaptability | Energy cost under dynamic maneuvers | Residual RL, hybrid reward functions | [44,59,61,67] |
| Hexapod/Octopod | High stability, redundancy | Complex phase coordination | GNN-based DRL, modular policy learning | [18,68,71] |
| Hybrid Morphology | Morphing capability, fast adaptation | Hardware complexity, sim-to-real gap | Meta-RL, policy distillation | [64,72,76] |
| Energy Condition | Average Reward | Stability Score | Energy Efficiency | Recovery Ability |
|---|---|---|---|---|
| Low | 0.75 | 0.70 | 0.88 | 0.72 |
| Medium | 0.85 | 0.78 | 0.82 | 0.80 |
| High | 0.92 | 0.87 | 0.75 | 0.88 |
| Actuator Constraint | Average Reward | Stability Score | Task Efficiency | Recovery Ability |
|---|---|---|---|---|
| Low | 0.90 | 0.88 | 0.85 | 0.87 |
| Medium | 0.82 | 0.80 | 0.78 | 0.80 |
| High | 0.75 | 0.72 | 0.70 | 0.73 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Sun, B.; Mohamed Haris, S.; Ramli, R. A Systematic Review of Deep Reinforcement Learning for Legged Robot Locomotion. Instruments 2026, 10, 8. https://doi.org/10.3390/instruments10010008
Sun B, Mohamed Haris S, Ramli R. A Systematic Review of Deep Reinforcement Learning for Legged Robot Locomotion. Instruments. 2026; 10(1):8. https://doi.org/10.3390/instruments10010008
Chicago/Turabian StyleSun, Bingxiao, Sallehuddin Mohamed Haris, and Rizauddin Ramli. 2026. "A Systematic Review of Deep Reinforcement Learning for Legged Robot Locomotion" Instruments 10, no. 1: 8. https://doi.org/10.3390/instruments10010008
APA StyleSun, B., Mohamed Haris, S., & Ramli, R. (2026). A Systematic Review of Deep Reinforcement Learning for Legged Robot Locomotion. Instruments, 10(1), 8. https://doi.org/10.3390/instruments10010008

