A Hybrid RRT–PPO Framework for Leg-Based Object Manipulation of Quadruped Robots
Abstract
1. Introduction
- A hybrid architecture integrating RRT for global path planning and PPO for local manipulation control, leveraging the strengths of model-based and learning-based approaches.
- A Bézier curve-parameterized leg trajectory representation that enables the RL agent to generate smooth, controllable pushing motions through a compact action space.
- A reward function design incorporating position error, angular error, and stability penalties that effectively guides the learning of box-pushing behaviors.
- A comprehensive robustness analysis demonstrating that the trained agent continues to push objects with masses up to eight times the training mass, with other physical parameters (shape, friction, and surface conditions) held constant.
- An extension of the learned straight-line pushing to curved paths: applying a fixed, asymmetric Bézier pushing action induces a consistent box rotation at a measurable mean rotational speed, and the resulting approximately circular arcs are composed to follow curved trajectories. Dubins path theory [13] provides the conceptual basis for composing straight segments and circular arcs.
2. Related Work
2.1. Sampling-Based Motion Planning with RRT
2.2. Reinforcement Learning for Robotic Manipulation
2.3. Hybrid Motion Planning and RL Approaches
3. Methodology
3.1. Robot Model
3.1.1. Coordinate Systems and Kinematics
3.1.2. Inverse Kinematics
3.1.3. Static Stability
3.1.4. Dynamic Stability During Pushing
3.2. RRT-Based Motion Planning
| Algorithm 1 RRT Path Planning for the Quadruped Robot |
|
3.2.1. Path Smoothing
3.2.2. Adaptation to Robot Geometry
3.3. Reinforcement Learning for Object Manipulation
3.3.1. Background
3.3.2. Proximal Policy Optimization
3.3.3. Bézier Curve Leg Trajectories
3.3.4. Reward Function Design
3.3.5. Episode Definition, Termination, and Reset
4. Implementation
4.1. Simulation Environment
4.2. Network Architecture
4.3. State Transition Dynamics
4.4. Observation and Action Spaces
5. Results
5.1. Motion Planning with RRT
5.2. Straight-Line Box Pushing
5.2.1. Fixed Entry Value
5.2.2. Variable Entry Value
5.3. Generalization Analysis
5.4. Straight-Line Pushing Illustration
5.5. Circular Path Pushing
5.6. Dubins Path Extension
6. Discussion
7. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CoM | center of mass |
| DDPG | deep deterministic policy gradient |
| DOF | degrees of freedom |
| DRL | deep reinforcement learning |
| IMU | inertial measurement unit |
| PPO | proximal policy optimization |
| RL | reinforcement learning |
| RRT | rapidly exploring random tree |
| SAC | soft actor–critic |
References
- Cong, Q.; Shi, X.; Wang, J.; Xiong, Y.; Su, B.; Xu, W.; Liu, H.; Zhou, K.; Jiang, L.; Tian, W. Stability Study and Simulation of Quadruped Robots with Variable Parameters. Appl. Bionics Biomech. 2022, 2022, 9968042. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, S.; Liu, M.; Yin, Y.; Rong, X.; Li, Y.; Hua, Z. Static Gait Planning Method for Quadruped Robot Walking on Unknown Rough Terrain. IEEE Access 2019, 7, 177651–177660. [Google Scholar] [CrossRef] [Scilit]
- Xin, G.; Wolfslag, W.; Lin, H.C.; Tiseo, C.; Mistry, M. An Optimization-Based Locomotion Controller for Quadruped Robots Leveraging Cartesian Impedance Control. Front. Robot. AI 2020, 7, 48. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, H.; Yu, W.; Zhang, T.; Wensing, P.M. Zero-Shot Retargeting of Learned Quadruped Locomotion Policies Using Hybrid Kinodynamic Model Predictive Control. In Proceedings of the 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Kyoto, Japan, 23–27 October 2022; pp. 11971–11977. [Google Scholar] [CrossRef] [Scilit]
- Shao, Y.; Jin, Y.; Liu, X.; He, W.; Wang, H.; Yang, W. Learning Free Gait Transition for Quadruped Robots Via Phase-Guided Controller. IEEE Robot. Autom. Lett. 2022, 7, 1230–1237. [Google Scholar] [CrossRef] [Scilit]
- LaValle, S.M. Rapidly-Exploring Random Trees: A New Tool for Path Planning; Technical report; Computer Science Department, Iowa State University: Ames, IA, USA, 1998. [Google Scholar]
- Karaman, S.; Frazzoli, E. Sampling-Based Algorithms for Optimal Motion Planning. Int. J. Robot. Res. 2011, 30, 846–894. [Google Scholar] [CrossRef] [Scilit]
- Gammell, J.D.; Srinivasa, S.S.; Barfoot, T.D. Informed RRT*: Optimal Sampling-Based Path Planning Focused via Direct Sampling of an Admissible Ellipsoidal Heuristic. In Proceedings of the 2014 IEEE/RSJ International Conference on Intelligent Robots and Systems, Chicago, IL, USA, 14–18 September 2014; pp. 2997–3004. [Google Scholar] [CrossRef] [Scilit]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018. [Google Scholar]
- Kaelbling, L.P.; Littman, M.L.; Moore, A.W. Reinforcement Learning: A Survey. J. Artif. Intell. Res. 1996, 4, 237–285. [Google Scholar] [CrossRef] [Scilit]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- OpenAI. Solving Rubik’s Cube with a Robot Hand. arXiv 2019, arXiv:1910.07113. [Google Scholar]
- Dubins, L.E. On Curves of Minimal Length with a Constraint on Average Curvature, and with Prescribed Initial and Terminal Positions and Tangents. Am. J. Math. 1957, 79, 497–516. [Google Scholar] [CrossRef] [Scilit]
- Ferguson, D.; Stentz, A. Anytime RRTs. In Proceedings of the 2006 IEEE/RSJ International Conference on Intelligent Robots and Systems, Beijing, China, 9–15 October 2006; pp. 5369–5375. [Google Scholar] [CrossRef] [Scilit]
- Kuffner, J.J.; LaValle, S.M. RRT-Connect: An Efficient Approach to Single-Query Path Planning. In Proceedings of the 2000 ICRA. IEEE International Conference on Robotics and Automation, San Francisco, CA, USA, 24–28 April 2000; Volume 2, pp. 995–1001. [Google Scholar] [CrossRef] [Scilit]
- Kober, J.; Bagnell, J.A.; Peters, J. Reinforcement Learning in Robotics: A Survey. Int. J. Robot. Res. 2013, 32, 1238–1274. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-Level Control Through Deep Reinforcement Learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rajeswaran, A.; Kumar, V.; Gupta, A.; Vezzani, G.; Schulman, J.; Todorov, E.; Levine, S. Learning Complex Dexterous Manipulation with Deep Reinforcement Learning and Demonstrations. In Proceedings of the Robotics: Science and Systems (RSS), Pittsburgh, PA, USA, 26–30 June 2018. [Google Scholar] [CrossRef] [Scilit]
- Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous Control with Deep Reinforcement Learning. arXiv 2019, arXiv:1509.02971. [Google Scholar]
- Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P.; et al. Soft Actor-Critic Algorithms and Applications. arXiv 2019, arXiv:1812.05905. [Google Scholar]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. arXiv 2018, arXiv:1801.01290. [Google Scholar]
- Han, J. From PID to Active Disturbance Rejection Control. IEEE Trans. Ind. Electron. 2009, 56, 900–906. [Google Scholar] [CrossRef] [Scilit]
- Vamvoudakis, K.G.; Lewis, F.L. Online Actor–Critic Algorithm to Solve the Continuous-Time Infinite Horizon Optimal Control Problem. Automatica 2010, 46, 878–888. [Google Scholar] [CrossRef] [Scilit]
- Chiang, H.T.L.; Hsu, J.; Fišer, M.; Xiao, L.; Faust, A. RL-RRT: Kinodynamic Motion Planning via Learning Reachability Estimators from RL Policies. IEEE Robot. Autom. Lett. 2019, 4, 4298–4305. [Google Scholar] [CrossRef] [Scilit]
- Faust, A.; Oslund, K.; Ramirez, O.; Francis, A.; Tapia, L.; Fiser, M.; Davidson, J. PRM-RL: Long-Range Robotic Navigation Tasks by Combining Reinforcement Learning and Sampling-Based Planning. In Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia, 21–25 May 2018; pp. 5113–5120. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Ji, C.; Cai, Y.; Yan, T.; Su, B. Deep Reinforcement Learning in Autonomous Car Path Planning and Control: A Survey. arXiv 2024, arXiv:2404.00340. [Google Scholar] [CrossRef] [Scilit]
- Zeng, A.; Song, S.; Welker, S.; Lee, J.; Rodriguez, A.; Funkhouser, T. Learning Synergies between Pushing and Grasping with Self-Supervised Deep Reinforcement Learning. In Proceedings of the 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Madrid, Spain, 1–5 October 2018; pp. 4238–4245. [Google Scholar] [CrossRef] [Scilit]
- LaValle, S.M.; Kuffner, J.J. Randomized Kinodynamic Planning. Int. J. Robot. Res. 2001, 20, 378–400. [Google Scholar] [CrossRef] [Scilit]
- Lynch, K.M.; Park, F.C. Modern Robotics: Mechanics, Planning, and Control; Cambridge University Press: Cambridge, UK, 2017. [Google Scholar]
- Craig, J.J. Introduction to Robotics: Mechanics and Control, 3rd ed.; Pearson/Prentice Hall: Upper Saddle River, NJ, USA, 2005. [Google Scholar]
- McGhee, R.B.; Frank, A.A. On the Stability Properties of Quadruped Creeping Gaits. Math. Biosci. 1968, 3, 331–351. [Google Scholar] [CrossRef]
- Pongas, D.; Mistry, M.; Schaal, S. A Robust Quadruped Walking Gait for Traversing Rough Terrain. In Proceedings of the 2007 IEEE International Conference on Robotics and Automation, Rome, Italy, 10–14 April 2007; pp. 1474–1479. [Google Scholar] [CrossRef] [Scilit]
- Queiroz, C.; Gonçalves, N.; Menezes, P. A Study on Static Gaits for a Four-Leg Robot. In Proceedings of the UKACC International Conference on Control (CONTROL 2000), Cambridge, UK, 4–7 September 2000; pp. 1–6. [Google Scholar]
- Ding, L.; Wang, G.; Gao, H.; Liu, G.; Yang, H.; Deng, Z. Footstep Planning for Hexapod Robots Based on 3D Quasi-Static Equilibrium Support Region. J. Intell. Robot. Syst. 2021, 103, 25. [Google Scholar] [CrossRef] [Scilit]
- Kajita, S.; Kanehiro, F.; Kaneko, K.; Fujiwara, K.; Harada, K.; Yokoi, K.; Hirukawa, H. Biped Walking Pattern Generation by Using Preview Control of Zero-Moment Point. In Proceedings of the 2003 IEEE International Conference on Robotics and Automation, Taipei, Taiwan, 14–19 September 2003; Volume 2, pp. 1620–1626. [Google Scholar] [CrossRef] [Scilit]
- Sutton, R.S.; McAllester, D.A.; Singh, S.P.; Mansour, Y. Policy Gradient Methods for Reinforcement Learning with Function Approximation. Adv. Neural Inf. Process. Syst. 2000, 12, 1057–1063. [Google Scholar]
- Silver, D.; Lever, G.; Heess, N.; Degris, T.; Wierstra, D.; Riedmiller, M. Deterministic Policy Gradient Algorithms. In Proceedings of the 31st International Conference on Machine Learning, Beijing, China, 22–24 June 2014. [Google Scholar]
- Rohmer, E.; Singh, S.P.N.; Freese, M. V-REP: A Versatile and Scalable Robot Simulation Framework. In Proceedings of the 2013 IEEE/RSJ International Conference on Intelligent Robots and Systems, Tokyo, Japan, 3–7 November 2013; pp. 1321–1326. [Google Scholar] [CrossRef] [Scilit]
- Hintjens, P. ZeroMQ: Messaging for Many Applications; O’Reilly Media, Inc.: Sebastopol, CA, USA, 2013. [Google Scholar]
- Raffin, A.; Hill, A.; Gleave, A.; Kanervisto, A.; Ernestus, M.; Dormann, N. Stable-Baselines3: Reliable Reinforcement Learning Implementations. J. Mach. Learn. Res. 2021, 22, 1–8. [Google Scholar]
- Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; Zaremba, W. OpenAI Gym. arXiv 2016, arXiv:1606.01540. [Google Scholar]
- Arulkumaran, K.; Deisenroth, M.P.; Brundage, M.; Bharath, A.A. Deep Reinforcement Learning: A Brief Survey. IEEE Signal Process. Mag. 2017, 34, 26–38. [Google Scholar] [CrossRef] [Scilit]
- Li, Y. Deep Reinforcement Learning: An Overview. arXiv 2018, arXiv:1701.07274. [Google Scholar]





















| Joint | Lower Limit [deg] | Upper Limit [deg] |
|---|---|---|
| Hip x | 45 | |
| Hip y | 100 | |
| Knee | 130 |
| Parameter | Symbol | Description |
|---|---|---|
| Displacement x | Horizontal displacement in the x-direction | |
| Displacement y | Horizontal displacement in the y-direction | |
| Step height | h | Maximum height of the parabolic swing arc |
| Swing duration | Time allocated for the swing phase |
| Parameter | Description |
|---|---|
| (start point) | Current position of the pushing leg tip |
| (end point) | Target contact point on the box surface |
| (control points) | Intermediate control points; each positioned by one scalar action |
| Scalar action positioning control point | |
| Scalar action positioning control point |
| Hyperparameter | Value | Description |
|---|---|---|
| Learning rate () | 0.001 | Gradient descent step size |
| Batch size | 512 | Samples per gradient update |
| Discount factor () | 0.99 | Future reward weighting |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Attias, Y.; Giladi, C. A Hybrid RRT–PPO Framework for Leg-Based Object Manipulation of Quadruped Robots. Machines 2026, 14, 773. https://doi.org/10.3390/machines14070773
Attias Y, Giladi C. A Hybrid RRT–PPO Framework for Leg-Based Object Manipulation of Quadruped Robots. Machines. 2026; 14(7):773. https://doi.org/10.3390/machines14070773
Chicago/Turabian StyleAttias, Yogev, and Chen Giladi. 2026. "A Hybrid RRT–PPO Framework for Leg-Based Object Manipulation of Quadruped Robots" Machines 14, no. 7: 773. https://doi.org/10.3390/machines14070773
APA StyleAttias, Y., & Giladi, C. (2026). A Hybrid RRT–PPO Framework for Leg-Based Object Manipulation of Quadruped Robots. Machines, 14(7), 773. https://doi.org/10.3390/machines14070773

