Research on Intelligent Path Planning and Management of X-Type Mecanum-Wheeled Mobile Robot Based on Improved Proximal Policy Optimization–Gated Recurrent Unit Model
Abstract
1. Introduction
- (1)
- A high-dimensional state space integrating environmental perception, goal guidance, and temporal memory is innovatively designed, and a composite reward function considering multiple influencing factors is also innovatively designed. Based on the kinematic modeling of the X-type Mecanum-wheeled mobile robot, the discrete mathematical model is transformed into a numerical simulation platform for interaction with reinforcement learning algorithms through Gym.
- (2)
- By embedding GRUs into the Actor–Critic framework, the PPO path planning model is innovatively improved, thereby obtaining a PPO-GRU-based path planning model with temporal perception capability.
- (3)
- The effectiveness and reliability of the improved PPO-GRU path planning model are comprehensively validated through ablation experiments, comparative experiments with multiple path planning algorithms, analysis of path planning characteristic evolution and path planning capability over multiple rounds in dynamic environments, and performance analysis of path planning considering key real-world phenomena.
2. Construction of the Agent Training Environment for the Mobile Robot
2.1. Kinematic Modeling of X-Type Mecanum-Wheeled Mobile Robot
2.2. Design of the State Space for the Mobile Robot
- (1)
- Polar coordinate representation and rotation invariance of navigation features
- (2)
- Sectorized compression and dimensionality reduction of perceptual information
- (3)
- Stability Design of the Robot’s Dynamic State
- (4)
- Recursive fusion and memory enhancement of the temporal features
2.3. Design of Composite Reward Function Considering Multiple Influencing Factors
- (2)
- (3)
- Dynamic Repulsion and Speed Limit Penalty Based on Velocity Projection () [32]
- (4)
- Heading Angle Dynamic Alignment Reward () [33]
- (5)
- Action Space Smoothness Constraint () [34]
- (6)
- Sparse Event Feedback () [35]
2.4. Construction of an Agent Training Environment Based on Gym
- (1)
- Physics and kinematics modeling module
- (2)
- State-space construction module
- (3)
- Composite reward function module
3. Research on Improvements to the Path Planning Model Based on PPO-GRU
3.1. Construction of the Proximal Policy Optimization (PPO) Path Planning Model
3.1.1. Design of the Actor–Critic Dual-Network Architecture and System Process
- (1)
- Policy Network
- (2)
- Value Network
3.1.2. Policy Updates Based on a Clipping Mechanism
3.1.3. Generalized Advantage Estimation and the Overall Loss Function
3.1.4. Construction of the Path Planning Model Based on the PPO Algorithm
3.2. Research on Improvements to the Path Planning Model Based on PPO-GRU
3.2.1. Gated Recurrent Unit (GRU) Temporal Feature Extraction Mechanism
- (1)
- Reset Gate (rt): This determines how much redundant information to discard from the historical memory, ht−1, aiming to capture abrupt changes characteristics in the environment.
- (2)
- Update Gate (zt): This controls the proportion by which the new candidate memory, , updates the final hidden state, ht, ensuring that the model can maintain long-term locking of the target’s bearing.
3.2.2. Design of the PPO-GRU Integrated Network Architecture
4. Performance Testing of the PPO-GRU Path Planning Model
4.1. Experimental Environment and Parameter Settings
4.2. Research on Path Planning Performance Based on Ablation Experiments
4.2.1. Training Convergence Performance and Stability Analysis
4.2.2. Robustness Testing in Dynamic Environments
4.3. Performance Validation of the PPO-GRU Path Planning Model Based on Comparative Experiments
4.3.1. Path Planning Performance Comparison Based on Classic APF
4.3.2. Path Planning Performance Comparison Based on A* and Dynamic Window Approach
4.4. Robustness Verification of the PPO-GRU Path Planning Model Based on Dynamic Environment Testing
4.4.1. Human-like High-Dynamic Scenario Construction
4.4.2. Spatiotemporal Evolution Analysis of the Dynamic-Obstacle Avoidance Process
4.4.3. Analysis of Kinematic Response Characteristics
4.4.4. Robustness Statistical Verification
4.5. Research on the Performance of the PPO-GRU Path Planning Model Considering Key Real-World Phenomena
4.5.1. Simulation of Key Real-World Phenomena
4.5.2. Research on PPO-GRU Path Planning Considering Key Realistic Phenomena
5. Conclusions and Prospects
5.1. Conclusions
- (1)
- The kinematic modeling of the X-type Mecanum-wheeled mobile robot was established. A state space was designed in polar coordinates, including goal features, environmental perception features, and body motion features. A dynamics-constrained composite reward function based on perception was designed using the concept of the Artificial Potential Field method. An interactive environment for reinforcement learning was built via Gym, providing a reliable numerical simulation platform for the subsequent training of PPO-GRU.
- (2)
- Based on the constructed PPO path planning model, a PPO-GRU-based path planning model with temporal perception ability was obtained by embedding GRUs into the Actor–Critic framework. Combined with the training environment built earlier using Gym, this ultimately provides a complete PPO-GRU reinforcement learning interactive update architecture enabling the robot’s autonomous path planning in dynamic environments.
- (3)
- Ablation experiments were conducted to compare the performance gap between the PPO path planning model and the improved PPO-GRU path planning model. By comparing their convergence curves and performance across multiple rounds in a high-density test environment, the improved PPO-GRU model demonstrated a 41.2% improvement in navigation success rate and an 82.7% reduction in average collision count compared to the PPO model. It performed better in terms of training convergence, stability, success rate, and robustness in dynamic environments than the PPO path planning model.
- (4)
- Comparative experiments were set up to test the path planning and escape capabilities of PPO-GRU and the APF method under identical physical constraints. By observing their performance in classic test scenarios and conducting batch comparative experiments on a test set with various shapes of non-convex obstacles, the planning success rate of the PPO-GRU model reached 96%. Compared to the 42% of the traditional APF method, this represents a 128.6% relative performance improvement, indicating that the improved PPO-GRU model has stronger generalization performance than the traditional APF method.
- (5)
- A human-like high-dynamic scenario was constructed, and the robustness of PPO-GRU was tested through multiple rounds of tests in a high-density obstacle environment with the same parameters. During navigation in the human-like dynamic scenario, the PPO-GRU model demonstrated the ability to avoid dynamic obstacles through its memory capability. In the high-density obstacle environment tests, the PPO-GRU model achieved a task success rate of 96%, with an average linear velocity of 2.57 m/s, reaching 84.5% of the maximum speed configured in the environment. The improved PPO-GRU model not only efficiently completes cruising tasks but also ensures a high planning success rate.
- (6)
- By adding a first-order low-pass filter, setting the friction coefficient, truncating the output of actions exceeding the acceleration limit, and adding normal distribution sampled noise to the speed output, the influence of key real-world phenomena such as inertia, friction, and skidding was simulated. Under the combined effect of inertia, friction, and skidding, the path planning success rate of the PPO-GRU model still remains above 80%, and even in extreme conditions such as ice surfaces, it can still remain above 63%, indicating that the improved PPO-GRU model can be applied to most road surfaces and different load conditions for obstacle avoidance navigation and can also provide navigation references in extreme road conditions.
5.2. Research Limitations and Future Prospects
- (1)
- This study only considered the motion obstacle avoidance of the Mecanum-wheeled mobile robot chassis on a two-dimensional plane. To better adapt to the real world, in future research, a 3D point cloud will be integrated to consider the obstacle avoidance effect in a stereoscopic environment.
- (2)
- The focus of this study is to find a more effective path planning model. In future research, the performance of the model proposed in this paper will be compared with more modern reinforcement learning architectures to further optimize and improve the proposed model, thereby achieving iterative upgrades in the research and obtaining more accurate path planning algorithms.
- (3)
- This study did not conduct in-depth research on the deployment issue of the proposed model. Due to the extremely high cost and risks of training reinforcement learning models in reality (initial training is bound to result in collisions), for migration to real machines, the next step is to use the sim2real technology for targeted training to adapt to real environments. Through techniques such as domain randomization (randomly perturbing various parameters in the environment, including speed, sensor readings, target coordinates, etc., to simulate possible disturbances in reality) and other technologies, the intelligent agent model can be made to face such relatively simple models of real environments, and end-to-end learning can be used to reinforce the existing model during real machine deployment.
- (4)
- Real phenomena such as wheel slippage, changes in surface friction, and complex inertial dynamics will all affect path planning. In future research, rigorous, complete, and rich experiments will be conducted to study the path planning algorithms under dynamic conditions such as wheel slippage, changes in surface friction, and complex inertial dynamics.
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Liu, Y.S.; Xu, Z.X.; He, N.; He, Y.L. Reinforcement learning-driven parameter tuning for mobile robot’s predictive control. J. Harbin Inst. Technol. 2026, 1–12. Available online: https://link.cnki.net/urlid/23.1235.T.20260129.1121.007 (accessed on 17 February 2026).
- Wang, Y.; Ye, Y.Y.; Zhong, W.; Gao, B.L.; Mu, C.Z.; Zhao, N. Micro-Platform Verification for LiDAR SLAM-Based Navigation of Mecanum-Wheeled Robot in Warehouse Environment. World Electr. Veh. J. 2025, 16, 571. [Google Scholar] [CrossRef]
- Lin, S.; Dai, J.; Song, Y.F.; Wang, H.G.; Yuan, B.B. Analysis of Curved Surface Motion Characteristics of Wheeled Mobile Robots with Reconfigurable Trunks. J. Mech. Eng. 2026, 1–14. Available online: https://link.cnki.net/urlid/11.2187.TH.20260106.1654.006 (accessed on 17 February 2026).
- He, C.; Wu, D.; Chen, K.; Liu, F.; Fan, N. Analysis of the Mecanum wheel arrangement of an omnidirectional vehicle. Part C J. Mech. Eng. Sci. 2019, 233, 12. [Google Scholar] [CrossRef]
- Dou, L.; Gao, Y.; He, Z.H.; Lv, A.; Ding, F.P. Simultaneous Localization and Mapping (SLAM) and Path Planning Algorithms Based on RadarVisual Sensor Fusion in Unknown Environments. Comput. Eng. Appl. 2026, 1–14. Available online: https://link.cnki.net/urlid/11.2127.tp.20260130.1042.004 (accessed on 17 February 2026).
- Setiadilaga, O.; Cahyadi, A.; Ataka, A. Mecanum-Wheeled Robot Control Based on Deep Reinforcement Learning. In 2023 15th International Conference on Information Technology and Electrical Engineering (ICITEE), Chiang Mai, Thailand, 2023; IEEE: New York, NY, USA, 2003; pp. 25–30. [Google Scholar] [CrossRef]
- Nguyen, C.T.; Pham, H.-A. Geometric Optimization Frameworks in Mobile Robot Path Planning. IEEE Access 2025, 13, 147127–147162. [Google Scholar] [CrossRef]
- Zhang, X.; Wu, W.; Li, X. A modified fruit fly optimization algorithm to active disturbance rejection control parameters tuning for trajectory tracking of omnidirectional mobile robotic chassis. Soft Comput. 2025, 29, 4401–4421. [Google Scholar] [CrossRef]
- Pérez-Juárez, J.G.; García-Martínez, J.R.; Santiago, A.M.; Cruz-Miguel, E.E.; Olmedo-García, L.F.; Barra-Vázquez, O.A.; Rojas-Hernández, M.A. Kinematic Fuzzy Logic-Based Controller for Trajectory Tracking of Wheeled Mobile Robots in Virtual Environments. Symmetry 2025, 17, 301. [Google Scholar] [CrossRef]
- Abut, T.; Salkim, E. Trajectory Tracking of a Mobile Robot with GWO-Based Type II Fuzzy Logic Controller. Bitlis Eren Üniversitesi Fen Balmier Derg. 2025, 14, 1523–1551. [Google Scholar] [CrossRef]
- Villalba-Aguilera, E.; Blesa, J.; Ponsa, P. Model-Based Predictive Control for Position and Orientation Tracking in a Multilayer Architecture for a Three-Wheeled Omnidirectional Mobile Robot. Robotics 2025, 14, 72. [Google Scholar] [CrossRef]
- Jia, J.; Xing, X.; Chang, D.E. GRU-Attention based TD3 Network for Mobile Robot Navigation. In 2022 22nd International Conference on Control, Automation and Systems (ICCAS), Jeju, Republic of Korea, 2022; IEEE: New York, NY, USA, 2022; pp. 1642–1647. [Google Scholar] [CrossRef]
- Qin, H.; Qiao, B.; Wu, W.; Deng, Y. A Path Planning Algorithm Based on Deep Reinforcement Learning for Mobile Robots in Unknown Environment. In 2022 IEEE 5th Advanced Information Management, Communicates, Electronic and Automation Control Conference (IMCEC), Chongqing, China, 2022; IEEE: New York, NY, USA, 2022; pp. 1661–1666. [Google Scholar] [CrossRef]
- Xu, H.; Terakawa, T.; Komori, M. Deep-reinforcement-learning-based trajectory tracking control for slidable-wheel omnidirectional mobile robot. Bull. JSME J. Adv. Mech. Des. Syst. Manuf. 2025, 19, JAMDSM0031. [Google Scholar] [CrossRef]
- Bie, T.; Zhu, X.; Li, X.; Ruan, X. A Mobile Robot Path Planning Method Based on Safe Pathfinding Guidance. In 2021 33rd Chinese Control and Decision Conference (CCDC), Kunming, China, 2021; IEEE: New York, NY, USA, 2021; pp. 3297–3303. [Google Scholar] [CrossRef]
- Li, H.; Zhong, P.; Liu, L.; Wang, X.; Liu, M.; Yuan, J. Robot Dynamic Path Planning Based on Prioritized Experience Replay and LSTM Network. IEEE Access 2025, 13, 22283–22299. [Google Scholar] [CrossRef]
- Zhou, S. PPO-Based Mobile Robot Path Planning with Dense Bootstrap Reward in Partially Observable Environments. In 2024 21st International Computer Conference on Wavelet Active Media Technology and Information Processing (ICCWAMTIP), Chengdu, China, 2024; IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef]
- Trojnacki, M. Tracking Control of a Four-Wheeled Skid-Steered Robot with Slip Compensation and Application of the Drive Unit Model. Electronics 2025, 14, 444. [Google Scholar] [CrossRef]
- Alcayaga, J.M.; Menéndez, O.A.; Torres-Torriti, M.A.; Vásconez, J.P.; Arévalo-Ramirez, T.; Romo, A.J.P. LSTM-Enhanced Deep Reinforcement Learning for Robust Trajectory Tracking Control of Skid-Steer Mobile Robots Under Terra-Mechanical Constraints. Robotics 2025, 14, 74. [Google Scholar] [CrossRef]
- Xing, X.; Ding, H.; Liang, Z.; Li, B.; Yang, Z. Robot path planner based on deep reinforcement learning and the seeker optimization algorithm. Mechatronics 2022, 88, 102918. [Google Scholar] [CrossRef]
- Magnusson, M.; Lilienthal, A.; Duckett, T. Scan registration for autonomous mining vehicles using 3D-NDT. J. Field Robot. 2007, 24, 803–827. [Google Scholar] [CrossRef]
- Wang, Q.; Wei, L.S. AGV dynamic obstacle avoidance path planning algorithm based on improved HLO and dynamic window. J. Electron. Meas. Instrum. 2025, 39, 213–221. [Google Scholar] [CrossRef]
- Jiang, H.; Fang, W.; Xu, T.F.; Chen, F.; Zhou, L.; Deng, Q. Optimal indoor evacuation path-planning model based on Dijkstra’s algorithm. J. Tsinghua Univ. (Nat. Sci. Ed.) 2025, 65, 742–749. [Google Scholar] [CrossRef]
- Wu, X.; Li, Y.; Zang, T.G.; Meng, Z.X.; Chen, J.Z.; Wang, C.T.; Xing, L.Y.W. Combined Path Planning Based on Voronoi Skeleton for Mobile Robots. J. Mech. Eng. 2025, 61, 165–177. [Google Scholar] [CrossRef]
- Cai, Y.; Du, P. Path planning of unmanned ground vehicle based on balanced whale optimization algorithm. Control. Decis. 2021, 36, 2647–2655. [Google Scholar] [CrossRef]
- Xiao, J.Z.; Yu, X.; Zhou, G.; Sun, K.; Zhou, Z. An improved ant colony algorithm for indoor AGV path planning. Chin. J. Sci. Instrum. 2022, 43, 277–285. [Google Scholar] [CrossRef]
- Li, Y.; Liao, Z.H.; Li, M.H. An Improved Ant Colony Optimization Algorithm Based on Reinforcement Learning for Mobile Robot Path Planning. Comput. Eng. Appl. 2026. Available online: https://link.cnki.net/urlid/11.2127.TP.20251205.1505.008 (accessed on 17 February 2026).
- Xie, M.; Yu, W.; Chen, M. End-to-end Autonomous Navigation Approach for Multiple mobile Robots Integrating Attention Mechanism and Velocity Obstacle Method. Robot 2026. [Google Scholar] [CrossRef]
- de Heuvel, J.; Zeng, X.; Shi, W.; Sethuraman, T.; Bennewitz, M. Spatiotemporal Attention Enhances Lidar-Based Robot Navigation in Dynamic Environments. arXiv 2023, arXiv:2310.19670. [Google Scholar] [CrossRef]
- Jiang, H.; Li, S.; Zhang, J.; Zhu, Y.; Xu, X.; Liu, D. Efficient state representation with artificial potential fields for reinforcement learning. Complex Intell. Syst. 2023, 9, 4911–4922. [Google Scholar] [CrossRef]
- Li, P.; Wang, Y.; Gao, Z. Path Planning of Mobile Robot Based on Improved TD3 Algorithm. In 2022 IEEE International Conference on Mechatronics and Automation (ICMA), Guilin, China, 2022; IEEE: New York, NY, USA, 2022; pp. 715–720. [Google Scholar] [CrossRef]
- Tao, Y.; Li, M.; Cao, X.; Lu, P. Mobile Robot Collision Avoidance Based on Deep Reinforcement Learning With Motion Constraints. IEEE Trans. Intell. Veh. 2025, 10, 2163–2173. [Google Scholar] [CrossRef]
- Yang, X.; Wang, Q.; Li, J.; Jiang, X. NM-TD3: A Hybrid Noise-Driven TD3 Algorithm With Long-Term Reward Propagation for Mobile Robot Path Planning. IEEE Access 2025, 13, 149921–149932. [Google Scholar] [CrossRef]
- Han, C.; Park, S.; Woo, J. Robust Collision Avoidance for ASVs Using Deep Reinforcement Learning with Sim2Real Methods in Static Obstacle Environments. J. Mar. Sci. Eng. 2025, 13, 1727. [Google Scholar] [CrossRef]
- Park, M.; Park, C.; Kwon, N.K. Autonomous Driving of Mobile Robots in Dynamic Environments Based on Deep Deterministic Policy Gradient: Reward Shaping and Hindsight Experience Replay. Biomimetics 2024, 9, 51. [Google Scholar] [CrossRef]
- Ren, J.; Zeng, Y.; Zhou, S.; Zhang, Y. An Experimental Study on State Representation Extraction for Vision-Based Deep Reinforcement Learning. Appl. Sci. 2021, 11, 10337. [Google Scholar] [CrossRef]
- Rojas, M.; Hermosilla, G.; Yunge, D.; Farias, G. An Easy to Use Deep Reinforcement Learning Library for AI Mobile Robots in Isaac Sim. Appl. Sci. 2022, 12, 8429. [Google Scholar] [CrossRef]
- Guo, B.; Wang, G.; Chen, Y.; Gao, Y.; Xie, Q. Risk-Aware Reinforcement Learning with Dynamic Safety Filter for Collision Risk Mitigation in Mobile Robot Navigation. Sensors 2025, 25, 5488. [Google Scholar] [CrossRef]
- Cheng, W.-C.; Ni, Z.; Zhong, X.; Wei, M. Autonomous Robot Goal Seeking and Collision Avoidance in the Physical World: An Automated Learning and Evaluation Framework Based on the PPO Method. Appl. Sci. 2024, 14, 11020. [Google Scholar] [CrossRef]
- Zhang, Q.; Ma, W.; Zheng, Q.; Zhai, X.; Zhang, W.; Zhang, T.; Wang, S. Path Planning of Mobile Robot in Dynamic Obstacle Avoidance Environment Based on Deep Reinforcement Learning. IEEE Access 2024, 12, 189136–189152. [Google Scholar] [CrossRef]
- Cho, K.; van Merriënboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation. In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Doha, Qatar, 25–29 October 2014. [Google Scholar] [CrossRef]
- Jiang, W.; Liu, J.; Wang, W. Global Path Planning for Land-Air Amphibious Biomimetic Robot Based on Improved PPO. Biomimetics 2026, 11, 25. [Google Scholar] [CrossRef] [PubMed]
- Nan, Z.; Nam, H. Multimodal Feature Fusion for Deep Reinforcement Learning-Based Mobile Robot Navigation. In 2025 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), Fukuoka, Japan, 2025; IEEE: New York, NY, USA, 2025; pp. 0893–0897. [Google Scholar] [CrossRef]
- Azizi, M.R.; Rastegarpanah, A.; Stolkin, R. Motion Planning and Control of an Omnidirectional Mobile Robot in Dynamic Environments. Robotics 2021, 10, 48. [Google Scholar] [CrossRef]
- Sun, Z.; Hu, S.; Miao, X.; Chen, B.; Zheng, J.; Man, Z.; Wang, T. Obstacle-avoidance trajectory planning and sliding mode-based tracking control of an omnidirectional mobile robot. Front. Control. Eng. 2023, 4, 1135258. [Google Scholar] [CrossRef]
- Wu, D.; Wei, L.; Wang, G.; Tian, L.; Dai, G. APF-IRRT*: An Improved Informed Rapidly-Exploring Random Trees-Star Algorithm by Introducing Artificial Potential Field Method for Mobile Robot Path Planning. Appl. Sci. 2022, 12, 10905. [Google Scholar] [CrossRef]
- Sun, Z.; Zhang, T.; Zhu, H.; Ma, T.; Bao, X.; Zhang, X. Path Planning for Mobile Robot Based on the Fusion Algorithm of Improved A* and APF. In 2024 6th International Symposium on Robotics & Intelligent Manufacturing Technology (ISRIMT), Changzhou, China, 2024; IEEE: New York, NY, USA, 2024; pp. 113–118. [Google Scholar] [CrossRef]
- Wang, X.; Li, G.; Bian, Z. Research on APF-Dijkstra Path Planning Fusion Algorithm Based on Steering Model and Volume Constraints. Algorithms 2025, 18, 403. [Google Scholar] [CrossRef]
- Kobayashi, M.; Zushii, H.; Nakamura, T.; Motoi, N. Local Path Planning: Dynamic Window Approach With Q-Learning Considering Congestion Environments for Mobile Robot. IEEE Access 2023, 11, 96733–96742. [Google Scholar] [CrossRef]
- Votion, J.; Cao, Y. Diversity-Based Cooperative Multivehicle Path Planning for Risk Management in Costmap Environments. IEEE Trans. Ind. Electron. 2019, 66, 6117–61274. [Google Scholar] [CrossRef]
























| Measurement radius | 20 m |
| Scanning frequency | 12 Hz |
| Sampling frequency | 20,000 Hz |
| Output content | Angle, distance |
| Angle resolution | 0.22° |
| Anti-environmental light intensity | 100 Klux |
| Category | Project | Configuration Details |
|---|---|---|
| Hardware | CPU | 13th Gen Intel(R) Core(TM) i5-13600KF |
| GPU | NVIDIA GeForce RTX 4070 SUPER (12 GB VRAM) | |
| AM | 32 GB DDR5 | |
| Software | Operating system | Microsoft Windows 11 professional edition |
| Programming language | Python 3.10/PyTorch 2.x |
| Parameters | Symbol | Value |
|---|---|---|
| Map size | 50.0 × 50.0 | |
| Number of target points | 3 | |
| Number of obstacles | 4 | |
| Obstacle moving speed | 1.4 m/s | |
| Obstacle radius | 1.5~2.5 m | |
| Maximum number of steps per round | 1500 | |
| Number of parallel environments | 20 | |
| Total number of training rounds | 5000 |
| Parameter Item | Symbol | Value |
|---|---|---|
| Policy network learning rate (Actor LR) | ||
| Value network learning rate (Critic LR) | ||
| Hidden layer dimension (Hidden Dim) | ||
| Update frequency (Steps per Update) | ||
| Batch size (Batch Size) | ||
| Discount factor/GAE parameter | ||
| Clipping coefficient/Entropy coefficient |
| Category | Parameter Item | Symbol/Name | Value/Configuration |
|---|---|---|---|
| Task event | Final goal reward | +1000.0 | |
| Stage goal reward | +400.0 | ||
| Task failure penalty | −1500.0 | ||
| Navigation guidance | Distance guidance coefficient | 1.0 | |
| Course alignment weight | 1.0 | ||
| Time cost | −0.5 | ||
| Obstacle avoidance safety | Alert threshold coefficient | 0.70 | |
| Danger threshold coefficient | 0.20 | ||
| Flexibility/rigidity repulsion force weight | 4.0/30.0 | ||
| Penalty points in the absolute penalty area | 150.0 | ||
| Speeding penalty coefficient | 20.0 | ||
| Control constraints | Action smoothness coefficient | 0.05 | |
| Yaw damping penalty | 0.01 |
| Model Framework | Average Round Reward | Navigation Success Rate | Average Number of Collisions |
|---|---|---|---|
| PPO path planning model | 1520.4 | 68.0% | 1.85 |
| PPO-GRU path planning model | 2510.6 | 96.0% | 0.32 |
| Optimization range | +65.1% | +41.2% | −82.7% |
| Performance Index | Statistic Value |
|---|---|
| Task Success Rate (SR) | 96.0% |
| Average linear velocity | 2.75 m/s |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
An, N.; Yang, S.; Kong, S. Research on Intelligent Path Planning and Management of X-Type Mecanum-Wheeled Mobile Robot Based on Improved Proximal Policy Optimization–Gated Recurrent Unit Model. Machines 2026, 14, 382. https://doi.org/10.3390/machines14040382
An N, Yang S, Kong S. Research on Intelligent Path Planning and Management of X-Type Mecanum-Wheeled Mobile Robot Based on Improved Proximal Policy Optimization–Gated Recurrent Unit Model. Machines. 2026; 14(4):382. https://doi.org/10.3390/machines14040382
Chicago/Turabian StyleAn, Ning, Songlin Yang, and Shihan Kong. 2026. "Research on Intelligent Path Planning and Management of X-Type Mecanum-Wheeled Mobile Robot Based on Improved Proximal Policy Optimization–Gated Recurrent Unit Model" Machines 14, no. 4: 382. https://doi.org/10.3390/machines14040382
APA StyleAn, N., Yang, S., & Kong, S. (2026). Research on Intelligent Path Planning and Management of X-Type Mecanum-Wheeled Mobile Robot Based on Improved Proximal Policy Optimization–Gated Recurrent Unit Model. Machines, 14(4), 382. https://doi.org/10.3390/machines14040382

