Adaptive Traversability Policy Optimization for an Unmanned Articulated Road Roller on Slippery, Geometrically Irregular Terrains
Abstract
1. Introduction
1.1. Low Level: Driveline and Traction Management
1.2. Mid-Level: Motion Planning and Search
1.3. High Level: Learning-Based Decision-Making
- Reward design. The task reward is naturally sparse and deceptive, easily causing credit-assignment difficulties and training stagnation. As shown by Devidze, R. et al. [14], a lack of effective reward shaping can markedly slow learning or even prevent convergence.
- Exploration–exploitation trade-off. Maximum-entropy RL algorithms (e.g., SAC [15]) encourage exploration via a temperature parameter. To improve adaptivity, several works learn or schedule this parameter, or otherwise modulate exploration, including meta-gradient temperature tuning and target-entropy annealing in SAC, value-conditional state-entropy exploration, uncertainty-driven bonuses for generalization, and adaptive entropy-regularization frameworks in multi-agent RL [16,17,18,19,20]. While effective on standard benchmarks, these schemes typically rely on a single generic signal (state entropy or value uncertainty) and do not explicitly incorporate a terrain’s physical difficulty or task-progress information into the entropy schedule, which limits their task-specificity in strongly heterogeneous, wet–geometry-coupled terrains.
- A component-based multi-physics simulation platform is developed for an articulated road roller operating on parameterized mud-pit terrains. The platform closely matches the real vehicle and generates a broad family of scenarios by varying adhesion, pit depth, and geometric shape, providing a physically grounded benchmark and training/test sets for traversability studies.
- A TICR function is designed that jointly encodes task-progress signals and physically grounded terrain attributes (adhesion, sinkage, and slope). This goes beyond existing terrain-aware reward shaping by explicitly targeting mud-pit traversability and alleviating credit-assignment difficulties under sparse and deceptive rewards.
- A context-aware adaptive entropy-regularization mechanism is developed, in which the policy temperature is an explicit state-dependent function of terrain physical difficulty, task-execution efficacy, and epistemic value uncertainty. This contrasts with meta-gradient, state-entropy, and uncertainty-only schemes that rely on a single generic signal and ignore task-specific terrain context, enabling more robust exploration–exploitation balancing on strongly heterogeneous, wet–geometry-coupled terrains.
2. Simulation Platform
2.1. Representative Case
- The low-adhesion property inherent to wet soil.
- The geometric irregularity is characterized by the localized, steep slopes formed by the depression.
2.2. Construction of the Simulation Platform
2.2.1. Overview of the Simulation Method
- Structured physical representation and high reusability. Each component encapsulates its internal constitutive laws and interacts solely via standardized power ports. This modular structure renders the assembly and reconfiguration of subsystems (e.g., articulated frame, drum/wheels) intuitive, akin to physical prototyping. Crucially, changes in system topology or parameters do not necessitate rewriting low-level solver code, yielding high flexibility.
- Unified solution of strongly coupled multi-physics. All components are numerically integrated within a single DAE solver. This intrinsically supports stable and consistent simulation of highly nonlinear, strongly coupled phenomena. Examples include valve-controlled hydraulics, line compressibility, and nonlinear wheel–ground contact/friction. Furthermore, it avoids interface errors common in co-simulation, which typically arise from cross-platform time-step synchronization and data interpolation, thereby improving numerical robustness and physical fidelity.
- Seamless co-simulation of control and physics. The physical model is co-simulated with control algorithms within the same environment (MATLAB/Simulink R20225b), forming a closed-loop testbed. Physical and control quantities share a unified time base and data space. This streamlines controller/decision parameter calibration, observer design, and sensitivity analysis, facilitating a smooth transition from Model-in-the-Loop (MIL) to Real-Time Simulation and Hardware-in-the-Loop (HIL) testing.
2.2.2. Overview of the Simulation Platform
Overall Model Architecture
Hydraulic Steering Subsystem
- Power supply and control valve. A plunger pump is modeled as the pressure source, emulating the engine output. A three-position, four-way proportional directional valve serves as the core control element; it meters and reverses the flow delivered to the left and right cylinders according to external command signals.
- Bilateral cylinder actuators. Two symmetrically arranged double-acting cylinders are modeled as the primary steering actuators. They convert hydraulic energy into mechanical work at the articulation joint. To ensure physical fidelity, key geometric and operating parameters—specifically the bore, rod diameter, and stroke/travel range—are calibrated using data from the real vehicle (see Table 2).
- Auxiliary and protection elements. The model also integrates auxiliary components essential for physical realism and numerical robustness. These include a relief valve to cap the maximum pressure, make-up valves to mitigate cavitation, and snubbers or accumulators to attenuate pressure spikes.
Wheel–Terrain Simulation Model
- (1)
- Wheel–ground contact model
- Terrain query. The world coordinates of the wheel center are used to query the local pit-surface height, the surface normal vector, and the position-dependent friction (adhesion) field.
- Local frame alignment. A local ground frame is instantiated based on the queried normal vector. The wheel–ground contact is then geometrically aligned with this frame to inject the terrain’s geometric excitation.
- Relative motion decomposition. The linear and angular velocities at the contact patch (provided by the multibody solver) are decomposed into normal and tangential components. Longitudinal slip and sideslip are subsequently computed from the tangential components.
- Effective adhesion evaluation. The terrain cells covered by the instantaneous contact patch are aggregated using area weighting to obtain a time-varying effective adhesion coefficient.
- Contact-force reconstruction. Normal support and tangential friction forces are calculated using equivalent normal stiffness/damping along with the effective adhesion. The tangential force component is constrained by the friction limit.
- Force–structure feedback. The reconstructed tri-axial forces and moments are applied back to the wheel and frame for the next multibody integration step. These forces are simultaneously exported to the virtual-sensor bus.
- (2)
- Parametric Mud-Pit Terrain
- Geometry
- Physical properties (spatially varying adhesion)
- (i)
- Nominal adhesion: defining the baseline friction level.
- (ii)
- Deterministic gradient: This term reflects the positional dependence (equivalently, depth) of the surface, where water accumulation toward the pit bottom lowers adhesion in deeper regions, and adhesion increases toward the rim.
- (iii)
- Feasibility constraints and discretization
- (i)
- Traction-feasibility boundary.
- (ii)
- Geometric no-interference boundary.
- Scenario set generationThe parameter space is discretized as follows:
- (i)
- Adhesion coefficient set : Eight nominal values are selected within with a step size of 0.05.
- (ii)
- Rim-size set : Nine representative short semi-axis values are sampled. These values are linked to the drum geometry (drum length , diameter ), .
- (iii)
- Depth set : The dimensionless sinkage coefficient is used to classify pits into shallow , medium , and deep ranges; Twelve representative depths are sampled, .
Validation of the Simulation Platform
3. Terrain-Adaptive Maximum-Entropy Policy Optimization
3.1. MDP Formulation of Traversability on Coupled Terrain
3.1.1. State Space
- Symmetric scaling to . Motion variables and the azimuth are linearly mapped to using their theoretical bounds. The radial distance is scaled to with respect to the maximum radial distance over all scenarios.
- Min–max scaling to . Terrain attributes are mapped to via min–max normalization. Likewise, the sinkage depth is normalized to using the maximum pit depth in the scenario set.
3.1.2. Action Space
3.1.3. Terrain-Interaction Critical Reward Function
- (1)
- Ego-behavior reward
- Progress reward.
- Speed reward.
- Smoothness reward.
- Time-efficiency reward.
- (2)
- Terrain-attribute exploration reward .
- Record-breaking condition.
- Effective-progress condition.
- (3)
- Terminal reward .
3.1.4. Overall Closed-Loop Control Architecture
3.2. TAMPO Framework
3.2.1. Maximum-Entropy Actor–Critic
3.2.2. Limitation of Entropy Regularization with a Global Target
3.2.3. Target Core Mechanisms and Update Flow
- (1)
- Context-aware adaptive entropy regularization:
- Terrain physical properties. The instantaneous difficulty index and its exponentially smoothed result are used to quantify traversal difficulty; both are defined in (15) and (17).
- Task-execution efficacy. A progress-based moving average is constructed as the advancement indicator ; its update is given in Equation (32):
- Epistemic uncertainty. The normalized sample standard deviation of the value network output, , is adopted as a proxy; the definition is given in Equation (33), and the online maintenance of its upper bound in Equation (34):
- (2)
- Value and advantage estimation based on GAE.
- (3)
- Entropy-regularized policy update.
3.2.4. Neural-Network Architecture and Hyperparameter Settings
- (1)
- Policy network (Actor)
- Mean head. A 2-D output provides the means for steering and longitudinal-velocity actions. A tanh activation is applied to normalize outputs to , matching the normalized action space.
- Log-standard-deviation head. A 2-D output provides . The standard deviation is obtained by exponentiation, guaranteeing positivity and numerical stability. This design keeps the sampling operation differentiable and supports efficient policy-gradient computation.
- (2)
- Value network (Critic)
4. Simulation Training and Generalization Evaluation
4.1. Experimental Setup
- (1)
- Training set :
- (2)
- Interpolation test set :
- (3)
- Extrapolation test set :
- Adhesion coefficient set: ;
- Depth set: ;
- Rim-size set: .
- (1)
- Success Rate (S.R.): The proportion of episodes in which the agent successfully escapes the pit. This serves as the primary metric of strategic effectiveness.
- (2)
- Average Escape Time (A.E.T.): The mean time from episode start to successful escape, computed only over successful episodes, reflecting policy efficiency.
- (3)
- Average Cumulative Reward (A.C.R.): The mean cumulative return over all episodes (successful or failed) on the test sets, reflecting the overall quality of the policy.
4.2. Training-Process Comparison
4.3. Quantitative Assessment of Generalization Performance
4.4. TAMPO Algorithm Ablation Study
- (1)
- Removing the exploration term driven by value-network uncertainty , with setting ;
- (2)
- Removing the progress-dependent feedback term , with setting ;
- (3)
- Removing the terrain difficulty feedback term , with setting ;
- (4)
- Removing the adaptive entropy-coefficient mechanism and replacing it with a finely tuned fixed constant.
4.5. Analysis of Typical Escape Scenarios
4.5.1. Micro-Behavioral Analysis of Typical Scenarios
- (1)
- Easy Scenario (, , )
- (2)
- Medium Scenario (, , )


- (3)
- Hard Scenario (, , )


4.5.2. Macroscopic Generalization Performance Evaluation
- (1)
- Narrow Pit Opening Scenarios ()
- (2)
- Medium Pit Opening Scenarios ()
- (3)
- Wide Pit Opening Scenarios ()
5. Conclusions
- A simulation platform based on a component-level physical network is constructed, comprising the road-roller system and parameterized mud-pit terrains. The vehicle model includes the drum, front frame, left and right rear wheels, rear frame, and hydraulic steering mechanism, with all subsystems built in accordance with the physical topology and assembly relationships of the real machine. The mud pit is modeled as a surface with elliptical contours and parabolic profiles. By combining the adhesion coefficient, pit depth, and the ellipse’s short semi-axis, parameterized scenarios are generated to form the training and test sets for the algorithms.
- A TICR function is designed. By combining dense rewards that encode task progress with sparse rewards that encourage terrain exploration, this function provides informative guidance signals that enable the agent to learn efficient and intelligent escape policies.
- The TAMPO algorithm is developed with a core context-aware adaptive entropy-regularization mechanism. By fusing, in real time, three types of information—terrain physical characteristics, task-execution efficacy, and model epistemic uncertainty—the mechanism dynamically adjusts policy entropy, thereby equipping the agent with the ability to intelligently balance exploration and exploitation in complex, heterogeneous environments.
- TAMPO exhibits clear advantages in both escape effectiveness and efficiency. On the 90 most challenging scenarios in the extrapolation test set, TAMPO attains an average success rate (S.R.) of 60.00% with an Average Escape Time (A.E.T.) of only 17.56 s. Compared with PPO, SAC, and DDPG, the S.R. is improved by 10%, 14.45%, and 22.22%, respectively, while the A.E.T. is reduced by 2.47 s, 4.73 s, and 5.73 s, respectively. These results highlight the algorithm’s strong decision-making capability and generalization performance.
6. Future Work
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| UARR | Unmanned Articulated Road Roller |
| TAMPO | Terrain-Adaptive Maximum-Entropy Policy Optimization |
| TICR | Terrain-Interaction Critical Reward |
| S.R. | Success rate |
| A.E.T. | Average Escape Time |
| A.C.R. | average cumulative reward |
| TCS | Traction control systems |
| RL | Reinforcement learning |
| DAE | Differential-Algebraic Equation |
| MIL | Model-in-the-Loop |
| HIL | Hardware-in-the-Loop |
| MDP | Markov decision process |
| AC | Actor–critic |
| TD | Temporal-difference |
| SAC | Soft actor–critic |
| GAE | Generalized advantage estimation |
Appendix A
| Algorithm | Parameter | Value |
|---|---|---|
| SAC | Discount factor | 0.99 |
| Actor learning rate | 0.0003 | |
| Critic learning rate | 0.001 | |
| Replay buffer size | 106 | |
| Batch size | 256 | |
| Target smoothing coefficient | 0.005 | |
| Initial temperature | 0.2 | |
| Entropy tuning | True (Auto-learned) | |
| Target entropy | −dim(A) = −2 |
| Algorithm | Parameter | Value |
|---|---|---|
| PPO | Discount factor | 0.99 |
| Learning rate | 0.0003 | |
| GAE parameter | 0.95 | |
| Clipping range | 0.2 | |
| Entropy coefficient | 0.01 | |
| Value function coefficient | 0.5 | |
| Update epochs | 10 | |
| Mini-batch size | 64 |
| Algorithms | Parameter | Value |
|---|---|---|
| DDPG | Discount factor | 0.99 |
| Actor learning rate | 0.0001 | |
| Critic learning rate | 0.001 | |
| Target smoothing coefficient | 0.005 | |
| Batch size | 128 | |
| Replay buffer size | 106 | |
| Exploration Noise Type | OU | |
| Noise Theta | 0.15 | |
| Noise Sigma | 0.20 |
References
- Ivanov, V.; Savitski, D.; Shyrokau, B. A survey of traction control and antilock braking systems of full electric vehicles with individually controlled electric motors. IEEE Trans. Veh. Technol. 2014, 64, 3878–3896. [Google Scholar] [CrossRef] [Scilit]
- Jin, L.-Q.; Ling, M.; Yue, W. Tire-road friction estimation and traction control strategy for motorized electric vehicle. PLoS ONE 2017, 12, e0179526. [Google Scholar] [CrossRef] [Scilit]
- Kučera, P.; Píštěk, V. Prototyping a system for truck differential lock control. Sensors 2019, 19, 3619. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jiang, C.; Hu, Z.; Mourelatos, Z.P.; Gorsich, D.; Jayakumar, P.; Fu, Y.; Majcher, M. R2-RRT*: Reliability-based robust mission planning of off-road autonomous ground vehicle under uncertain terrain environment. IEEE Trans. Autom. Sci. Eng. 2021, 19, 1030–1046. [Google Scholar] [CrossRef] [Scilit]
- Sakayori, G.; Ishigami, G. Modeling of slip rate-dependent traversability for path planning of wheeled mobile robot in sandy terrain. Front. Robot. AI 2024, 11, 1320261. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Urvina, R.P.; Guevara, C.L.; Vásconez, J.P.; Prado, A.J. An integrated route and path planning strategy for skid–steer mobile robots in assisted harvesting tasks with terrain traversability constraints. Agriculture 2024, 14, 1206. [Google Scholar] [CrossRef] [Scilit]
- Gargano, I.E.; von Ellenrieder, K.D.; Vivolo, M. A Survey of Trajectory Planning Algorithms for Off-Road Uncrewed Ground Vehicles. In Proceedings of the International Conference on Modelling and Simulation for Autonomous Systems; Springer: Cham, Switzerland, 2025; pp. 120–148. [Google Scholar]
- Wang, N.; Li, X.; Zhang, K.; Wang, J.; Xie, D. A survey on path planning for autonomous ground vehicles in unstructured environments. Machines 2024, 12, 31. [Google Scholar] [CrossRef] [Scilit]
- Uwano, F.; Tajima, Y.; Murata, A.; Takadama, K. Recovery system based on exploration-biased genetic algorithm for stuck rover in planetary exploration. J. Robot. Mechatron. 2017, 29, 877–886. [Google Scholar] [CrossRef] [Scilit]
- Xiao, X.; Xu, Z.; Wang, Z.; Song, Y.; Warnell, G.; Stone, P.; Zhang, T.; Ravi, S.; Wang, G.; Karnan, H. Autonomous ground navigation in highly constrained spaces: Lessons learned from the benchmark autonomous robot navigation challenge at icra 2022 [competitions]. IEEE Robot. Autom. Mag. 2022, 29, 148–156. [Google Scholar] [CrossRef] [Scilit]
- Sánchez, M.; Morales, J.; Martínez, J.L. Reinforcement and curriculum learning for off-road navigation of an UGV with a 3D LiDAR. Sensors 2023, 23, 3239. [Google Scholar] [CrossRef] [Scilit]
- Xu, T.; Pan, C.; Xiao, X. Reinforcement learning for wheeled mobility on vertically challenging terrain. In Proceedings of the 2024 IEEE International Symposium on Safety Security Rescue Robotics (SSRR), New York, NY, USA, 12–14 November 2024; pp. 125–130. [Google Scholar]
- Siva, S.; Wigness, M.; Rogers, J.G.; Quang, L.; Zhang, H. Self-reflective terrain-aware robot adaptation for consistent off-road ground navigation. Int. J. Robot. Res. 2024, 43, 1003–1023. [Google Scholar] [CrossRef] [Scilit]
- Devidze, R.; Kamalaruban, P.; Singla, A. Exploration-guided reward shaping for reinforcement learning under sparse rewards. Adv. Neural Inf. Process. Syst. 2022, 35, 5829–5842. [Google Scholar]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; pp. 1861–1870. [Google Scholar]
- Jiang, Y.; Kolter, J.Z.; Raileanu, R. On the importance of exploration for generalization in reinforcement learning. Adv. Neural Inf. Process. Syst. 2023, 36, 12951–12986. [Google Scholar]
- Kim, D.; Shin, J.; Abbeel, P.; Seo, Y. Accelerating reinforcement learning with value-conditional state entropy exploration. Adv. Neural Inf. Process. Syst. 2023, 36, 31811–31830. [Google Scholar]
- Kim, W.; Sung, Y. An adaptive entropy-regularization framework for multi-agent reinforcement learning. In Proceedings of the 40th International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; pp. 16829–16852. [Google Scholar]
- Wang, Y.; Ni, T. Meta-sac: Auto-tune the entropy temperature of soft actor-critic via metagradient. arXiv 2020, arXiv:2007.01932. [Google Scholar]
- Xu, Y.; Hu, D.; Liang, L.; McAleer, S.; Abbeel, P.; Fox, R. Target entropy annealing for discrete soft actor-critic. arXiv 2021, arXiv:2112.02852. [Google Scholar] [CrossRef] [Scilit]
- Dumitriu, D.N.; Chiroiu, V.; Munteanu, L. Car vertical dynamics simulations using both an in-house 7 DOF model simulator and Carsim commercial software. UPB Sci. Bull. Ser. D Mech. Eng. 2015, 77, 77–84. [Google Scholar]
- Kinjawadekar, T.; Dixit, N.; Heydinger, G.J.; Guenther, D.A.; Salaani, M.K. Vehicle Dynamics Modeling and Validation of the 2003 Ford Expedition with ESC Using CarSim; 0148-7191; SAE Technical Paper: Warrendale, PA, USA, 2009. [Google Scholar]
- Takács, D.; Zelei, A. Performance Optimization of a Formula Student Racing Car Using the IPG CarMaker, Part 1: Lap Time Convergence and Sensitivity Analysis. Eng. Proc. 2024, 79, 86. [Google Scholar]
- Takács, D.; Zelei, A. Performance Optimization of a Formula Student Racing Car Using IPG CarMaker—Part 2: Aiding Aerodynamics and Drag Reduction System Package Design. Eng. Proc. 2024, 79, 77. [Google Scholar]
- Zhang, S.; Liu, Y.; Li, P.; Wang, L.; Zhou, X.; Li, Y.; Luo, Y. Study of instability mechanisms of trucks turning right at long downhill T-junctions based on Trucksim simulation. PLoS ONE 2023, 18, e0282779. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, J.; Oh, T.; Yoo, J. Adaptive Longitudinal Speed Control for Heavy-Duty Vehicles Considering Actuator Constraints and Disturbances Using Simulation Validation. Appl. Sci. 2025, 15, 7327. [Google Scholar] [CrossRef] [Scilit]
- Xu, F.; Liu, X.; Chen, W.; Zhou, C.; Cao, B. Modeling and co-simulation based on Adams and AMESim of pivot steering system. J. Eng. 2019, 2019, 392–396. [Google Scholar] [CrossRef] [Scilit]
- Backhuijs, S. Development and Validation of a Multibody Model of the Lupo 3L and EL; Eindhoven University of Technology: Eindhoven, The Netherlands, 2020. [Google Scholar]
- Tomasikova, M.; Sojcak, D.; Nieoczym, A.; Brumercik, F. Experimental data in vehicle modeling. LOGI Sci. J. Transp. Logist. 2017, 8, 82–87. [Google Scholar] [CrossRef] [Scilit]
- Achterhold, J.; Guttikonda, S.; Kreber, J.U.; Li, H.; Stueckler, J. Learning a Terrain-and Robot-Aware Dynamics Model for Autonomous Mobile Robot Navigation. arXiv 2024, arXiv:2409.11452. [Google Scholar]
- Etherington, T.R. Perlin noise as a hierarchical neutral landscape model. Web Ecol. 2022, 22, 1–6. [Google Scholar] [CrossRef] [Scilit]
- Puterman, M.L. Markov Decision Processes: Discrete Stochastic Dynamic Programming; John Wiley & Sons: New York, NY, USA, 2014. [Google Scholar]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction, 2nd ed.; MIT Press: Cambridge, MA, USA, 2018; Volume 1, p. 25. [Google Scholar]
- Konda, V.R.; Tsitsiklis, J.N. Onactor-critic algorithms. SIAM J. Control. Optim. 2003, 42, 1143–1166. [Google Scholar] [CrossRef] [Scilit]
- Mnih, V.; Badia, A.P.; Mirza, M.; Graves, A.; Lillicrap, T.; Harley, T.; Silver, D.; Kavukcuoglu, K. Asynchronous methods for deep reinforcement learning. In Proceedings of the 33rd International Conference on Machine Learning, New York, NY, USA, 19–24 June 2016; pp. 1928–1937. [Google Scholar]
- Watkins, C.J.; Dayan, P. Q-learning. Mach. Learn. 1992, 8, 279–292. [Google Scholar] [CrossRef] [Scilit]
- Williams, R.J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Mach. Learn. 1992, 8, 229–256. [Google Scholar] [CrossRef] [Scilit]
- Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P. Soft actor-critic algorithms and applications. arXiv 2018, arXiv:1812.05905. [Google Scholar]
- Greensmith, E.; Bartlett, P.L.; Baxter, J. Variance reduction techniques for gradient estimates in reinforcement learning. J. Mach. Learn. Res. 2004, 5, 1471–1530. [Google Scholar]
- Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.; Abbeel, P. High-dimensional continuous control using generalized advantage estimation. arXiv 2015, arXiv:1506.02438. [Google Scholar]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Duan, Y.; Chen, X.; Houthooft, R.; Schulman, J.; Abbeel, P. Benchmarking deep reinforcement learning for continuous control. In Proceedings of the 33rd International Conference on Machine Learning, New York, NY, USA, 19–24 June 2016; pp. 1329–1338. [Google Scholar]
- Schulman, J.; Levine, S.; Abbeel, P.; Jordan, M.; Moritz, P. Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 6–11 July 2015; pp. 1889–1897. [Google Scholar]
- Glorot, X.; Bordes, A.; Bengio, Y. Deep sparse rectifier neural networks. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, Fort Lauderdale, FL, USA, 11–13 April 2011; pp. 315–323. [Google Scholar]
- Lillicrap, T.P.; Hunt, J.J.; Pritzel, A.; Heess, N.; Erez, T.; Tassa, Y.; Silver, D.; Wierstra, D. Continuous control with deep reinforcement learning. arXiv 2015, arXiv:1509.02971. [Google Scholar]
- Silver, D.; Lever, G.; Heess, N.; Degris, T.; Wierstra, D.; Riedmiller, M. Deterministic policy gradient algorithms. In Proceedings of the 31st International Conference on Machine Learning, Beijing, China, 21–26 June 2014; pp. 387–395. [Google Scholar]


































| Component | Parameter | Value | Unit |
|---|---|---|---|
| Front frame | Mass | 5200 | kg |
| CG height (relative to ground) | 780 | mm | |
| CG longitudinal position (relative to articulation) | 1425 | mm | |
| CG lateral position (relative to articulation) | 0 | mm | |
| Moment of inertia about x-axis | 1841 | kg·m2 | |
| Moment of inertia about y-axis | 2358 | kg·m2 | |
| Moment of inertia about z-axis | 3695 | kg·m2 | |
| Rear frame | Mass | 6750 | kg |
| CG height (relative to ground) | 880 | mm | |
| CG longitudinal position (relative to articulation) | 1530 | mm | |
| CG lateral position (relative to articulation) | 0 | mm | |
| Moment of inertia about x-axis | 2731 | kg·m2 | |
| Moment of inertia about y-axis | 3158 | kg·m2 | |
| Moment of inertia about z-axis | 5233 | kg·m2 |
| Parameter | Value | Unit |
|---|---|---|
| Steering pump displacement | 80~500 | cc/r |
| Inlet pressure | 20 | MPa |
| Cylinder stroke | 455 | mm |
| Cylinder bore | 100 | mm |
| Piston-rod travel | 265 | mm |
| Piston-rod diameter | 57.3 | mm |
| Hydraulic line inner diameter | 24 | mm |
| Component | Parameter | Value | Unit |
|---|---|---|---|
| Front drum | Diameter | 1600 | mm |
| Steel width | 2130 | mm | |
| Mass | 9400 | kg | |
| CG height (with respect to ground) | 800 | mm | |
| CG longitudinal position (with respect to articulation) | 1500 | mm | |
| CG lateral position (with respect to articulation) | 0 | mm | |
| Rear wheels | Diameter | 1500 | mm |
| Tire width | 595 | mm | |
| Track width | 1980 | mm | |
| Mass (per wheel assembly) | 650 | kg | |
| CG height (with respect to ground) | 750 | mm | |
| CG longitudinal position (with respect to articulation) | 1680 | mm | |
| CG lateral position (with respect to articulation) | 0 | mm |
| Turn Index | Measured Peak Y | Simulated Peak Y | Normalized Error (% of Range) |
|---|---|---|---|
| 1 | 870.18 | 870.25 | 1.56% |
| 2 | 873.20 | 873.31 | 2.44% |
| 3 | 870.59 | 870.75 | 3.56% |
| 4 | 872.80 | 872.87 | 1.56% |
| 5 | 870.51 | 870.56 | 1.11% |
| 6 | 871.75 | 871.68 | 1.56% |
| 7 | 869.52 | 869.43 | 2.00% |
| 8 | 871.03 | 870.78 | 5.56% |
| 9 | 868.70 | 868.51 | 4.22% |
| 10 | 871.56 | 871.40 | 3.56% |
| 11 | 870.48 | 870.43 | 1.11% |
| 12 | 872.78 | 872.87 | 2.00% |
| 13 | 871.90 | 872.01 | 2.44% |
| Parameter | Symbol | Value | Description |
|---|---|---|---|
| Discount factor | γ | 0.99 | Trades off long- vs. short-term return |
| Actor learning rate | ηθ | 0.0003 | Adam step size for policy network |
| Critic learning rate | ηψ | 0.001 | Adam step size for value network |
| GAE parameter | λ | 0.95 | Bias-variance trade-off in generalized advantage estimation |
| Initial entropy coefficient | α0 | 0.2 | Initial value for adaptive entropy |
| Entropy decay rate | βH | 0.0001 | Exponential time-decay factor for |
| Minimum entropy coefficient | αmin | 0.01 | Lower bound to preserve exploration |
| Weight: terrain difficulty | λD | 0.5 | Weight in adaptive exploration scheduler |
| Weight: progress efficacy | λη | 0.2 | Weight in adaptive exploration scheduler |
| Weight: value uncertainty | λV | 0.3 | Weight in adaptive exploration scheduler |
| Polyak averaging factor | τ | 0.005 | Soft-update rate for target value network |
| Max steps per episode | Tmax | 256 | Maximum environment interactions per episode |
| DRL Algorithm | Interpolation Test Set | Extrapolation Test Set | ||||
|---|---|---|---|---|---|---|
| S.R. (%) | A.E.T. (s) | A.C.R. | S.R. (%) | A.E.T. (s) | A.C.R. | |
| TAMPO | 70.83 ± 1.74 | 14.92 ± 0.59 | 428.76 | 60.00 ± 1.86 | 17.56 ± 0.59 | 389.15 |
| SAC | 61.46 ± 3.16 | 16.77 ± 0.82 | 346.53 | 50.00 ± 2.53 | 20.03 ± 1.74 | 290.73 |
| R.P. | 9.37 ↓ | 1.85 ↑ | 82.23 ↓ | 10.00 ↓ | 2.47 ↑ | 98.42 ↓ |
| PPO | 54.17 ± 3.55 | 18.46 ± 1.35 | 269.27 | 45.55 ± 2.81 | 22.29 ± 2.21 | 215.48 |
| R.P. | 16.66 ↓ | 3.54 ↑ | 159.49 ↓ | 14.45 ↓ | 4.73 ↑ | 173.67 ↓ |
| DDPG | 43.75 ± 4.32 | 21.37 ± 1.70 | 251.82 | 37.78 ± 2.90 | 23.29 ± 2.07 | 196.92 |
| R.P. | 27.08 ↓ | 6.45 ↑ | 176.94 ↓ | 22.22 ↓ | 5.73 ↑ | 192.23 ↓ |
| Pit Opening Dimension (m) | Depth Setpoint (m) | Actual Climb Height (m) | ||
|---|---|---|---|---|
| 1.80 | = 0.37 | = 0.52 | = 0.68 | |
| −1.00 | 0.258 | 0.413 | 0.558 | |
| −0.90 | 0.279 | 0.405 | 0.603 | |
| −0.82 | 0.252 | 0.478 | 0.614 | |
| −0.70 | 0.334 | 0.492 | 0.674 | |
| −0.60 | 0.336 | 0.579 | 0.600 | |
| −0.50 | 0.429 | 0.500 | 0.500 | |
| −0.42 | 0.420 | 0.420 | 0.420 | |
| −0.30 | 0.300 | 0.300 | 0.300 | |
| −0.20 | 0.200 | 0.200 | 0.200 | |
| −0.10 | 0.100 | 0.100 | 0.100 | |
| Pit Opening Dimension (m) | Depth Setpoint (m) | Actual Climb Height (m) | ||
|---|---|---|---|---|
| 3.30 | = 0.37 | = 0.52 | = 0.68 | |
| −1.00 | 0.379 | 0.518 | 0.665 | |
| −0.90 | 0.387 | 0.569 | 0.706 | |
| −0.82 | 0.443 | 0.598 | 0.778 | |
| −0.70 | 0.474 | 0.668 | 0.700 | |
| −0.60 | 0.542 | 0.600 | 0.600 | |
| −0.50 | 0.500 | 0.500 | 0.500 | |
| −0.42 | 0.420 | 0.420 | 0.420 | |
| −0.30 | 0.300 | 0.300 | 0.300 | |
| −0.20 | 0.200 | 0.200 | 0.200 | |
| −0.10 | 0.100 | 0.100 | 0.100 | |
| Pit Opening Dimension (m) | Depth Setpoint (m) | Actual Climb Height (m) | ||
|---|---|---|---|---|
| 3.30 | = 0.37 | = 0.52 | = 0.68 | |
| −1.00 | 0.379 | 0.518 | 0.665 | |
| −0.90 | 0.387 | 0.569 | 0.706 | |
| −0.82 | 0.443 | 0.598 | 0.778 | |
| −0.70 | 0.474 | 0.668 | 0.700 | |
| −0.60 | 0.542 | 0.600 | 0.600 | |
| −0.50 | 0.500 | 0.500 | 0.500 | |
| −0.42 | 0.420 | 0.420 | 0.420 | |
| −0.30 | 0.300 | 0.300 | 0.300 | |
| −0.20 | 0.200 | 0.200 | 0.200 | |
| −0.10 | 0.100 | 0.100 | 0.100 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Qiang, W.; Xu, Q.; Xie, H. Adaptive Traversability Policy Optimization for an Unmanned Articulated Road Roller on Slippery, Geometrically Irregular Terrains. Machines 2026, 14, 79. https://doi.org/10.3390/machines14010079
Qiang W, Xu Q, Xie H. Adaptive Traversability Policy Optimization for an Unmanned Articulated Road Roller on Slippery, Geometrically Irregular Terrains. Machines. 2026; 14(1):79. https://doi.org/10.3390/machines14010079
Chicago/Turabian StyleQiang, Wei, Quanzhi Xu, and Hui Xie. 2026. "Adaptive Traversability Policy Optimization for an Unmanned Articulated Road Roller on Slippery, Geometrically Irregular Terrains" Machines 14, no. 1: 79. https://doi.org/10.3390/machines14010079
APA StyleQiang, W., Xu, Q., & Xie, H. (2026). Adaptive Traversability Policy Optimization for an Unmanned Articulated Road Roller on Slippery, Geometrically Irregular Terrains. Machines, 14(1), 79. https://doi.org/10.3390/machines14010079

