Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (2,242)

Search Parameters:
Keywords = reinforcement learning for policy

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
24 pages, 5175 KB  
Article
InfRA-FL: Information-Driven Robust Federated Learning via Saturation-Aware Reinforcement Learning
by Jiao Tian, Jinlin He and Liejun Wang
Mach. Learn. Knowl. Extr. 2026, 8(8), 247; https://doi.org/10.3390/make8080247 - 14 Aug 2026
Abstract
Client selection is a critical mechanism for ensuring robust convergence in Federated Learning (FL) systems, yet it remains vulnerable to Non-IID data distributions and Byzantine attacks. Deep Reinforcement Learning (DRL) has shown promise for automated client selection, yet existing methods suffer from three [...] Read more.
Client selection is a critical mechanism for ensuring robust convergence in Federated Learning (FL) systems, yet it remains vulnerable to Non-IID data distributions and Byzantine attacks. Deep Reinforcement Learning (DRL) has shown promise for automated client selection, yet existing methods suffer from three structural deficiencies: observation ambiguity, where scalar states cannot distinguish malicious updates from benign heterogeneity; reward fragility, whereby attackers exploit unbounded feedback to hijack policy updates; and risk blindness, as risk-neutral agents overlook the inherent variance in client contributions. To address these deficiencies, we propose InfRA-FL, a robust adaptive framework. A mutual-information-based state construction extracts high-utility features, resolving observation ambiguity. A Saturation-Aware Robust Reward (SARR) mechanism applies soft-clipping to bound each client’s influence on the policy gradient, provably neutralizing reward poisoning. URA-PPO, an uncertainty-aware algorithm with a dual-head Critic, optimizes a risk-penalized objective that shifts the agent from risk-neutral to risk-averse decision-making. Experiments on MNIST and CIFAR-10 under 20% Byzantine adversaries show that InfRA-FL outperforms state-of-the-art baselines by 5–10% in accuracy while accelerating convergence, establishing that principled information-theoretic observation and robust reward design suffice to secure RL-driven federated learning against targeted poisoning. Full article
25 pages, 15084 KB  
Article
Preference-Conditioned Sequential Optimization for Multi-Target Cleanup in Action-Dependent Risk Fields
by Fengyou Wu, Chenyang Li, Wenjun Kang, Boxin Liu, Peng Zhi, Rui Zhou, Ling-Huey Li, Qingguo Zhou and Kuan-Ching Li
Symmetry 2026, 18(8), 1369; https://doi.org/10.3390/sym18081369 - 14 Aug 2026
Abstract
Radioactive source cleanup in nuclear environments is an important task for facility decommissioning and autonomous radiation management. In multi-target cleanup scenarios, each source removal action alters the subsequent radiation risk distribution. This action-dependent evolution invalidates the static planning assumptions commonly adopted in conventional [...] Read more.
Radioactive source cleanup in nuclear environments is an important task for facility decommissioning and autonomous radiation management. In multi-target cleanup scenarios, each source removal action alters the subsequent radiation risk distribution. This action-dependent evolution invalidates the static planning assumptions commonly adopted in conventional routing and path planning methods. To address this challenge, this work formulates multi-target radioactive hotspot cleanup as a preference-conditioned sequential combinatorial optimization problem. We propose a neural sequential optimization framework that integrates hotspot map encoding and radiation field image encoding through cross-modal interaction. The proposed framework generates preference-conditioned risk-aware cleanup policies, enabling effective coordination between target selection and safe navigation under different risk–efficiency trade-offs. The nuclear hot-cell simulation environment is constructed using the Robot Operating System (ROS) and the Gazebo physics simulation engine. Compared with classical heuristic methods, evolutionary optimization methods, and reinforcement learning-based baselines, the proposed method achieves lower cumulative radiation exposure across different problem scales while maintaining competitive task efficiency. The results demonstrate that the proposed framework provides an effective task-level sequencing strategy for radiation-aware robotic cleanup in action-dependent risk fields. Full article
Show Figures

Figure 1

33 pages, 21275 KB  
Article
Egocentric Constraint Corridor: Deep Reinforcement Learning for Fixed-Wing UAV Navigation in Vertically Constrained Airspace
by Yuhao Gong, Jinfu Lin, Jiaqiang Zhang and Han Wang
Drones 2026, 10(8), 622; https://doi.org/10.3390/drones10080622 - 14 Aug 2026
Abstract
Fixed-wing UAVs operating in long-range missions often fly through airspace subject to heterogeneous multi-source constraints that vertically compress the flyable space into a constraint corridor of continuously varying thickness. Conventional path planning methods incur high online computational costs in such scenarios. Deep reinforcement [...] Read more.
Fixed-wing UAVs operating in long-range missions often fly through airspace subject to heterogeneous multi-source constraints that vertically compress the flyable space into a constraint corridor of continuously varying thickness. Conventional path planning methods incur high online computational costs in such scenarios. Deep reinforcement learning can generate reactive decisions from local observations, yet existing approaches predominantly target multirotor obstacle avoidance and rely on observations designed for discrete obstacles, lacking a unified representation for corridor constraints. Moreover, constraint conditions vary across mission scenarios, demanding cross-scenario policy generalization. This paper proposes the Egocentric Constraint Corridor (ECC), which fuses multi-source constraints into upper and lower boundary surfaces defining the corridor, then egocentrically encodes the surrounding corridor relative to the vehicle into a margin field serving as structured policy input. A deep reinforcement learning framework built on ECC is trained end-to-end, with its multi-branch network and composite reward function following from the structure of the corridor encoding. Experiments show that ECC-DRL achieves path efficiency approaching that of globally informed A*, and that it is the only one of the compared methods that computes its decisions online within the decision interval. Ablation studies confirm the margin field is necessary for reliable navigation, and the ECC encoding enables zero-shot transfer to scenarios with unseen terrains and radar deployments without retraining. Hardware-in-the-loop experiments on an embedded platform verify real-time closed-loop feasibility. Full article
(This article belongs to the Section Artificial Intelligence in Drones (AID))
Show Figures

Figure 1

25 pages, 8647 KB  
Article
Network Dynamic Spatiotemporal Dispatch Based on Multi-Head Graph Attention Reinforcement Learning and Balanced Responsibility
by Hucheng Li, Lifei Sun, Haifeng Fan, Hongjin Pan, Fei Xu and Ling Hao
Electronics 2026, 15(16), 3622; https://doi.org/10.3390/electronics15163622 - 14 Aug 2026
Abstract
The rapid integration of distributed renewable energy and flexible loads significantly intensifies supply and demand uncertainty in active distribution networks (ADNs), threatening economic and secure grid operations. Existing deep reinforcement learning (DRL) dispatch methods fail to extract spatial features properly, leading to a [...] Read more.
The rapid integration of distributed renewable energy and flexible loads significantly intensifies supply and demand uncertainty in active distribution networks (ADNs), threatening economic and secure grid operations. Existing deep reinforcement learning (DRL) dispatch methods fail to extract spatial features properly, leading to a local optimal solution. To address these limitations, this paper proposes a state-adaptive topology-aware continuous-dispatch framework via multi-head graph attention network and deep deterministic policy gradient (GAT-DDPG). A multi-head graph attention network is embedded within a centralized Actor–Critic training paradigm to adaptively update spatial message-passing weights based on operational states. Extensive simulations on a modified IEEE 33-bus ADN over a 125-day unseen test set demonstrate that the proposed framework achieves lower comprehensive operating costs and fewer voltage violations compared with representative DRL-based dispatch baselines. Visualizations of state-dependent attention shifts confirm the model’s capability to track dynamically shifting network vulnerabilities, providing physical interpretability. Based on this, and combined with the optimized dispatch method of balancing responsibility, the ability of different flexible resources to support safe and stable operation and the ideal dispatch results under the temporary reduction in new energy output are further simulated. Full article
Show Figures

Figure 1

28 pages, 7749 KB  
Article
Imagine to Ensure Safety in Hierarchical Reinforcement Learning
by Gregory Gorbov, Artem Latyshev and Aleksandr Panov
Mach. Learn. Knowl. Extr. 2026, 8(8), 243; https://doi.org/10.3390/make8080243 - 13 Aug 2026
Abstract
This work investigates the safe exploration problem in reinforcement learning, where an agent must maximize cumulative performance while simultaneously satisfying safety constraints. This challenge becomes particularly difficult in long-horizon tasks because safe reinforcement learning methods often become overly conservative, restricting exploration and preventing [...] Read more.
This work investigates the safe exploration problem in reinforcement learning, where an agent must maximize cumulative performance while simultaneously satisfying safety constraints. This challenge becomes particularly difficult in long-horizon tasks because safe reinforcement learning methods often become overly conservative, restricting exploration and preventing agents from reaching distant goals, while estimation errors accumulate over extended horizons. We propose Imagine To Ensure Safety in Hierarchical Reinforcement Learning (ITES), which combines a learnable world model with high-level and low-level policies to promote safety at both hierarchical levels. The novelty of ITES lies in jointly ensuring safe subgoal generation and safe subgoal execution within a hierarchical framework. The high-level policy generates intermediate subgoals that guide exploration toward safe regions, while the low-level policy uses imagined rollouts in the learned world model to reduce unsafe behavior during subgoal execution. We evaluate ITES on the long-horizon SafeAntMaze C-shape, SafeAntMaze W-shape, and SafePusher benchmarks, as well as on short-horizon Safety Gym tasks, using task performance and episodic cost under prescribed safety budgets. The results show that ITES achieves substantially higher task success and constraint-compliant success rates on the long-horizon benchmarks while maintaining mean episodic costs below the prescribed budgets. Full article
(This article belongs to the Section Learning)
Show Figures

Figure 1

32 pages, 4771 KB  
Article
ODARRL: Obstacle- and Disturbance-Aware End-to-End Residual Reinforcement Learning for Underwater Robot Trajectory Tracking with Obstacle Avoidance
by Linghan Meng, Zebin Huang, Qingfeng Yao, Yunxiu Zhang and Qifeng Zhang
J. Mar. Sci. Eng. 2026, 14(16), 1501; https://doi.org/10.3390/jmse14161501 - 13 Aug 2026
Abstract
ROVs are essential for marine exploration and underwater operations, yet conventional teleoperation relies heavily on skilled human operators, and many autonomous methods stop at high-level planning rather than low-level actuation, limiting robustness in disturbed and cluttered environments. This paper proposes ODARRL, an obstacle- [...] Read more.
ROVs are essential for marine exploration and underwater operations, yet conventional teleoperation relies heavily on skilled human operators, and many autonomous methods stop at high-level planning rather than low-level actuation, limiting robustness in disturbed and cluttered environments. This paper proposes ODARRL, an obstacle- and disturbance-aware sensor-to-thruster (ST) end-to-end residual reinforcement learning framework for safe trajectory execution of underwater robots. Using a three-stage curriculum, ODARRL first acquires a basic policy from MPC demonstrations in a static obstacle-free environment, then improves disturbance-robust tracking under random currents, and finally extends to scenarios involving both currents and obstacles. A Dual-Horizon Attention Disturbance Encoder is further designed to capture current-related features from long- and short-term histories, which are fused with robot states and reference information as the input to the ST end-to-end policy. Experiments in Marine Gym with BlueROV2 Heavy demonstrate that ODARRL achieves more stable and robust trajectory tracking under random currents, reducing the mean total tracking error by 69.3%, 31.9%, 45.8%, 73.0% and 25.8% relative to the MPC-imitation policy, PPO, SAC, A2C and VNRS-SAC, respectively. With obstacles introduced, curriculum-initialized policies also exhibit higher path progress and more stable task completion during obstacle-avoidance training. Full article
(This article belongs to the Special Issue Advanced Modeling and Intelligent Control of Marine Vehicles)
Show Figures

Figure 1

26 pages, 1385 KB  
Article
Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation
by Yitan Wang, Yangyang Zhang and Zhenxing Gao
Aerospace 2026, 13(8), 721; https://doi.org/10.3390/aerospace13080721 - 13 Aug 2026
Viewed by 62
Abstract
Low-altitude micro downbursts pose a severe threat to aviation safety, yet conventional control approaches and standard deep reinforcement learning (DRL) often fail due to explicit modeling difficulties and sparse reward constraints. To address these challenges, this study proposes an explainable, data-driven framework integrating [...] Read more.
Low-altitude micro downbursts pose a severe threat to aviation safety, yet conventional control approaches and standard deep reinforcement learning (DRL) often fail due to explicit modeling difficulties and sparse reward constraints. To address these challenges, this study proposes an explainable, data-driven framework integrating active-reward proximal policy optimization (AR-PPO). A bilevel optimization architecture driven by meta-gradients is developed to dynamically discover optimal reward functions without human intervention. Furthermore, a policy distillation pipeline utilizing wavelet-multivariate singular spectrum analysis (W-MSSA) and classification and regression trees (CART) is proposed to translate high-frequency continuous neural outputs into discrete, pilot-readable rules. Simulation results on a B737-800 model demonstrate that AR-PPO effectively overcomes the “stall trap” by autonomously learning to trade altitude for airspeed, outperforming static-reward baselines and empirical human pilots in extreme, zero-shot windshear encounters (22.0 m/s downdraft). Ultimately, the proposed framework successfully distills black-box AI strategies into verifiable, physics-informed standard operating procedures (SOPs), providing a highly transparent and robust solution for autonomous windshear escape and future competency-based flight training. Full article
Show Figures

Figure 1

31 pages, 22125 KB  
Article
Carbon-Aware Dynamic Human–Robot Collaborative Flexible Job Shop Scheduling Under Safety-Proximity Disruption
by Fan Wu, Yufan Zheng and Wenkang Zhang
Machines 2026, 14(8), 931; https://doi.org/10.3390/machines14080931 - 12 Aug 2026
Viewed by 84
Abstract
Human–robot collaborative flexible job shop scheduling (HRC-FJSP) must coordinate heterogeneous capabilities, mode-dependent processing times, safety feasibility, and carbon constraints. The problem becomes harder when a collaboration mode that is attractive during planning becomes infeasible after a human enters the robot safety separation zone. [...] Read more.
Human–robot collaborative flexible job shop scheduling (HRC-FJSP) must coordinate heterogeneous capabilities, mode-dependent processing times, safety feasibility, and carbon constraints. The problem becomes harder when a collaboration mode that is attractive during planning becomes infeasible after a human enters the robot safety separation zone. Unlike conventional dynamic disturbances such as machine breakdown or order insertion, this event changes the feasible collaboration mode of the unfinished operation remainder rather than only delaying a resource or adding a job. This study formulates a carbon-aware dynamic HRC-FJSP and evaluates a carbon-aware multi-agent deep reinforcement learning scheduler (CA-MADRL) with local recovery after safety-proximity-induced collaboration disruption. The objective combines normalized makespan, carbon emission, and human workload imbalance with carbon accounting based on operation energy and time-varying grid carbon intensity. Across the benchmark cases, CA-MADRL obtains the best average global criterion (0.7235), wins nine of 12 cases, and achieves the lowest average carbon emissions among the compared policies (48.991 kg CO2e). Sensitivity analysis shows that stronger carbon preference reduces emissions but increases makespan and tardiness, while adaptive collaboration outperforms fixed human–robot, human-only, and robot-only regimes. The results indicate that dynamic mode adaptation and local rescheduling improve carbon-aware collaborative schedules under safety disruption. Full article
(This article belongs to the Special Issue Human-Centred Manufacturing Towards Industry 5.0)
Show Figures

Figure 1

53 pages, 820 KB  
Systematic Review
Applications of Reinforcement Learning for Autonomous Surgical Robotics: A Systematic Review
by Muhammad Shahid, Abdullah, Zulaikha Fatima, Wasif Feroze, Miguel Jesús Torres Ruiz, Magdalena Saldaña-Pérez, Carlos Guzmán Sánchez-Mejorada and Rolando Quintero Tellez
Biomimetics 2026, 11(8), 577; https://doi.org/10.3390/biomimetics11080577 - 12 Aug 2026
Viewed by 95
Abstract
Reinforcement learning (RL) has emerged as a promising approach for autonomous surgical robotic subtasks. Recent advances include deep reinforcement learning (DRL), imitation learning (IL), and vision–language–action (VLA) models. However, current evidence remains fragmented across simulation benchmarks, task-specific demonstrations, and limited clinical studies. Existing [...] Read more.
Reinforcement learning (RL) has emerged as a promising approach for autonomous surgical robotic subtasks. Recent advances include deep reinforcement learning (DRL), imitation learning (IL), and vision–language–action (VLA) models. However, current evidence remains fragmented across simulation benchmarks, task-specific demonstrations, and limited clinical studies. Existing reviews primarily focus on RL algorithms, while the broader pathway from algorithm development to clinically deployable surgical autonomy has not been comprehensively synthesised. This PRISMA 2020-guided systematic review examines RL, IL, safe RL, simulation-to-real (sim-to-real) transfer, foundation models, VLA systems, and regulatory readiness in surgical robotics. We searched IEEE Xplore, PubMed/MEDLINE, Embase, Scopus, Web of Science, the Cochrane Library, ACM Digital Library, arXiv, and medRxiv for studies published between January 2015 and March 2026, with additional studies identified through backward citation tracing. Eligible studies proposed novel RL, imitation learning, or foundation-model approaches for surgical robotics with empirical validation in simulation or on physical robotic platforms. Two reviewers independently extracted data using a predefined coding scheme, and a third reviewer resolved disagreements. Owing to substantial heterogeneity in platforms, tasks, and outcome measures, a quantitative meta-analysis was not feasible; therefore, the evidence was synthesised narratively using a comparative framework. A total of 220 studies met the inclusion criteria, covering eleven active surgical RL platforms, seven paired sim-to-real studies, emerging foundation-model architectures, and three FDA-cleared robotic systems exhibiting Level 3 autonomy. Available comparative studies suggest that hierarchical approaches can outperform flat policies in long-horizon tasks, while language-conditioned models demonstrated promising multi-step surgical capabilities. Seven paired simulation-to-real studies were identified, encompassing tissue retraction, guidewire navigation, and surgical cutting tasks. Sim-to-real performance gaps varied substantially by task and metric, with success-rate gaps ranging from −10 to 50 percentage points (negative values indicating better real-world than simulated performance), while paired mean spatial errors differed by at most 0.61 mm. Most studies employed domain randomization or visual domain adaptation; hierarchical reinforcement learning demonstrated advantages over flat policies in multi-step surgical tasks. Explicit safety-constrained methods (CPO, CBF, and SER), formal verification, and regulatory-aligned evaluation were reported in fewer than 3% of applied studies. Most evidence remained simulation-based, with no reported autonomous RL execution in vivo in humans. Overall, RL-based surgical robotics appears mature at the simulation stage but remains preclinical for autonomous clinical deployment. Future progress requires stronger sim-to-real validation, multimodal safety-aware architectures, alignment with IEC 62304, ISO 14971, FDA guidance, and the EU AI Act, and open benchmarks that jointly evaluate performance, safety, and surgeon trust. Full article
(This article belongs to the Section Locomotion and Bioinspired Robotics)
Show Figures

Graphical abstract

22 pages, 1042 KB  
Article
Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning in Long-Horizon AntMaze Navigation
by Lidong Sun, Ye Wang, Zheheng Fan and Fuchun Sun
Machines 2026, 14(8), 928; https://doi.org/10.3390/machines14080928 - 12 Aug 2026
Viewed by 121
Abstract
Long-horizon navigation requires a high-level policy to select locally reachable subgoals, yet a scalar task reward provides little information about how an unsuitable proposal should be changed. We introduce Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning (SCS-HRL), a two-level method in which [...] Read more.
Long-horizon navigation requires a high-level policy to select locally reachable subgoals, yet a scalar task reward provides little information about how an unsuitable proposal should be changed. We introduce Simulated Corrective Subgoal Supervision for Hierarchical Reinforcement Learning (SCS-HRL), a two-level method in which a topology- and clearance-aware programmatic supervisor evaluates each proposed subgoal and returns both a scalar score and a continuous target in the same subgoal space. The score trains the high-level critic, and the target enters a masked regression term for the high-level actor. Primitive actions are always conditioned on the actor’s subgoal; the supervisor is inactive during learned-policy evaluation. In AntMaze, using 6000 training episodes, five seeds, and 100 deterministic evaluation episodes per seed, SCS-HRL attained an 88.4±7.8% final success rate (mean ± sample standard deviation; 95% Student-t confidence interval [78.7%,98.1%]). The matched scalar-only condition and HIRO attained 0% rates. Applying the same route rule directly to the SCS-HRL low-level controllers yielded 82.2±9.9% success; the paired difference favored the learned high-level policy by 6.2 percentage points (95% confidence interval [2.1,10.3], p=0.013). Across three matched seeds, nonzero corrective weights of 0.5, 1.0, and 2.0 remained stable, whereas 0.25 was seed-sensitive. Term-level ablations further show that the continuous target, rather than the exact scalar-shaping formula, was the principal additional signal. Separate fixed-policy tests obtained 0% success rates on two unseen maze layouts. These results indicate that continuous subgoal targets can encode task-specific route information in the source maze, while cross-layout transfer remains unresolved. Full article
(This article belongs to the Special Issue Machine Learning Application in Robots)
Show Figures

Figure 1

23 pages, 1919 KB  
Article
Artificial Intelligence, Tourism Development, and Ecological Footprint in Advanced Economies: Evidence from MMQR and PQQKRLS Approaches
by Muhammad Sonail, Deyi Xu, Zohaib Hassan, Farrukh Fazal and Mojawir Ahmad Sadat
Economies 2026, 14(8), 338; https://doi.org/10.3390/economies14080338 - 12 Aug 2026
Viewed by 116
Abstract
Achieving environmental sustainability, particularly the targets outlined in Sustainable Development Goal 13 (Climate Action), is a critical global imperative. This investigation analyzes the heterogeneous effects of artificial intelligence (AI), tourism intensity, tourism expenditure, the Gross Domestic Product (GDP) share contributed by tourism, natural [...] Read more.
Achieving environmental sustainability, particularly the targets outlined in Sustainable Development Goal 13 (Climate Action), is a critical global imperative. This investigation analyzes the heterogeneous effects of artificial intelligence (AI), tourism intensity, tourism expenditure, the Gross Domestic Product (GDP) share contributed by tourism, natural resource rents, and environmental policy stringency on the ecological footprints of advanced countries from 2000 to 2022. Using a robust analytical framework featuring advanced econometric methods, specifically the Method of Moments Quantile Regression (MMQR) and an innovative machine learning approach—Panel Quantile-on-Quantile Kernel-Based Regularized Least Squares (PQQKRLS)—the research elucidates complex, nonlinear interdependencies. Key empirical results show that AI adoption significantly mitigates ecological footprints across all quantile distributions. Conversely, heightened tourism intensity and increased tourism expenditure are associated with greater environmental degradation. The analysis further indicates a U-shaped tourism–ecological footprint relationship, suggesting that tourism’s economic contribution may initially reduce ecological pressure but may increase it again beyond a certain expansion threshold. These conclusions underscore the necessity for advanced nations to adopt synergistic policy frameworks that strategically leverage AI technologies, promote sustainable tourism practices, and reinforce rigorous environmental governance to advance climate action and ecological sustainability. Full article
(This article belongs to the Special Issue Advances in Applied Economics: Trade, Growth and Policy Modeling)
Show Figures

Figure 1

29 pages, 1976 KB  
Article
Physics-Constrained Dual-Attention Reinforcement Learning for Semi-Active Lateral Vibration Control of High-Speed Trains
by Dongyu Fan, Lei Gao, Runliang Tian, Yiwei Zhao and Zhaoyang Xing
Actuators 2026, 15(8), 437; https://doi.org/10.3390/act15080437 - 12 Aug 2026
Viewed by 73
Abstract
High-speed trains are prone to severe lateral vibrations induced by track irregularity excitations under complex operating conditions, which deteriorate ride comfort, reduce running stability, and accelerate wheel–rail wear. Existing lateral suspension vibration control methods mainly rely on fixed parameters and expert experience. To [...] Read more.
High-speed trains are prone to severe lateral vibrations induced by track irregularity excitations under complex operating conditions, which deteriorate ride comfort, reduce running stability, and accelerate wheel–rail wear. Existing lateral suspension vibration control methods mainly rely on fixed parameters and expert experience. To further optimize and improve the control performance of semi-active lateral suspension systems for high-speed trains, this paper proposes the Physics-Constrained Dual-Attention Reinforcement Learning for Semi-Active Lateral Vibration Control of High-Speed Trains (SA-PRL). The proposed method is built upon a semi-active Twin Delayed Deep Deterministic Policy Gradient algorithm (SATD3). To ensure the physical realizability of control actions, the Logical Constraints of Skyhook Control (LCSC) are introduced to map the controller output into physically feasible damping commands. Furthermore, to enhance the perception of critical state features and improve value estimation capability, a Critic Network with Dual-Head Self-Attention (CDHA) is developed. Based on a railway vehicle lateral dynamic model, a series of simulation experiments are conducted under four excitation conditions, including single-peak sinusoidal excitation and multi-peak sinusoidal excitation, as well as the Chinese high-speed railway track irregularity (CHSRTI) and German low-interference track irregularity (GLITI) excitations. The results demonstrate the lateral vibration suppression capability of the proposed method over passive suspension, skyhook damping control, and existing reinforcement learning methods, with reductions of 47.5% and 47.7% in the RMS carbody lateral acceleration under the CHSRTI and GLITI spectra, respectively. The proposed method provides an effective solution for semi-active lateral vibration control of high-speed trains. Full article
(This article belongs to the Special Issue Vibration Control Based on Intelligent Actuators and Sensors)
Show Figures

Figure 1

23 pages, 2504 KB  
Article
Game-Theoretic Reinforcement Learning Framework for Local Multi-Vehicle Interactive Guided Trajectory Generation in Representative Traffic Scenarios
by Chagen Luo, Weifu Wang, Yadong Wang and Di Zhang
Vehicles 2026, 8(8), 187; https://doi.org/10.3390/vehicles8080187 - 12 Aug 2026
Viewed by 75
Abstract
The generation of guided trajectories for autonomous vehicles in multi-vehicle interaction scenarios remains challenging because surrounding vehicles continuously adapt their actions under uncertainty. This paper proposes a game-theoretic reinforcement learning (GT-RL) framework that combines a posterior-weighted rolling-horizon local game with proximal policy optimization [...] Read more.
The generation of guided trajectories for autonomous vehicles in multi-vehicle interaction scenarios remains challenging because surrounding vehicles continuously adapt their actions under uncertainty. This paper proposes a game-theoretic reinforcement learning (GT-RL) framework that combines a posterior-weighted rolling-horizon local game with proximal policy optimization (PPO) trajectory refinement. The revision makes the incomplete-information cost explicit through a normalized Bayesian posterior, a posterior expected cost, and a certainty-equivalent numerical approximation used by the SQP best-response solver. Using the supplied run-level logs (500 runs per method and scenario), GT-RL achieved mean safety scores of 94.1% (95% bootstrap CI: 93.91–94.28) in the intersection scenario and 91.5% (91.33–91.66) in the highway-merging scenario, with zero collisions in both sets of 500 runs. Paired comparisons with DQN, PPO, MPC-only, and potential-field baselines were significant after Holm correction (p < 0.001) for safety, traversal time, and comfort; DQN and PPO were faster in some cases but had lower safety and comfort. At the intersection, the rule-based method had higher safety but required 4.2 s more traversal time and had a 7.8-point lower comfort score than GT-RL. The findings support game-theoretic reasoning as a strategic prior for local interaction-aware trajectory generation. Claims are restricted to the tested low-to-moderate-speed, non-limit-handling simulations; high-fidelity vehicle dynamics and empirical large-scale timings remain areas of future study. Full article
(This article belongs to the Special Issue Trajectory Tracking of Autonomous Vehicles)
Show Figures

Figure 1

26 pages, 675 KB  
Article
Reinforcement Learning-Based Day-to-Day Route Choice Models Under Full and Partial Information Conditions: A Theoretical and Experimental Research
by Yuance Yang, Ning Jia, Nianlu Ren, Hongye Fan and Zhanghao Lei
Mathematics 2026, 14(16), 2910; https://doi.org/10.3390/math14162910 - 12 Aug 2026
Viewed by 179
Abstract
Learning behavior plays an important role in the day-to-day route choice process, which is essentially a repeated decision-making process. This paper proposes two learning models for the day-to-day route choice process, one for the full-information (FI) condition and one for the partial-information (PI) [...] Read more.
Learning behavior plays an important role in the day-to-day route choice process, which is essentially a repeated decision-making process. This paper proposes two learning models for the day-to-day route choice process, one for the full-information (FI) condition and one for the partial-information (PI) condition. Both models are developed based on attraction-based reinforcement learning theory but with different learning mechanisms. In our FI model, travelers learn by comparing different paths, while in our PI model, travelers’ route choice decision is modeled by long-term cost minimization and a policy-based reinforcement algorithm. Both models are rooted in individual strategy updating and are theoretically linked to classical network flow assignment theory: we establish that individual-level invariance of choice probabilities is sufficient for aggregate stationarity consistent with Wardrop equilibrium, and provide an explicit stability condition for the FI dynamics. To validate the models, laboratory experiments under FI and PI conditions were designed and conducted. The experimental data confirm the presence of learning behavior and show that the proposed models achieve better predictive performance than several behavioral benchmarks, including EWA, Q-learning, mixed logit, and a Selten variant. Moreover, an interesting phenomenon was found in the FI experimental dataset: even though FI was provided, some subjects still made decisions hinging on their own experience. Our findings would benefit traffic information services and traffic administration. Full article
(This article belongs to the Special Issue AI, Machine Learning and Optimization)
Show Figures

Figure 1

27 pages, 23489 KB  
Article
Toward Self-Evolving Lunar Robotic Autonomy Through Contract-Governed Skill Registration
by Bingqi Huang, Bingchuan Wei, Yingkai Cai and Zhaokui Wang
Astronautics 2026, 1(3), 15; https://doi.org/10.3390/astronautics1030015 - 11 Aug 2026
Viewed by 98
Abstract
Permanent lunar habitation will require robotic systems that can maintain infrastructure, recover from local failures, and acquire new operational capabilities under limited Earth supervision. Existing planetary robots are largely fixed-function specialists, while end-to-end foundation-model policies remain difficult to validate and extend for safety-critical [...] Read more.
Permanent lunar habitation will require robotic systems that can maintain infrastructure, recover from local failures, and acquire new operational capabilities under limited Earth supervision. Existing planetary robots are largely fixed-function specialists, while end-to-end foundation-model policies remain difficult to validate and extend for safety-critical surface operations. We present SELENE (Self-Evolving Lunar Embodied ageNt Ecosystem), an architectural proposal for contract-governed lunar robotic autonomy centered on a shared Atomic Action Library A. The key abstraction is the Atomic Action Contract: a typed skill interface that specifies parameters, preconditions, goal predicates, execution bindings, safety envelopes, runtime reports, and validation metadata. Through this contract, a VLM-driven Cognitive Agent plans over executable skills, a multi-modal Execution Agent realizes them through optimization-based controllers, Vision–Language–Action (VLA) policies, Vision–Language–Navigation (VLN) policies, or reinforcement-learned policies, and an offline Evolutionary Agentic Framework synthesizes and registers new candidate contracts without modifying the planner or the execution interface. This paper presents an architecture-level validation of that contract mechanism. We instantiate SELENE across two heterogeneous pathways on LunarBot and its simulation counterpart, with optimization-based control supported as a third execution modality. A pre-trained VLA policy adapted from 100 teleoperated demonstrations achieves 29/30 task success (96.7 percent) in in-domain trials on the physical LunarBot. A curriculum–RL policy instantiates the traversal pathway in simulated lunar-gravity terrain. Together, these results show that the Atomic Action Contract can serve as a common registration and dispatch interface across heterogeneous control modalities. The same contract layer also defines the path toward runtime gap-triggered self-evolution, mission-grade admission, and lunar-environment validation in subsequent system-level studies. Full article
Show Figures

Figure 1

Back to TopTop