Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (44)

Search Parameters:
Keywords = randomness level (RL)

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
20 pages, 4604 KB  
Article
Enhancing Sim-to-Real Transfer for a High-Gear-Ratio Quadruped Robot via Extended Actuator Dynamics Identification
by Hansol Kang, Hyunyong Lee, Jiman Park, Seongwon Nam, Yeongwoo Son, Bumsu Yi, Jaeyoung Oh, Hyeonwoo Yu and Hyouk Ryeol Choi
Machines 2026, 14(9), 1031; https://doi.org/10.3390/machines14091031 - 9 Sep 2026
Viewed by 326
Abstract
Reinforcement learning (RL) has become a powerful tool for quadrupedal locomotion, and a sim-to-real approach is widely adopted to avoid hardware damage during training. However, the “sim-to-real gap” remains a critical challenge, particularly for robots driven by high-gear-ratio actuators, in which nonlinear friction [...] Read more.
Reinforcement learning (RL) has become a powerful tool for quadrupedal locomotion, and a sim-to-real approach is widely adopted to avoid hardware damage during training. However, the “sim-to-real gap” remains a critical challenge, particularly for robots driven by high-gear-ratio actuators, in which nonlinear friction effects are strongly amplified. Conventional methods, such as actuator networks or heuristic domain randomization, often require specialized sensors or extensive trial-and-error to tune appropriate randomization ranges. Building on a recent system-identification framework for actuator dynamics, we extend it with an augmented friction model that incorporates the Stribeck effect to capture the low-velocity nonlinearities characteristic of high-gear-ratio actuators. The physical parameters are identified from real-robot trajectory data using an evolutionary algorithm, and the resulting simulation is used to train a locomotion policy that is transferred zero-shot to a 55 kg quadruped without additional fine-tuning or base- or controller-level dynamics randomization. On our platform, adding the Stribeck term lowers the actuator identification error by 16% relative to a Coulomb–Viscous model on the trajectory used for identification, and this advantage generalizes to an unseen trajectory not used for identification. It also lowers the simulation-to-reality mean-velocity degradation from 40.2% and 26.8% for the Coulomb–Viscous model to 27.2% and 21.6% for our method at the 0.3 and 1.0 m/s commands, respectively. The trained policy achieves stable locomotion on flat ground as well as rough terrain including steps and stairs. These results indicate that explicitly modeling low-velocity friction is beneficial for high-fidelity sim-to-real transfer in high-reduction systems. Full article
(This article belongs to the Special Issue The Future of Mobility: Exploring Wheeled–Legged Robot Systems)
Show Figures

Figure 1

27 pages, 882 KB  
Article
A Center-Guided Reinforcement Learning Method for Hyperparameter Optimization and Its Application to Relation Extraction
by Yangbin Tan, Liping Mo and Yu Yan
Mach. Learn. Knowl. Extr. 2026, 8(9), 264; https://doi.org/10.3390/make8090264 - 28 Aug 2026
Viewed by 249
Abstract
Hyperparameter optimization (HPO) aims to identify high-quality model configurations under a limited evaluation budget. To address mixed search spaces, sparse feedback, and low sample efficiency in reinforcement learning (RL)-based HPO, a Center-Guided Reinforcement Learning (CGRL) method is proposed. In CGRL, the policy output [...] Read more.
Hyperparameter optimization (HPO) aims to identify high-quality model configurations under a limited evaluation budget. To address mixed search spaces, sparse feedback, and low sample efficiency in reinforcement learning (RL)-based HPO, a Center-Guided Reinforcement Learning (CGRL) method is proposed. In CGRL, the policy output is reformulated from a configuration to be directly evaluated into a search center that defines a promising region, decoupling region-level guidance from exact configuration selection. A mixed candidate pool is generated around the center, and a promising candidate for real evaluation is selected by a Random Forest surrogate model. Meanwhile, a process-aware reward provides dense and informative feedback for policy learning. Experiments on 20 Yet Another Hyperparameter Optimization (YAHPO) Gym environments validate the effectiveness of CGRL. Compared with random search (RS), Tree-structured Parzen Estimator (TPE), Sequential Model-based Algorithm Configuration 3 (SMAC3), a Proximal Policy Optimization baseline (PPO-basic), Hyperparameter Optimization by Reinforcement Learning (Hyp-RL), and Q-Learning for Hyperparameter Optimization (HyperQ-Opt), CGRL achieves the best average rank of 1.800 in terms of the final best objective value, versus 6.000, 3.600, 2.200, 4.450, 6.350, and 3.600, respectively. For Low-Rank Adaptation (LoRA) HPO for relation extraction (RE) from ancient Chinese historical documents, CGRL improves Macro-F1 by 8.66%, 3.10%, 3.18%, and 5.13% on the validation set relative to RS, TPE, SMAC3, and PPO, respectively, and by 11.15%, 2.03%, 3.11%, and 9.59% on the test set. These results demonstrate the effectiveness of CGRL for limited-budget HPO and its applicability to practical RE tasks. Full article
Show Figures

Figure 1

53 pages, 820 KB  
Systematic Review
Applications of Reinforcement Learning for Autonomous Surgical Robotics: A Systematic Review
by Muhammad Shahid, Abdullah, Zulaikha Fatima, Wasif Feroze, Miguel Jesús Torres Ruiz, Magdalena Saldaña-Pérez, Carlos Guzmán Sánchez-Mejorada and Rolando Quintero Tellez
Biomimetics 2026, 11(8), 577; https://doi.org/10.3390/biomimetics11080577 - 12 Aug 2026
Viewed by 1149
Abstract
Reinforcement learning (RL) has emerged as a promising approach for autonomous surgical robotic subtasks. Recent advances include deep reinforcement learning (DRL), imitation learning (IL), and vision–language–action (VLA) models. However, current evidence remains fragmented across simulation benchmarks, task-specific demonstrations, and limited clinical studies. Existing [...] Read more.
Reinforcement learning (RL) has emerged as a promising approach for autonomous surgical robotic subtasks. Recent advances include deep reinforcement learning (DRL), imitation learning (IL), and vision–language–action (VLA) models. However, current evidence remains fragmented across simulation benchmarks, task-specific demonstrations, and limited clinical studies. Existing reviews primarily focus on RL algorithms, while the broader pathway from algorithm development to clinically deployable surgical autonomy has not been comprehensively synthesised. This PRISMA 2020-guided systematic review examines RL, IL, safe RL, simulation-to-real (sim-to-real) transfer, foundation models, VLA systems, and regulatory readiness in surgical robotics. We searched IEEE Xplore, PubMed/MEDLINE, Embase, Scopus, Web of Science, the Cochrane Library, ACM Digital Library, arXiv, and medRxiv for studies published between January 2015 and March 2026, with additional studies identified through backward citation tracing. Eligible studies proposed novel RL, imitation learning, or foundation-model approaches for surgical robotics with empirical validation in simulation or on physical robotic platforms. Two reviewers independently extracted data using a predefined coding scheme, and a third reviewer resolved disagreements. Owing to substantial heterogeneity in platforms, tasks, and outcome measures, a quantitative meta-analysis was not feasible; therefore, the evidence was synthesised narratively using a comparative framework. A total of 220 studies met the inclusion criteria, covering eleven active surgical RL platforms, seven paired sim-to-real studies, emerging foundation-model architectures, and three FDA-cleared robotic systems exhibiting Level 3 autonomy. Available comparative studies suggest that hierarchical approaches can outperform flat policies in long-horizon tasks, while language-conditioned models demonstrated promising multi-step surgical capabilities. Seven paired simulation-to-real studies were identified, encompassing tissue retraction, guidewire navigation, and surgical cutting tasks. Sim-to-real performance gaps varied substantially by task and metric, with success-rate gaps ranging from −10 to 50 percentage points (negative values indicating better real-world than simulated performance), while paired mean spatial errors differed by at most 0.61 mm. Most studies employed domain randomization or visual domain adaptation; hierarchical reinforcement learning demonstrated advantages over flat policies in multi-step surgical tasks. Explicit safety-constrained methods (CPO, CBF, and SER), formal verification, and regulatory-aligned evaluation were reported in fewer than 3% of applied studies. Most evidence remained simulation-based, with no reported autonomous RL execution in vivo in humans. Overall, RL-based surgical robotics appears mature at the simulation stage but remains preclinical for autonomous clinical deployment. Future progress requires stronger sim-to-real validation, multimodal safety-aware architectures, alignment with IEC 62304, ISO 14971, FDA guidance, and the EU AI Act, and open benchmarks that jointly evaluate performance, safety, and surgeon trust. Full article
(This article belongs to the Section Locomotion and Bioinspired Robotics)
Show Figures

Graphical abstract

22 pages, 700 KB  
Article
Cross-Layer Resource Optimization for Ultra-Low-Power TinyML Inference on ARM Cortex-M Microcontrollers
by Abdulaziz G. Alanazi, Haifa A. Alanazi and Nasser S. Albalawi
Electronics 2026, 15(13), 2918; https://doi.org/10.3390/electronics15132918 - 3 Jul 2026
Viewed by 633
Abstract
Running neural networks on battery-powered Internet of Things (IoT) sensor nodes is difficult because flash memory, SRAM, latency, and energy per inference are limited at the same time. Existing TinyML co-design methods usually improve model size or memory use, but runtime voltage–frequency control [...] Read more.
Running neural networks on battery-powered Internet of Things (IoT) sensor nodes is difficult because flash memory, SRAM, latency, and energy per inference are limited at the same time. Existing TinyML co-design methods usually improve model size or memory use, but runtime voltage–frequency control is often handled as a separate step. This separation limits energy saving because the power policy does not use the layer-wise compute profile of the final compressed model. We propose the Cross-Layer Resource Optimizer (CLRO), a three-stage resource optimization pipeline for TinyML inference on an ARM Cortex-M7 target. The first stage, Mixed-Precision Aware Pruning and Distillation (MPAD), assigns per-layer bit widths and pruning ratios using calibration-set sensitivity scores. The second stage, consisting of the Activation Lifetime-Aware Tensor Scheduler (ALTS), uses the compressed graph to find an execution order that reduces peak live static random-access memory (SRAM). The third stage, Reinforcement Learning-Based Dynamic Voltage and Frequency Scaling (DVFS-RL), trains a tabular Q-learning policy from the multiply–accumulate (MAC) utilization profile of the compressed and scheduled model. The learned voltage–frequency policy is stored as a small flash lookup table, so it adds no runtime decision cost during inference. We evaluate the CLRO on all four MLPerf Tiny tasks using an STM32H743ZI microcontroller with 512 kB SRAM and 2 MB flash. The CLRO reaches 91.7% image classification accuracy, 95.4% keyword-spotting accuracy, 89.6% visual wake words accuracy, and 0.913 anomaly detection AUC. The final deployment uses 198 kB flash and 174 kB peak SRAM, with 387 μJ energy per inference and 38 ms latency. Compared with the MCUNet baseline, the CLRO reduces energy by 58.1% and peak SRAM by 39% while keeping the same accuracy level. Full article
Show Figures

Figure 1

24 pages, 2593 KB  
Article
Regional Strategy Composition: A Hierarchical-Action Reinforcement Learning Framework for Dynamic Smart-Meter Association over 5G NR mMTC Networks
by Muhammed Al-Ali, Esteban Inga, Juan Inga and Elias Yaacoub
Future Internet 2026, 18(7), 337; https://doi.org/10.3390/fi18070337 - 25 Jun 2026
Viewed by 594
Abstract
Advanced Metering Infrastructure (AMI) over 5G New Radio (NR) massive machine-type communication (mMTC) networks require efficient and adaptive communication mechanisms to support reliable data delivery for large numbers of smart meters under dynamic traffic and channel conditions. In this work, we propose a [...] Read more.
Advanced Metering Infrastructure (AMI) over 5G New Radio (NR) massive machine-type communication (mMTC) networks require efficient and adaptive communication mechanisms to support reliable data delivery for large numbers of smart meters under dynamic traffic and channel conditions. In this work, we propose a framework in which each smart meter chooses, at runtime, whether to transmit directly to the base station (BS) or via a nearby Data Aggregation Point (DAP). The optimal choice is dynamic and depends on DAP buffer occupancy, periodic congestion, channel quality, and packet deadline pressure. Formulating this as a per-meter binary decision yields an action space of size 2N for N meters, which is intractable for reinforcement learning (RL). We reformulate the problem as regional strategy composition: the RL agent selects one parameterized association strategy for each DAP region from a small library of interpretable rules, and a deterministic mapping expands the regional choice into per-meter modes. It reduces the policy action space from 2N to KD, where D is the number of DAPs and K the number of strategies, while preserving meter-level control granularity. We evaluate Proximal Policy Optimization (PPO) and Deep Q-Network (DQN) controllers against eight meter-level baselines on a 5G NR-calibrated simulator with 1500 m, six DAPs, deadline-bounded delivery, stale channel-state information, and phase-offset congestion cycles. Across three traffic regimes and five random seeds, PPO improves packet delivery ratio (PDR) over the strongest heuristic by +0.63, +2.41, and +2.66 percentage points under baseline, high-load, and bursty-cycle conditions, respectively; all gains are statistically significant (paired t-test, p<0.001; Cohen’s d up to 5.12), and the advantage grows with traffic stress. The results show that learned regional composition of classical heuristics outperforms any single fixed heuristic precisely when no individual rule is globally optimal. Full article
(This article belongs to the Special Issue Artificial Intelligence in Smart Grids)
Show Figures

Graphical abstract

22 pages, 32308 KB  
Article
Mastering the Twin–Game: Hierarchical Reinforcement Learning in a Digital Twin Sandbox for Adaptive Urban Healthcare Optimization—A Case Study of Wuhan
by Yuxuan Hu, Shaohua Wang and Haojian Liang
ISPRS Int. J. Geo-Inf. 2026, 15(6), 273; https://doi.org/10.3390/ijgi15060273 - 16 Jun 2026
Cited by 2 | Viewed by 842
Abstract
Urban healthcare systems are fundamentally constrained by the mismatch between static resource configurations and dynamically evolving patient demand. Under the tiered healthcare system, traditional static planning methods struggle to capture the complexity and randomness of patient flows. While recent reinforcement learning (RL) approaches [...] Read more.
Urban healthcare systems are fundamentally constrained by the mismatch between static resource configurations and dynamically evolving patient demand. Under the tiered healthcare system, traditional static planning methods struggle to capture the complexity and randomness of patient flows. While recent reinforcement learning (RL) approaches enable adaptive decision-making, they suffer from dimensionality explosion and unstable convergence due to massive action spaces and delayed spatiotemporal credit assignment in city-scale environments. To address this gap, we propose Twin–Game: a digital twin-driven hierarchical reinforcement learning (HRL) framework that formulates adaptive healthcare resource optimization as a “Twin Game” between a simulation-based game environment (Strategic Sandbox) and a hierarchical decision policy. First, we construct the “first twin”—an offline digital twin that serves as the Strategic Sandbox parameterized with Wuhan’s observed facility, population, and transportation data, while patient arrivals and disease profiles are generated synthetically under documented assumptions because individual-level clinical flow data are not publicly available. This environment integrates a dynamic gravity model with a two-way referral mechanism to represent the nonlinear coupling between hospital attractiveness, crowding levels, and patient choice behaviors. Second, we build the “second twin”—an Option-based HRL policy. The Manager (Macro-level Strategic Layer) uses a Deep Q-Network (DQN) for discrete spatial attention allocation; the Worker (Micro-level Execution Layer) uses Proximal Policy Optimization (PPO) for continuous, fine-grained controls such as bed expansion ratios and personnel scheduling. The two twins interact in a closed-loop game, performing strategy search and game evolution under complex constraints to optimize allocation. Experimental results from the Wuhan case indicate that the Twin–Game framework outperforms static baselines and single-layer RL in reducing average travel times, enhancing resource utilization, and improving tiered diagnosis and treatment within the simulation setting. The results should be interpreted as simulation-based decision-support evidence rather than direct clinical validation. This study provides a data-driven, game-theoretic decision support tool for building resilient urban healthcare systems. Full article
Show Figures

Figure 1

24 pages, 5986 KB  
Article
Multi-Scale Quality Evaluation of Road Markings from Glass Bead Distribution: A Novel GSMTA Framework
by Xiaosong Lu, Haoqin Guo, Rui He, Hao Wu, Liangliang Li and Jianrong Hu
Materials 2026, 19(11), 2244; https://doi.org/10.3390/ma19112244 - 26 May 2026
Viewed by 655
Abstract
Road markings constitute essential traffic control elements that ensure safety and traffic flow efficiency. The nighttime visibility of road markings, quantified through retroreflective luminance (RL), is fundamentally governed by the distribution characteristics of embedded glass beads (GBs) within the marking [...] Read more.
Road markings constitute essential traffic control elements that ensure safety and traffic flow efficiency. The nighttime visibility of road markings, quantified through retroreflective luminance (RL), is fundamentally governed by the distribution characteristics of embedded glass beads (GBs) within the marking matrix. Yet, three persistent limitations hinder reliable GB distribution evaluation: measurement variability, oversimplified model assumptions (fixed 50% embedment depth vs. observed 50–60% variations), and fragmented correlations between GB morphology and RL metrics. This study proposes a granulometric–spatial–morphological triad assessment (GSMTA) framework, integrating instance segmentation with hierarchical performance analytics. The GSMTA framework achieves 15% higher segmentation accuracy over Otsu/Fast Random Forest methods, quantifying GB distribution via granulometric (size gradation), spatial (homogeneity index), and morphological (shape factor) descriptors. Through principal component analysis, the derived L3D performance indices establish statistically robust retroreflective luminance (RL) prediction models, with PC1 and PC2 capturing 91.7% of the variance and validation errors. The model remained below 8% error across a 750 mcd·m−2·lx−1 RL range, ensuring reliable and precise performance evaluation. Field validation demonstrates the framework’s capacity of transforming pixel-level segmentation data into practical quality control metrics. This advancement supports lifecycle management through standardized GB distribution evaluation, overcoming prior incompatibility issues between microscopic morphology analysis and macroscale RL measurements. Full article
Show Figures

Figure 1

16 pages, 2877 KB  
Article
Red Ginseng Extract Intake and Changes in Metabolite Profiles, Gut Microbiota, and Immune Responses of Healthy Rats
by Madhuri Sangar, Seong-Hwa Song, Saoraya Chanmuang, Dong-Shin Kim, Gwang-Ju Jang, Hyeon-Jeong Lee, Young Kyoung Rhee, Hee-Do Hong, Chang-Won Cho and Hyun-Jin Kim
Nutrients 2026, 18(9), 1462; https://doi.org/10.3390/nu18091462 - 2 May 2026
Viewed by 1947
Abstract
Background: Red ginseng (RG) exhibits enhanced bioactivity compared to white ginseng. Although the beneficial effects of RG have been well investigated in disease models, its impacts on the metabolome, gut microbiota, and immune response under normal physiological conditions remain poorly understood. Methods: Rats [...] Read more.
Background: Red ginseng (RG) exhibits enhanced bioactivity compared to white ginseng. Although the beneficial effects of RG have been well investigated in disease models, its impacts on the metabolome, gut microbiota, and immune response under normal physiological conditions remain poorly understood. Methods: Rats were randomized into three groups: control (normal diet), RL (low-dose RGE at 100 mg/kg body weight), and RH (high-dose RGE at 200 mg/kg body weight). After five weeks, metabolite profiles of the blood, liver, kidney, and large intestinal contents were analyzed and the gut microbiota was assessed. Splenocytes were isolated and treated with or without ethanol-precipitated carbohydrate fractions isolated from RGE or from intestinal contents, and IL-12 secretion was measured. Additionally, the correlations among biochemical characteristics, metabolites, gut microbiota, and immune markers were analyzed. Results: RGE intake decreased plasma triglycerides, liver function biomarkers, and epididymal adipose tissue weight. It also altered metabolite profiles for plasma, liver, kidney, and intestinal contents and increased the hepatic NAD+/NADH ratio. RGE intake reduced the populations of harmful bacteria, whereas it increased Lachnospiraceae. RGE intake enhanced IL-12 production in splenocytes. Furthermore, splenocytes treated with carbohydrates isolated from the small and large intestinal contents of RGE-fed rats secreted higher IL-12 levels than those of the control group. Conclusions: RGE modulated the gut microbiota, metabolism, and immune responses in healthy rats under normal physiological conditions, warranting further investigation into the underlying mechanisms. Full article
Show Figures

Graphical abstract

32 pages, 85093 KB  
Article
Modeling Seismic Resilience and Hospital Evacuation: A Comparative Analysis of Multi-Agent Reinforcement Learning and Classical Evacuation Models
by Chunlin Bian, Yonghao Guo, Gang Meng, Liuyang Li, Hua Chen, Fuhong Lv and Xiaofeng Chai
Buildings 2026, 16(8), 1538; https://doi.org/10.3390/buildings16081538 - 14 Apr 2026
Viewed by 734
Abstract
Hospitals in earthquake-prone regions must evacuate heterogeneous occupants rapidly while preserving operational continuity under disrupted conditions. However, many hospital-evacuation studies still rely on static routing assumptions or narrowly defined behavioral rules, which limits their value for building-level resilience planning. This paper develops a [...] Read more.
Hospitals in earthquake-prone regions must evacuate heterogeneous occupants rapidly while preserving operational continuity under disrupted conditions. However, many hospital-evacuation studies still rely on static routing assumptions or narrowly defined behavioral rules, which limits their value for building-level resilience planning. This paper develops a comparative hospital-campus evacuation framework that combines GIS-based geodesic routing, heterogeneous agent-based modeling, and reinforcement-learning-based decision policies. Puge County People’s Hospital in Sichuan, China, is used as the case study. Six algorithms are evaluated: three rule-based baselines—Shortest Path (SP), Random Walk (RW), and the Social Force Model (SFM)—together with a training-free density-aware heuristic, Density-Aware Gradient Routing (DAGR), and two reinforcement-learning approaches, Density-Aware Q-Learning (DAQL) and SARSA. Experiments cover three population scales (N{50,100,200}), normal daytime conditions, staffing-variation scenarios, and a blocked-exit disruption scenario, with 30 independent runs for each main condition. The results show that the rule-based and training-free methods remain the most reliable under full multi-agent evaluation: the SFM and RW achieve the highest completion ratios (approximately 100% and 93.5%, respectively), while DAGR provides the strongest balance between completion and evacuation efficiency among the non-trained methods. In contrast, the trained RL agents perform substantially worse in direct multi-agent deployment with DAQL reaching approximately 37% completion and SARSA approximately 17%, highlighting a train–evaluation distribution shift associated with independent Q-learning. The ablation analysis further shows that collision avoidance is the most critical reward component, whereas density-avoidance shaping can unintentionally induce collective deadlock when all agents execute the learned policy simultaneously. Among the enhanced variants, DAQL_RoleAware yields the best overall improvement, increasing the completion ratio to approximately 52% and reducing the 90th-percentile evacuation time to approximately 363 s. Overall, this paper clarifies both the promise and the present limitations of density-aware reinforcement learning for hospital evacuation while providing a more building-centred and reproducible basis for future coordination-aware evacuation design and emergency-planning research. Full article
(This article belongs to the Special Issue Innovative Solutions for Enhancing Seismic Resilience of Buildings)
Show Figures

Figure 1

37 pages, 2896 KB  
Article
Energy-Efficient Resilience Scheduling for Elevator Group Control via Queueing-Based Planning and Safe Reinforcement Learning
by Tingjie Zhang, Tiantian Zhang, Hao Zou, Chuanjiang Li and Jun Huang
Machines 2026, 14(3), 352; https://doi.org/10.3390/machines14030352 - 21 Mar 2026
Viewed by 967
Abstract
High-rise elevator group control systems operate under pronounced nonstationarity during commuting peaks, post-event surges, and capacity degradation, where the waiting time distribution becomes right-tail heavy and stresses service-level agreements (SLAs) defined by coverage and high-quantile targets. At the same time, the time-of-use tariffs [...] Read more.
High-rise elevator group control systems operate under pronounced nonstationarity during commuting peaks, post-event surges, and capacity degradation, where the waiting time distribution becomes right-tail heavy and stresses service-level agreements (SLAs) defined by coverage and high-quantile targets. At the same time, the time-of-use tariffs and carbon constraints sharpen the tension between peak-power control, energy savings, and service capacity. This paper proposes a two-layer resilience scheduling framework that integrates queueing-based planning with safe reinforcement learning (RL) fine-tuning. In the planning layer, parsimonious queueing approximations and scenario-based evaluation construct a finite set of implementable mode cards and emergency switching cards; Sample Average Approximation (SAA) combined with Conditional Value-at-Risk (CVaR) constraints filter candidates to enforce tail-risk-aware service limits while keeping power demand within a prescribed envelope. In the execution layer, online dispatch is formulated as a constrained Markov decision process; within the planning layer limits, action masking and Lagrangian safe RL learn small adaptive adjustments to suppress tail-waiting risk and improve recovery dynamics without increasing peak-power commitments. The experiments under morning peaks and post-event surges confirm tail risk reduction and accelerated recovery. For partial outages, the framework prioritizes SLA coverage and recovery speed, accepting a bounded increase in tail risk as a manageable trade-off. Throughout all tests, peak power remains within the prescribed limits. Improvements persist across random seeds and demand fluctuations, indicating distributional robustness and cross-scenario generalization. Ablation studies further reveal complementary roles: removing the planning layer CVaR screening worsens tail performance, while removing the execution layer action masking increases constraint violations and destabilizes recovery. Full article
Show Figures

Figure 1

28 pages, 1825 KB  
Article
Combinatorial Game Theory and Reinforcement Learning in Cumulative Tic-Tac-Toe via Evaluation Functions
by Kai Li and Wei Zhu
Stats 2026, 9(2), 28; https://doi.org/10.3390/stats9020028 - 10 Mar 2026
Viewed by 1940
Abstract
We introduce cumulative tic-tac-toe, a novel variant of the classic 3×3 tic-tac-toe game in which play continues until the board is completely filled. Each player’s final score is determined by the total number of three-in-a-row sequences they form. Using combinatorial game [...] Read more.
We introduce cumulative tic-tac-toe, a novel variant of the classic 3×3 tic-tac-toe game in which play continues until the board is completely filled. Each player’s final score is determined by the total number of three-in-a-row sequences they form. Using combinatorial game theory (CGT), we establish that under optimal play, the game is a draw, and we characterize its theoretical properties. To empirically validate and optimize practical play, we develop a reinforcement learning (RL) framework based on temporal-difference (TD) learning, which is enhanced with a domain-informed evaluation function to accelerate convergence. The experimental results show that our triplet-coverage difference (TCD) evaluation function reduces the average number of training episodes by approximately 23.1% compared with a random-initialization baseline, a statistically significant improvement at the 5% significance level. These results demonstrate the efficiency of our CGT–RL approach for cumulative tic-tac-toe and suggest that similar methods may be useful for analyzing related combinatorial games. We also discuss potential analogies in domains such as competitive resource allocation and coalition formation, illustrating how cumulative-scoring games connect abstract game-theoretic ideas to practical sequential decision problems. Full article
Show Figures

Figure 1

28 pages, 34395 KB  
Article
Container Slot Allocation with Empty Container Repositioning: A Multi-Objective Optimization Approach
by Lei Huang, Mei Sha, Wenwen Guo and Yinping Gao
J. Mar. Sci. Eng. 2026, 14(5), 424; https://doi.org/10.3390/jmse14050424 - 25 Feb 2026
Viewed by 991
Abstract
Trade imbalances and equipment shortages are making it increasingly important to coordinate container slot allocation with empty container repositioning on liner services. This paper develops an integrated bi-objective mixed-integer model for voyage-level slot planning on a fixed cyclic route. The model jointly decides [...] Read more.
Trade imbalances and equipment shortages are making it increasingly important to coordinate container slot allocation with empty container repositioning on liner services. This paper develops an integrated bi-objective mixed-integer model for voyage-level slot planning on a fixed cyclic route. The model jointly decides booking acceptance, inter-voyage shipment, and empty repositioning with port-level empty-inventory dynamics and leg-based vessel capacity constraints. We optimize two conflicting objectives: maximizing operational profit and minimizing empty container TEU-miles. To solve the model at practical scales, we propose a hybrid evolutionary framework, NSGA-II-RL, which uses a lightweight Q-learning controller to adapt operator and repair choices during NSGA-II evolution. Computational experiments on representative service route instances show that NSGA-II-RL produces diverse Pareto-efficient solutions and improves hypervolume relative to fixed-operator and random-control variants, revealing clear trade-offs between profitability and repositioning intensity. Full article
(This article belongs to the Section Ocean Engineering)
Show Figures

Figure 1

23 pages, 2236 KB  
Technical Note
SmartBuildSim: An Open-Source Synthetic-Twin Framework for Reproducible AI Benchmarking in Smart-Building Analytics
by Tymoteusz Miller, Irmina Durlik, Agnieszka Nowy and Ewelina Kostecka
Sensors 2025, 25(23), 7263; https://doi.org/10.3390/s25237263 - 28 Nov 2025
Cited by 1 | Viewed by 1542
Abstract
This paper introduces SmartBuildSim, an open-source synthetic-twin framework that generates configurable and reproducible multi-sensor building streams using lightweight statistical models with tunable trend, seasonality, correlation, delays, and anomaly mechanisms. Deterministic seeding ensures experiment-level reproducibility, while modular pipelines support unified evaluation across forecasting, anomaly [...] Read more.
This paper introduces SmartBuildSim, an open-source synthetic-twin framework that generates configurable and reproducible multi-sensor building streams using lightweight statistical models with tunable trend, seasonality, correlation, delays, and anomaly mechanisms. Deterministic seeding ensures experiment-level reproducibility, while modular pipelines support unified evaluation across forecasting, anomaly detection, and RL tasks. A comprehensive validation against an ASHRAE Great Energy Predictor III reference signal demonstrates that the synthetic data capture realistic magnitude and variability (KS ≈ 0.32; DTW ≈ 9.69), while preserving interpretable and controllable temporal structure. Benchmark results show that simple linear models achieve strong forecasting performance (RMSE ≈ 21.27), IsolationForest reliably outperforms LOF in anomaly detection (F1 ≈ 0.17 vs. 0.10), and Soft-Q Learning provides substantially more stable RL convergence than tabular Q-learning (variance reduced by >95%). Scenario-level analyses further illustrate reproducible daily cycles, zone-specific differences, and the scalability of model behaviour across building configurations. By combining declarative YAML configurations, deterministic randomness management, and an extensible scenario engine, SmartBuildSim provides a transparent and lightweight alternative to high-fidelity building simulators. The framework offers a practical, reproducible testbed for smart-building AI research, bridging the gap between simplistic synthetic datasets and complex physical digital twins. All code, tables, figures, and a Google Colab workflow are openly available to ensure full replicability. Full article
(This article belongs to the Special Issue Smart Sensing Technology for Industry and Environmental Applications)
Show Figures

Figure 1

21 pages, 4678 KB  
Article
Impact of Beacon Feedback on Stabilizing RL-Based Power Optimization in SLM-Controlled FSO Uplinks Under Turbulence
by Erfan Seifi and Peter LoPresti
Photonics 2025, 12(10), 979; https://doi.org/10.3390/photonics12100979 - 1 Oct 2025
Viewed by 2046
Abstract
Atmospheric turbulence severely limits the stability and reliability of free-space optical (FSO) uplinks by inducing wavefront distortions and random intensity fluctuations. This study investigates the use of reinforcement learning (RL) with beacon-based feedback for adaptive beam shaping in a spatial light modulator (SLM)-controlled [...] Read more.
Atmospheric turbulence severely limits the stability and reliability of free-space optical (FSO) uplinks by inducing wavefront distortions and random intensity fluctuations. This study investigates the use of reinforcement learning (RL) with beacon-based feedback for adaptive beam shaping in a spatial light modulator (SLM)-controlled FSO link. The RL agent dynamically adjusts phase patterns to maximize received signal strength, while the beacon channel provides turbulence estimates that guide the optimization process. Experiments under low, moderate, and high turbulence levels demonstrate that incorporating beacon feedback can enhance link stability in severe conditions, reducing signal variability and suppressing extreme fluctuations. In low-turbulence scenarios, the performance is comparable to non-feedback operation, whereas under high turbulence, beacon-assisted control consistently achieves lower coefficients of variation and improved bit error rate (BER) performance. Under high turbulence replay experiments—where the best-performing RL-learned phase patterns are reapplied without learning—further show that configurations trained with feedback retain robustness, even without real-time turbulence measurements under high turbulence. These results highlight the potential of integrating contextual feedback with RL to achieve turbulence-resilient and stable optical uplinks in dynamic atmospheric environments. Full article
Show Figures

Graphical abstract

43 pages, 1895 KB  
Article
Bi-Level Dependent-Chance Goal Programming for Paper Manufacturing Tactical Planning: A Reinforcement-Learning-Enhanced Approach
by Yassine Boutmir, Rachid Bannari, Abdelfettah Bannari, Naoufal Rouky, Othmane Benmoussa and Fayçal Fedouaki
Symmetry 2025, 17(10), 1624; https://doi.org/10.3390/sym17101624 - 1 Oct 2025
Cited by 2 | Viewed by 853
Abstract
Tactical production–distribution planning in paper manufacturing involves hierarchical decision-making under hybrid uncertainty, where aleatory randomness (demand fluctuations, machine variations) and epistemic uncertainty (expert judgments, market trends) simultaneously affect operations. Existing approaches fail to address the bi-level nature under hybrid uncertainty, treating production and [...] Read more.
Tactical production–distribution planning in paper manufacturing involves hierarchical decision-making under hybrid uncertainty, where aleatory randomness (demand fluctuations, machine variations) and epistemic uncertainty (expert judgments, market trends) simultaneously affect operations. Existing approaches fail to address the bi-level nature under hybrid uncertainty, treating production and distribution decisions independently or using single-paradigm uncertainty models. This research develops a bi-level dependent-chance goal programming framework based on uncertain random theory, where the upper level optimizes distribution decisions while the lower level handles production decisions. The framework exploits structural symmetries through machine interchangeability, symmetric transportation routes, and temporal symmetry, incorporating symmetry-breaking constraints to eliminate redundant solutions. A hybrid intelligent algorithm (HIA) integrates uncertain random simulation with a Reinforcement-Learning-enhanced Arithmetic Optimization Algorithm (RL-AOA) for bi-level coordination, where Q-learning enables adaptive parameter tuning. The RL component utilizes symmetric state representations to maintain solution quality across symmetric transformations. Computational experiments demonstrate HIA’s superiority over standard metaheuristics, achieving 3.2–7.8% solution quality improvement and 18.5% computational time reduction. Symmetry exploitation reduces search space by approximately 35%. The framework provides probability-based performance metrics with optimal confidence levels (0.82–0.87), offering 2.8–4.5% annual cost savings potential. Full article
Show Figures

Figure 1

Back to TopTop