Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (54)

Search Parameters:
Keywords = conditional imitation learning

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
33 pages, 11051 KB  
Article
Which Training-Data Axes Matter for Conditional Imitation Learning in CARLA? A Leave-One-Out Ablation UnderPure and Guardrailed Deployment
by Laurentiu Carabulea and Claudiu Pozna
Appl. Sci. 2026, 16(17), 8587; https://doi.org/10.3390/app16178587 (registering DOI) - 28 Aug 2026
Abstract
Conditional imitation learning (CIL) for CARLA depends on diverse expert data spanning map, weather, traffic, and control-perturbation axes, yet it remains unclear which axes actually drive closed-loop behavior, and whether ablation conclusions survive deployment guardrails. We train a matched baseline and four equal-budget [...] Read more.
Conditional imitation learning (CIL) for CARLA depends on diverse expert data spanning map, weather, traffic, and control-perturbation axes, yet it remains unclear which axes actually drive closed-loop behavior, and whether ablation conclusions survive deployment guardrails. We train a matched baseline and four equal-budget leave-one-axis-out (LOO) variants of a fixed CIL architecture (v14) and evaluate each under three nested tiers: pure policy rollout, minimal traffic-rule shields, and a fully deployed stack with route blending and recovery. The factorial design comprises 18×5×3 scenario-variant-tier cells, each repeated under n = 5 traffic-seed replicates (1350 closed-loop episodes) to estimate NPC-seed variance on every eval stack. No single withheld axis dominates pooled outcomes. LOO effects are tag- (scenario-category) and spawn- (vehicle starting location) specific: removing perturbation-labeled recovery data costs 285 m on geometry_stress but can gain distance on in-distribution spawns; removing multi-town data changes held-out Town05 mobility on some routes while depressing others. Axis-importance rankings reorder across tiers; Kendall τ between pure and full rankings is 0.0, and guardrails compress or erase pure-tier gaps (e.g., drop_perturbation pooled distance Δ from 93 m to 0 m). Pure-tier seed replicates show that geometry-driven lane-tracking degradation under drop_perturbation is seed-stable; traffic-axis and full-tier drop_traffic loads are more seed-sensitive. We release the evaluation ledgers, parsing scripts, and analysis tooling with the paper. Training-data ablation claims should report per-scenario or tag-stratified LOO metrics under pure evaluation; guardrailed tiers are supplementary deployment checks, not substitutes for isolating what the policy learned. Full article
(This article belongs to the Section Transportation and Future Mobility)
53 pages, 820 KB  
Systematic Review
Applications of Reinforcement Learning for Autonomous Surgical Robotics: A Systematic Review
by Muhammad Shahid, Abdullah, Zulaikha Fatima, Wasif Feroze, Miguel Jesús Torres Ruiz, Magdalena Saldaña-Pérez, Carlos Guzmán Sánchez-Mejorada and Rolando Quintero Tellez
Biomimetics 2026, 11(8), 577; https://doi.org/10.3390/biomimetics11080577 - 12 Aug 2026
Viewed by 455
Abstract
Reinforcement learning (RL) has emerged as a promising approach for autonomous surgical robotic subtasks. Recent advances include deep reinforcement learning (DRL), imitation learning (IL), and vision–language–action (VLA) models. However, current evidence remains fragmented across simulation benchmarks, task-specific demonstrations, and limited clinical studies. Existing [...] Read more.
Reinforcement learning (RL) has emerged as a promising approach for autonomous surgical robotic subtasks. Recent advances include deep reinforcement learning (DRL), imitation learning (IL), and vision–language–action (VLA) models. However, current evidence remains fragmented across simulation benchmarks, task-specific demonstrations, and limited clinical studies. Existing reviews primarily focus on RL algorithms, while the broader pathway from algorithm development to clinically deployable surgical autonomy has not been comprehensively synthesised. This PRISMA 2020-guided systematic review examines RL, IL, safe RL, simulation-to-real (sim-to-real) transfer, foundation models, VLA systems, and regulatory readiness in surgical robotics. We searched IEEE Xplore, PubMed/MEDLINE, Embase, Scopus, Web of Science, the Cochrane Library, ACM Digital Library, arXiv, and medRxiv for studies published between January 2015 and March 2026, with additional studies identified through backward citation tracing. Eligible studies proposed novel RL, imitation learning, or foundation-model approaches for surgical robotics with empirical validation in simulation or on physical robotic platforms. Two reviewers independently extracted data using a predefined coding scheme, and a third reviewer resolved disagreements. Owing to substantial heterogeneity in platforms, tasks, and outcome measures, a quantitative meta-analysis was not feasible; therefore, the evidence was synthesised narratively using a comparative framework. A total of 220 studies met the inclusion criteria, covering eleven active surgical RL platforms, seven paired sim-to-real studies, emerging foundation-model architectures, and three FDA-cleared robotic systems exhibiting Level 3 autonomy. Available comparative studies suggest that hierarchical approaches can outperform flat policies in long-horizon tasks, while language-conditioned models demonstrated promising multi-step surgical capabilities. Seven paired simulation-to-real studies were identified, encompassing tissue retraction, guidewire navigation, and surgical cutting tasks. Sim-to-real performance gaps varied substantially by task and metric, with success-rate gaps ranging from −10 to 50 percentage points (negative values indicating better real-world than simulated performance), while paired mean spatial errors differed by at most 0.61 mm. Most studies employed domain randomization or visual domain adaptation; hierarchical reinforcement learning demonstrated advantages over flat policies in multi-step surgical tasks. Explicit safety-constrained methods (CPO, CBF, and SER), formal verification, and regulatory-aligned evaluation were reported in fewer than 3% of applied studies. Most evidence remained simulation-based, with no reported autonomous RL execution in vivo in humans. Overall, RL-based surgical robotics appears mature at the simulation stage but remains preclinical for autonomous clinical deployment. Future progress requires stronger sim-to-real validation, multimodal safety-aware architectures, alignment with IEC 62304, ISO 14971, FDA guidance, and the EU AI Act, and open benchmarks that jointly evaluate performance, safety, and surgeon trust. Full article
(This article belongs to the Section Locomotion and Bioinspired Robotics)
Show Figures

Graphical abstract

101 pages, 20860 KB  
Review
AI-Enhanced Evolutionary Game Theory for Intelligent Coordination and Adaptive Optimization in Low-Carbon Energy Systems: A Multi-Scale Review from Smart Grids to Carbon Markets
by Guorui Wang, Liang Zhong and Yixuan Zeng
Processes 2026, 14(16), 2568; https://doi.org/10.3390/pr14162568 - 11 Aug 2026
Viewed by 475
Abstract
The modern energy transition has outpaced the control and optimization frameworks built to govern it. As power and energy systems fragment into webs of renewable generators, storage operators, flexible loads, and carbon-constrained firms, the deterministic, single-optimizer models that once sufficed buckle against nonlinearity, [...] Read more.
The modern energy transition has outpaced the control and optimization frameworks built to govern it. As power and energy systems fragment into webs of renewable generators, storage operators, flexible loads, and carbon-constrained firms, the deterministic, single-optimizer models that once sufficed buckle against nonlinearity, bounded rationality, and strategic conflict among parties who learn and revise as they go. Evolutionary game theory (EGT), which traces how strategies propagate through populations by imitation and selection rather than instantaneous optimization, offers a route through this difficulty—one this review develops across three scales of low-carbon coordination central to cleaner production: enterprise-level industrial symbiosis, system-level smart energy operation, and market-level carbon governance. We synthesize three decades of theory alongside the recent fusion of EGT with artificial intelligence, where deep reinforcement learning approximates high-dimensional payoffs, federated learning lets rival firms co-train models without surrendering proprietary data, and blockchain underwrites decentralized mechanism execution. The synthesis is accompanied by two illustrative numerical case studies, constructed for this review rather than drawn from the surveyed literature, whose quantitative outputs are reported below as demonstrations of modeled behavior rather than as empirical measurements. In the first of these, cooperative emergence in industrial symbiosis hinges on critical thresholds that travel from 0.15 to 0.75 as subsidies and transaction costs vary, with anchor-enterprise targeting accelerating cooperation 2.4-fold while cutting outcome variance 3-fold. In smart energy coordination, AI-enhanced learning buys 32 to 41% faster convergence, yet pays 25 to 39% larger oscillations—a speed–stability tension whose resolution lives in a narrow learning-rate band near 0.08 to 0.12, outside which either sluggishness or instability takes hold. Carbon-market behavior turns on price thresholds: emitters switch abruptly from buying quotas toward investing in abatement once the clearing price clears firm-specific triggers, a discrete state switch that smooth equilibrium analysis misses entirely. Across all three domains, fragmented data, path dependence, and regime-switching dynamics recur as the binding constraints on modeling and on governance alike. Four mechanisms prove invariant to scale—the decisive weight of initial conditions, the catalytic leverage of well-positioned anchor agents, the equilibrium-shaping force of institutional design, and the computational reach added by AI integration—which suggests that insight earned in one domain transfers to the others. We close by mapping open problems in heterogeneity modeling, verification under deep uncertainty, and the still-unrealized coupling of digital twins with privacy-preserving learning. EGT emerges not as retrospective description but as prospective guidance for the cooperative transitions on which credible decarbonization depends. Full article
Show Figures

Figure 1

24 pages, 955 KB  
Review
Sensor Fusion and Perception for Autonomous Driving: A Critical Review of Modalities, AI Models, Algorithms, and Industry Configurations
by Esraa Khatab, Fares Fathy, Abdallah AlKholy and Omar Shalash
Mach. Learn. Knowl. Extr. 2026, 8(7), 199; https://doi.org/10.3390/make8070199 - 7 Jul 2026
Cited by 1 | Viewed by 1291
Abstract
Autonomous driving systems rely on a sophisticated pipeline of artificial intelligence models to perceive, predict, and plan in dynamic environments. This review presents a systematic analysis of the machine learning and deep learning models underpinning vehicle autonomy, spanning classical convolutional neural networks (CNNs) [...] Read more.
Autonomous driving systems rely on a sophisticated pipeline of artificial intelligence models to perceive, predict, and plan in dynamic environments. This review presents a systematic analysis of the machine learning and deep learning models underpinning vehicle autonomy, spanning classical convolutional neural networks (CNNs) for object detection and semantic segmentation to recurrent and Transformer-based architectures for trajectory prediction and motion planning. It also provides a critical examination of the autonomous vehicle sensor stack, including cameras, LiDAR, radar, ultrasonics, and GNSS/IMU as data acquisition systems, highlighting modality-specific AI challenges such as monocular depth estimation, 3D point cloud processing, and radar Doppler interpretation. The evolution of perception and decision-making pipelines is reviewed, contrasting modular architectures with end-to-end learning paradigms that directly map raw sensor data to control commands, and discussing their trade-offs in interpretability, safety assurance, and robustness to rare edge cases. We further survey specialized hardware accelerators and heterogeneous automotive SoCs designed to meet stringent real-time and power constraints. Industrial strategies are compared, including multi-modal sensor fusion and vision-centric approaches based on large-scale imitation learning. Finally, we identify open challenges related to robustness under adverse conditions, domain shift, causal ambiguity, and the need for interpretable and certifiable AI in safety-critical autonomous driving systems. Full article
Show Figures

Figure 1

26 pages, 3759 KB  
Article
Prediction-Regularized Spatio-Temporal Transformer Framework for Offline Multi-Intersection Traffic Signal Control
by Yueting Deng, Huale Li, Tong Xia, Zhaobin Wang and Ruoming Lei
Appl. Sci. 2026, 16(10), 5156; https://doi.org/10.3390/app16105156 - 21 May 2026
Viewed by 483
Abstract
Multi-intersection traffic signal control must jointly address local coordination and delayed traffic propagation under strongly time-varying conditions. Existing offline sequence-imitation methods mainly recover actions from historical trajectories and make limited use of short-term future traffic evolution in shared-representation learning. To address this issue, [...] Read more.
Multi-intersection traffic signal control must jointly address local coordination and delayed traffic propagation under strongly time-varying conditions. Existing offline sequence-imitation methods mainly recover actions from historical trajectories and make limited use of short-term future traffic evolution in shared-representation learning. To address this issue, we propose PR-STLight, a prediction-regularized spatio-temporal extension of TransformerLight for offline multi-intersection traffic signal control. PR-STLight introduces short-term future inbound-queue evolution as structural supervision for shared representation learning. The model combines neighborhood-constrained spatial self-attention, causal temporal self-attention, and a Topology-Recurrent Queue Predictor (TRQP) to capture topology-aware spatio-temporal dependencies and near-future congestion dynamics. Training adopts a two-stage strategy, namely queue-prediction pretraining followed by joint control-prediction optimization, to improve optimization stability on a fixed offline replay buffer. In experiments on the adopted CityFlow benchmarks, PR-STLight obtains average travel times of 274.39 s on Jinan 3×4 and 288.09 s on Hangzhou 4×4, corresponding to 1.14% and 2.82% lower travel times than the strongest non-PR baseline, and 21.27% and 22.54% lower travel time than the TransformerLight backbone, respectively. It also achieves the lowest average inbound queue on Hangzhou and remains competitive on Jinan. These results show that PR-STLight provides an effective offline spatio-temporal sequence framework for coordinated multi-intersection signal control. Full article
(This article belongs to the Special Issue Advances in Intelligent Decision-Making Systems)
Show Figures

Figure 1

26 pages, 21250 KB  
Article
Social Modulation of Imitation in Children with Autism Spectrum Disorder: Evidence from EEG and Reciprocal Imitation Training
by Yonggu Wang, Zihan Wang, Guohao Li and Zhou Jin
Appl. Sci. 2026, 16(9), 4297; https://doi.org/10.3390/app16094297 - 28 Apr 2026
Viewed by 724
Abstract
Imitation is crucial for social learning, yet children with autism spectrum disorder (ASD) often show atypical imitation abilities. To probe the neural dynamics that precede overt imitation, electroencephalography (EEG)—with a focus on α (8–12 Hz) and β (13–30 Hz) activity commonly linked to [...] Read more.
Imitation is crucial for social learning, yet children with autism spectrum disorder (ASD) often show atypical imitation abilities. To probe the neural dynamics that precede overt imitation, electroencephalography (EEG)—with a focus on α (8–12 Hz) and β (13–30 Hz) activity commonly linked to action observation and sensorimotor processing—was used to index pre-imitation processing in preschool-aged children with ASD. Grounded in the social motivation framework, this study combined an EEG experiment and a naturalistic behavioral intervention. In Study 1, 11 preschool children with ASD completed an action-observation (pre-imitation) task under low- versus high-sociality video conditions. Time–frequency and spectral analyses were conducted to compare α- and β-band responses across conditions. In Study 2, four children received a six-week Reciprocal Imitation Training (RIT) program, and imitation and social-communication outcomes were assessed pre-, mid-, and post-intervention. The results showed that low-sociality stimuli elicited stronger frontal and prefrontal power increases in both α and β bands, whereas high-sociality stimuli elicited more temporally dynamic β-band responses but with lower overall power engagement. Although inferential support was limited by sample size, behavioral trends suggested improvements following RIT in imitation and related social functioning, with larger gains in children with mild-to-moderate ASD. Together, these findings suggest that social context modulates pre-imitation neural activity in ASD and that socially grounded imitation training may support broader social development. Full article
Show Figures

Figure 1

20 pages, 4199 KB  
Article
Parkour Learning for Quadrupeds via Terrain-Conditional Adversarial Motion Priors
by Shuomo Zhang, Wei Zou and Hu Su
Appl. Sci. 2026, 16(7), 3448; https://doi.org/10.3390/app16073448 - 2 Apr 2026
Viewed by 1192
Abstract
Agile parkour in unstructured environments poses significant challenges for quadruped robots, requiring both dynamic motion generation and terrain adaptability. Recent advances such as Adversarial Motion Priors (AMP) have shown promise in learning dynamic behaviors through motion imitation, but the resulting policies are typically [...] Read more.
Agile parkour in unstructured environments poses significant challenges for quadruped robots, requiring both dynamic motion generation and terrain adaptability. Recent advances such as Adversarial Motion Priors (AMP) have shown promise in learning dynamic behaviors through motion imitation, but the resulting policies are typically specialized and struggle to generalize across varying terrains. However, existing AMP-based approaches largely lack explicit environmental awareness, leading to limited adaptability and revealing a clear gap in achieving general agile locomotion. To address this limitation, we propose a novel terrain-conditional AMP framework that extends adversarial motion priors by conditioning the discriminator on explicit terrain features, enabling the learning of terrain-aware motion representations adaptable to diverse environments. To improve practical applicability, we further leverage a vision-based policy distillation scheme, where a teacher policy with privileged terrain height information supervises a student policy operating only on forward-looking depth images. This enables agile, perception-driven locomotion in real time. To the best of our knowledge, this is the first work to integrate environmental information into adversarial motion priors and jointly learn a vision-based policy through policy distillation for agile quadruped locomotion. Experiments on terrains such as platforms, gaps, stairs, slopes, and debris show that the proposed method achieves more efficient training convergence and higher success rates compared to pure AMP-based and RL-based methods. These results highlight the effectiveness of the proposed framework and represent a step toward perception-driven agile locomotion for quadruped robots in complex environments. Full article
(This article belongs to the Special Issue Intelligent Control of Robotic System)
Show Figures

Figure 1

32 pages, 10936 KB  
Article
PLM-Net: Perception Latency Mitigation Network for Vision-Based Lateral Control of Autonomous Vehicles
by Aws Khalil and Jaerock Kwon
Sensors 2026, 26(6), 1798; https://doi.org/10.3390/s26061798 - 12 Mar 2026
Cited by 1 | Viewed by 758
Abstract
This study introduces the Perception Latency Mitigation Network (PLM-Net), a modular deep learning framework designed to mitigate perception latency in vision-based imitation-learning lane-keeping systems. Perception latency, defined as the delay between visual sensing and steering actuation, can degrade lateral tracking performance and steering [...] Read more.
This study introduces the Perception Latency Mitigation Network (PLM-Net), a modular deep learning framework designed to mitigate perception latency in vision-based imitation-learning lane-keeping systems. Perception latency, defined as the delay between visual sensing and steering actuation, can degrade lateral tracking performance and steering stability. While delay compensation has been extensively studied in classical predictive control systems, its treatment within vision-based imitation-learning architectures under constant and time-varying perception latency remains limited. Rather than reducing latency itself, PLM-Net mitigates its effect on control performance through a plug-in architecture that preserves the original control pipeline. The framework consists of a frozen Base Model (BM), representing an existing lane-keeping controller, and a Timed Action Prediction Model (TAPM), which predicts future steering actions corresponding to discrete latency conditions. Real-time mitigation is achieved by interpolating between model outputs according to the measured latency value, enabling adaptation to both constant and time-varying latency. The framework is evaluated in a closed-loop deterministic simulation environment under fixed-speed conditions to isolate the impact of perception latency. Results demonstrate significant reductions in steering error under multiple latency settings, achieving up to 62% and 78% reductions in Mean Absolute Error (MAE) for constant and time-varying latency cases, respectively. These findings demonstrate the architectural feasibility of modular latency mitigation for vision-based lateral control under controlled simulation settings. The project page including video demonstrations, code, and dataset is publicly released. Full article
(This article belongs to the Special Issue Intelligent Control Systems for Autonomous Vehicles)
Show Figures

Figure 1

26 pages, 10726 KB  
Article
PI-VLA: Adaptive Symmetry-Aware Decision-Making for Long-Horizon Vision–Language–Action Manipulation
by Yina Jian, Di Tian, Xuan-Jing Chen, Zhen-Yuan Wei, Chen-Wei Liang and Mu-Jiang-Shan Wang
Symmetry 2026, 18(3), 394; https://doi.org/10.3390/sym18030394 - 24 Feb 2026
Cited by 2 | Viewed by 2719
Abstract
Vision–language–action (VLA) models often suffer from limited robustness in long-horizon manipulation tasks—where robots must execute extended sequences of actions over multiple time steps to achieve complex goals—due to their inability to explicitly exploit structural symmetries and to react adaptively when such symmetries are [...] Read more.
Vision–language–action (VLA) models often suffer from limited robustness in long-horizon manipulation tasks—where robots must execute extended sequences of actions over multiple time steps to achieve complex goals—due to their inability to explicitly exploit structural symmetries and to react adaptively when such symmetries are violated by environmental uncertainty. To address this limitation, this paper proposes PI-VLA, a symmetry-aware predictive and interactive VLA framework for robust robotic manipulation. PI-VLA is built upon three key symmetry-driven principles. First, a Cognitive–Motor Synergy (CMS) module jointly generates discrete and continuous action chunks together with predictive world-model features in a single forward pass, enforcing cross-modal action consistency as an implicit symmetry constraint across heterogeneous action representations. Second, a unified training objective integrates imitation learning, reinforcement learning, and state prediction, encouraging invariance to task-relevant transformations while enabling adaptive symmetry breaking when long-horizon deviations emerge. Third, an Active Uncertainty-Resolving Decider (AURD) explicitly monitors action consensus discrepancies and state prediction errors as symmetry-breaking signals, dynamically adjusting the execution horizon through closed-loop replanning. Extensive experiments on long-horizon benchmarks demonstrate that PI-VLA achieves state-of-the-art performance, attaining a 73.2% average success rate on the LIBERO benchmark (with particularly strong gains on the Long-Horizon suite) and an 88.3% success rate in real-world manipulation tasks under visual distractions and unseen conditions. Ablation studies confirm that symmetry-aware action consensus and uncertainty-triggered replanning are critical to robust execution. These results establish PI-VLA as a principled framework that leverages symmetry preservation and controlled symmetry breaking to enable reliable and interactive robotic manipulation. Full article
(This article belongs to the Section A: Computer Science)
Show Figures

Figure 1

31 pages, 461 KB  
Systematic Review
Techniques Applied to Autonomous Liquid Pouring: A Scoping Review
by Jeeangh Jennessi Reyes-Montiel, Ericka Janet Rechy-Ramirez and Antonio Marin-Hernandez
Math. Comput. Appl. 2026, 31(1), 30; https://doi.org/10.3390/mca31010030 - 14 Feb 2026
Viewed by 1281
Abstract
In recent years, autonomous liquid pouring systems have gained more relevance, with applications from daily service tasks to complex industrial operations. While seemingly simple for humans, this task poses major challenges for automated systems, as it requires precise control and adaptation to varying [...] Read more.
In recent years, autonomous liquid pouring systems have gained more relevance, with applications from daily service tasks to complex industrial operations. While seemingly simple for humans, this task poses major challenges for automated systems, as it requires precise control and adaptation to varying container geometries, liquid properties, and environmental conditions. This review examines the state-of-the-art on liquid pouring through five research questions: (1) What are the characteristics of the liquids used in the experiments? (2) What are the characteristics of the containers used in the experiments and how do they affect the performance of the pouring tasks? (3) What techniques are used to control liquid pouring (i.e., to control the robotic arm or device)? (4) What metrics are used to assess the methods for pouring liquid? (5) What devices are used to measure poured volume? This scoping review follows the Arksey and O’Malley framework, and uses the PRISMA-ScR protocol to filter the articles. A total of 285 studies published between 2018 and 2025 were screened from IEEE Xplore, SpringerLink, ScienceDirect, Web of Science, and EBSCOhost, of which 23 met the inclusion criteria. Results showed that the most widely used methods for autonomous liquid pouring were classical control methods—PID, PD (30.4% of the studies). Conversely, the least widely used methods for autonomous liquid pouring were learning, imitation learning, and probabilistic models (15% of the studies). Full article
(This article belongs to the Special Issue New Trends in Computational Intelligence and Applications 2025)
Show Figures

Figure 1

21 pages, 3516 KB  
Article
Diffusion-Guided Model Predictive Control for Signal Temporal Logic Specifications
by Jonghyuck Choi and Kyunghoon Cho
Electronics 2026, 15(3), 551; https://doi.org/10.3390/electronics15030551 - 27 Jan 2026
Viewed by 822
Abstract
We study control synthesis under Signal Temporal Logic (STL) specifications for driving scenarios where strict rule satisfaction is not always feasible and human experts exhibit context-dependent flexibility. We represent such behavior using robustness slackness—learned rule-wise lower bounds on STL robustness—and introduce sub-goals that [...] Read more.
We study control synthesis under Signal Temporal Logic (STL) specifications for driving scenarios where strict rule satisfaction is not always feasible and human experts exhibit context-dependent flexibility. We represent such behavior using robustness slackness—learned rule-wise lower bounds on STL robustness—and introduce sub-goals that encode intermediate intent in the state/output space (e.g., lane-level waypoints). Prior learning-based MPC–STL methods typically infer slackness with VAE priors and plug it into MPC, but these priors can underrepresent multimodal and rare yet valid expert behaviors and do not explicitly model intermediate intent. We propose a diffusion-guided MPC–STL framework that jointly learns slackness and sub-goals from demonstrations and integrates both into STL-constrained MPC. A conditional diffusion model generates pairs of (rule-wise slackness, sub-goal) conditioned on features from the ego vehicle, surrounding traffic, and road context. At run time, a few denoising steps produce samples for the current situation; slackness values define soft STL margins, while sub-goals shape the MPC objective via a terminal (optionally stage) cost, enabling context-dependent trade-offs between rule relaxation and task completion. In closed-loop simulations on held-out highD track-driving scenarios, our method improves task success and yields more realistic lane-changing behavior compared to imitation-learning baselines and MPC–STL variants using CVAE slackness or strict rule enforcement, while remaining computationally tractable for receding-horizon MPC in our experimental setting. Full article
(This article belongs to the Special Issue Real-Time Path Planning Design for Autonomous Driving Vehicles)
Show Figures

Figure 1

23 pages, 2058 KB  
Article
On the Evolutionary Dynamics and Optimal Control of a Tripartite Game in the Pharmaceutical Procurement Supply Chain with Regulatory Participation
by Zhao Li and Yumu Wang
Mathematics 2026, 14(1), 56; https://doi.org/10.3390/math14010056 - 24 Dec 2025
Viewed by 837
Abstract
This study involves the construction of a dynamic evolutionary game model involving three key participants, including the Group Purchasing Organization (GPO), medical institutions, and pharmaceutical suppliers, while comprehensively considering critical factors such as benefit compensation, bad debt risk, and fiscal costs. The model [...] Read more.
This study involves the construction of a dynamic evolutionary game model involving three key participants, including the Group Purchasing Organization (GPO), medical institutions, and pharmaceutical suppliers, while comprehensively considering critical factors such as benefit compensation, bad debt risk, and fiscal costs. The model characterizes the strategy evolution of each participant under bounded rationality and imitation learning mechanisms. Based on the replicator dynamics equations, the evolutionary trajectories and equilibrium conditions of the three parties’ strategies are systematically derived. The Jacobian matrix is then used to analyze the local stability of eight boundary equilibria and potential internal mixed equilibria. Furthermore, to capture the optimal adjustment process of the compensation mechanism, the GPO’s compensation level is introduced into an optimal control framework. A controlled evolutionary system is formulated, and the dynamic optimal relationship between compensation intensity and system state is described using the Hamilton–Jacobi–Bellman (HJB) equation. Through analytical linearization and numerical simulations, the optimal feedback compensation law and its closed-loop evolutionary trajectory are obtained, allowing for a comparative analysis between the “fixed compensation” and “optimal compensation” scenarios. The results reveal that an appropriately designed dynamic compensation mechanism can significantly enhance system cooperation stability and overall social welfare. This provides a quantitative theoretical foundation and methodological tool for the refined design and dynamic regulation of pharmaceutical group purchasing policies. Full article
(This article belongs to the Special Issue Dynamic Analysis and Decision-Making in Complex Networks)
Show Figures

Figure 1

18 pages, 1678 KB  
Article
Body Knowledge and Emotion Recognition in Preschool Children: A Comparative Study of Human Versus Robot Tutors
by Alice Araguas, Arnaud Blanchard, Sébastien Derégnaucourt, Adrien Chopin and Bahia Guellai
Behav. Sci. 2026, 16(1), 29; https://doi.org/10.3390/bs16010029 - 23 Dec 2025
Viewed by 1298
Abstract
Social robots are increasingly integrated into early childhood education, yet limited research exists examining preschoolers’ learning from robotic versus human demonstrators across embodied tasks. This study investigated whether children (aged between 3 and 6) demonstrate comparable performance when learning body-centered tasks from a [...] Read more.
Social robots are increasingly integrated into early childhood education, yet limited research exists examining preschoolers’ learning from robotic versus human demonstrators across embodied tasks. This study investigated whether children (aged between 3 and 6) demonstrate comparable performance when learning body-centered tasks from a humanoid robot compared to a human demonstrator. Sixty-two typically developing children were randomly assigned to a robot or a human condition. Participants completed three tasks: body part comprehension and production, body movement imitation, and emotion recognition from body postures. Performance was measured using standardized protocols. No significant main effects of demonstrator type emerged across most tasks. However, age significantly predicted performance across all measures, with systematic improvements between 3 and 6. A significant age × demonstrator interaction was observed for sequential motor imitation, with stronger age effects for the human demonstrator condition. Preschool children demonstrate comparable performance when interacting with a humanoid robot versus a human in body-centered tasks, though motor imitation shows differential developmental trajectories. These findings suggest appropriately designed social robots may serve as supplementary pedagogical tools for embodied learning in early childhood education under specific conditions. The primacy of developmental effects highlights the importance of age-appropriate design in both traditional and technology-enhanced educational contexts. Full article
Show Figures

Figure 1

19 pages, 11024 KB  
Article
Contact-Aware Diffusion Sampling for RRT-Based Manipulation
by Kyoungho Lee and Kyunghoon Cho
Electronics 2025, 14(24), 4837; https://doi.org/10.3390/electronics14244837 - 8 Dec 2025
Cited by 1 | Viewed by 814
Abstract
Rapidly exploring Random Trees (RRT) provide probabilistic completeness but often explore inefficiently in high-DOF manipulation tasks. We address this by proposing a contact-aware, two-level planner that couples a learned toggle–subgoal predictor with a conditional diffusion sampler in joint space under a completeness-preserving mixture [...] Read more.
Rapidly exploring Random Trees (RRT) provide probabilistic completeness but often explore inefficiently in high-DOF manipulation tasks. We address this by proposing a contact-aware, two-level planner that couples a learned toggle–subgoal predictor with a conditional diffusion sampler in joint space under a completeness-preserving mixture with uniform sampling. An upper ResNet-based network predicts task-relevant milestones from RGB images: grasp/release “toggle” configurations and intermediate joint-space subgoals that serve as phase-wise, receding-horizon targets between consecutive contact events. Conditioned on these predictions and the current state, a lower-level diffusion model samples tree-extension segments—joint-space directions and step lengths—instead of absolute configurations. These proposals act as a drop-in replacement for uniform sampling in standard RRT/RRT-Connect, while a nonzero fraction of uniform samples preserves probabilistic completeness. By biasing growth toward contact-relevant regions, the planner concentrates the search near feasible approach manifolds without altering nearest-neighbor, steering, or collision-checking primitives. In mug pick-and-place simulations, the proposed method achieves higher success rates than diffusion and other sequence-based policies trained by imitation learning, and requires fewer RRT expansions than uniform and goal-biased RRT as well as prior learning-guided samplers based on CVAE and conditional GAN, under identical collision checking and iteration limits. Full article
(This article belongs to the Special Issue Intelligent Perception and Control for Robotics)
Show Figures

Figure 1

24 pages, 1158 KB  
Article
Symbolic Imitation Learning: From Black-Box to Explainable Driving Policies
by Iman Sharifi, Mustafa Yildirim and Saber Fallah
Appl. Sci. 2025, 15(23), 12464; https://doi.org/10.3390/app152312464 - 24 Nov 2025
Cited by 2 | Viewed by 1394
Abstract
Current imitation learning approaches, predominantly based on deep neural networks (DNNs), offer efficient mechanisms for learning driving policies from real-world datasets. However, they suffer from inherent limitations in interpretability and generalizability—issues of critical importance in safety-critical domains such as autonomous driving. In this [...] Read more.
Current imitation learning approaches, predominantly based on deep neural networks (DNNs), offer efficient mechanisms for learning driving policies from real-world datasets. However, they suffer from inherent limitations in interpretability and generalizability—issues of critical importance in safety-critical domains such as autonomous driving. In this paper, we introduce Symbolic Imitation Learning (SIL), a novel framework that leverages Inductive Logic Programming (ILP) to derive explainable and generalizable driving policies from synthetic datasets. We evaluate SIL on real-world HighD and NGSim datasets, comparing its performance with state-of-the-art neural imitation learning methods using metrics such as collision rate, lane change efficiency, and average speed. The results indicate that SIL significantly enhances policy transparency while maintaining strong performance across varied driving conditions. These findings highlight the potential of integrating ILP into imitation learning to promote safer and more reliable autonomous systems. Full article
(This article belongs to the Special Issue Intelligent Vehicle Collaboration and Positioning)
Show Figures

Figure 1

Back to TopTop