Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (341)

Search Parameters:
Keywords = soft actor-critic

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
43 pages, 51585 KB  
Article
Adaptive Control of Lower-Limb Assistive Exoskeleton for Rehabilitation Using Deep Reinforcement Learning
by Ali Foroutannia, Masoud Mohammadian and Kumudu Munasinghe
Sensors 2026, 26(16), 5217; https://doi.org/10.3390/s26165217 - 17 Aug 2026
Viewed by 335
Abstract
Lower-limb rehabilitation exoskeletons have the potential to improve gait recovery after stroke by providing intensive and repetitive training. However, conventional control strategies often rely on fixed control parameters and exhibit limited adaptability to patient-specific characteristics, sensor noise, and dynamic uncertainties. This paper proposes [...] Read more.
Lower-limb rehabilitation exoskeletons have the potential to improve gait recovery after stroke by providing intensive and repetitive training. However, conventional control strategies often rely on fixed control parameters and exhibit limited adaptability to patient-specific characteristics, sensor noise, and dynamic uncertainties. This paper proposes an adaptive control framework that combines deep reinforcement learning (RL) with model-based impedance control for personalised lower-limb exoskeleton assistance. Patient-specific biological parameters are incorporated into the simulation environment and reward formulation to improve adaptability and robustness. Three state-of-the-art deep RL algorithms, Deep Deterministic Policy Gradient (DDPG), Twin Delayed Deep Deterministic Policy Gradient (TD3), and Soft Actor-Critic (SAC), are evaluated in a continuous control environment under varying signal-to-noise ratio (SNR) conditions ranging from 5 dB to noise-free conditions. Results demonstrate that TD3 achieves the most stable learning performance, obtaining a mean reward of −354.24 under noise-free conditions, while DDPG provides the highest joint-angle tracking accuracy with an RMSE of 0.0369 rad. SAC exhibits superior robustness in noisy environments, achieving the highest learning ratio of 0.51 at 5 dB SNR. Furthermore, the proposed personalised framework reduces tracking errors by up to 27% compared with non-personalised baseline approaches. The findings indicate that integrating patient-specific information with RL-based adaptive control can significantly enhance robustness, tracking performance, and personalisation in exoskeleton-assisted gait rehabilitation, providing a promising direction for future intelligent rehabilitation systems. Full article
(This article belongs to the Section Wearables)
Show Figures

Figure 1

37 pages, 8462 KB  
Article
A Nonlinear Model Predictive Controller for 4WID Electric Vehicles Incorporating a Hierarchical Architecture
by Minghui Ye, Meng Zhang, Bowen Li, Wen He and Mengna Li
Vehicles 2026, 8(8), 193; https://doi.org/10.3390/vehicles8080193 - 16 Aug 2026
Viewed by 137
Abstract
In light of the advancement of vehicle electrification and intelligence, four-wheel independent drive (4WID) electric vehicles (EVs) have garnered significant attention as a promising platform. Integrating advanced torque-vectoring (TV) strategies into 4WID EVs can effectively optimize the synergistic performance between handling stability and [...] Read more.
In light of the advancement of vehicle electrification and intelligence, four-wheel independent drive (4WID) electric vehicles (EVs) have garnered significant attention as a promising platform. Integrating advanced torque-vectoring (TV) strategies into 4WID EVs can effectively optimize the synergistic performance between handling stability and energy efficiency of the over-actuated system across various driving conditions. In this paper, a hierarchical Combined Sliding Mode Control–Adaptive Nonlinear Model Predictive Control (cSMC-ANMPC) TV strategy is proposed to enhance the comprehensive performance of 4WID EVs and ensure adaptive control across diverse driving conditions. Firstly, a hierarchical control architecture is developed to decouple the complex multi-objective problem. The upper layer performs robust stability decision-making by observing the vehicle’s state errors. The lower layer determines the optimal torque distribution throughout the powertrain. Secondly, a Combined Sliding Mode Controller (cSMC) is developed for the upper layer to promptly generate a robust stability command. By co-regulating both yaw rate and sideslip angle into a single command, it simplifies the lower layer’s task and enhances overall stability. Thirdly, a Soft Actor-Critic (SAC) intelligent tuner is integrated into the lower-layer NMPC to mitigate the effects of varying conditions on the stability–economy trade-off and strengthen the adaptability of the controller. Finally, co-simulation evaluations on the MATLAB R2023b/CarSim 2020.0platform demonstrate that the proposed cSMC-ANMPC strategy can improve comprehensive performance for the studied 4WID EV. Compared with other baselines, the stability enhancement in extreme maneuvers and the long-term energy-saving capability are remarkable, showcasing its promising performance. Full article
(This article belongs to the Special Issue Computer Vision Applications in Autonomous Vehicles)
Show Figures

Figure 1

14 pages, 2487 KB  
Article
CM-FuseNet: An Attention-Augmented Hybrid EEG–EMG Cognitive–Motor Fusion Network with Soft Actor-Critic Reinforcement Learning for Adaptive Lower-Limb Exoskeleton Control
by Yong-Deok Park, Dae-seob Shin and Hun-kee Kim
Appl. Sci. 2026, 16(16), 8042; https://doi.org/10.3390/app16168042 - 12 Aug 2026
Viewed by 163
Abstract
Population aging and the rising prevalence of motor disorders are driving demand for assistive lower-limb robotic systems capable of decoding user intention rather than merely providing mechanical support. We present CM-FuseNet, an attention-augmented hybrid Brain–Computer–Muscle Interface (BCMI) that simultaneously fuses cortical concentration indices [...] Read more.
Population aging and the rising prevalence of motor disorders are driving demand for assistive lower-limb robotic systems capable of decoding user intention rather than merely providing mechanical support. We present CM-FuseNet, an attention-augmented hybrid Brain–Computer–Muscle Interface (BCMI) that simultaneously fuses cortical concentration indices extracted from electroencephalography (EEG) and lower-limb intention patterns derived from electromyography (EMG) to adaptively control a 4-DOF assistive lower-limb exoskeleton. To eliminate the burden of human-subject ethics review and to ensure reproducibility of the proposed methodology, all validation is performed exclusively on (i) permissively licensed open-access biomedical datasets, (ii) high-fidelity OpenSim 4.5 and MuJoCo 3.1 musculoskeletal–exoskeleton co-simulation, and (iii) limited self-experimentation by the corresponding author with non-invasive consumer-grade devices. Three components are introduced: (i) a log-tanh normalized concentration index CI in (0, 1) derived from the (PSMR+PMidBeta)/PTheta ratio; (ii) a bidirectional Cross-Modal Transformer (CMT) with eight-head self- and cross-attention; and (iii) a Soft Actor-Critic (SAC) reinforcement-learning controller that adaptively tunes four servo PID gains using a concentration-weighted state. Experiments on the PhysioNet EEGMMIDB, Ninapro DB2/DB7, HuMoD and WAY-EEG-GAL datasets (combining N = 162 trial sessions, 47,520 windows, and five-fold cross-validation) yield a gait-phase classification accuracy of 96.84 ± 1.18%, torque-tracking RMSE of 0.072 ± 0.008 N·m, information transfer rate of 38.6 bits/min, end-to-end latency of 9.4 ms, and a 27.4% reduction in simulated metabolic cost over an EMG-only PID baseline (one-way ANOVA: F(4, 75) = 47.83, p < 0.001; Tukey HSD: p < 0.01 against all baselines). Under high cognitive load, CM-FuseNet preserves accuracy with only a 4.63 percentage-point degradation versus 13.22 percentage points for the EMG-only baseline. Full article
(This article belongs to the Section Robotics and Automation)
Show Figures

Figure 1

34 pages, 3143 KB  
Article
Multi-Objective Optimization for Data Center HVAC Systems Based on Edge–Cloud Collaborative Deep Reinforcement Learning
by Shichao Huang, Yibing Zhou and Yuan Liu
Sensors 2026, 26(16), 5031; https://doi.org/10.3390/s26165031 - 7 Aug 2026
Viewed by 432
Abstract
The sustained growth of cloud computing and AI training workloads drives data center expansion. Optimizing their control is therefore critical for reducing operational costs. Edge real-time control is indispensable for guaranteeing thermal safety, data sovereignty, and offline availability. Yet deploying Deep Reinforcement Learning [...] Read more.
The sustained growth of cloud computing and AI training workloads drives data center expansion. Optimizing their control is therefore critical for reducing operational costs. Edge real-time control is indispensable for guaranteeing thermal safety, data sovereignty, and offline availability. Yet deploying Deep Reinforcement Learning (DRL) in production Heating, Ventilation, and Air Conditioning (HVAC) environments confronts cold-start risks, edge–cloud computational asymmetry, and multi-objective conflicts spanning energy efficiency, electricity cost, and thermal safety. To address these challenges, this paper proposes an edge-cloud collaborative physics-informed reinforcement learning framework for production data center HVAC control. The framework integrates a physics-informed cold-start solution using Adaptive Particle Swarm Optimization (APSO) to generate physically constrained initial policies on a gray-box digital twin without expert demonstration data, a three-time-scale edge–cloud architecture coordinating minute-level edge Soft Actor-Critic (SAC) real-time inference, weekly edge APSO online model identification, daily cloud Non-dominated Sorting Genetic Algorithm III (NSGA-III) thermal storage scheduling, and a constraint-aware safe projection layer that embeds thermal safety hard constraints directly into the neural network policy. The framework is validated through a seven-month production deployment spanning the complete summer-to-winter transition, comprising approximately 3.2 million sensor records and evaluated with rigorous statistical methods. Full article
(This article belongs to the Special Issue Edge Computing for Beyond 5G and Wireless Sensor Networks)
Show Figures

Figure 1

30 pages, 1256 KB  
Article
Multimodal History-Window Gated-Attention Soft Actor-Critic for Urban Low-Altitude UAV Navigation
by Xi You and Wenjun Yi
Drones 2026, 10(8), 605; https://doi.org/10.3390/drones10080605 - 5 Aug 2026
Viewed by 540
Abstract
Urban low-altitude unmanned aerial vehicle (UAV) navigation combines partial observability, building occlusion, wind disturbance, and continuous control. This study develops and evaluates HW-GA-SAC, a multimodal history-window Soft Actor-Critic (SAC) policy for procedurally generated three-dimensional MuJoCo cities. A Gated Transformer-XL (GTrXL)-inspired gated-attention encoder processes [...] Read more.
Urban low-altitude unmanned aerial vehicle (UAV) navigation combines partial observability, building occlusion, wind disturbance, and continuous control. This study develops and evaluates HW-GA-SAC, a multimodal history-window Soft Actor-Critic (SAC) policy for procedurally generated three-dimensional MuJoCo cities. A Gated Transformer-XL (GTrXL)-inspired gated-attention encoder processes a fixed eight-step navigation history, while a current-frame safety branch supplies vertical clearance, sparse Light Detection and Ranging (LiDAR)-like range sectors, and handcrafted safety cues directly to the actor and critic. The policy uses obstacle-related observations and reward shaping to support collision avoidance; it does not include constrained policy optimization or a separate runtime safety filter. In a seven-method comparison using five training seeds and five evaluation layouts, HW-GA-SAC achieved a 96% ± 3% success rate, 207 ± 16 average return, and 3% ± 4% timeout rate. Feedforward SAC achieved 92% ± 11% success and a 7% ± 10% timeout rate, but its successful paths were more direct. Five-seed learning curves, city-split evaluation, wind sensitivity, sensing perturbations, inference profiling, and ablation studies further characterize the method. Within this simulation protocol, HW-GA-SAC provides the strongest completion-oriented performance, with a measurable trade-off between task completion and path directness. Full article
(This article belongs to the Section Innovative Urban Mobility)
Show Figures

Figure 1

38 pages, 5594 KB  
Article
A Cooperative Game-Based Low-Carbon Optimal Operation Strategy for Multi-Microgrids Based on Multi-Agent Deep Reinforcement Learning
by Pengfei Zhang, Pan Liu, Li Jiang and Dong Han
Energies 2026, 19(15), 3683; https://doi.org/10.3390/en19153683 - 5 Aug 2026
Viewed by 317
Abstract
Distributed integrated energy microgrids support the low-carbon transition of regional energy systems. However, multi-agent trading among microgrids still faces insufficient cross-market coordination, weak low-carbon incentives, and difficulties in fair benefit allocation. To address these issues, this paper proposes a cooperative-game-based low-carbon optimal operation [...] Read more.
Distributed integrated energy microgrids support the low-carbon transition of regional energy systems. However, multi-agent trading among microgrids still faces insufficient cross-market coordination, weak low-carbon incentives, and difficulties in fair benefit allocation. To address these issues, this paper proposes a cooperative-game-based low-carbon optimal operation strategy for multi-microgrids using multi-agent deep reinforcement learning. First, an energy-carbon-green certificate peer-to-peer coordinated trading mechanism and a green-carbon offsetting-based dual-incentive model are developed to link energy exchange, carbon quota adjustment, and green certificate circulation. Second, a Nash bargaining-based cooperative game model is formulated for multi-commodity P2P trading to maximize coalition benefits and ensure a fair allocation of surplus. Finally, the cooperative game is transformed into a Markov decision process, and a centralized training and decentralized execution framework with homogeneous agents is constructed based on the multi-agent soft actor-critic algorithm. Case studies using data from the Yangtze River Delta region of China show that the proposed method achieves a 1.25% optimality gap compared with the centralized MILP benchmark and reduces the coalition operating cost by 8.19% relative to independent operation. The carbon trading costs of the three microgrids are reduced by 37.27%, 40.13%, and 33.82%, respectively, verifying the economic applicability of the proposed method. Full article
(This article belongs to the Section B: Energy and Environment)
Show Figures

Figure 1

34 pages, 17014 KB  
Article
Hierarchical Model Selection and Control for Latency-Energy Optimization in MEC-Assisted Vehicular Networks
by Inseok Song, Seungwoo Kang, Seyha Ros and Seokhoon Kim
Sensors 2026, 26(15), 4969; https://doi.org/10.3390/s26154969 - 5 Aug 2026
Viewed by 233
Abstract
Multi-access edge computing (MEC) enables computation-intensive perception and decision-making tasks in vehicular networks to be offloaded to nearby edge servers. Existing approaches usually fix the artificial intelligence (AI) inference model, overlooking how model selection jointly affects latency, energy consumption, and service reliability. We [...] Read more.
Multi-access edge computing (MEC) enables computation-intensive perception and decision-making tasks in vehicular networks to be offloaded to nearby edge servers. Existing approaches usually fix the artificial intelligence (AI) inference model, overlooking how model selection jointly affects latency, energy consumption, and service reliability. We propose a hierarchical model selection and control (HMSC) framework based on deep reinforcement learning (DRL) for MEC-assisted vehicular networks. The framework couples a vehicle-layer MAPPO component that provides a communication interface representation for subchannel assignment and energy accounting with a centralized MEC-layer soft actor-critic (SAC) agent that, under SDN orchestration, adaptively selects lightweight or high-fidelity AI models and allocates computational resources. Accordingly, the core contribution of this paper lies in MEC-side model-aware computation control under an explicitly defined subchannel-contention abstraction, rather than in physical-layer transmit-power optimization. Both layers are guided by a composite objective that integrates normalized end-to-end (E2E) latency, normalized energy consumption, and a deadline-violation penalty. Using a discrete-time simulation framework, HMSC reduces E2E latency compared with static inference and non-hierarchical DRL baselines and sustains a higher deadline satisfaction ratio (DSR) under constrained uplink throughput and varying traffic loads. The learned policy is load-aware, favoring high-fidelity inference under light load and lightweight inference under congestion; a post hoc analysis using YOLOv5-family accuracy reference further quantifies the inference-quality implications of this adaptive selection behavior. These results show that coordinated MEC-side control of AI model selection and computation, under a shared deadline-aware objective, provides a robust latency–energy trade-off for MEC-assisted vehicular networks. Full article
(This article belongs to the Special Issue Edge Computing for Resource Sharing and Sensing in IoT Systems)
Show Figures

Figure 1

36 pages, 3449 KB  
Article
Joint Task Offloading and Resource Allocation with Data Caching in UAV-Aided Mobile Edge Computing Networks for Latency-Sensitive Applications
by Tanmay Baidya and Sangman Moh
Sensors 2026, 26(15), 4966; https://doi.org/10.3390/s26154966 - 5 Aug 2026
Viewed by 289
Abstract
The rapid growth of computing-intensive and latency-sensitive applications, including augmented reality, virtual reality, and self-driving systems, has increased the demand for low-latency and energy-efficient processing solutions. Mobile edge computing (MEC) has evolved as a transformative paradigm by relocating computation to the network edge, [...] Read more.
The rapid growth of computing-intensive and latency-sensitive applications, including augmented reality, virtual reality, and self-driving systems, has increased the demand for low-latency and energy-efficient processing solutions. Mobile edge computing (MEC) has evolved as a transformative paradigm by relocating computation to the network edge, closer to end users. Unmanned aerial vehicles (UAVs) further strengthen MEC by offering flexible deployment, mobility, and reliable line-of-sight communication, making them suitable for temporary high-demand scenarios. Moreover, such latency-sensitive applications often generate numerous repetitive tasks and, thus, storing the results of these tasks can reduce both communication overhead and computational workload. However, jointly addressing the caching of task-results alongside offloading and resource allocation decisions in UAV-aided MEC networks remains a non-trivial challenge. In this study, an integrated task offloading and resource allocation with data caching (JORC) framework is proposed to address these challenges. The offloading and resource allocation problems are formulated as a Markov decision process and solved using the soft actor–critic reinforcement learning algorithm. In addition, dynamic and adaptive caching manages limited storage and reduces redundant computations by using a hybrid strategy that integrates the least-frequently used and least-recently used policies to reduce computational redundancy. Simulation results confirm that the proposed JORC framework substantially reduces latency, energy consumption, and overall system cost, while increasing the successful task completion ratio compared to existing baseline approaches. Full article
(This article belongs to the Special Issue Feature Papers in the ‘Sensor Networks’ Section 2026)
Show Figures

Figure 1

20 pages, 3654 KB  
Article
Distribution Network Optimization with Aggregation and Reinforcement Learning Under Massive Distributed Resources Integration
by Peng Yu, Jiawei Xing, Xinbin Zuo, Yan Cheng, Yu Yi, Shunmin Sun, Xiao Wei, Zhigang Zhang, Jianxiu Li and Yunpeng Zhang
Energies 2026, 19(15), 3664; https://doi.org/10.3390/en19153664 - 4 Aug 2026
Viewed by 298
Abstract
The integration of large-scale distributed energy resources (DERs) into distribution networks (DNs) brings challenges to the effective control of DNs. In traditional approaches, mathematical or reinforcement learning (RL)-based solution algorithms are commonly used. However, the exponential increase in the number of DERs reduces [...] Read more.
The integration of large-scale distributed energy resources (DERs) into distribution networks (DNs) brings challenges to the effective control of DNs. In traditional approaches, mathematical or reinforcement learning (RL)-based solution algorithms are commonly used. However, the exponential increase in the number of DERs reduces the effectiveness of these strategies. Mathematical methods struggle to cope with the dynamic uncertainty caused by the high penetration of renewable energy, while RL algorithms relying on global data training may violate multi-agent privacy protocols. This paper proposes a DNs cooperative optimization method based on resource aggregation and RL. To reduce optimization dimensionality and ensure the privacy of resource data, a dynamic aggregation strategy is employed to aggregate a large number of distributed energy resources into aggregated entities, and the adjustable active–reactive power boundaries of each aggregated entity are derived. To fully exploit the regulation capability of DNs, data centers (DCs), as novel devices, are considered as flexible loads. To improve the convergence speed of model training and decision-making accuracy, evolution strategies (ES) and prioritized experience replay (PER) are integrated into the Soft Actor-Critic (SAC) algorithm, respectively. The proposed method is validated on the IEEE 33-bus and IEEE 123-bus systems. The results demonstrate the effectiveness and superiority of the proposed method in ensuring the secure operation of DNs. Full article
Show Figures

Figure 1

36 pages, 40887 KB  
Article
RL-Augmented Dual Robust Adaptive Propagated Interval Observer for Actuator and Residual-Framed Sensor Fault Detection and Isolation in Underactuated AUVs
by Ishaq Ahmed, Jun Lu, Talha Younas, Ghulam Farid, Muhammad Bilal and Sohaib Tahir Chauhdary
Drones 2026, 10(8), 598; https://doi.org/10.3390/drones10080598 - 3 Aug 2026
Viewed by 235
Abstract
Reliable fault detection and isolation (FDI) for actuators and sensors in underactuated autonomous underwater vehicles (AUVs) is challenging because nonlinear hydrodynamic coupling redistributes fault signatures, and persistent ocean currents can trigger false alarms. This paper presents a reinforcement learning (RL)-augmented dual-layer FDI framework [...] Read more.
Reliable fault detection and isolation (FDI) for actuators and sensors in underactuated autonomous underwater vehicles (AUVs) is challenging because nonlinear hydrodynamic coupling redistributes fault signatures, and persistent ocean currents can trigger false alarms. This paper presents a reinforcement learning (RL)-augmented dual-layer FDI framework for actuator and sensor faults in underactuated AUVs. The actuator layer uses a robust adaptive propagated interval observer (RAPIO) that evaluates thruster and control-surface residuals against a calibrated dynamics-consistency tube. The sensor layer forms estimator-consistency residuals for Doppler velocity log (DVL), depth, and inertial measurement unit (IMU) measurements against a reference-separated finite-time extended state observer (FTESO). An offline-trained soft actor–critic (SAC) policy schedules bounded actuator uncertainty margins and sensor alarm thresholds according to operating confidence. The scheduled actuator error dynamics remain Metzler and Hurwitz, preserving positive interval propagation and center-error input-to-state stability (ISS) independent of policy convergence. A Schmitt-trigger alarm and signal-space disambiguation rule classify healthy, actuator-only, sensor-only, and simultaneous-fault conditions under explicit residual-separation and persistence conditions. Across 72 simultaneous-fault episodes over a 4×6 uncertainty–current grid, the proposed method achieved 100% detection coverage for both actuator and sensor faults with only five false-alarm events, retaining full coverage in the severe-current, high-uncertainty subset where the selected actuator and sensor baselines achieved only 88.9% and 70.4% detection, respectively, with more false alarms. These results indicate that the proposed bounded RL scheduler can deliver reliable, certifiable actuator and sensor fault diagnosis under significant operational uncertainty. Full article
(This article belongs to the Section Unmanned Surface and Underwater Drones)
Show Figures

Figure 1

27 pages, 12252 KB  
Article
Dynamic Energy-Efficient Path Planning for Unmanned Surface Vehicles Based on SAC-BSTFN
by Zhaohui Liu and Qing Li
Eng 2026, 7(8), 376; https://doi.org/10.3390/eng7080376 - 1 Aug 2026
Viewed by 251
Abstract
To address the limited endurance of Unmanned Surface Vehicles (USVs) in time-varying sea conditions, this study investigates global energy-efficient path planning that balances obstacle avoidance and energy efficiency. First, based on ship seakeeping theory, an energy consumption model incorporating wave height, wave period, [...] Read more.
To address the limited endurance of Unmanned Surface Vehicles (USVs) in time-varying sea conditions, this study investigates global energy-efficient path planning that balances obstacle avoidance and energy efficiency. First, based on ship seakeeping theory, an energy consumption model incorporating wave height, wave period, speed, and wave-encounter angle is constructed to represent the impact of dynamic sea states on resistance. The model is assessed through formula–program consistency verification, a multi-factor input ablation on a physics-constrained benchmark, and external trend validation against published towing-tank added-resistance data. Second, the planning problem is modeled as a Partially Observable Markov Decision Process (POMDP). Built upon the Soft Actor-Critic (SAC) algorithm, a Bimodal Spatio-Temporal Feature Fusion Network (BSTFN) is proposed to achieve deep fusion of spatial perception information and historical temporal sea state sequences for decision-making. Furthermore, a composite reward function is designed, integrating energy consumption penalties, heading guidance, and smoothness constraints. Simulation results demonstrate that the proposed method effectively utilizes favorable encounter angles to avoid high sea state regions. While maintaining high task success rates, it significantly reduces average energy consumption, effectively enhancing the endurance and robustness of USVs in complex dynamic environments. Full article
Show Figures

Figure 1

32 pages, 8955 KB  
Article
Sensor-Informed Motion-Continuity Control of Shared-Return Electro-Hydraulic Actuator Networks Under Neighboring-Branch Disturbances
by Tiangu Wu, Lijuan Zhao, Guocong Lin and Shutian Gong
Sensors 2026, 26(15), 4739; https://doi.org/10.3390/s26154739 - 26 Jul 2026
Viewed by 200
Abstract
This study focuses on the development of a sensor-informed motion-continuity control method for shared-return electro-hydraulic actuator networks subject to neighboring-branch disturbances. The objective is to reduce the local velocity fluctuations induced by return-line pressure transients while retaining explicit hydraulic and valve constraints. A [...] Read more.
This study focuses on the development of a sensor-informed motion-continuity control method for shared-return electro-hydraulic actuator networks subject to neighboring-branch disturbances. The objective is to reduce the local velocity fluctuations induced by return-line pressure transients while retaining explicit hydraulic and valve constraints. A control-oriented shared-return disturbance model is established to map neighboring-valve action, T-port replenishment, accumulator buffering, common return-line pressure, net driving pressure difference, and local actuator motion. On this basis, sensor-derived motion and pressure states together with neighboring-action prior information are used to reconstruct the objective of a constrained predictive controller according to the disturbance stage. Soft Actor–Critic is restricted to bounded objective-weight inference, whereas the valve command remains generated by locally linearized receding-horizon optimization; bounded mapping, smoothing update, and soft pressure constraints preserve positive weighting matrices and online quadratic programming solvability. Co-simulation, simulation-based ablation and baseline comparisons, timing evaluation, and scaled dual-branch experiments show that the proposed framework improves motion continuity, reduces disturbance-induced pressure-difference excursions, maintains smoother valve execution, and completes each tested online update within the sampling period. These findings support feasibility-preserving sensor-driven objective reconstruction under the investigated shared-return disturbance scenarios. Full article
(This article belongs to the Section Industrial Sensors)
Show Figures

Figure 1

30 pages, 10197 KB  
Article
Autonomous Approach and Stable Tracking of Dynamic Target for Intelligent Ship Using Recurrent Soft Actor-Critic
by Zixuan Qiu, Shaosong Min and Cong Liu
J. Mar. Sci. Eng. 2026, 14(14), 1345; https://doi.org/10.3390/jmse14141345 - 22 Jul 2026
Viewed by 235
Abstract
Dynamic target tracking is a challenging task for autonomous ships due to the continuous variation of relative motion states and the requirement for coordinated control of position, heading, and speed. This paper proposes a task-oriented continuous decision-making framework based on Soft Actor-Critic (SAC) [...] Read more.
Dynamic target tracking is a challenging task for autonomous ships due to the continuous variation of relative motion states and the requirement for coordinated control of position, heading, and speed. This paper proposes a task-oriented continuous decision-making framework based on Soft Actor-Critic (SAC) reinforcement learning for the autonomous approach and stable following of dynamic target vessels. A finite-history-enhanced SAC framework is developed by incorporating LSTM-based sequence encoding into the policy and value networks to capture recent evolution patterns of target motion and own-ship maneuvering responses. Furthermore, a sector-annular tracking region defined by distance and relative bearing constraints is constructed and a multi-component reward function is designed to integrate distance convergence, heading adjustment, region maintenance, speed matching, and control smoothness into policy learning. Simulation experiments under straight-line motion, curved motion, and randomized initial conditions demonstrate that the proposed SAC-LSTM method achieves improved task completion capability and control quality compared with SAC, PPO, and DDPG under the same task settings. Compared with standard SAC, SAC-LSTM improves the average success rate by 7.5 percentage points, reduces the average episode length by approximately 22.0%, and decreases the average heading error by approximately 35.5%. Additional sequence-length analysis, reward-component ablation, and multi-level disturbance tests further validate the effectiveness of the proposed design. The results indicate that the proposed method provides an effective solution for continuous decision-making in dynamic target tracking tasks. Full article
(This article belongs to the Section Ocean Engineering)
Show Figures

Figure 1

28 pages, 14867 KB  
Article
Dynamic Uplink Power Control for Cell-Free Massive MIMO
by Hussein A. Jasim, Mohd Fadlee A. Rasid, Fazirulhisyam Hashim and Syamsiah Mashohor
Eng 2026, 7(7), 357; https://doi.org/10.3390/eng7070357 - 22 Jul 2026
Viewed by 324
Abstract
Dynamic uplink power allocation is a critical challenge in cell-free massive MIMO (CF-mMIMO) networks, where distributed access points (APs) jointly serve multiple user equipment (UEs) under mobility, time-varying propagation conditions, and strong inter-user interference. Conventional optimization-based methods can improve fairness or spectral efficiency, [...] Read more.
Dynamic uplink power allocation is a critical challenge in cell-free massive MIMO (CF-mMIMO) networks, where distributed access points (APs) jointly serve multiple user equipment (UEs) under mobility, time-varying propagation conditions, and strong inter-user interference. Conventional optimization-based methods can improve fairness or spectral efficiency, but they often require repeated numerical solving and are usually designed for a specific objective. Learning-based approaches can reduce online decision time after training; however, their effectiveness depends strongly on the reward design and the selected operating objective. In response to these challenges, we propose a Deep Hybrid Intelligent (DHI) architecture designed to evaluate dynamic uplink power management within cell-free massive MIMO environments. The framework uses Soft Actor-Critic (SAC) learning to generate continuous uplink transmit-power decisions and evaluates objective-specific configurations for fairness, signal-to-interference-plus-noise ratio (SINR) improvement, and spectral-efficiency enhancement. In addition, three optimization-based strategies, namely max-min fairness, max-product SINR optimization, and max-sum-rate maximization, are incorporated to analyze the trade-off among fairness, signal quality, throughput, and computational cost. Limited-memory Broyden-Fletcher-Goldfarb-Shanno with bound constraints (L-BFGS-B) optimization is employed for the max-product and max-sum-rate objectives, while the max-min strategy is evaluated through a fairness-oriented feasibility procedure. Simulation results show that the fairness-oriented configuration achieves the highest Jain’s fairness index, reaching 0.989 at 120 access points, whereas the sum-rate-oriented configuration provides stronger SINR and user-rate performance. The results also indicate execution-time reductions of 51.6%, 83.7%, and 85.0% for the evaluated max-min, max-product, and max-sum-rate strategies, respectively, compared with conventional optimization-based implementations. These execution-time gains are accompanied by a clear performance trade-off: the max-min strategy provides the strongest fairness behavior, the max-sum-rate strategy improves total spectral efficiency and user-rate performance, and the max-product strategy offers a balanced operating point between collective SINR improvement and user-service balance. Therefore, the proposed framework does not optimize only computational speed, but also clarifies the trade-off among execution time, SINR, spectral efficiency, and fairness under dynamic uplink CF-mMIMO conditions. These results indicate that this architecture serves as an adaptable platform to evaluate dynamic uplink power distribution across CF-mMIMO networks. Full article
(This article belongs to the Special Issue Signal Processing Challenges and Solutions in Mobile Communications)
Show Figures

Figure 1

32 pages, 6063 KB  
Article
Reinforcement Learning-Based Adaptive Control for a Permanent Magnet Synchronous Generator Connected to a Hybrid AC/DC Grid with Virtual Inertia Support
by Islam A. Zenhom, Mostafa I. Marei and Ahmed M. I. Mohamad
Sustainability 2026, 18(14), 7404; https://doi.org/10.3390/su18147404 - 20 Jul 2026
Viewed by 484
Abstract
The increasing penetration of renewable energy sources has increased the need for advanced control strategies capable of maintaining stability under low-inertia, converter-dominated operating conditions. In grid-connected wind energy conversion systems (WECSs), constant power loads (CPLs) exhibit negative incremental impedance characteristics that can amplify [...] Read more.
The increasing penetration of renewable energy sources has increased the need for advanced control strategies capable of maintaining stability under low-inertia, converter-dominated operating conditions. In grid-connected wind energy conversion systems (WECSs), constant power loads (CPLs) exhibit negative incremental impedance characteristics that can amplify DC-link oscillations and complicate the coordination between the electrical and mechanical subsystems. The main contribution of this work is a Soft Actor–Critic (SAC) reinforcement learning algorithm that tunes the outer proportional-integral gains of the machine-side DC-voltage-squared control loop together with the active damping gain, allowing online adaptation of the controller according to the operating condition and disturbance level, thereby improving energy system sustainability. The proposed control framework includes a two-mass shaft model, virtual inertia control, and DC-link load uncertainty in the form of both resistive loads and CPLs. The system is modeled and evaluated using MATLAB/Simulink, and its performance is compared with that of a conventional fixed-gain controller under AC load disturbances and wind speed variations. It has been found that for a 25% load disturbance, the maximum DC-link voltage deviation is reduced by 1.2% under resistive loading and 6.5% under CPL operation. For a 1 m/s reduction in wind speed, the corresponding reductions are 0.8% and 0.9%, respectively. The proposed controller also provides smoother output power and improved damping of the rotor speed and system frequency responses. Full article
(This article belongs to the Special Issue Driving Electric Power Solutions for a Sustainable Energy Transition)
Show Figures

Figure 1

Back to TopTop