Next Article in Journal
Dynamic Tribological Behavior of Surface-Textured Bushings in External Gear Pumps: A CFD Investigation
Previous Article in Journal
Learning Nonlinear Motor Control: How Integrating Machine Learning and Nonlinear Dynamics Reveals Structure, Adaptation, and Control in Human Movement
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators

1
Faculty of Electrical and Electronic Engineering, Hung Yen University of Technology and Education, Hung Yen 17000, Vietnam
2
Faculty of Information Technology, Ton Duc Thang University, Ho Chi Minh City 700000, Vietnam
*
Authors to whom correspondence should be addressed.
Actuators 2026, 15(3), 167; https://doi.org/10.3390/act15030167
Submission received: 12 January 2026 / Revised: 8 March 2026 / Accepted: 9 March 2026 / Published: 16 March 2026
(This article belongs to the Section Actuators for Robotics)

Abstract

Reinforcement learning controllers for robot manipulators depend strongly on reward tuning, and fixed weights may yield poor trade-offs under uncertainty and disturbances. This paper proposes a disturbance observer-based actor–critic RL (DOB–ACRL) with adaptive multi-objective reward shaping for a torque-saturated 2-DOF manipulator, where the reward weights are updated online using normalized indicators of tracking error, control energy, and effort. A Lyapunov analysis guarantees the uniform ultimate boundedness of closed-loop signals. The simulations show improved learning and performance over a static reward actor–critic baseline, reducing the RMS tracking error by up to 22.8%, the control energy by ~4.6%, the control effort by 1.9%, and the settling time by up to 29.2%.
Keywords: reward shaping; actor–critic reinforcement learning; disturbance observer; multi-objective control; Lyapunov stability; energy-efficient robotics reward shaping; actor–critic reinforcement learning; disturbance observer; multi-objective control; Lyapunov stability; energy-efficient robotics

Share and Cite

MDPI and ACS Style

Tam, L.T.M.; Ngu, N.V.; Pham, D.H.; Mai, V.T. Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators. Actuators 2026, 15, 167. https://doi.org/10.3390/act15030167

AMA Style

Tam LTM, Ngu NV, Pham DH, Mai VT. Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators. Actuators. 2026; 15(3):167. https://doi.org/10.3390/act15030167

Chicago/Turabian Style

Tam, Le Thi Minh, Nguyen Viet Ngu, Duc Hung Pham, and V. T. Mai. 2026. "Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators" Actuators 15, no. 3: 167. https://doi.org/10.3390/act15030167

APA Style

Tam, L. T. M., Ngu, N. V., Pham, D. H., & Mai, V. T. (2026). Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators. Actuators, 15(3), 167. https://doi.org/10.3390/act15030167

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop