Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators
Abstract
1. Introduction
- Adaptive Reward Shaping: Unlike existing methods that rely on fixed gain adjustments [44], the reward function in this study is decomposed into tracking error, energy, and control effort components, with their weights updated online via normalized performance indicators.
- Lyapunov-based Stability Guarantee: Sufficient conditions are established to ensure the uniform ultimate boundedness (UUB) of all closed-loop signals. This analysis explicitly accounts for bounded disturbances, DOB estimation errors, and actor–critic approximation errors, extending the robustness frameworks discussed in the recent literature [45].
- Comparative Evaluation: The simulation results on a 2-DOF manipulator demonstrate superior performance in both the accuracy and efficiency compared to two baselines: a static reward actor–critic and an adaptive reward actor–critic without a DOB.
2. Related Work
2.1. Actor–Critic and ADP with Stability Guarantees
2.2. Reward Shaping and Adaptive and Evolving Rewards
2.3. Disturbance Observers, RL, and Energy-Efficient Manipulator Control
3. System Model and Control Architecture
Disturbance Observer and Estimation Error
4. Adaptive Reward Design and Lyapunov-Based Analysis
4.1. Adaptive Multi-Objective Reward Shaping
4.2. Assumptions and Lyapunov Function
4.3. Stability Results and Finite-Time Remark
5. Simulation Evaluation
5.1. Simulation Setup and Implementation Details
- (a)
- Actor network: The actor is implemented as a feedforward neural network with two hidden layers. The input vector consists of the tracking error and its derivatives and . The proposed neural network architecture consists of four distinct layers. The network begins with an input layer comprising four neurons, which receive the normalized error states of the system. Following the input, there are two consecutive hidden layers. Hidden layer 1 contains 64 neurons and utilizes the ReLU activation function and, similarly, hidden layer 2 also has 64 neurons and employs the ReLU activation function. Finally, the network concludes with the output layer, which consists of two neurons using a linear activation function corresponding to the adaptive torque components . The final torque command is obtained by combining the nominal control, the DOB compensation and the actor output, and then applying saturation to respect the actuator limits.
- (b)
- Critic network: The critic network serves to approximate the state-value function . It utilizes the same input vector as the actor network, which comprises the system’s state information. The critic’s structure is defined by its layers: it features an input layer of four neurons, followed by two consecutive hidden layers, each containing 64 neurons and employing the ReLU activation function. The network concludes with an output layer consisting of a single neuron with a linear activation, which provides the scalar value estimate . The crucial TD error is then computed by comparing this critic output with the reward provided by the adaptive reward module. This TD error is subsequently used as the training signal to update the parameters of both the critic and the actor networks.
- (c)
- Learning rates and discount factor: The critic and actor parameters are updated online using gradient-based rules driven by the TD error. The learning rates are chosen as the critic learning rate and the actor learning rate , and the discount factor is set to . These values are selected to ensure a compromise between the convergence speed and numerical stability of the learning process.
- (d)
- Simulation horizon and episodes: All simulations are carried out in discrete time with a fixed sampling period of . Each episode has a duration of corresponding to 10,000 time steps per episode. Unless otherwise stated, the results reported in Section 5 were obtained after 200 training episodes, which were sufficient for the TD error and performance indices to reach a steady regime in all tested configurations.
- (e)
- Numerical integration: The continuous-time manipulator dynamics are integrated using a fourth-order Runge–Kutta (RK4) scheme with the step size matching the control update period. At each integration step, the torque command is held constant, and the DOB, actor and critic are updated based on the most recent state and reward.
- (f)
- Network initialization: All network weights and biases for the actor and critic are initialized using a zero-mean Gaussian distribution with a standard deviation of 0.01, and the biases are initialized to zero. Before training, the input states are normalized by fixed scaling factors corresponding to the maximum expected ranges of , and the tracking errors. No pre-training or offline learning is used; all the parameter updates are performed online during interaction with the simulated manipulator.
5.2. Simulation Process
5.2.1. Numerical Simulation Case 1: Mode A—AC (Static Reward)
5.2.2. Numerical Simulation Case 2: Mode B—AC (Actor–Critic with the Adaptive Composite Reward Without Disturbance Observer)
5.2.3. Numerical Simulation Case 3: Mode C-DOB–ACRL (Adaptive Reward + DOB)
5.3. Results
6. Discussion and Implications
6.1. Tracking Optimality and Adaptive Weight Scheduling
6.2. Energy Efficiency, Sustainability and Trade-Offs
6.3. Relationship to Existing Adaptive Reward and DOB–RL Methods
6.4. Limitations and Choice of Actor–Critic Structure
7. Conclusions and Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Nomenclature
| Symbol | Definition | Units |
| Joint angle vector of the 2-DOF manipulator | rad | |
| Joint velocity vector | rad/s | |
| Inertia matrix of the manipulator | kg·m2 | |
| Coriolis/centrifugal matrix | kg·m2/s | |
| Friction matrix | N·m·s | |
| Joint disturbance over time | N·m | |
| Control torque applied to joints | N·m | |
| Reference trajectory for joint i | rad | |
| Frequency of reference trajectory | rad/s | |
| Tracking and filtered errors | — | |
| Composite reward at time t | — | |
| Tracking accuracy weight | — | |
| Energy efficiency weight | — | |
| Control effort smoothness weight | — | |
| Normalized performance index | — | |
| Adaptation gain | — | |
| Lyapunov candidate function | — | |
| Design constant in | — | |
| Critic learning rate | — | |
| TD loss | Temporal difference error of critic | — |
References
- Zhang, Z.; Chen, G.; Chen, W.; Jia, R.; Chen, G.; Zhang, L.; Zhou, P. A Joint Learning of Force Feedback of Robotic Manipulation and Textual Cues for Granular Materials Classification. IEEE Robot. Autom. Lett. 2025, 10, 7166–7173. [Google Scholar] [CrossRef] [Scilit]
- Jin, J.; Zhao, L.; Chen, L.; Chen, W. A robust zeroing neural network and its applications to dynamic complex matrix equation solving and robotic manipulator trajectory tracking. Front. Neurorobotics 2022, 16, 1065256. [Google Scholar] [CrossRef] [Scilit]
- Li, R.; Jin, J.; Zhang, D.; Chen, C. A Segmented Activation Function-Based Zeroing Neural Network Model for Dynamic Sylvester Equation Solving and Robotic Manipulator Control. Concurr. Comput.-Pract. Exp. 2025, 37, e70243. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Yang, C.; He, W. Integrated disturbance observer and reinforcement learning control for uncertain nonlinear systems with application to robot manipulators. IEEE/ASME Trans. Mechatron. 2023, 28, 558–569. [Google Scholar]
- Wang, Y.; Zhang, H.; Wang, J.; Chen, Z. Adaptive dynamic programming for robotic ma-nipulators with disturbance observer: A robust and optimal approach. IEEE Trans. Cybern. 2023, 53, 3214–3225. [Google Scholar]
- Xiong, J.; Chen, Y. RBFNN-Based Parameter Adaptive Sliding Mode Control for an Uncertain TQUAV with Time-Varying Mass. Int. J. Robust Nonlinear Control 2025, 35, 4658–4668. [Google Scholar] [CrossRef] [Scilit]
- Liu, X.; Wu, C.; Zhen, S.; Sun, H.; Sun, C.; Chen, Y. Robust Control under Servo Constraint Following via Nash Equilibrium Theory for Bimanual Humanoid Manipulation. IEEE Trans. Fuzzy Syst. 2025, 33, 4069–4082. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Wang, S.; Zhang, Y.; Wang, X.; Zhang, Y.; Wu, W.; Mu, R. Distributed Bearing-based Fault-tolerant Formation Control of Fixed-wing UAV Swarm with Prescribed Performance. Aerosp. Sci. Technol. 2025, 168, 110897. [Google Scholar] [CrossRef] [Scilit]
- Lu, Q.; Wu, X.; She, J.; Guo, F.; Yu, L. Disturbance Rejection for Systems with Uncertainties Based on Fixed-Time Equivalent-Input-Disturbance Approach. IEEE-CAA J. Automatica Sin. 2024, 11, 2384–2395. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Zhao, X.; Zheng, G.; Zhu, F.; Dinh, T.N. On Distributed Prescribed-Time Unknown Input Observers. IEEE Trans. Autom. Control 2025, 70, 4743–4750. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Song, Y.; Zheng, G. Prescribed-time observer for descriptor systems with unknown input. Automatica 2025, 172, 111999. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Tan, C.P.; Zheng, G.; Wang, Y. On sliding mode observers for non-infinitely observable descriptor systems. Automatica 2023, 147, 110676. [Google Scholar] [CrossRef] [Scilit]
- Xiong, J.; Wang, X.; Li, C. Recurrent Neural Network Based Sliding Mode Control for an Uncertain Tilting Quadrotor UAV. Int. J. Robust Nonlinear Control 2025, 35, 8030–8046. [Google Scholar] [CrossRef] [Scilit]
- Li, G.; Liang, X.; Zhang, J.; Su, T.; Hou, Z. A Stiffness-Enhanced Extensible Continuum Surgical Robot: Design, Modeling, and Evaluation. IEEE-ASME Trans. Mechatron. 2025, 31, 413–424. [Google Scholar] [CrossRef] [Scilit]
- Liu, Q.; Chen, P.; Lin, K.; Zhao, K.; Ding, J.; Li, Y. Sample-efficient backtrack temporal difference deep reinforcement learning. Knowl.-Based Syst. 2025, 330, 114613. [Google Scholar] [CrossRef] [Scilit]
- Jian, G.; Yin, W. A Constrained Reinforcement Learning Based Approach for Cooperative Control of Multi-UAV in Dense Obstacle Environments. Sci. China-Technol. Sci. 2025, 69, 1120601. [Google Scholar] [CrossRef] [Scilit]
- Fan, Q.Y.; Yang, G.H. Adaptive actor–critic design-based integral sliding-mode control for partially unknown nonlinear systems with input disturbances. IEEE Trans. Neural Netw. Learn. Syst. 2016, 27, 165–177. [Google Scholar] [CrossRef] [Scilit]
- Jiang, Y.; Jiang, Z.-P. Approximate dynamic programming for stochastic nonlinear systems with continuous state and action spaces. IEEE Trans. Neural Netw. Learn. Syst. 2018, 29, 3816–3828. [Google Scholar]
- Kamalapurkar, R.; Rosenfeld, J.A.; Dixon, W.E. Efficient model-based reinforcement learning for approximate online optimal control. Automatica 2016, 74, 247–258. [Google Scholar] [CrossRef] [Scilit]
- Kiumarsi, B.; Lewis, F.L.; Modares, H. Actor–critic-based optimal tracking for partially unknown nonlinear discrete-time systems. IEEE Trans. Neural Netw. Learn. Syst. 2017, 28, 69–82. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, S.; Liu, Z.; Li, Z. Observer-based adaptive optimal control for nonlinear systems via reinforcement learning: A dual-heuristic programming approach. Automatica 2023, 156, 111198. [Google Scholar]
- Li, X.; Zhang, Q.; Chai, T. Robust actor–critic learning for continuous-time nonlinear systems with matched uncertainties and disturbances. IEEE Trans. Neural Netw. Learn. Syst. 2023, 34, 7821–7832. [Google Scholar]
- Gu, S.; Lilleicrap, T.; Sutskever, I.; Levine, S. Deep Reinforcement Learning for Robotic Manipulation with Asynchronous off-Policy Updates. In Proceedings of the 2017 IEEE International Conference on Robotics and Automation (ICRA), Singapore, 29 May–3 June 2017; pp. 3389–3396. [Google Scholar]
- Theodorou, E.; Buchli, J.; Schaal, S. A generalized path integral control approach to re-enforcements learning. J. Mach. Learn. Res. 2010, 11, 3137–3181. [Google Scholar]
- Vu, M.T.; Pham, D.H.; Nguyen, V.T.; Do, Q.T.; Alanazi, A.K.; Nguyen, T.H. Adaptive nonlinear integral-backstepping control for frequency stabilization in cyber-physical shipboard microgrids using double deep Q-learning. Eng. Appl. Artif. Intell. 2025, 160, 111943. [Google Scholar] [CrossRef] [Scilit]
- Vamvoudakis, K.G.; Lewis, F.L. Online actor–critic algorithm to solve the continuous-time infinite horizon optimal control problem. Automatica 2010, 46, 878–888. [Google Scholar] [CrossRef] [Scilit]
- Lewis, F.L.; Vrabie, D.; Vamvoudakis, K.G. Reinforcement learning and adaptive dynamic programming for feedback control. IEEE Circuits Syst. Mag. 2011, 11, 76–105. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Luo, Y.; Liu, D. Adaptive Dynamic Programming for Control: Algorithms and Stability; Springer: Singapore, 2017. [Google Scholar]
- Chen, D.; Wang, H.; Cheng, L.; Gong, S. Stability enhancement in reinforcement learning via adaptive control Lyapunov function. arXiv 2025, arXiv:2504.19473. [Google Scholar] [CrossRef] [Scilit]
- Han, H.; Zhang, L.; Wang, J.; Pan, W. Actor–critic reinforcement learning for control with stability guarantee. IEEE Robot. Autom. Lett. 2020, 5, 6217–6224. [Google Scholar] [CrossRef] [Scilit]
- Yang, P.; Zhang, S.; Yu, X.; He, W. Reinforcement-learning-based finite-time fault-tolerant control for a manipulator with actuator faults. IEEE Trans. Cybern. 2025, 55, 2621–2632. [Google Scholar] [CrossRef] [Scilit]
- Bertsekas, D.P. Dynamic Programming and Optimal Control; Athena Scientific: Belmont, MA, USA, 2017; Volumes 1–2. [Google Scholar]
- Wang, B.; Sun, J.; Peng, B.; Cui, X.; Cheng, L.; Zheng, X. Optimal Event-Triggered Neural Learning Tracking Control for Pneumatic Muscle Antagonistic Joint with Asymmetric Constraints. IEEE Trans. Ind. Electron. 2025, 72, 14677–14687. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Z.; Wang, Y.; Liu, X.; Li, Z.; Wu, M.; Zhou, G. Hybrid of Neural Network and Physics-Based Estimator for Vehicle Longitudinal Dynamics Modeling Using Limited Driving Data. IEEE Trans. Intell. Transp. Syst. 2025, 26, 16735–16746. [Google Scholar] [CrossRef] [Scilit]
- Todorov, E. Linearly-solvable Markov decision problems. In Advances in Neural Information Processing Systems (NeurIPS); NIPS: Grenada, Spain, 2006; pp. 1369–1376. [Google Scholar]
- Guo, W.; Liu, J.; Qin, W.; Lan, X.; Bai, H.; Li, X. Robust adaptive dynamic programming for morphing air-breathing hypersonic vehicles under unmatched uncertainty. Sci. China-Inf. Sci. 2026, 69, 122205. [Google Scholar] [CrossRef] [Scilit]
- Hu, R.; Chen, Y.; Huang, L. Finite-time convergence analysis of actor–critic with evolving reward. arXiv 2025, arXiv:2510.12334. [Google Scholar] [CrossRef] [Scilit]
- Levine, S.; Popovic, Z.; Koltun, V. Feature construction for inverse reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS); NIPS: Grenada, Spain, 2011. [Google Scholar]
- Wang, G.; Feng, Z.; Qu, Y.; Sun, H. Event-triggered adaptive predefined-time anti-unwinding attitude tracking control for spacecraft. PLoS ONE 2025, 20, e0333700. [Google Scholar] [CrossRef] [Scilit]
- Ma, H.; Luo, Z.; Vo, T.V.; Sima, K.; Leong, T.-Y. Highly efficient self-adaptive reward shaping for reinforcement learning. arXiv 2024, arXiv:2408.03029. [Google Scholar] [CrossRef] [Scilit]
- Nohooji, H.R.; Zaraki, A.; Voos, H. Actor–Critic Learning-Based PID Control for Robotic Manipulators. SSRN 2024. Available online: https://ssrn.com/abstract=4409551 (accessed on 13 January 2026).
- Vu, M.T.; Nguyen, V.T.; Do, Q.T.; Youn, W.; Nguyen, T.H. Robust non-integer predictive control for wind turbine pitch angle regulation in full load regions using deep on-policy learning. Eng. Appl. Artif. Intell. 2025, 156, 111156. [Google Scholar] [CrossRef] [Scilit]
- Romero, A.; Aljalbout, E.; Song, Y.; Scaramuzza, D. Actor–critic model predictive control: Differentiable optimization meets reinforcement learning. arXiv 2025, arXiv:2306.09852. [Google Scholar] [CrossRef] [Scilit]
- Lee, H.; Choi, K.; Kim, W. Using Deep Reinforcement Learning for Dynamic Gain Adjustment of a Disturbance Observer. In Proceedings of the 2025 25th International Conference on Control, Automation and Systems (ICCAS), Jeju, Republic of Korea, 4–7 November 2025. [Google Scholar] [CrossRef] [Scilit]
- Lee, D.; Ahn, H.; Lee, J.; Bang, H. UAV with Parametric Uncertainty and Unmodeled Dynamics. In Proceedings of the AIAA SCITECH 2023 Forum, National Harbor, MD, USA, 23–27 January 2023; AIAA: Reston, VA, USA, 2023; p. 2357. [Google Scholar] [CrossRef] [Scilit]











| Parameter | Symbol | Value | Unit |
|---|---|---|---|
| Inertia coefficient 1 | 3.473 | Kg·m2 | |
| Inertia coefficient 2 | 0.196 | kg·m2 | |
| Inertia coefficient 3 | 0.242 | kg·m2 | |
| Joint 1 friction | 5.3 | N·m·s | |
| Joint 2 friction | 1.1 | N·m·s | |
| Link length 1 | 0.5 | m | |
| Link length 2 | 0.5 | m |
| Mode | Scenario | Notes | ||||||
|---|---|---|---|---|---|---|---|---|
| A | Baseline (fixed) | 1.0 | 1.0 | 1.0 | 0.60 | 0.25 | 0.15 | Adaptive reward only |
| B | Adaptive (no DOB) | 1.0 | 1.0 | 1.0 | 0.70 | 0.20 | 0.10 | Default configuration |
| C | Proposed (DOB + ACRL) | +2.0 | 0.7 | 0.7 | 0.80 | 0.15 | 0.05 | High accuracy, higher energy |
| C | Tracking heavy | 0.7 | 2.0 | 1.0 | 0.55 | 0.35 | 0.10 | Energy savings, mild error |
| C | Energy heavy | 0.8 | 1.0 | 2.0 | 0.55 | 0.20 | 0.25 | Smoother torques |
| C | Effort heavy | 2.5 | 2.5 | 2.5 | 0.70 | 0.20 | 0.10 | Fastest learning, risk of oscillation |
| C | Aggressive adaptation | 0.5 | 0.5 | 0.5 | 0.70 | 0.20 | 0.10 | Slower but stable |
| C | Conservative adaptation | 1.0 | 1.0 | 1.0 | 0.60 | 0.25 | 0.15 | Adaptive reward only |
| Case | RMS e1 | RMS e2 | Avg. Control Effort | Max EE Error (m) | Settling Time (s) |
|---|---|---|---|---|---|
| Case 1: AC (static reward) | 0.0139 | 0.0809 | 10.4380 | 0.2525 | 0.65 |
| Case 2: AC (actor–critic with an adaptive composite reward without a disturbance observer) | 0.0132 | 0.0358 | 10.8923 | 0.2399 | 0.54 |
| Case 3: DOB–ACRL (adaptive reward + DOB) | 0.0137 | 0.0274 | 10.9185 | 0.2299 | 0.34 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Tam, L.T.M.; Ngu, N.V.; Pham, D.H.; Mai, V.T. Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators. Actuators 2026, 15, 167. https://doi.org/10.3390/act15030167
Tam LTM, Ngu NV, Pham DH, Mai VT. Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators. Actuators. 2026; 15(3):167. https://doi.org/10.3390/act15030167
Chicago/Turabian StyleTam, Le Thi Minh, Nguyen Viet Ngu, Duc Hung Pham, and V. T. Mai. 2026. "Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators" Actuators 15, no. 3: 167. https://doi.org/10.3390/act15030167
APA StyleTam, L. T. M., Ngu, N. V., Pham, D. H., & Mai, V. T. (2026). Disturbance Observer-Based Actor–Critic Reinforcement Learning with Adaptive Reward for Energy-Efficient Control of Robotic Manipulators. Actuators, 15(3), 167. https://doi.org/10.3390/act15030167

