Next Article in Journal
Research on Stewart Platform Control Method for Wave Compensation Based on BiLSTM Prediction and ADRC
Next Article in Special Issue
A Hybrid Enhanced Harris Hawks Optimization Algorithm for AGV Path Planning in Smart Warehousing
Previous Article in Journal
Fault-Tolerant Vertical Load Redistribution of an Active Suspension Under Yaw-Rate and Roll-Rate Sensor Faults
Previous Article in Special Issue
Analysis of Time Drift and Real-Time Challenges in Programmable Logic Controller-Based Industrial Automation Systems: Insights from 24-Hour and 14-Day Tests
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Adaptive Actor–Critic Optimal Tracking Control for a Class of High-Order Nonlinear Systems with Partially Unknown Dynamics

1
School of Automation, Guangxi University of Science and Technology, Liuzhou 545616, China
2
Guangxi Key Laboratory of Logistics Unmanned Aircraft Technology for Transportation Industry, Liuzhou 545616, China
*
Author to whom correspondence should be addressed.
Actuators 2026, 15(3), 138; https://doi.org/10.3390/act15030138
Submission received: 20 January 2026 / Revised: 18 February 2026 / Accepted: 19 February 2026 / Published: 2 March 2026

Abstract

Optimal tracking control for high-order partially unknown nonlinear systems poses significant challenges, particularly in deriving tractable solutions without requiring persistent excitation (PE) conditions or precise system models. This study develops an adaptive optimal tracking control law using neural network (NN)-based reinforcement learning (RL) for high-order partially unknown nonlinear systems. By designing a cost function associated with the sliding mode variable (SMV), the original tracking control problem is equivalently transformed into solving the optimal control problem related to the tracking Hamilton–Jacobi–Bellman (HJB) equation. Since the analytical solution of the HJB equation is generally intractable, we employ a policy iteration algorithm derived from the HJB equation, where both the partial derivative of the optimal tracking cost function and the optimal control law are approximated by NNs. The proposed RL framework achieves simplification through actor–critic training laws derived under the condition that a simple function is zero. Finally, both a numerical example and a single-link robotic arm application are provided to demonstrate the effectiveness and advantages of the proposed adaptive optimal tracking control method.

1. Introduction

In recent decades, optimal tracking control for nonlinear systems has remained one of the most significant research topics in control theory, with extensive applications in industrial manufacturing [1], aerospace [2], and robotics [3,4]. However, the majority of existing controller designs fail to incorporate formal optimization frameworks for performance cost functions, despite energy/time efficiency being critical in practical engineering. This study addresses this gap by developing an optimal tracking control law for partially unknown continuous-time nonlinear systems. Using an actor–critic reinforcement learning strategy derived from the HJB equation, the proposed method enables system states to achieve asymptotic tracking of the desired reference trajectory with guaranteed convergence.
Nonlinear tracking control problems can be effectively addressed through nonlinear control techniques including feedback linearization [5,6], the backstepping approach [7,8,9,10], and sliding mode control (SMC) [3,11,12,13], among others. Feedback linearization, a mature methodology in nonlinear control theory, transforms nonlinear systems into partially or fully linear equivalents through coordinated nonlinear state feedback and coordinate transformations (including dynamic compensation). This enables the application of established linear control design methods. Unlike feedback linearization, the backstepping approach provides an alternative framework for nonlinear systems unbounded by linear constraints. This methodology synthesizes feedback control laws through an iterative recursive procedure. Its principal advantage lies in preserving beneficial nonlinearities while guaranteeing asymptotic stability for regulation and tracking tasks. However, these methods are mainly applicable to strictly feedback nonlinear systems. To enhance transient response performance, SMC has garnered significant research attention. As a robust control strategy, SMC’s distinctive feature lies in its adaptive control structure switching based on state-dependent sliding surfaces. Its multiple variants have been developed, including terminal SMC [11], integral SMC [3,12], and hierarchical SMC [13]. Recent years have witnessed increasingly in-depth research on adaptive optimal control integrated with sliding mode surface (SMS). Zhang et al. addressed the problem of SMS-based adaptive optimal control for a class of switched continuous-time nonlinear systems with average dwell time by employing an actor–critic RL strategy [14]. Zhao et al. investigated the SMS-based approximate optimal control problem in the context of nonlinear multi-player Stackelberg–Nash games [15]. Furthermore, Zhang et al. studied the adaptive optimal control scheme based on a hierarchical SMS for a class of switched continuous-time nonlinear systems subject to unknown disturbances, utilizing an actor–critic NN architectures [16]. However, all of these are aimed at the problem of optimal regulation.
Designing optimal control laws for nonlinear systems typically requires solving the HJB equation. However, the inherent nonlinearity of the HJB equation makes analytical solutions challenging to derive via conventional methods. To address this, RL algorithms have been integrated into the optimal control framework, enabling feasible solutions. As a machine learning paradigm, RL allows agents to learn optimal policies through environmental interactions [17,18]. This RL-based function approximation approach has successfully enabled adaptive optimal control, emerging as a prominent methodology for complex nonlinear control in recent decades.
Current research extensively explores optimal tracking for nonlinear systems using adaptive dynamic programming (ADP) [19,20]. This method combines dynamic programming (DP) and RL. Notably, Radac et al. address the key challenge of interleaving real-time interaction with online learning in RL-based control. They propose a near real-time online RL framework for model reference tracking, which employs an extended moving window for state encoding and online backpropagation training [21]. Modares et al. proposed a novel integral RL formulation for continuous-time nonlinear systems. This approach adapts methodologies originally developed for optimal regulation problems to address optimal tracking control. Recent advancements in RL-based tracking control have demonstrated significant progress across diverse nonlinear systems [18]. Jiang et al. investigated the adaptive optimal tracking control problem for networked discrete-time linear systems by directly utilizing the data transmitted through the communication network. This problem can be addressed by solving two sub-problems: an adaptive optimal control problem and an adaptive output regulation problem. In the case of random packet loss in the network, the tracking error of the closed-loop control system has also been proven to be asymptotically stable [22]. Radac et al. proposed a novel virtual state-feedback reference feedback tuning and approximate iterative value iterative reinforcement learning method for learning the linear reference model’s output tracking control of observable systems with unknown dynamic characteristics. The approximate iterative value iterative reinforcement learning can effectively reduce the learning complexity; the virtual state-feedback reference feedback tuning can learn stable controllers even in an open-loop environment with a low exploration degree, demonstrating its superiority in tracking control learning [23]. Wen et al. developed an optimized tracking control framework employing neural network-based actor–critic reinforcement learning for a class of nonlinear dynamic systems [24]. Wang et al. subsequently proposed a novel actor–critic RL scheme to achieve optimal tracking control for unmanned surface vehicles, explicitly addressing complex unknowns including dead-zone input nonlinearities, uncertain system dynamics, and external disturbances [25]. Shi et al. proposed a robust predictive fault-tolerant control method based on signal compensation. This method established a new model for nonlinear industrial processes that contain fault factors and unknown concentrated dynamic nonlinear terms. A sub-compensator was designed to eliminate the influence of unmodeled dynamics on the control effect. This method effectively reduced the conservatism of traditional fault-tolerant control and significantly improved the tracking performance of the system [26]. Further innovations include the fuzzy integral RL-based fault-tolerant control algorithm, which integrates RL techniques with fuzzy-augmented models to handle partially unknown systems with actuator faults [27]. Specialized applications have been demonstrated, including robust online tracking control for space manipulators and the nearly optimal trajectory tracking of autonomous surface vehicles [28,29]. While the backstepping technique has been widely applied in optimal tracking control of strict-feedback nonlinear systems [30,31,32,33], the same problem for high-order canonical nonlinear systems remains largely unexplored.
Motivated by the above discussions, this paper addresses a class of canonical-form high-order nonlinear system by proposing a novel adaptive optimal tracking control scheme. The main contributions comprise three aspects. First, an adaptive optimal tracking control scheme is proposed via an SMS for high-order canonical nonlinear systems. By constructing a cost function specifically related to the SMS, the original control problem is equivalently transformed into one that seeks an optimal control strategy. The designed SMS constrains the states of the error dynamic system, thereby forcing the tracking error to converge to zero with predefined dynamics. Second, in most of the existing literature, such as [18,20,24,25], the requirement of persistent excitation is necessary for training RL optimal control with adaptive parameters. The proposed optimization method can avoid this requirement. Third, an adaptive optimal tracking law is developed using an actor–critic NN RL framework. Compared with method in [34], our method achieves higher tracking accuracy and computational efficiency with fewer conditional constraints.
The rest of this paper is organized as follows: Section 2 describes the optimal tracking control problem of a class of canonical-form high-order nonlinear system. In Section 3, we present an SMS-based adaptive optimal tracking control using actor–critic RL. Based on the Lyapunov theory, it is proven that the error signals are semi-globally uniformly ultimately bounded (SGUUB), and the tracking error can be steered into a small neighborhood of zero in Section 4. Simulations are provided to demonstrate the effectiveness of the proposed method in Section 5. Finally, the conclusion summarized in Section 6.

2. Problem Description

A canonical form nonlinear system with relative degree n is given by [35]:
x ˙ i ( t ) = x i + 1 ( t ) , i = 1 , , n 1 x ˙ n ( t ) = f ( x ¯ ) + g u ,
where x ¯ ( t ) = x 1 ( t ) , x 2 ( t ) , , x n ( t ) T R n is the system state vector, u R is the control input, f ( x ¯ ) R is the unknown and bounded nonlinear dynamic function, and g R is the known control gain constant.
Assumption 1.
The system (1) is stabilizable, meaning there exists an admissible control policy u Ψ ( Ω ) that ensures the global stability of the closed-loop system.
Assumption 2.
The reference trajectory y d ( t ) and its i-th order derivatives y d ( i ) ( t ) ( i = 1 , , n ) are assumed to be bounded.
Definition 1
([36]). For a nonlinear system ζ ˙ = f ( ζ , t ) , the solution ζ ( t ) R n is called SGUUB if for any initial state ζ ( t 0 ) Ω , where Ω is a compact set, there exists constants σ > 0 and T ( σ , ζ ( t 0 ) ) > 0 such that ζ ( t )   σ for all t t 0 + T ( σ , ζ ( t 0 ) ) .
The control objective of this paper is to design an optimized tracking controller for partially unknown nonlinear dynamic system (1) using an NN-based actor–critic RL strategy. The proposed control law not only ensures all closed-loop error signals are SGUUB, but also guarantees that the output state x 1 ( t ) and its derivative signals rapidly track the reference trajectory y d ( t ) and its corresponding derivatives, respectively.

3. Optimal Tracking Control Design

3.1. Tracking HJB Equation

The tracking errors are defined as z 1 ( t ) = x 1 ( t ) y d ( t ) , z 2 ( t ) = x 2 ( t ) y ˙ d ( t ) , , and z n ( t ) = x n ( t ) y d ( n 1 ) ( t ) . Based on the dynamic model (1) and the above tracking error definitions, the error dynamical system is derived as follows:
z ˙ i ( t ) = z i + 1 ( t ) , i = 1 , , n 1 z ˙ n ( t ) = f ( z ¯ , y ¯ d ) + g u y d ( n ) ( t )
where z ¯ = [ z 1 ( t ) , z 2 ( t ) , , z n ( t ) ] T R n , and y ¯ d = [ y d ( t ) , y ˙ d ( t ) , , y d ( n 1 ) ( t ) ] T R n .
Based on the tracking errors, the sliding mode variable [37] is constructed as
s ( t ) = K 0 T z ¯ ( t ) = k 1 z 1 ( t ) + + k n 1 z n 1 ( t ) + z n ( t ) ,
where s ( t ) R , K 0 = [ k 1 , , k n 1 , 1 ] T R n , and k j > 0 ( j = 1 , , n 1 ) are design parameters.
Taking the time derivative of s ( t ) along the error dynamical system (2) yields
s ˙ ( t ) = K 1 T z ¯ ( t ) + f ( z ¯ , y ¯ d ) + g u y d ( n ) ( t ) = k 1 z 2 ( t ) + + k n 1 z n ( t ) + f ( z ¯ , y ¯ d ) + g u y d ( n ) ( t ) ,
where K 1 = [ 0 , k 1 , , k n 1 ] T R n .
The design parameters k j are chosen such that the polynomial ρ n 1 + k n 1 ρ n 2 + + k 2 ρ + k 1 is Hurwitz, i.e., all its roots lie in the open left-half plane. Note that when the tracking errors z 1 ( t ) , , z n ( t ) are constrained to the SMS s ( t ) = 0 , they will converge to a small neighborhood of zero.
Define an infinite-horizon cost function with respect to the SMS s as follows:
V ( s ) = t r ( s ( τ ) , u ( s ) ) d τ ,
where r ( s , u ) = s 2 + u 2 denotes the utility function.
Definition 2.
For the error dynamical system (2), the control policy u ( s ) is defined as admissible with respect to V ( s ) on the compact set Ω, denoted by u ( s ) Ψ ( Ω ) ; if u ( s ) is continuous, u ( 0 ) = 0 , u ( s ) can stabilize (2) on Ω and V ( s ) is finite for s Ω .
The optimal control objective for the error dynamical system (2) is to find an admissible control law that minimizes the cost function (5) over an infinite horizon, thus achieving the desired control task. Accordingly, the optimal cost function is defined as
V * ( s ) = min u Ψ ( Ω ) t r ( s ( τ ) , u ( s ) ) d τ = t r ( s ( τ ) , u * ( s ) ) d τ ,
where V * ( s ) and u * ( s ) denote the optimal cost function and optimal control law, respectively, and Ψ ( Ω ) represents the set of admissible controls over the domain Ω .
By calculating the time derivative of the cost function (6) along the sliding mode dynamic (4) and applying the optimality condition, the following HJB equation is derived:
H ( s , u * , V * ( s ) ) = r ( s , u * ) + V * ( s ) s ˙ = s 2 + u * 2 + V * ( s ) ( K 1 T z ¯ ( t ) + f ( z ¯ , y ¯ d ) + g u * y d ( n ) ( t ) ) = 0
where V * ( s ) = V * ( s ) / s denotes the partial derivative of V * ( s ) with respect to s.
Assuming the solution of (7) exists and is unique, the optimal control u * can be obtained by applying the stationarity condition H ( s , u * , V * ( s ) u * = 0 as
u * = g 2 V * ( s ) .
Substituting (8) into (7) yields
s 2 g 2 4 V * ( s ) 2 + V * ( s ) K 1 T z ¯ ( t ) + f ( z ¯ , y ¯ d ) y d ( n ) ( t ) = 0 .
The partial derivative V * ( s ) can be obtained by solving the HJB Equation (9). The optimal control protocol can then be derived by combining the result with (8). However, solving (9) analytically is difficult due to the inherent system nonlinearity. Furthermore, the lack of complete knowledge about the system dynamics f ( z ¯ , y ¯ d ) further complicates the solution of (9). To overcome these challenges, this paper employs a reinforcement learning method with an actor–critic architecture.

3.2. NNs Approximation in Actor–Critic RL

To achieve the optimal tracking control, the term V * ( s ) in (9) is decomposed as
V * ( s ) = 2 γ u s ( t ) + 2 g 2 K 1 T z ¯ ( t ) + 1 g 2 V 0 ( s , z ¯ ) ,
where γ u > 0 is the designed parameter and 1 g 2 V 0 ( s , z ¯ ) = 2 γ u s ( t ) 2 g 2 K 1 T z ¯ ( t ) + V * ( s ) is the unknown function in this dual architecture.
Remark 1.
In (10), the terms 2 γ u s ( t ) and 2 g 2 K 1 T z ¯ ( t ) are designed to ensure the boundedness and stability of the system signals.
According to the Weierstrass approximation theorem [38], the continuous function V 0 ( s , z ¯ ) can be approximated by the weighted sum of basis functions. Given the relationship s ( t ) = K 0 T z ¯ ( t ) between s and z ¯ , the approximation takes the form as
V 0 ( s , z ¯ ) = W * T Φ ( z ¯ ) + ε ( z ¯ ) ,
where W * = [ w 1 , w 2 , , w N ] R N is the ideal weight vector and bounded by a positive constant W ¯ such that W *   W ¯ . Moreover, Φ = [ ϕ 1 , ϕ 2 , , ϕ N ] R N and ε ( z ¯ ) R are, respectively, the activation function vector and the approximation error. With the number of neurons in the hidden layer N , the NN approximation error ε ( z ¯ ) 0 .
Putting (11) into (10) yields
V * ( s ) = 2 γ u s ( t ) + 2 g 2 K 1 T z ¯ ( t ) + 1 g 2 W * T Φ ( z ¯ ) + 1 g 2 ε ( z ¯ ) .
Since the ideal weight vector W * is unknown, (12) is approximated by a critic NN as
V ^ * ( s ) = 2 γ u s ( t ) + 2 g 2 K 1 T z ¯ ( t ) + 1 g 2 W ^ c T Φ ( z ¯ ) ,
where W ^ c denotes the current estimate of W * . The weight estimation error is then defined as W ˜ c = W ^ c W * .
According to (8) and (13), the optimal control to the system is approximated by an additional actor NN as
u ^ * = g γ u s ( t ) 1 g K 1 T z ¯ ( t ) 1 2 g W ^ a T Φ ( z ¯ ) ,
where W ^ a is the current estimate of W * . The actor NN weight estimation error is defined as W ˜ a = W ^ a W * .
The critic and actor adaptive training laws are designed as
W ^ ˙ c ( t ) = γ c Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ c ( t ) ,
W ^ ˙ a ( t ) = Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m γ a W ^ a ( t ) W ^ c ( t ) γ c W ^ c ( t ) ,
where γ c > 0 and γ a > 0 are the learning rates for the critic and actor networks, respectively. Here, σ > 0 is a regularization parameter, and I m denotes the m × m identity matrix.
Inserting both (13) and (14) into (7), the approximated HJB equation yields
H ^ ( s , u ^ * , V ^ * ( s ) ) = s 2 + g 2 γ u s ( t ) 1 g 2 K 1 T z ¯ ( t ) 1 2 g 2 W ^ a T Φ ( z ¯ ) 2 + 2 γ u s ( t ) + 2 g 2 K 1 T z ¯ ( t ) + 1 g 2 W ^ c T ( t ) Φ ( z ¯ ) × g 2 γ u s ( t ) 1 2 W ^ a T Φ ( z ¯ ) + f ( z ¯ , y ¯ d ) y d ( n ) ( t ) .
The Bellman residual error E is introduced, with the following expression:
E = H ^ ( s , u ^ * , V ^ * ( s ) ) H ( s , u * , V * ( s ) ) = H ^ ( s , u ^ * , V ^ * ( s ) )
According to the previous analysis and the optimal control theory, the approximate optimized control u ^ * should ensure E 0 . If the condition E = 0 is met and has a unique solution, it is equivalent to
H ^ ( s , u ^ * , V ^ * ( s ) ) W ^ a = 1 2 g 2 Φ ( z ¯ ) Φ T ( z ¯ ) W ^ a ( t ) W ^ c ( t ) = 0 .
To make (19) hold, construct the following function Γ ( t ) such that
Γ ( t ) = W ^ a ( t ) W ^ c ( t ) T W ^ a ( t ) W ^ c ( t ) .
Obviously, Γ ( t ) = 0 ensures that (19) holds. When the weight vectors of both networks are trained to satisfy W ^ a ( t ) = W ^ c ( t ) , Equations (13) and (14) fulfill the relation (8), and the control action approaches the optimal solution.
Theorem 1.
Under the adaptive laws given by (15) and (16), the function Γ ( t ) converges to zero asymptotically.
Proof. 
The partial derivatives of Γ ( t ) with respect to the weight estimates are
Γ ( t ) W ^ a ( t ) = Γ ( t ) W ^ c ( t ) = 2 W ^ a ( t ) W ^ c ( t ) .
The time derivative of Γ ( t ) along the trajectories of the system is
Γ ˙ ( t ) = Γ ( t ) W ^ a ( t ) W ^ ˙ a ( t ) + Γ ( t ) W ^ c ( t ) W ^ ˙ c ( t ) = Γ ( t ) W ^ a ( t ) W ^ ˙ a ( t ) W ^ ˙ c ( t ) .
Substituting the adaptive laws (15) and (16) into (22) yields
Γ ˙ ( t ) = 2 W ^ a W ^ c T γ a Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ a W ^ c = 2 γ a W ^ a W ^ c T Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ a W ^ c .
Noting that Γ ( t ) = W ^ a W ^ c 2 , the expression in (23) can be rewritten as a function of Γ ( t ) itself. Since the matrix Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m is positive definite, it follows that
Γ ˙ ( t ) = 2 γ a W ^ a W ^ c T Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ a W ^ c 2 γ a λ min W ^ a W ^ c 2 = 2 γ a λ min Γ ( t ) ,
where λ min > 0 is the minimum eigenvalue of the positive definite matrix Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m .
This inequality, Γ ˙ ( t ) 2 γ a λ min Γ ( t ) , implies that Γ ( t ) converges to zero exponentially. □
Remark 2.
The critic NN in (14) may be used to determine the actor without using another NN for the actor. However, to obtain the form of (19) and to facilitate the derivation of weight update laws and subsequent stability analysis, separate NNs are necessary for the actor and critic.
Remark 3.
This paper presents an online reinforcement learning framework based on an actor–critic architecture. The actor and critic networks are employed to approximate the optimal control law and the value function gradient, respectively. Critically, the update laws for the network weights are derived from the convergence condition of a constructed scalar function Γ ( t ) to zero, rather than from the minimization of the Bellman residual error. This constitutes a key methodological departure. Conventional adaptive dynamic programming methods typically design update laws by minimizing the squared Bellman residual. However, solving this nonlinear equation online is computationally complex. The proposed approach, by leveraging the simple convergence condition of Γ ( t ) , circumvents this complexity, thereby significantly reducing the online computational burden of the control design.
Remark 4.
As demonstrated in this subsection, the proposed RL approach not only features a simple structure but also eliminates the need for explicit knowledge of the system dynamic f ( z ¯ , y ¯ d ) when solving the optimal control problem.

4. Main Results

Lemma 1
([39]). X ( t ) R is a positive continuous function and has the bounded initial value X ( 0 ) . If the inequality X ˙ ( t ) α X ( t ) + β holds, where α > 0 and β > 0 are constants, then
X ( t ) e α t X ( 0 ) + β α ( 1 e α t ) .
Theorem 2.
Consider the partially unknown nonlinear system (1) under Assumptions 1 and 2, with a bounded initial state. The optimal control law is derived from the actor–critic RL framework, where the critic and actor are updated according to (13) and (14), using the learning rules given in (15) and (16), respectively. If the parameters satisfy γ u > 1 2 + 1 g 2 , γ c > γ a 2 and γ a > 1 , then the tracking error vector z ¯ ( t ) and the weight estimation errors W ˜ c ( t ) and W ˜ a ( t ) are SGUUB. Moreover, the tracking errors converge to an arbitrarily small neighborhood of zero.
Proof. 
Choose the Lyapunov candidate function as
L ( t ) = 1 2 s 2 + 1 2 W ˜ c T ( t ) W ˜ c ( t ) + 1 2 W ˜ a T ( t ) W ˜ a ( t ) .
Calculating the time derivative (26) along (4), (15) and (16) yields
L ˙ ( t ) = s g 2 γ u s 1 2 W ^ a T ( t ) Φ ( z ¯ ) + f ( z ¯ , y ¯ d ) y d ( n ) ( t ) γ c W ˜ c T ( t ) ( Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m ) W ^ c ( t ) γ a W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ a ( t ) W ^ c ( t ) + γ c W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ c ( t ) .
It follows from W ˜ c = W ^ c W * and W ˜ a = W ^ a W * that
W ^ a T Φ ( z ¯ ) = W ˜ a T Φ ( z ¯ ) + W * T Φ ( z ¯ )
and
W ^ a ( t ) W ^ c ( t ) = W ^ a ( t ) W * + W * W ^ c ( t ) = W ˜ a W ˜ c .
Using (28) and (29), (27) can be rewritten as
L ˙ ( t ) = g 2 γ u s 2 1 2 s W ˜ a T ( t ) Φ ( z ¯ ) 1 2 s W * T Φ ( z ¯ ) + s f ( z ¯ , y ¯ d ) s y d ( n ) ( t ) γ c W ˜ c T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ c ( t ) γ a W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m × W ˜ a ( t ) W ˜ c ( t ) + γ c W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ c ( t ) .
From Cauchy–Schwartz and Young’s inequalities, we obtain
s f ( z ¯ , y ¯ d ) 1 2 s 2 + 1 2 f 2 ( z ¯ , y ¯ d ) , s y d ( n ) ( t ) 1 2 s 2 + 1 2 y d ( n ) ( t ) 2 , s W ˜ a T ( t ) Φ ( z ¯ ) 1 2 s 2 + 1 2 W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ a ( t ) , s W * T Φ ( z ¯ ) 1 2 s 2 + 1 2 W * T Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W * .
Under these inequalities, (30) reduces to
L ˙ ( t ) g 2 γ u 1 2 1 s 2 + 1 2 f 2 ( z ¯ , y ¯ d ) + 1 2 y d ( n ) ( t ) 2 γ c W ˜ c T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m × W ^ c ( t ) γ a 1 4 W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ a ( t ) + γ a W ˜ a T ( t ) ( Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m ) W ˜ c ( t ) + γ c W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ c ( t ) + 1 4 W * T Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W * .
From the definition W ˜ c = W ^ c W * , the following equation can be derived:
W ˜ c T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ c ( t ) = W ˜ c T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ c ( t ) + W ^ c T ( t ) ( Φ ( z ¯ ) × Φ T ( z ¯ ) + σ I m ) W ^ c ( t ) W * T Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W * .
Moreover, the following inequality can be derived according to Young’s inequality:
W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ c ( t ) 1 2 W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ a ( t ) + 1 2 W ˜ c T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ c ( t ) , W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ c ( t ) 1 2 W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ a ( t ) + 1 2 W ^ c T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ^ c ( t ) .
According to the above Equation (33) and inequality (34), the following inequality can be derived from (32):
L ˙ ( t ) g 2 γ u 1 2 1 s 2 γ c γ a 2 W ˜ c T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ c ( t ) 1 2 γ a γ c 1 2 W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ a ( t ) γ c 2 W ^ c T ( t ) ( Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m ) W ^ c ( t ) + γ c + 1 4 W * T Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W * + 1 2 f 2 ( z ¯ , y ¯ d ) + 1 2 y d ( n ) ( t ) 2 .
This inequality (35) can be simplified to
L ˙ ( t ) g 2 γ u 1 2 1 s 2 γ c γ a 2 W ˜ c T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ c ( t ) 1 2 γ a γ c 1 2 W ˜ a T ( t ) Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W ˜ a ( t ) + C ( t ) ,
where C ( t ) = ( γ c + 1 4 ) W * T Φ ( z ¯ ) Φ T ( z ¯ ) + σ I m W * + 1 2 f 2 ( z ¯ , y ¯ d ) + 1 2 y d ( n ) ( t ) 2 can be bounded by a constant β , i.e., C ( t ) β .
Let α = min g 2 γ u 1 2 1 , ( γ c γ a 2 ) σ , 1 2 ( γ a γ c 1 2 ) σ ; then, (36) can be rewritten as
L ˙ ( t ) α L ( t ) + β
Applying Lemma 1, we rigorously obtain the following inequality:
L ( t ) e α t L ( 0 ) + β α 1 e α t .
The above inequality implies that the signals z ¯ ( t ) , W ˜ c ( t ) , and W ˜ a ( t ) are SGUUB. Moreover, by selecting the design parameters to make α sufficiently large, the tracking errors z 1 ( t ) , , z n ( t ) can be made to converge to an arbitrarily small neighborhood of zero. This completes the proof. □
Remark 5.
The designed parameter γ u and the learning rates γ c and γ a must satisfy the inequalities in Theorem 2 to guarantee stability. An excessively large γ u enhances robustness but leads to high control effort, while an excessively small γ u results in a sluggish response and poor disturbance rejection. Excessively large learning rates γ c and γ a accelerate convergence but risk numerical instability and oscillatory updates; excessively small rates, though ensuring update smoothness, significantly slow down the learning process.

5. Simulation Experiment

In this section, the effectiveness and advantages of the proposed adaptive optimal tracking control method are demonstrated through two simulation examples.
Example 1.
Consider a second-order nonlinear dynamic model as presented in [34]:
x ˙ 1 ( t ) = x 2 ( t ) , x ˙ 2 ( t ) = 1 x 1 ( t ) 2 x 2 ( t ) + g u ,
where x 1 ( t ) and x 2 ( t ) are the system states, u is the control input, and g = 1 is the control gain, with initial conditions x 1 ( 0 ) = 2 and x 2 ( 0 ) = 3 .
Based on the reference trajectory y d ( t ) = 5 s i n ( t ) , the tracking errors can be obtained as
z 1 ( t ) = x 1 ( t ) 5 sin ( t ) , z 2 ( t ) = x 2 ( t ) 5 cos ( t ) .
The error dynamics derived in (39) are
z ˙ 1 ( t ) = z 2 ( t ) , z ˙ 2 ( t ) = 1 z 1 ( t ) 5 s i n ( t ) 2 z 2 ( t ) 5 c o s ( t ) + u + 5 sin ( t ) .
Following from (3), the sliding mode variable is designed as
s ( t ) = 2 z 1 ( t ) + z 2 ( t ) .
The critic and actor network weight vectors are defined as W ^ c = W ^ c , 1 , W ^ c , 2 , W ^ c , 3 T and W ^ a = W ^ a , 1 , W ^ a , 2 , W ^ a , 3 T , respectively. The activation function vector is given by Φ ( z ¯ ) = z 1 2 , z 1 z 2 , z 2 2 T . The learning rates and regularization parameters for the training laws of both the critic and actor (15) and (16) are set to γ c = 3 , γ a = 5 , and σ = 1 . The weight vectors are initialized to W ^ c ( 0 ) = 1 , 0 , 1 T and W ^ a ( 0 ) = 0 , 1 , 0 T .
The simulation results are presented in Figure 1, Figure 2, Figure 3 and Figure 4. Specifically, Figure 1 shows that the states x 1 ( t ) and x 2 ( t ) closely follow the reference trajectories y d and its derivative y ˙ d , respectively. The corresponding tracking errors z 1 and z 2 , shown in Figure 2, converge to zero. Furthermore, the neural network weights for the actor and critic remain bounded, as illustrated in Figure 3. Finally, the evolution of the utility function is depicted in Figure 4. Compared with the simulation results in [34], the method proposed in this paper achieves a superior tracking performance.
The performance of our proposed algorithm is first evaluated in Example 1 through a comparative study with the algorithm in [34]. Its adaptability is then examined in Example 2 using a more complex, practical robotic arm system. This system is integrated with commercial servo actuators, which are responsible for converting electrical control signals into joint movements.
Example 2.
This example considers a single-link robotic arm system with flexible joints. In the case of ignoring damping, its nonlinear dynamic equations are given by [37]:
I q ¨ 1 + M g L s i n ( q 1 ) + k ( q 1 + q 2 ) = 0 , J q ¨ 2 k ( q 1 q 2 ) = u ,
where q 1 and q 2 represent angular positions ( q 1 is the system output), I and J are moments of inertia, k is the spring constant, M is the total mass, L is the distance, and u is the torque input.
The above Equation (43) can be transformed to the standard form:
x ˙ 1 ( t ) = x 2 ( t ) , x ˙ 2 ( t ) = x 3 ( t ) , x ˙ 3 ( t ) = x 4 ( t ) , x ˙ 4 ( t ) = M g L I c o s ( x 1 ( t ) ) + k I + k J x 3 ( t ) + M g L I x 2 2 ( t ) k J s i n ( x 1 ( t ) ) + k I J u ,
where x 1 = q 1 , x 2 = q ˙ 1 , x 3 = M g L I s i n ( q 1 ) k I ( q 1 q 2 ) , and x 4 = M g L I q ˙ 1 c o s ( q 1 ) k I ( q ˙ 1 q ˙ 2 )
The parameter values of the robotic arm system are set to k = 31 [N m/rad], I = 0.031 [kg m2], J = 0.004 [kg m2], g = 9.81 [m/s2], M = 0.4 [kg], and L = 0.2 [m].
Given the reference trajectory y d ( t ) = s i n ( 2.5 t ) , the tracking errors can be calculated:
z 1 ( t ) = x 1 ( t ) s i n ( 2.5 t ) , z 2 ( t ) = x 2 ( t ) 2.5 c o s ( 2.5 t ) , z 3 ( t ) = x 3 ( t ) + 6.25 s i n ( 2.5 t ) , z 4 ( t ) = x 4 ( t ) + 15.625 c o s ( 2.5 t ) .
The error dynamics derived in (44) are
z ˙ 1 ( t ) = z 2 ( t ) , z ˙ 2 ( t ) = z 3 ( t ) , z ˙ 3 ( t ) = z 4 ( t ) , z ˙ 4 ( t ) = M g L I c o s ( z 1 ( t ) + s i n ( 2.5 t ) ) + k I + k J ( z 3 ( t ) 6.25 s i n ( 2.5 t ) ) + M g L I ( z 2 ( t ) + 2.5 c o s ( 2.5 t ) ) 2 k J s i n ( z 1 ( t ) + s i n ( 2.5 t ) ) + k I J u 39.0625 s i n ( 2.5 t )
Based on (3), the sliding mode variable is designed as
s ( t ) = z 1 ( t ) + z 2 ( t ) + 2 z 3 ( t ) + z 4 ( t ) .
The activation function vector is given by
Φ ( z ¯ ) = [ z 1 3 , z 1 2 z 2 , z 1 2 z 3 , z 1 2 z 4 , z 1 z 2 2 , z 1 z 3 2 , z 1 z 4 2 , z 1 z 2 z 3 , z 1 z 2 z 4 , z 1 z 3 z 4 , z 2 3 , z 2 2 z 3 , z 2 2 z 4 , z 2 z 3 2 , z 2 z 4 2 , z 2 z 3 z 4 , z 3 3 , z 3 2 z 4 , z 3 z 4 2 , z 4 3 ] T .
The initial states are x 1 ( 0 ) = 5 , x 2 ( 0 ) = 1 , x 3 ( 0 ) = 2 and x 4 ( 0 ) = 0.3 . The weight vectors are initialized to
W ^ c ( 0 ) = 1 , 0 , 1 , 2 , 0 , 1 , 0 , 0 , 1 , 1 , 1 , 1 , 0 , 0 , 1 , 1 , 2 , 0 , 1 , 0 T
and
W ^ a ( 0 ) = 0 , 5 , 0 , 5 , 0 , 2 , 0 , 8 , 0 , 3 , 4 , 0 , 2 , 0 , 6 , 2 , 2 , 3 , 0 , 5 T .
The design parameters are set to γ u = 30 , γ c = 3 , γ a = 5 , and σ = 1 .
The simulation results of Example 2 are presented in Figure 5, Figure 6, Figure 7 and Figure 8. As shown in Figure 5, the states x 1 ( t ) , x 2 ( t ) , x 3 ( t ) , and x 4 ( t ) closely follow the reference trajectory y d and its derivatives of various orders y ˙ d , y ¨ d , and y d ( 3 ) , respectively. The corresponding tracking errors z 1 , z 2 , z 3 , and z 4 , depicted in Figure 6, converge to a small neighborhood of zero. Furthermore, Figure 7 demonstrates that the neural network weights for both the actor and critic remain bounded. Finally, the evolution of the utility function is shown in Figure 8. These results demonstrate that the proposed high-order tracking control scheme for canonical nonlinear systems maintains the boundedness of the error signal while ensuring accurate tracking of each system state to its corresponding reference signal.

6. Conclusions

This paper proposes an SMS-based adaptive optimal tracking control scheme, utilizing an NN-based RL approach, for a class of high-order nonlinear systems with partial unknowns. A cost function defined in terms of the SMS is constructed to derive the optimal control law. Accordingly, the HJB equation is addressed within an actor–critic NN framework. Leveraging the inherent relationship between the SMS and the tracking errors significantly simplifies the computational process. Under the premise of ensuring the stability of the closed-loop system, the update law of the actor–critic network is indirectly derived under the condition that a simple function is zero, without the need for the persistence excitation condition. Furthermore, the proposed design requires no prior knowledge of the internal system dynamics. Meanwhile, the proposed method offers a significant advantage over [34] by eliminating the need for a Riccati-like equation. The comparative simulation results demonstrate the superior tracking performance of the proposed scheme over existing methods.
The inadequacy of this method lies in two of its aspects: (1) The control gain of the research system is constant. However, in most practical engineering applications, the control gain often changes with the state, so we will study the case where the control gain changes with the state, even in the case of unknown dynamics. (2) The performance of the closed-loop system and the learning efficiency are hig y dependent on the selection of the SMS parameter k j . Although these parameters satisfy the Hurwitz condition to ensure stability, their specific values need to be empirically balanced among the convergence speed, control energy consumption, and robustness, lacking a systematic tuning method.

Author Contributions

D.X.: investigation, methodology, and writing—review; X.L.: investigation and software; F.L.: formal analysis; J.T.: conceptualization and writing—original draft preparation. All authors have read and agreed to the published version of the manuscript.

Funding

This research was supported by the National Natural Science Foundation of China under Grant No. 61463002 and the Doctoral Foundation of Guangxi University of Science and Technology Grant No. Xiaokebo 22Z04.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are already in the article. For further consultation, please contact the corresponding author.

Acknowledgments

The authors thank the journal editors and reviewers for their valuable suggestions and opinions.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Li, P.; Duan, G.; Zhang, B.; Wang, P.; Wang, Y. High-Order Fully Actuated Approach for Output Tracking Control of Flexible Servo Systems Subject to Uncertainties and Disturbances. IEEE Trans. Ind. Electron. 2025, 72, 9433–9443. [Google Scholar] [CrossRef]
  2. Bu, X.; Qi, Q. Fuzzy optimal tracking control of hypersonic flight vehicles via single-network adaptive critic design. IEEE Trans. Fuzzy Syst. 2020, 30, 270–278. [Google Scholar] [CrossRef]
  3. Lee, J.; Chang, P.H.; Jin, M. Adaptive integral sliding mode control with time-delay estimation for robot manipulators. IEEE Trans. Ind. Electron. 2017, 64, 6796–6804. [Google Scholar] [CrossRef]
  4. Wang, C.; Liu, Z.; Sun, S.; Wang, Z.; Ma, K.; Mao, Q.; Xue, X.; Chen, X.; Zhao, K.; Hu, T. Trajectory tracking of a mobile robot in underground roadways based on hierarchical model predictive control. Actuators 2026, 15, 47. [Google Scholar] [CrossRef]
  5. Voos, H. Nonlinear control of a quadrotor micro-UAV using feedback-linearization. In Proceedings of the 2009 IEEE International Conference on Mechatronics, Malaga, Spain, 14–17 April 2009; pp. 1–6. [Google Scholar]
  6. Bonna, R.; Camino, J.F. Trajectory tracking control of a quadrotor using feedback linearization. In Proceedings of the International Symposium on Dynamic Problems of Mechanics, Natal, Brazil, 22–27 February 2015; Volume 1. [Google Scholar]
  7. Yousefizadeh, S.; Bendtsen, J.D.; Vafam, N.; Khooban, M.H.; Blaabjerg, F.; Dragičević, T. Tracking control for a DC microgrid feeding uncertain loads in more electric aircraft: Adaptive backstepping approach. IEEE Trans. Ind. Electron. 2018, 66, 5644–5652. [Google Scholar] [CrossRef]
  8. Zhao, K.; Song, Y. Removing the feasibility conditions imposed on tracking control designs for state-constrained strict-feedback systems. IEEE Trans. Autom. Control 2018, 64, 1265–1272. [Google Scholar] [CrossRef]
  9. Yu, J.; Ma, Y.; Yu, H.; Lin, C. Adaptive fuzzy dynamic surface control for induction motors with iron losses in electric vehicle drive systems via backstepping. Inf. Sci. 2017, 376, 172–189. [Google Scholar] [CrossRef]
  10. Song, Y.D.; Zhou, S. Tracking control of uncertain nonlinear systems with deferred asymmetric time-varying full state constraints. Automatica 2018, 98, 314–322. [Google Scholar] [CrossRef]
  11. Wang, Y.; Gu, L.; Xu, Y.; Cao, X. Practical tracking control of robot manipulators with continuous fractional-order nonsingular terminal sliding mode. IEEE Trans. Ind. Electron. 2016, 63, 6194–6204. [Google Scholar] [CrossRef]
  12. Qiao, L.; Zhang, W. Adaptive non-singular integral terminal sliding mode tracking control for autonomous underwater vehicles. IET Control Theory Appl. 2017, 11, 1293–1306. [Google Scholar] [CrossRef]
  13. Hwang, C.L.; Chiang, C.C.; Yeh, Y.W. Adaptive fuzzy hierarchical sliding-mode control for the trajectory tracking of uncertain underactuated nonlinear dynamic systems. IEEE Trans. Fuzzy Syst. 2013, 22, 286–299. [Google Scholar] [CrossRef]
  14. Zhang, H.; Wang, H.; Niu, B.; Zhang, L.; Ahmad, A.M. Sliding-mode surface-based adaptive actor–critic optimal control for switched nonlinear systems with average dwell time. Inf. Sci. 2021, 580, 756–774. [Google Scholar] [CrossRef]
  15. Zhao, H.; Zhao, N.; Zong, G.; Zhao, X.; Xu, N. Sliding-mode surface-based approximate optimal control for nonlinear multiplayer Stackelberg-Nash games via adaptive dynamic programming. Commun. Nonlinear Sci. Numer. Simul. 2024, 132, 107928. [Google Scholar] [CrossRef]
  16. Zhang, H.; Zhao, X.; Wang, H.; Zong, G.; Xu, N. Hierarchical sliding-mode surface-based adaptive actor–critic optimal control for switched nonlinear systems with unknown perturbation. Trans. Neural Netw. Learn. Syst. 2022, 35, 1559–1571. [Google Scholar] [CrossRef]
  17. Huang, J.; Xu, D.; Li, Y.; Zhang, X.; Zhao, J. Inverse reinforcement learning for discrete-time linear systems based on inverse optimal control. ISA Trans. 2025, 163, 108–119. [Google Scholar] [CrossRef] [PubMed]
  18. Modares, H.; Lewis, F.L. Optimal tracking control of nonlinear partially-unknown constrained-input systems using integral reinforcement learning. Automatica 2014, 50, 1780–1792. [Google Scholar] [CrossRef]
  19. Xu, D.; Wang, Q.; Li, Y. Optimal guaranteed cost tracking of uncertain nonlinear systems using adaptive dynamic programming with concurrent learning. Int. J. Control Autom. Syst. 2020, 18, 1116–1127. [Google Scholar] [CrossRef]
  20. Na, J.; Lv, Y.; Zhang, K.; Zhao, J. Adaptive identifier-critic-based optimal tracking control for nonlinear systems with experimental validation. IEEE Trans. Syst. Man Cybern. Syst. 2020, 52, 459–472. [Google Scholar] [CrossRef]
  21. Radac, M.B.; Chirla, D.P. Near real-time online reinforcement learning with synchronous or asynchronous updates. Sci. Rep. 2025, 15, 17158. [Google Scholar] [CrossRef] [PubMed]
  22. Jiang, Y.; Liu, L.; Feng, G. Adaptive optimal tracking control of networked linear systems under two-channel stochastic dropouts. Automatica 2024, 165, 111690. [Google Scholar] [CrossRef]
  23. Radac, M.B.; Borlea, A.I. Virtual state feedback reference tuning and value iteration reinforcement learning for unknown observable systems control. Energies 2021, 14, 1006. [Google Scholar] [CrossRef]
  24. Wen, G.; Chen, C.P.; Ge, S.S.; Yang, H.; Liu, X. Optimized adaptive nonlinear tracking control using actor–critic reinforcement learning strategy. IEEE Trans. Ind. Inform. 2019, 15, 4969–4977. [Google Scholar] [CrossRef]
  25. Wang, N.; Gao, Y.; Zhao, H.; Ahn, C.K. Reinforcement learning-based optimal tracking control of an unknown unmanned surface vehicle. IEEE Trans. Neural Netw. Learn. Syst. 2020, 32, 3034–3045. [Google Scholar] [CrossRef]
  26. Shi, H.; Yang, C.; Peng, B.; Su, C.; El-Sherbeeny, A.M.; Li, Z. Robust predictive fault-tolerant control based on signal compensation for nonlinear industrial processes with partial actuator failures. Int. J. Control 2026, 314, 131655. [Google Scholar] [CrossRef]
  27. Zhang, H.; Zhang, K.; Cai, Y.; Han, J. Adaptive fuzzy fault-tolerant tracking control for partially unknown systems with actuator faults via integral reinforcement learning method. IEEE Trans. Fuzzy Syst. 2019, 27, 1986–1998. [Google Scholar] [CrossRef]
  28. Zhuang, H.; Zhou, H.; Shen, Q.; Wu, S.; Razoumny, V.Y.; Razoumny, Y.N. Optimal robust online tracking control for space manipulator in task space using off-policy reinforcement learning. Aerosp. Sci. Technol. 2024, 153, 109446. [Google Scholar] [CrossRef]
  29. Wu, T.; Zhang, Y.; Yang, X.; Ye, H.; Xiang, Z. Predefined-time nearly optimal trajectory tracking control for autonomous surface vehicles with unknown dynamics. Ocean. Eng. 2025, 327, 121021. [Google Scholar] [CrossRef]
  30. Wen, G.; Ge, S.S.; Tu, F. Optimized backstepping for tracking control of strict-feedback systems. IEEE Trans. Neural Netw. Learn. Syst. 2018, 29, 3850–3862. [Google Scholar] [CrossRef] [PubMed]
  31. Liu, Y.; Zhu, Q.; Wen, G. Adaptive tracking control for perturbed strict-feedback nonlinear systems based on optimized backstepping technique. IEEE Trans. Neural Netw. Learn. Syst. 2020, 33, 853–865. [Google Scholar] [CrossRef]
  32. Huang, Z.; Bai, W.; Li, T.; Long, Y.; Chen, C.P.; Liang, H.; Yang, H. Adaptive reinforcement learning optimal tracking control for strict-feedback nonlinear systems with prescribed performance. Inf. Sci. 2023, 621, 407–423. [Google Scholar] [CrossRef]
  33. Yuan, H.; Cao, L.; Lin, W.; Xiao, W.; Li, X. Learning-observer-based fixed-time optimal tracking control for robotic manipulators with full-state constraints. Neurocomputing 2025, 642, 130186. [Google Scholar] [CrossRef]
  34. Wen, G.; Niu, B. Optimized tracking control based on reinforcement learning for a class of high-order unknown nonlinear dynamic systems. Inf. Sci. 2022, 606, 368–379. [Google Scholar] [CrossRef]
  35. Song, Y.; Wang, Y.; Holloway, J.; Krstic, M. Time-varying feedback for regulation of normal-form nonlinear systems in prescribed finite time. Automatica 2017, 83, 243–251. [Google Scholar] [CrossRef]
  36. Zhu, J.; Wen, G.; Veluvolu, K.C. Optimized backstepping consensus control using adaptive observer-critic–actor reinforcement learning for strict-feedback multi-agent systems. J. Frankl. Inst. 2024, 361, 106693. [Google Scholar] [CrossRef]
  37. Hušek, P. Adaptive sliding mode control with moving sliding surface. Appl. Soft Comput. 2016, 42, 178–183. [Google Scholar] [CrossRef]
  38. Dao, P.N.; Phung, M.H. Nonlinear robust integral based actor–critic reinforcement learning control for a perturbed three-wheeled mobile robot with mecanum wheels. Comput. Electr. Eng. 2025, 121, 109870. [Google Scholar] [CrossRef]
  39. Lin, J.; Wang, M.; Yan, H.; Yang, W. Prescribed-Time Optimal Tracking Control for a Class of Stochastic Systems Using Reinforcement Learning. J. Frankl. Inst. 2025, 362, 107881. [Google Scholar] [CrossRef]
Figure 1. The tracking performances: (a) The state x 1 tracking performance; (b) The state x 2 tracking performance.
Figure 1. The tracking performances: (a) The state x 1 tracking performance; (b) The state x 2 tracking performance.
Actuators 15 00138 g001
Figure 2. The tracking errors: (a) The tracking error z 1 ; (b) The tracking error z 2 .
Figure 2. The tracking errors: (a) The tracking error z 1 ; (b) The tracking error z 2 .
Actuators 15 00138 g002
Figure 3. The actor and critic network weights: (a) The actor network weight W ^ a ; (b) The critic network weight W ^ c .
Figure 3. The actor and critic network weights: (a) The actor network weight W ^ a ; (b) The critic network weight W ^ c .
Actuators 15 00138 g003
Figure 4. The utility function r.
Figure 4. The utility function r.
Actuators 15 00138 g004
Figure 5. The tracking performances: (a) The state x 1 tracking performance; (b) The state x 2 tracking performance; (c) The state x 3 tracking performance; (d) The state x 4 tracking performance.
Figure 5. The tracking performances: (a) The state x 1 tracking performance; (b) The state x 2 tracking performance; (c) The state x 3 tracking performance; (d) The state x 4 tracking performance.
Actuators 15 00138 g005aActuators 15 00138 g005b
Figure 6. The tracking errors: (a) The tracking error z 1 ; (b) The tracking error z 2 ; (c) The tracking error z 3 ; (d) The tracking error z 4 .
Figure 6. The tracking errors: (a) The tracking error z 1 ; (b) The tracking error z 2 ; (c) The tracking error z 3 ; (d) The tracking error z 4 .
Actuators 15 00138 g006
Figure 7. The actor and critic network weights: (a) The actor network weight W ^ a ; (b) The critic network weight W ^ c .
Figure 7. The actor and critic network weights: (a) The actor network weight W ^ a ; (b) The critic network weight W ^ c .
Actuators 15 00138 g007
Figure 8. The utility function r.
Figure 8. The utility function r.
Actuators 15 00138 g008
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Xu, D.; Li, X.; Li, F.; Tian, J. Adaptive Actor–Critic Optimal Tracking Control for a Class of High-Order Nonlinear Systems with Partially Unknown Dynamics. Actuators 2026, 15, 138. https://doi.org/10.3390/act15030138

AMA Style

Xu D, Li X, Li F, Tian J. Adaptive Actor–Critic Optimal Tracking Control for a Class of High-Order Nonlinear Systems with Partially Unknown Dynamics. Actuators. 2026; 15(3):138. https://doi.org/10.3390/act15030138

Chicago/Turabian Style

Xu, Dengguo, Xinsuo Li, Fapeng Li, and Jingbei Tian. 2026. "Adaptive Actor–Critic Optimal Tracking Control for a Class of High-Order Nonlinear Systems with Partially Unknown Dynamics" Actuators 15, no. 3: 138. https://doi.org/10.3390/act15030138

APA Style

Xu, D., Li, X., Li, F., & Tian, J. (2026). Adaptive Actor–Critic Optimal Tracking Control for a Class of High-Order Nonlinear Systems with Partially Unknown Dynamics. Actuators, 15(3), 138. https://doi.org/10.3390/act15030138

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop