Next Article in Journal
Analysis of Dimension Dependence in Quasi-Vertical GaN Schottky Barrier Diodes
Previous Article in Journal
PrivFuzz: Privacy-Preserving Distributed Fuzzing for CPS-Facing Parsing Components on Untrusted Clients
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Zero-Sum Game-Based Practical Predefined-Time Reinforcement Learning for Robust Tracking Control of Highly Flexible Aircraft

College of Intelligence Science and Technology, National University of Defense Technology, Changsha 410000, China
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(17), 3838; https://doi.org/10.3390/electronics15173838
Submission received: 15 July 2026 / Revised: 19 August 2026 / Accepted: 22 August 2026 / Published: 26 August 2026
(This article belongs to the Section Systems & Control Engineering)

Abstract

This paper develops a zero-sum game-based practical predefined-time reinforcement learning method for robust tracking control of highly flexible aircraft. First, the disturbed tracking problem of highly flexible aircraft is transformed into a min–max optimal control problem, where the controller minimizes the infinite-horizon performance index and the adversarial disturbance maximizes it. Then, a critic neural network is used to approximate the solution of the Hamilton–Jacobi–Isaacs equation online, and a fractional-power critic-update law is constructed so that the critic weight estimation error converges to a bounded neighborhood within a user-predefined time. The boundedness of closed-loop and practical predefined-time convergence of critic weight estimation error are rigorously proven. Several simulations are conducted to verify the advancement, effectiveness and robustness of this control algorithm. Overall, the results demonstrate that the proposed method provides an effective practical predefined-time learning framework for robust tracking control of highly flexible aircraft.

1. Introduction

Highly flexible aircraft (HFA) serve as a representative benchmark for advanced flight control due to their lightweight high-aspect-ratio wings and strong aeroelastic coupling. Nonlinear aeroelasticity and flight dynamics must be treated together for HFA, since flexible deformation changes both the structural response and rigid-body motion [1]. Large elastic deformation in flight induces strong rigid–flexible coupling in HFA, resulting in pronounced nonlinearities and modeling uncertainties. These characteristics make controller design for HFA particularly challenging [2].
A variety of nonlinear control methods have been developed for highly flexible aircraft. Gregory used modified nonlinear dynamic inversion to improve flexible-aircraft control, but this type of approach still relies on a carefully constructed dynamic model [3]. Patil and Hodges studied output-feedback control for nonlinear aeroelastic response, showing that feedback design can suppress flexible response when the aeroelastic model is available [4]. Shearer and Cesnik proposed a trajectory-control architecture for very flexible aircraft, and Raghavan and Patil extended trajectory control ideas to high-aspect-ratio flying wings [5,6]. Gust and load-alleviation studies then introduced robust control, model-predictive control, and LQG-based designs to improve performance under atmospheric excitation [7]. Cook et al. emphasized robust gust alleviation and stabilization for highly flexible aircraft, while Haghighat et al. used model-predictive control for gust-load alleviation [8,9]. Liu et al. showed that LQG-based model-predictive control can be effective for gust-load alleviation in flexible aircraft [10]. However, Qu and Annaswamy pointed out that trim drift, unmeasured flexible modes, and actuator anomalies can exceed the robustness margin of conventional linear output-feedback designs, motivating adaptive output-feedback control with closed-loop reference models [11].
Reinforcement learning (RL) and approximate dynamic programming (ADP) provide another route because they learn value functions and policies from system data rather than requiring a completely known drift model. Lewis and Vrabie connected reinforcement learning and ADP to feedback control, and Lewis et al. later summarized how natural decision methods can be used to design optimal adaptive controllers [12,13]. Kiumarsi et al. reviewed optimal and autonomous control using reinforcement learning, showing the breadth of ADP/RL in modern feedback-control design [14]. For continuous-time systems, Vrabie et al. introduced policy-iteration-based adaptive optimal control, and Jiang and Jiang developed computational adaptive optimal control for systems with completely unknown dynamics [15,16]. Zhu et al. further showed that integral reinforcement learning can be used for suboptimal output-feedback control of partially unknown linear continuous-time systems [17]. Then, Peng and Ma used online integral RL to stabilize an uncertain HFA using full-state and output-feedback information, reducing the dependence on exact drift dynamics [2]. Nevertheless, RL controllers for HFA still face two important difficulties: the learning convergence time is not explicitly assigned by the designer, and external disturbances are not always treated as worst-case strategic inputs during the learning process.
Zero-sum differential games provide a natural framework for robust RL because the controller minimizes the performance index while the disturbance attempts to maximize it. Abu-Khalaf et al. used neurodynamic programming and zero-sum games to handle constrained control, which established an important link between Hamiltonian function equations and robust learning-based control [18]. An efficient adaptive synchronization criteria was developed in [19] for fractional-order fuzzy neural networks with uncertain parameters, providing a useful perspective on adaptive control of uncertain nonlinear systems. Vamvoudakis and Lewis then developed an online solution for nonlinear two-player zero-sum games using synchronous policy iteration [20]. More recently, Wang et al. combined safe reinforcement learning, fixed-time stability, and zero-sum games to address nonlinear systems with external disturbances and obstacle-avoidance awareness [21]. A novel ESO-enhanced actor–critic RL method for robust marine-vessel tracking was developed in [22]. However, fixed-time bounds can still be conservative, and the designer may not be able to set the actual convergence time directly as a mission parameter.
To directly set the convergence time, Yuan et al. developed neural adaptive fixed-time control for nonlinear systems with full-state constraints, illustrating the value of neural approximation in time-constrained control [23]. An efficient fixed-time event-triggered control was proposed for distributed underwater vehicles in [24]. Wang, Cao, and Liu then proposed adaptive fuzzy control with both predefined time and predefined accuracy, emphasizing that convergence time and convergence accuracy can be assigned in advance for engineering systems [25]. A new standard was developed in [26] to measure the lowest convergence rate for leader–follower consensus under communication delays. Efficient cooperative consensus tracking control was investigated in [27] for multi-agent systems on cooperation–competition networks with asynchronous information exchange. Ni and Shi developed predefined-time adaptive neural control with output constraints, and Liu et al. studied predefined-time backstepping for strict-feedback nonlinear systems [28,29]. Related predefined-time stability results further support the use of user-specified time parameters in nonlinear systems [30]. However, a zero-sum game-based predefined-time reinforcement learning method has not been designed for highly flexible aircraft.
Motivated by the above discussions, this work develops a zero-sum game-based practical predefined-time reinforcement learning method for robust tracking control of highly flexible aircraft. The main contributions are summarized as follows:
(1) Compared with existing RL-based control methods for highly flexible aircraft [2], where external disturbances are treated as exogenous signals, a zero-sum differential-game architecture is constructed to explicitly characterize the worst-case interaction between the controller and disturbance, thereby converting the robust tracking problem into a min–max optimal control problem.
(2) Compared with conventional zero-sum ADP/RL methods [18,20] and the recent fixed-time zero-sum RL method [21], the critic weight-estimation error is guaranteed to converge to a bounded neighborhood within the predefined time.
(3) In contrast with predefined-time adaptive neural/fuzzy and backstepping control methods [25,28,29], the proposed predefined-time mechanism is incorporated into the critic learning process of the zero-sum game. The practical predefined-time convergence of the critic weight-estimation error and uniform ultimate boundedness of the overall closed-loop tracking system are rigorously established.

2. Problem Statement

2.1. Preliminaries

Consider the following autonomous nonlinear system:
x ˙ ( t ) = f ( x ( t ) ) , x ( 0 ) = x 0 , t 0 ,
where x ( t ) R n is the state vector, x 0 is the initial state, t is the time, and the nonlinear function expressed as f : R n R n is continuous and satisfies f ( 0 ) = 0 .
Definition 1.
Let T ( x 0 ) be the physical time required for the state starting from x 0 to reach the origin. If sup x 0 R n T ( x 0 ) T c holds with a given time constant ( T c > 0 ), then the origin is globally predefined-time stable.
Lemma 1.
For the above nonlinear system, suppose that there exists a continuously differentiable, positive-definite, and radially unbounded Lyapunov function ( V : R n R 0 ), and suppose that for a given time constant ( T c > 0 ) and a parameter ( η ( 0 , 1 ) ), the derivative of V along the system trajectory satisfies [25]
V ˙ π η T c V ( x ( t ) ) 1 η / 2 + V ( x ( t ) ) 1 + η / 2 + b .
Then, the origin is globally practical predefined-time, the convergence-time upper bound is determined by the physical constant ( 2 T c ), and V converges to the region expressed as V η b T c π .
Lemma 2.
For any real sequence ( { y j } j = 1 m R ) and any η ( 0 , 1 ) , the following inequalities hold [23]:
j = 1 m y j 2 2 η 2 j = 1 m | y j | 2 η ,
m η / 2 j = 1 m y j 2 2 + η 2 j = 1 m | y j | 2 + η .

2.2. Dynamics of Highly Flexible Aircraft

The aircraft under consideration is shown in Figure 1 [31]. It comprises three identical rigid wing bodies connected through hinges. Each rigid panel includes a propeller for thrust production, an aileron along the aft of the main wing, and an elevator attached at the end of the boom. In comparison with conventional flexible aerial vehicles, elastic modal behaviors exhibited by HFA are simplified and represented by dihedral angle parameters [31]. The dynamic modeling methodology adopted for HFA aligns fully with that utilized for standard rigid-body aircraft platforms. Concentrated substantial structural deflection occurs at the hinge interface between dual wing panels, rendering the model capable of reproducing flexible structural deformation responses.
The longitudinal control-oriented nonlinear model is written as [32]
V ˙ = T cos α D m g sin γ , α ˙ = T sin α + L m V + q + g cos γ V , h ˙ = V sin γ , θ ˙ = q , q ˙ = M 2 c 2 sin ( η ) cos ( η ) η ˙ q c 1 + c 2 sin 2 ( η ) , η ¨ = H κ c η ˙ κ k η + d 1 d 2 d 3 ,
where V is velocity, γ = θ α is the flight-path angle, h is the altitude, α is the angle of attack, θ is the pitch angle, q is the pitch rate, η is the dihedral angle describing the elastic model, T is total thrust, D is drag, L is lift, and M and H are the total moment and angular moment related to control inputs. The coefficients are
c 1 = 3 I y y , c 2 = 2 I z z 2 I y y + m s 2 6 , d 1 = s 2 m ( V ˙ sin α + V cos α α ˙ ) cos η V sin α sin η η ˙ 2 s 3 cos η sin η η ˙ 2 , d 2 = I y y I z z m s 2 12 sin η cos η q 2 m s 2 cos η V cos α q , d 3 = I x 3 x 3 + m s 2 4 + s 2 6 cos 2 η .
The compact nonlinear HFA model is
X ˙ = f ( X , U ) , X = [ V , α , h , θ , q , η , η ˙ ] T ,
with the five-input vector expressed as
U = [ δ t , δ a , c , δ a , o , δ e , c , δ e , o ] T .
After linearization around a trim point ( X 0 , U 0 ) satisfying f ( X 0 , U 0 ) = 0 , the local model is
x ˙ = A x + B u ,
where x = X X 0 and u = U U 0 . In the simulation of this paper, the altitude deviation is not controlled explicitly, and the six-state vector is
x ( t ) = [ V ( t ) , α ( t ) , θ ( t ) , q ( t ) , η ( t ) , η ˙ ( t ) ] T R 6 ,
with the five-input vector ( u ( t ) R 5 ) inherited from U .

2.3. Disturbed Tracking-Error Dynamics

Consider the following linear system subject to unknown external disturbance:
x ˙ ( t ) = A x ( t ) + B u ( t ) + d ( t ) ,
where u ( t ) R 5 is the control input vector, d ( t ) R 6 is the unknown external bounded disturbance, A R n × n is the system matrix, and B R n × m is the input coupling matrix.
The error dynamics model is
z ˙ ( t ) = F ( t ) + B u ( t ) + d ( t ) ,
where
z ( t ) = x ( t ) x d ( t ) , F ( t ) = A x ( t ) x ˙ d ( t ) ,
x d ( t ) is the the desired reference signal, and x ˙ d ( t ) denotes the derivative of the reference signal. In this paper, · denotes the vector 2-norm or the induced matrix 2-norm.
The control objective is to design an RL algorithm and a parameter-update strategy such that: (i) the neural-network parameter estimation error vector ( W ˜ ( t ) ) converges to a bounded neighborhood within the strict predefined time ( T w ) and (ii) under the worst external disturbance d ( z ( t ) ) , the tracking error can achieve bounded convergence.

3. Zero-Sum Game-Based Practical Predefined-Time RL Controller Design

3.1. Zero-Sum-Game Optimal Control

For the error system (12), we define the nonlinear utility function ( J : R n × R m × R n R ) to evaluate the dynamic-system performance:
J ( z ( t ) , u ( t ) , d ( t ) ) = z T ( t ) Q z ( t ) + u T ( t ) R u ( t ) δ 2 d T ( t ) d ( t ) ,
where Q R n × n is the tracking-error weight, R R m × m is the control-input weight, and δ > 0 is the disturbance weight.
The optimal value function ( V c ( t ) ) is defined as the min–max functional of the utility index over the infinite horizon:
V c ( t ) = min u max d 0 J ( z ( t ) , u ( t ) , d ( t ) ) d t .
For V c ( t ) , the differential-game Hamiltonian ( H ( z ( t ) , V c ( z ( t ) ) , u ( t ) , d ( t ) ) ) is constructed as
H = z T ( t ) Q z ( t ) + u T ( t ) R u ( t ) δ 2 d T ( t ) d ( t ) + V c T ( z ( t ) ) F ( t ) + B u ( t ) + d ( t ) .
Taking the extremum derivatives with respect to the control and disturbance variables gives
H u ( t ) = 2 R u ( t ) + B T V c ( t ) = 0 ,
H d ( t ) = 2 δ 2 d ( t ) + V c ( z ( t ) ) = 0 .
Therefore, the optimal feedback Nash strategies are
u ( t ) = 1 2 R 1 B T V c ( t ) , d ( t ) = 1 2 δ 2 V c ( t ) .
According to the extremum conditions, the optimal feedback strategies satisfy the steady-state HJI partial differential equation, i.e.,
H ( z ( t ) , V c ( t ) , u ( z ( t ) ) , d ( z ( t ) ) ) = 0 .

3.2. Critic Neural-Network Approximation

A critic neural network is adopted to parameterize and reconstruct the optimal value function ( V c ( z ( t ) ) ):
V c ( z ( t ) ) = W T ψ ( z ( t ) ) + ε ( z ( t ) ) .
Taking the gradient with respect to the state variable yields
V c ( z ( t ) ) = ψ T ( z ( t ) ) W + ε ( z ( t ) ) .
Here, W R M is the ideal network weight vector, M is the number of hidden nodes, and ψ : R n R M is the activation basis vector. The approximation residual satisfies
ε ( z ( t ) ) ε c , ε ( z ( t ) ) ε d c .
Here, ε c > 0 and ε d c > 0 denote the upper bounds of the approximation residual ( ε ( z ( t ) ) ) and its gradient ( ε ( z ( t ) ) ), respectively. Thus, the optimal control and disturbance strategies can be written as
u ( t ) = 1 2 R 1 B T ψ T ( z ( t ) ) W + ε ( z ( t ) ) ,
d ( t ) = 1 2 δ 2 ψ T ( z ( t ) ) W + ε ( z ( t ) ) .
The critic neural network approximates the optimal value function as
V ^ c ( z , t ) = W ^ T ( t ) ψ ( z ) ,
whose state gradient is directly obtained as
V ^ c ( z , t ) = ψ T ( z ) W ^ ( t ) .
Substituting this learned value-function gradient into the analytical saddle-point mappings yields the implementable control policy:
u ^ c ( z ( t ) ) = 1 2 R 1 B T ψ T ( z ( t ) W ^ ( t ) ) , d ^ c ( z ( t ) ) = 1 2 δ 2 ψ T ( z ( t ) W ^ ( t ) ) .
σ ( t ) R M is defined as
σ ( t ) = ψ ( z ( t ) ) F ( t ) + B u ^ ( t ) + d ^ ( z ( t ) ) .
Then, the Bellman residual equation ( e ( t ) R ) is
e ( t ) = W ^ T ( t ) σ ( t ) + z T ( t ) Q z ( t ) + u ^ T ( t ) R u ^ ( t ) δ 2 d ^ T ( t ) d ^ ( t ) .
The online parameter-estimation error os defined as
W ˜ ( t ) = W ^ ( t ) W .
By algebraic subtraction, the residual equation is converted into the explicit mapping with respect to the parameter error:
e ( t ) = W ˜ T ( t ) σ ( t ) + ε H ( t ) ,
where
ε H ( t ) = W T ( t ) σ ( t ) + z T ( t ) Q z ( t ) + u ^ T ( t ) R u ^ ( t ) δ 2 d ^ T ( t ) d ^ ( t )
represents the residual caused by the limited approximation capability of the network, bounded by
sup z ( t ) Z | ε H ( t ) | ε H d .

3.3. Practical Predefined-Time Adaptive Critic-Update Law Design

To make the sufficient-excitation assumption valid in the learning law, the critic update is constructed by using both the current Hamiltonian estimation error and a finite stack of recorded data. Let
D h ( t ) = z ( t j ) , u ^ ( t j ) , d ^ ( t j ) , σ ( t j ) j = 1 N h
denote the recorded dataset collected before the current time (t), where N h M . For each stored sample ( t j ), the historical Bellman residual is evaluated by the current critic weight ( W ^ ( t ) ) as
e ( t j , t ) = W ^ T ( t ) σ ( t j ) + z T ( t j ) Q z ( t j ) + u ^ T ( t j ) R u ^ ( t j ) δ 2 d ^ T ( t j ) d ^ ( t j ) , j = 1 , 2 , , N h ,
that is, the state, policy and regressor information at t j is stored, while the unknown ideal parameter is still replaced by the current online estimate ( W ^ ( t ) ). The normalized current and recorded regressors are defined as
σ ¯ ( t ) = σ ( t ) 1 + σ T ( t ) σ ( t ) , e ¯ ( t ) = e ( t ) 1 + σ T ( t ) σ ( t ) , σ ¯ ( t j ) = σ ( t j ) 1 + σ T ( t j ) σ ( t j ) , e ¯ ( t j , t ) = e ( t j , t ) 1 + σ T ( t j ) σ ( t j ) .
Then, the current and historical normalized Bellman residuals can be written as
e ¯ ( t ) = σ ¯ T ( t ) W ˜ ( t ) + ε ¯ H ( t ) ,
e ¯ ( t j , t ) = σ ¯ T ( t j ) W ˜ ( t ) + ε ¯ H ( t j ) , j = 1 , 2 , , N h ,
where ε ¯ H ( t ) and ε ¯ H ( t j ) are the normalized approximation residuals. On the compact set expressed as Z , there exists a positive constant ( ε ¯ H d ) such that
max | ε ¯ H ( t ) | , | ε ¯ H ( t 1 ) | , , | ε ¯ H ( t N h ) | ε ¯ H d .
Let the recorded-data matrix be
Π = σ ¯ ( t 1 ) , σ ¯ ( t 2 ) , , σ ¯ ( t N h ) R M × N h ,
and define
Σ = Π Π T = j = 1 N h σ ¯ ( t j ) σ ¯ T ( t j ) .
Assumption 1.
The recorded historical dataset satisfies the sufficient-excitation condition—namely, the data stack ( { σ ¯ ( t j ) } j = 1 N h ) is N h -sufficiently rich, and
rank ( Π ) = M ,
or, equivalently, there exists a constant ( λ 0 > 0 ) such that
λ min ( Σ ) λ 0 > 0 .
Remark 1.
Assumption 1 imposes an excitation condition on the historical data, i.e., rank ( Π ) = M , and persistent excitation is not required. Without a sufficient-excitation condition, the practical predefined-time convergence of the critic weight-estimation error cannot be guaranteed. Once rank ( Π ) = M , persistent excitation of σ ¯ ( t ) is not required after the data stack has been established.
Before online critic learning starts, a finite historical dataset ( D h ) is collected in a preliminary data-collection stage. The stored samples are selected such that the normalized regressor matrix ( Π ) satisfies rank ( Π ) = M . Therefore, the online critic learning starts with Assumption 1 already satisfied. The practical predefined learning-time bound ( T w ) is consequently measured from the beginning of the online learning process.
The critic update is obtained from the fractional-power minimization of both current and historical Bellman residuals. Specifically, we define
E ( W ^ ( t ) ) = 1 2 η | e ¯ ( t ) | 2 η + j = 1 N h | e ¯ ( t j , t ) | 2 η + 1 2 + η | e ¯ ( t ) | 2 + η + j = 1 N h | e ¯ ( t j , t ) | 2 + η ,
where η ( 0 , 1 ) . To guarantee that the learning convergence time is strictly constrained within T w , the following current-and-history network parameter update law is designed:
W ^ ˙ ( t ) = ( α 1 + 1 ) Γ σ ¯ ( t ) sig 1 η e ¯ ( t ) ( α 1 + 1 ) Γ j = 1 N h σ ¯ ( t j ) sig 1 η e ¯ ( t j , t ) ( α 2 + 1 ) Γ σ ¯ ( t ) sig 1 + η e ¯ ( t ) ( α 2 + 1 ) Γ j = 1 N h σ ¯ ( t j ) sig 1 + η e ¯ ( t j , t ) .
The sign-power operator is defined by
sig p ( x ) = | x | p sgn ( x ) .
The gain matrix ( Γ R M × M ) is positive-definite, symmetric, and diagonal, satisfying Γ = Γ T > 0 . We introduce an analytical constant ( θ ( 0 , 1 ) ) and define
c 1 = 2 λ min ( Γ ) λ 0 1 η / 2 , c 2 = N h η / 2 2 λ min ( Γ ) λ 0 1 + η / 2 .
The learning factors ( α 1 and α 2 ) are selected as
α 1 = π θ T w η c 1 , α 2 = π θ T w η c 2 ,
where T w = 0.5 T w is the auxiliary learning-time constant, so the predefined upper bound for the critic-weight learning process is T w = 2 T w .
The structural diagram of the algorithm proposed in this article is shown in Figure 2.

4. Stability Analysis

The stability analysis is developed in two stages. First, the learning dynamics of the critic network are analyzed to establish practical predefined-time convergence of the critic weight-estimation error. This result further provides boundedness of the learned value-function approximation and the corresponding policy reconstruction errors. Second, the critic-learning dynamics and the physical tracking dynamics are incorporated into a composite Lyapunov analysis to establish uniform ultimate boundedness of the overall closed-loop system. Therefore, Theorem 1 provides the learning-layer guarantee required for the closed-loop stability analysis in Theorem 2.
Theorem 1.
Under Assumption 1, if the current-and-history fractional-power network update law (46) is applied, the practical predefined-time stability of the critic approximation error can be guaranteed.
Proof. 
Construct the Lyapunov functional reflecting the critic error energy:
V w ( W ˜ ( t ) ) = 1 2 W ˜ T ( t ) Γ 1 W ˜ ( t ) .
Taking its derivative and substituting the current-and-history update law (46) gives
V ˙ w ( t ) = ( α 1 + 1 ) σ ¯ T ( t ) W ˜ ( t ) sig 1 η e ¯ ( t ) ( α 1 + 1 ) j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) sig 1 η e ¯ ( t j , t ) ( α 2 + 1 ) σ ¯ T ( t ) W ˜ ( t ) sig 1 + η e ¯ ( t ) ( α 2 + 1 ) j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) sig 1 + η e ¯ ( t j , t ) .
Using (38) and (39), Equation (51) can be rewritten as
V ˙ w ( t ) = ( α 1 + 1 ) σ ¯ T ( t ) W ˜ ( t ) sig 1 η σ ¯ T ( t ) W ˜ ( t ) + ε ¯ H ( t ) ( α 1 + 1 ) j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) sig 1 η σ ¯ T ( t j ) W ˜ ( t ) + ε ¯ H ( t j ) ( α 2 + 1 ) σ ¯ T ( t ) W ˜ ( t ) sig 1 + η σ ¯ T ( t ) W ˜ ( t ) + ε ¯ H ( t ) ( α 2 + 1 ) j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) sig 1 + η σ ¯ T ( t j ) W ˜ ( t ) + ε ¯ H ( t j ) .
For the ideal approximation case—namely, ε ¯ H ( t ) = 0 and ε ¯ H ( t j ) = 0 , j = 1 , , N h —one obtains
V ˙ w ( t ) = ( α 1 + 1 ) σ ¯ T ( t ) W ˜ ( t ) 2 η ( α 2 + 1 ) σ ¯ T ( t ) W ˜ ( t ) 2 + η ( α 1 + 1 ) j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 η ( α 2 + 1 ) j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 + η .
Since the current term is non-positive, it can be retained in the learning law but discarded in the upper-bound estimate. Hence,
V ˙ w ( t ) α 1 j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 η α 2 j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 + η .
According to Lemmas 2 and 3, one has
j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 η j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 1 η / 2 ,
j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 + η N h η / 2 j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 1 + η / 2 .
According to the definition of Π and Assumption 1,
j = 1 N h σ ¯ T ( t j ) W ˜ ( t ) 2 = W ˜ T ( t ) Π Π T W ˜ ( t ) = W ˜ T ( t ) Σ W ˜ ( t ) λ 0 W ˜ ( t ) 2 .
Moreover, since
V w ( t ) = 1 2 W ˜ T ( t ) Γ 1 W ˜ ( t ) 1 2 λ min ( Γ ) W ˜ ( t ) 2 ,
one obtains
W ˜ ( t ) 2 2 λ min ( Γ ) V w ( t ) .
Substituting (55)–(59) into (54) gives
V ˙ w ( t ) α 1 c 1 V w 1 η / 2 ( t ) α 2 c 2 V w 1 + η / 2 ( t ) .
According to the definitions of α 1 and α 2 in (49),
V ˙ w ( t ) π θ T w η V w ( t ) 1 η / 2 + V w ( t ) 1 + η / 2 π T w η V w ( t ) 1 η / 2 + V w ( t ) 1 + η / 2 .
According to Lemma 1, the ideal critic weight-estimation error converges to zero within the predefined upper bound ( T w = 2 T w ).
For the nonideal case, let
x 0 ( t ) = σ ¯ T ( t ) W ˜ ( t ) , x j ( t ) = σ ¯ T ( t j ) W ˜ ( t ) , j = 1 , , N h ,
and
r 0 ( t ) = ε ¯ H ( t ) , r j ( t ) = ε ¯ H ( t j ) , j = 1 , , N h .
Outside the residual-dominated region, i.e., when | x i ( t ) | | r i ( t ) | , i = 0 , 1 , , N h , it holds that
sgn ( x i ( t ) ) = sgn ( x i ( t ) + r i ( t ) ) , i = 0 , 1 , , N h .
Therefore, Equation (52) yields
V ˙ w ( t ) = ( α 1 + 1 ) i = 0 N h | x i ( t ) | x i ( t ) + r i ( t ) 1 η ( α 2 + 1 ) i = 0 N h | x i ( t ) | x i ( t ) + r i ( t ) 1 + η .
Using the absolute-value inequalities, i.e.,
| x i ( t ) | 1 η | x i ( t ) + r i ( t ) | 1 η + | r i ( t ) | 1 η ,
| x i ( t ) | 1 + η 2 η | x i ( t ) + r i ( t ) | 1 + η + | r i ( t ) | 1 + η ,
Equation (65) is upper-bounded by
V ˙ w ( t ) ( α 1 + 1 ) i = 0 N h | x i ( t ) | 2 η 2 η ( α 2 + 1 ) i = 0 N h | x i ( t ) | 2 + η + ( α 1 + 1 ) i = 0 N h | x i ( t ) | | r i ( t ) | 1 η + ( α 2 + 1 ) i = 0 N h | x i ( t ) | | r i ( t ) | 1 + η .
Because ε ¯ H ( t ) and ε ¯ H ( t j ) are bounded and the normalized regressors are bounded, there exists a positive constant ( σ ¯ d ) satisfying
max σ ¯ ( t ) , σ ¯ ( t 1 ) , , σ ¯ ( t N h ) σ ¯ d .
Then, the residual-induced terms in (68) satisfy
( α 1 + 1 ) i = 0 N h | x i ( t ) | | r i ( t ) | 1 η + ( α 2 + 1 ) i = 0 N h | x i ( t ) | | r i ( t ) | 1 + η ( N h + 1 ) σ ¯ d ( α 1 + 1 ) ε ¯ H d 1 η + ( α 2 + 1 ) ε ¯ H d 1 + η W ˜ ( t ) .
Let
ϱ w = ( N h + 1 ) σ ¯ d ( α 1 + 1 ) ε ¯ H d 1 η + ( α 2 + 1 ) ε ¯ H d 1 + η .
Dropping the non-positive current term and using the same historical-data excitation argument as in the ideal case gives
V ˙ w ( t ) α 1 λ 0 1 η / 2 W ˜ ( t ) 2 η α 2 N h η / 2 λ 0 1 + η / 2 W ˜ ( t ) 2 + η + ϱ w W ˜ ( t ) .
The compact set is defined as
Ω w = W ˜ ( t ) : W ˜ ( t ) ρ w ,
where
ρ w = max ϱ w ( 1 θ ) α 1 λ 0 1 η / 2 1 1 η , ϱ w ( 1 θ ) α 2 N h η / 2 λ 0 1 + η / 2 1 1 + η .
For W ˜ ( t ) Ω w , inequality (72) implies
V ˙ w ( t ) θ α 1 λ 0 1 η / 2 W ˜ ( t ) 2 η θ α 2 N h η / 2 λ 0 1 + η / 2 W ˜ ( t ) 2 + η .
Using (59) further yields
V ˙ w ( t ) θ α 1 c 1 V w 1 η / 2 ( t ) θ α 2 c 2 V w 1 + η / 2 ( t ) .
According to (49), one has
V ˙ w ( t ) π T w η V w ( t ) 1 η / 2 + V w ( t ) 1 + η / 2 , W ˜ ( t ) Ω w .
Thus, according to Lemma 1, W ˜ ( t ) converges to the compact set ( Ω w ) within the predefined upper bound ( T w = 2 T w ) and remains therein thereafter. This completes the proof. □
Remark 2.
Theorem 1 provides a practical predefined-time learning guarantee for the critic network. In the ideal approximation case, the critic weight-estimation error converges to zero within the predefined upper bound ( T w ). In the nonideal case, critic error converges to the compact set ( Ω w ) within the prescribed learning time. Therefore, T w characterizes the learning-time requirement, whereas the size of Ω w reflects the effects of approximation accuracy.
Based on Theorem 1, the critic weight-estimation error converges to a bounded region within the predefined time, which further ensures bounded errors in the learned value-function gradient and the reconstructed control policy. This learning-layer result provides the basis for analysis of the overall closed-loop stability in Theorem 2.
Theorem 2.
For the disturbed physical system (12), apply the RL control policy in (28) and the current-and-history network parameter update law (46). If Assumption 1 holds, then the closed-loop system is bounded.
Proof. 
The Lyapunov function is defined as
V ( t ) = V c ( z ( t ) ) + V w ( W ˜ ( t ) ) .
Taking the derivative yields
V ˙ ( t ) = V ˙ c ( t ) + V ˙ w ( t ) .
Here,
V ˙ c ( t ) = V c T ( t ) F ( t ) + B u ^ ( t ) + d ^ ( t ) .
By further mathematical manipulation,
V ˙ c ( t ) = V c T ( t ) F ( t ) + B u ( t ) + d ( t ) + V c T ( t ) B ( u ^ ( t ) u ( t ) ) + V c T ( t ) ( d ^ ( t ) d ( t ) ) .
According to the optimality condition,
V c T ( z ( t ) ) F ( t ) + B u ( t ) + d ( t ) = z T ( t ) Q z ( t ) u T ( t ) R u ( t ) + δ 2 d T ( t ) d ( t ) .
Moreover, according to the extremum conditions of the optimal control and disturbance laws, i.e.,
V c T ( t ) B = 2 u T ( t ) R , V c T ( t ) = 2 δ 2 d T ( t ) ,
one has
V ˙ c ( t ) = z T ( t ) Q z ( t ) u T ( t ) R u ( t ) + δ 2 d T ( t ) d ( t ) 2 u T ( t ) R ( u ^ ( t ) u ( t ) ) + 2 δ 2 d T ( t ) ( d ^ ( t ) d ( t ) ) .
Applying Young’s inequality to the cross terms gives
2 u T ( t ) R ( u ^ ( t ) u ( t ) ) u T ( t ) R u ( t ) + ( u ^ ( t ) u ( t ) ) T R ( u ^ ( t ) u ( t ) ) ,
2 δ 2 d T ( t ) ( d ^ ( t ) d ( t ) ) δ 2 d T ( t ) d ( t ) + δ 2 ( d ^ ( t ) d ( t ) ) T ( d ^ ( t ) d ( t ) ) .
Substituting the bounds into V ˙ c ( t ) and eliminating the u T ( t ) R u ( t ) term yields
V ˙ c ( t ) z T ( t ) Q z ( t ) + 2 δ 2 d T ( t ) d ( t ) + ( u ^ ( t ) u ( t ) ) T R ( u ^ ( t ) u ( t ) ) + δ 2 ( d ^ ( t ) d ( t ) ) T ( d ^ ( t ) d ( t ) ) .
Combining the critic neural-network approximation property, the errors between the practical and optimal policies are expressed by the parameter reconstruction error as
u ^ ( t ) u ( t ) = 1 2 R 1 B T W ˜ T ( t ) ψ ( z ( t ) ) ε T ( z ( t ) ) ,
d ^ ( t ) d ( t ) = 1 2 δ 2 W ˜ T ( t ) ψ ( z ( t ) ) ε T ( z ( t ) ) .
Theorem 1 has already shown that the network weight-estimation error ( W ˜ ( t ) ) converges to a bounded region within the predefined time. On a compact set, ψ ( z ( t ) ) , ε ( z ( t ) ) , and d ( t ) are bounded. Therefore, we define a positive constant ( D M ) such that the residual terms induced by the network approximation error and external bounded disturbance satisfy
2 δ 2 d T ( t ) d ( t ) + ( u ^ ( t ) u ( t ) ) T R ( u ^ ( t ) u ( t ) ) + δ 2 ( d ^ ( t ) d ( t ) ) T ( d ^ ( t ) d ( t ) ) D M .
The quadratic form ( z T ( t ) Q z ( t ) ) satisfies
V ˙ c ( t ) λ min ( Q ) z ( t ) 2 + D M .
Substituting this result into the derivative of the composite Lyapunov functional and using Theorem 1 gives, for W ˜ ( t ) Ω w ,
V ˙ ( t ) λ min ( Q ) z ( t ) 2 + D M π T w η V w ( t ) 1 η / 2 + V w ( t ) 1 + η / 2 .
When W ˜ ( t ) Ω w , the critic error is already bounded. Therefore, when
z ( t ) > D M λ min ( Q ) ,
one has V ˙ ( t ) < 0 . According to Lyapunov stability theory, the coupled closed-loop system is uniformly ultimately bounded. This completes the proof. □

5. Simulation

To comprehensively validate the efficacy of the proposed algorithm, this article conducts a series of experiments designed to evaluate its generality, performance, and robustness. Specifically, the experimental validation consists of three parts. Firstly, different trim points are selected to verify the generality and broad applicability of the proposed algorithm. Secondly, the proposed algorithm is compared with traditional Integral Reinforcement Learning (IRL) algorithms to demonstrate its performance superiority. Finally, model uncertainty and measurement noise are introduced into the experimental setting to verify the robustness of the proposed algorithm.

5.1. Different Trim Points

The applicable range of this linear model is when the dihedral angle is no more than 15 degrees [31]. In order to verify that this method has a good generalization ability for this model, the paper selects two equilibrium points with dihedral angles of 15 degrees and 5 degrees, which are relatively large and small, respectively. The HFA is trimmed under the following conditions: V = 68 ft/s, h = 40,000 ft, α = 2 . 8 , θ = 2 . 8 , η = 15 or 5 , and η ˙ = 0 .
The state-space matrix and input matrix ( ( A 1 , B 1 ) and ( A 2 , B 2 ) ) are shown as follows:
A 1 = 0.0428 4.1764 32.2 0.0010 0.2557 0.0292 0.0139 8.7716 0 1.0045 0.1618 0.1564 0 0 0 1.0000 0 0 0 71.3347 0 0.0052 0.8590 8.2122 0 0 0 0 0 1.0000 0.0015 0.3995 0 1.9228 0.0308 7.5587 ,
B 1 = 0.0357 2.6729 6.4431 0 0 0 0.9144 1.7666 0.1795 0.3469 0 0 0 0 0 0 2.2479 3.9797 16.3906 33.4039 0 0 0 0 0 0 0.8925 1.0479 0.1749 0.2053 ,
A 2 = 0.0417 4.1722 32.2 0.0008 0.0835 0.1788 0.0139 9.1315 0 1.0009 0.0541 0.0191 0 0 0 1.0000 0 0 0 446.4040 0 0.1753 2.2393 38.4633 0 0 0 0 0 1 0.0005 0.0480 0 1.9991 0.1384 7.4605 ,
B 2 = 0.0357 2.6729 6.2501 0 0 0 0.9144 1.8219 0.1795 0.3577 0 0 0 0 0 0 12.8245 24.5028 91.5144 192.2739 0 0 0 0 0 0 0.9196 0.9375 0.1802 0.1836 .
The bounded disturbance used for plant propagation is
d ( t ) = 1.1 [ 0.1 sin ( 0.1 t ) , 0.2 sin ( 0.1 t ) , 0.2 cos ( 0.1 t ) , 0.1 sin ( 0.1 t ) , 0.1 cos ( 0.1 t ) , 0.1 sin ( 0.1 t ) ] T .
The zero-sum game weights and critic-learning parameters are
Q = 10 diag ( 100 , 100 , 0.001 , 0.001 , 100 , 0.001 ) , R = 10 diag ( 10 , 40 , 40 , 40 , 5 ) , δ = 200 .
T w in (46) is set as T w = 5 s , and the predefined learning time is T w = 2 T w = 10 s . The fractional parameter is η = 0.8 , the number of basis functions is M = 25 , the critic gain matrix is Γ = 0.01 I M , and the initial critic weight is W ^ ( 0 ) = 100 1 M . The simulation step is 10 4 s, the final time is 20 s, and the initial state is x ( 0 ) = [ 5 , 0.05 , 0 , 0 , 0 , 0 ] T . The reference command and its derivative are
x d ( t ) = [ 7 + 0.01 t + 2 sin ( 0.5 t ) , 0.1 , 0 , 0 , 0.5 , 0 ] T ,
x ˙ d ( t ) = [ 0.01 + cos ( 0.5 t ) , 0 , 0 , 0 , 0 , 0 ] T .
Figure 3, Figure 4 and Figure 5 compare the tracking performance of the two parameter sets for velocity, angle of attack, and dihedral angle. Both cases show a short initial adjustment followed by bounded tracking. The second parameter set contains stronger pitch-rate and dihedral angle coupling, especially in the fourth row of A 2 and in the fourth row of B 2 , but the controller still keeps the main rigid and elastic responses close to their commands. This indicates that the zero-sum RL policy addsa degree of robustness to the considered model-parameter variation.
Figure 6 shows that the five control channels remain smooth. Figure 7, Figure 8, Figure 9 and Figure 10 describe the critic-learning process. The critic weights change rapidly during the early learning interval, then remain bounded. The weight increments decrease after the main adaptation phase, showing that the learning process is concentrated around the prescribed learning window. The Bellman residual is not forced exactly to zero because the basis functions and disturbance estimate are approximate, but it stays bounded in both cases. The weight-update norm also decreases after the initial transient, which supports the practical predefined-time critic convergence stated in Theorem 1.
Overall, the two-case simulations verify three points. First, the zero-sum controller suppresses bounded disturbances. Second, the predefined-time critic update concentrates learning activity within the prescribed time window and keeps the Bellman residual bounded afterward. Third, the same controller parameters work for both HFA parameter sets, which indicates robustness to the considered model-parameter variation.

5.2. Comparison with Traditional Methods

In this section, the proposed algorithm is compared with traditional Integral Reinforcement Learning (IRL) algorithms introduced by Ma [32]. The linearized state matrix and input matrix is ( A 2 , B 2 ) , which is introduced in the previous section. T w is 5   s , and other simulation conditions are the same as the previous section.
Based on the simulation results presented in the figures, a comparative analysis between the proposed method and the conventional nominal Integral Reinforcement Learning (IRL) approach reveals significant performance advantages of the former. The superiority of the proposed method is primarily evidenced by its rapid convergence characteristics and high-precision tracking capabilities.
As illustrated in the state trajectory plots in Figure 11, the proposed controller demonstrates a markedly faster transient response compared to the nominal IRL method. Specifically, for states α and η , the proposed method achieves convergence to the desired reference values within approximately 4 s . In contrast, the nominal IRL method exhibits a sluggish response, failing to reach the steady-state target, even after 20 s of operation. This indicates that the prescribed-time framework effectively overcomes the asymptotic convergence limitations often associated with traditional adaptive or learning-based controllers, guaranteeing system stabilization within a user-defined timeframe independent of initial conditions.
The tracking-error dynamics further corroborate the robustness of the proposed scheme in Figure 12. While the nominal IRL method maintains non-negligible steady-state errors, the proposed method drives tracking errors z 1 , z 2 , and z 5 to zero with high precision. The error trajectories for the proposed method remain virtually flat at zero after the settling time, whereas the nominal IRL shows continuous drift or offset, highlighting the superior disturbance rejection and parameter-estimation accuracy of the proposed algorithm.
Regarding the control input activity in Figure 13, although the proposed method initially utilizes higher control authority to enforce the fast convergence, it settles to a stable equilibrium rapidly within 5 s. Conversely, the nominal IRL method applies significantly smaller control magnitudes but fails to stabilize the system effectively, resulting in prolonged transient periods. This trade-off suggests that the proposed method efficiently utilizes available control energy to achieve superior dynamic performance, whereas the nominal method’s conservative input leads to inadequate system regulation.
In summary, the comparative results validate that the proposed control method offers substantial improvements over the nominal IRL method, particularly in terms of convergence rate, steady-state accuracy, and transient response quality.

5.3. Model Uncertainty

In this section, model uncertainty is introduced to assess the robustness of the algorithm proposed in this paper. The linearized state matrix and input matrix is ( A 2 , B 2 ) , which is introduced in the previous section. The linearized system has eigenvalues with λ 1 = 7.11 , λ 2 = 0.019 , λ 3 , 4 = 0.016 ± 0.64 i , and λ 5 , 6 = 4.82 ± 18.66 i [32]. The system-matched uncertainties are caused by flexible effects of dihedral angle [33].
In the simulation environment, A r e a l replaces A in the calculation of environmental information feedback, when the proposed algorithm has no information about it and still uses A in the calculation of controller, where
A r e a l = A + Δ A
Different Δ A values are settled as follows:
Case I: no model uncertainty, Δ A = 0 ;
Case II: low model uncertainty, Δ A = 0.1 A ;
Case III: high model uncertainty, Δ A = 0.3 A
Case IV: negative model uncertainty, Δ A = 0.2 A .
Other simulation conditions are the same as in the previous section.
Figure 14, Figure 15 and Figure 16 demonstrate that the proposed controller preserves essentially invariant tracking performance across all tested uncertainty levels. The commanded and measured trajectories overlap closely after a short transient of approximately 4–5 s, with negligible steady-state error and nearly identical rise and settling times, irrespective of the sign or magnitude of Δ A .
Under the same uncertainty cases, the control inputs remain bounded, smooth, and free of saturation or sustained oscillations in Figure 17. The transient profiles and peak amplitudes are nearly indistinguishable across Δ A , and no chattering is observed; input rates stay within acceptable limits, indicating benign actuation demands and strong robustness of the control law.
Figure 18, Figure 19, Figure 20 and Figure 21 characterize the learning process. The critic weights converge to bounded constants, the weight-update increments rapidly decay to small values, and the Bellman residual drops by several orders of magnitude and stays near zero. The norm of the weight-update vector exhibits an early peak, then quickly vanishes. Closed-loop performance is preserved, even after learning is frozen at T w , evidencing stable adaptation and insensitivity to modeling errors up to every tested Δ A .

5.4. Measurement Noise

The measurement noiseis represented as
x m ( t ) = x ( t ) + n ( t ) ,
where n ( t ) = 0.02 cos ( 0.1 t ) , 0.005 sin ( 0.5 t ) , 0.005 sin ( 0.1 t ) , 0.005 cos ( 0.5 t ) , 0.005 sin ( 0.1 t ) , 0.005 cos ( 0.5 t ) denotes the measurement-noise vector. The resulting tracking errors and critic-weight responses are given below.
Figure 22 and Figure 23 present the results in the presence of measurement noise. The tracking errors remain bounded and decrease to a small neighborhood of zero after a short transient, while the critic weights exhibit bounded learning responses and gradually approach steady values before T w . Therefore, the introduced measurement noise does not significantly degrade the tracking performance across tested measurement-noise levels.

6. Conclusions

This paper presented a practical predefined-time reinforcement learning method for robust tracking control of highly flexible aircraft. A zero-sum differential-game formulation was used to model the interaction between the tracking controller and the worst-case disturbance. A critic neural network was introduced to approximate the Hamilton–Jacobi–Isaacs equation solution online, and a fractional-power update law was developed to make the critic weight-estimation error converge to a bounded neighborhood within a prescribed learning time. A composite Lyapunov analysis established uniform ultimate boundedness of the coupled closed-loop system. Different trim-point experiments based on two HFA parameter sets demonstrated that the proposed method can achieve bounded tracking under external sinusoidal disturbances. Comparison with traditional IRL methods demonstrated the superiority of the proposed method. Model-uncertainty experiments demonstrated the method proposed in this paper performs well under different model uncertainties. Zero-sum game-based practical predefined-time reinforcement learning control for highly flexible aircraft with actuator saturation and time delay is an important direction for future work.

Author Contributions

Conceptualization, H.Z. and M.W.; methodology, H.Z.; software, H.Z.; validation, H.Z. and Y.Z.; formal analysis, H.Z. and J.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The data presented in this study are available upon request from the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Patil, M.; Hodges, D.; Cesnik, C. Nonlinear aeroelasticity and flight dynamics of high-altitude long-endurance aircraft. J. Aircr. 2001, 38, 88–94. [Google Scholar] [CrossRef] [Scilit]
  2. Peng, C.; Ma, J. Online integral reinforcement learning control for an uncertain highly flexible aircraft using state and output feedback. Aerosp. Sci. Technol. 2021, 109, 106442. [Google Scholar] [CrossRef] [Scilit]
  3. Gregory, I. Modified dynamic inversion to control large flexible aircraft: What’s going on? In Proceedings of the AIAA Guidance, Navigation, and Control Conference and Exhibit, Portland, OR, USA, 9–11 August 1999. [Google Scholar]
  4. Patil, M.; Hodges, D. Output feedback control of the nonlinear aeroelastic response of a slender wing. J. Guid. Control Dyn. 2002, 25, 302–308. [Google Scholar] [CrossRef] [Scilit]
  5. Shearer, C.; Cesnik, C. Trajectory control for very flexible aircraft. J. Guid. Control Dyn. 2008, 31, 340–357. [Google Scholar] [CrossRef] [Scilit]
  6. Raghavan, B.; Patil, M. Flight control for flexible high-aspect-ratio flying wings. J. Guid. Control Dyn. 2010, 33, 64–74. [Google Scholar] [CrossRef] [Scilit]
  7. Dillsaver, M.; Cesnik, C.; Kolmanovsky, I. Trajectory control of very flexible aircraft with gust disturbance. In Proceedings of the AIAA Atmospheric Flight Mechanics Conference, Reston, VA, USA, 19–22 August 2013. [Google Scholar]
  8. Cook, R.; Palacios, R.; Goulart, P. Robust gust alleviation and stabilization of very flexible aircraft. AIAA J. 2013, 51, 330–340. [Google Scholar] [CrossRef] [Scilit]
  9. Haghighat, S.; Liu, H.; Martins, J. Model-predictive gust load alleviation controller for a highly flexible aircraft. J. Guid. Control Dyn. 2012, 35, 1751–1766. [Google Scholar] [CrossRef] [Scilit]
  10. Liu, X.; Sun, Q.; Cooper, J. LQG based model predictive control for gust load alleviation. Aerosp. Sci. Technol. 2017, 71, 499–509. [Google Scholar] [CrossRef] [Scilit]
  11. Qu, Z.; Annaswamy, A. Adaptive output-feedback control with closed-loop reference models for very flexible aircraft. J. Guid. Control Dyn. 2016, 39, 873–888. [Google Scholar] [CrossRef] [Scilit]
  12. Lewis, F.; Vrabie, D. Reinforcement learning and adaptive dynamic programming for feedback control. IEEE Circuits Syst. Mag. 2009, 9, 32–50. [Google Scholar] [CrossRef] [Scilit]
  13. Lewis, F.; Vrabie, D.; Vamvoudakis, K. Reinforcement learning and feedback control: Using natural decision methods to design optimal adaptive controllers. IEEE Control Syst. Mag. 2012, 32, 76–105. [Google Scholar] [CrossRef] [Scilit]
  14. Kiumarsi, B.; Vamvoudakis, K.; Modares, H.; Lewis, F. Optimal and autonomous control using reinforcement learning: A survey. IEEE Trans. Neural Netw. Learn. Syst. 2018, 29, 2042–2062. [Google Scholar] [CrossRef] [Scilit]
  15. Vrabie, D.; Pastravanu, O.; Abu-Khalaf, M.; Lewis, F. Adaptive optimal control for continuous-time linear systems based on policy iteration. Automatica 2009, 45, 477–484. [Google Scholar] [CrossRef] [Scilit]
  16. Jiang, Y.; Jiang, Z.P. Computational adaptive optimal control for continuous-time linear systems with completely unknown dynamics. Automatica 2012, 48, 2699–2704. [Google Scholar] [CrossRef] [Scilit]
  17. Zhu, L.; Modares, H.; Peen, G.; Lewis, F.; Yue, B. Adaptive suboptimal output-feedback control for linear systems using integral reinforcement learning. IEEE Trans. Control Syst. Technol. 2015, 23, 264–273. [Google Scholar] [CrossRef] [Scilit]
  18. Abu-Khalaf, M.; Lewis, F.; Huang, J. Neurodynamic programming and zero-sum games for constrained control systems. IEEE Trans. Neural Netw. 2008, 19, 1243–1252. [Google Scholar] [CrossRef] [Scilit]
  19. Zhou, A.; Fan, H. Novel adaptive synchronization criteria of fractional-order fuzzy neural networks with parameter uncertainties and information interactions. AIMS Math. 2026, 11, 9166–9190. [Google Scholar] [CrossRef] [Scilit]
  20. Vamvoudakis, K.; Lewis, F. Online solution of nonlinear two-player zero-sum games using synchronous policy iteration. Int. J. Robust Nonlinear Control 2012, 22, 1460–1483. [Google Scholar] [CrossRef] [Scilit]
  21. Wang, P.; Yu, C.; Lv, M.; Duan, G.R. Safe fixed-time reinforcement learning for nonlinear zero-sum games with obstacle avoidance awareness. Automatica 2026, 183, 112673. [Google Scholar] [CrossRef] [Scilit]
  22. Liang, X.; Li, J. ESO-enhanced actor–critic reinforcement learning-optimised trajectory tracking control for 3-DOF marine vessels. Mathematics 2026, 14, 867. [Google Scholar] [CrossRef] [Scilit]
  23. Yuan, X.; Chen, B.; Lin, C. Neural adaptive fixed-time control for nonlinear systems with full-state constraints. IEEE Trans. Cybern. 2023, 53, 3048–3059. [Google Scholar] [CrossRef] [Scilit]
  24. Liang, X.; Li, J.; Bao, D. Fixed-time event-triggered control for distributed unmanned underwater vehicles. J. Mar. Sci. Eng. 2026, 14, 202. [Google Scholar] [CrossRef] [Scilit]
  25. Wang, Q.; Cao, J.; Liu, H. Adaptive fuzzy control of nonlinear systems with predefined time and accuracy. IEEE Trans. Fuzzy Syst. 2022, 30, 5152–5165. [Google Scholar] [CrossRef] [Scilit]
  26. Shi, L.; Yan, S.; Li, W. Consensus and products of substochastic matrices: Convergence rate with communication delays. IEEE Trans. Syst. Man Cybern. Syst. 2025, 55, 4752–4761. [Google Scholar] [CrossRef] [Scilit]
  27. Li, W.; Yan, S.; Shi, L.; Yue, J.; Shi, M.; Lin, B.; Qin, K. Multiagent consensus tracking control over asynchronous cooperation–competition networks. IEEE Trans. Cybern. 2025, 55, 4347–4360. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Ni, J.; Shi, P. Global predefined time and accuracy adaptive neural network control for uncertain strict-feedback systems with output constraint and dead zone. IEEE Trans. Syst. Man Cybern. Syst. 2021, 51, 7903–7918. [Google Scholar] [CrossRef] [Scilit]
  29. Liu, B.; Hou, M.; Wu, C.; Wang, W.; Wu, Z.; Huang, B. Predefined-time backstepping control for a nonlinear strict-feedback system. Int. J. Robust Nonlinear Control 2021, 31, 3354–3372. [Google Scholar] [CrossRef] [Scilit]
  30. Sánchez-Torres, J.; Gómez-Gutiérrez, D.; López, E.; Loukianov, A. A class of predefined-time stable dynamical systems. IMA J. Math. Control Inf. 2018, 35, i1–i29. [Google Scholar] [CrossRef] [Scilit]
  31. Gibson, T.; Annaswamy, A.; Lavretsky, E. Modeling for Control of Very Flexible Aircraft. In Proceedings of the AIAA Guidance, Navigation, and Control Conference, Portland, OR, USA, 8–11 August 2011. [Google Scholar]
  32. Ma, J.; Peng, C. Adaptive model-free fault-tolerant control based on integral reinforcement learning for a highly flexible aircraft with actuator faults. Aerosp. Sci. Technol. 2021, 119, 107204. [Google Scholar] [CrossRef] [Scilit]
  33. Qu, Z.; Thomsen, B.; Annaswamy, A. Adaptive control for a class of multi-input multi-output plants with arbitrary relative degree. IEEE Trans. Autom. Control 2020, 65, 3023–3038. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Three-rigid-wingmodel for highly flexible aircraft.
Figure 1. Three-rigid-wingmodel for highly flexible aircraft.
Electronics 15 03838 g001
Figure 2. Structural diagram of the proposed algorithm.
Figure 2. Structural diagram of the proposed algorithm.
Electronics 15 03838 g002
Figure 3. Velocity tracking. (Left): Case I. (Right): Case II.
Figure 3. Velocity tracking. (Left): Case I. (Right): Case II.
Electronics 15 03838 g003
Figure 4. Angle-of-attack tracking. (Left): Case I. (Right): Case II.
Figure 4. Angle-of-attack tracking. (Left): Case I. (Right): Case II.
Electronics 15 03838 g004
Figure 5. Dihedral angle tracking. (Left): Case I. (Right): Case II.
Figure 5. Dihedral angle tracking. (Left): Case I. (Right): Case II.
Electronics 15 03838 g005
Figure 6. Control inputs. (Left): Case I. (Right): Case II.
Figure 6. Control inputs. (Left): Case I. (Right): Case II.
Electronics 15 03838 g006
Figure 7. Critic weights. (Left): Case I. (Right): Case II.
Figure 7. Critic weights. (Left): Case I. (Right): Case II.
Electronics 15 03838 g007
Figure 8. Critic-weight increments. (Left): Case I. (Right): Case II.
Figure 8. Critic-weight increments. (Left): Case I. (Right): Case II.
Electronics 15 03838 g008
Figure 9. Bellman residuals. (Left): Case I. (Right): Case II.
Figure 9. Bellman residuals. (Left): Case I. (Right): Case II.
Electronics 15 03838 g009
Figure 10. Critic weight-update norms. (Left): Case I. (Right): Case II.
Figure 10. Critic weight-update norms. (Left): Case I. (Right): Case II.
Electronics 15 03838 g010
Figure 11. State Comparison.
Figure 11. State Comparison.
Electronics 15 03838 g011
Figure 12. Error Comparison.
Figure 12. Error Comparison.
Electronics 15 03838 g012
Figure 13. Control Comparison.
Figure 13. Control Comparison.
Electronics 15 03838 g013
Figure 14. Comparison of model uncertainties regarding velocity tracking.
Figure 14. Comparison of model uncertainties regarding velocity tracking.
Electronics 15 03838 g014
Figure 15. Comparison of model uncertainties regarding angle-of-attack tracking.
Figure 15. Comparison of model uncertainties regarding angle-of-attack tracking.
Electronics 15 03838 g015
Figure 16. Comparison of model uncertainties regarding dihedral angle tracking.
Figure 16. Comparison of model uncertainties regarding dihedral angle tracking.
Electronics 15 03838 g016
Figure 17. Comparison of model uncertainties regarding control inputs.
Figure 17. Comparison of model uncertainties regarding control inputs.
Electronics 15 03838 g017
Figure 18. Comparison of model uncertainties regarding critic weights.
Figure 18. Comparison of model uncertainties regarding critic weights.
Electronics 15 03838 g018
Figure 19. Comparison of model uncertainties regarding critic-weight increments.
Figure 19. Comparison of model uncertainties regarding critic-weight increments.
Electronics 15 03838 g019
Figure 20. Comparison of model uncertainties regarding Bellman residuals.
Figure 20. Comparison of model uncertainties regarding Bellman residuals.
Electronics 15 03838 g020
Figure 21. Comparison of model uncertainties regarding critic weight-update norms.
Figure 21. Comparison of model uncertainties regarding critic weight-update norms.
Electronics 15 03838 g021
Figure 22. Time response of tracking errors.
Figure 22. Time response of tracking errors.
Electronics 15 03838 g022
Figure 23. Time response of critic weights.
Figure 23. Time response of critic weights.
Electronics 15 03838 g023
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, H.; Ma, J.; Zhang, Y.; Wu, M. Zero-Sum Game-Based Practical Predefined-Time Reinforcement Learning for Robust Tracking Control of Highly Flexible Aircraft. Electronics 2026, 15, 3838. https://doi.org/10.3390/electronics15173838

AMA Style

Zhang H, Ma J, Zhang Y, Wu M. Zero-Sum Game-Based Practical Predefined-Time Reinforcement Learning for Robust Tracking Control of Highly Flexible Aircraft. Electronics. 2026; 15(17):3838. https://doi.org/10.3390/electronics15173838

Chicago/Turabian Style

Zhang, Hanwen, Jianjun Ma, Yuxin Zhang, and Meiping Wu. 2026. "Zero-Sum Game-Based Practical Predefined-Time Reinforcement Learning for Robust Tracking Control of Highly Flexible Aircraft" Electronics 15, no. 17: 3838. https://doi.org/10.3390/electronics15173838

APA Style

Zhang, H., Ma, J., Zhang, Y., & Wu, M. (2026). Zero-Sum Game-Based Practical Predefined-Time Reinforcement Learning for Robust Tracking Control of Highly Flexible Aircraft. Electronics, 15(17), 3838. https://doi.org/10.3390/electronics15173838

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop