1. Introduction
Highly flexible aircraft (HFA) serve as a representative benchmark for advanced flight control due to their lightweight high-aspect-ratio wings and strong aeroelastic coupling. Nonlinear aeroelasticity and flight dynamics must be treated together for HFA, since flexible deformation changes both the structural response and rigid-body motion [
1]. Large elastic deformation in flight induces strong rigid–flexible coupling in HFA, resulting in pronounced nonlinearities and modeling uncertainties. These characteristics make controller design for HFA particularly challenging [
2].
A variety of nonlinear control methods have been developed for highly flexible aircraft. Gregory used modified nonlinear dynamic inversion to improve flexible-aircraft control, but this type of approach still relies on a carefully constructed dynamic model [
3]. Patil and Hodges studied output-feedback control for nonlinear aeroelastic response, showing that feedback design can suppress flexible response when the aeroelastic model is available [
4]. Shearer and Cesnik proposed a trajectory-control architecture for very flexible aircraft, and Raghavan and Patil extended trajectory control ideas to high-aspect-ratio flying wings [
5,
6]. Gust and load-alleviation studies then introduced robust control, model-predictive control, and LQG-based designs to improve performance under atmospheric excitation [
7]. Cook et al. emphasized robust gust alleviation and stabilization for highly flexible aircraft, while Haghighat et al. used model-predictive control for gust-load alleviation [
8,
9]. Liu et al. showed that LQG-based model-predictive control can be effective for gust-load alleviation in flexible aircraft [
10]. However, Qu and Annaswamy pointed out that trim drift, unmeasured flexible modes, and actuator anomalies can exceed the robustness margin of conventional linear output-feedback designs, motivating adaptive output-feedback control with closed-loop reference models [
11].
Reinforcement learning (RL) and approximate dynamic programming (ADP) provide another route because they learn value functions and policies from system data rather than requiring a completely known drift model. Lewis and Vrabie connected reinforcement learning and ADP to feedback control, and Lewis et al. later summarized how natural decision methods can be used to design optimal adaptive controllers [
12,
13]. Kiumarsi et al. reviewed optimal and autonomous control using reinforcement learning, showing the breadth of ADP/RL in modern feedback-control design [
14]. For continuous-time systems, Vrabie et al. introduced policy-iteration-based adaptive optimal control, and Jiang and Jiang developed computational adaptive optimal control for systems with completely unknown dynamics [
15,
16]. Zhu et al. further showed that integral reinforcement learning can be used for suboptimal output-feedback control of partially unknown linear continuous-time systems [
17]. Then, Peng and Ma used online integral RL to stabilize an uncertain HFA using full-state and output-feedback information, reducing the dependence on exact drift dynamics [
2]. Nevertheless, RL controllers for HFA still face two important difficulties: the learning convergence time is not explicitly assigned by the designer, and external disturbances are not always treated as worst-case strategic inputs during the learning process.
Zero-sum differential games provide a natural framework for robust RL because the controller minimizes the performance index while the disturbance attempts to maximize it. Abu-Khalaf et al. used neurodynamic programming and zero-sum games to handle constrained control, which established an important link between Hamiltonian function equations and robust learning-based control [
18]. An efficient adaptive synchronization criteria was developed in [
19] for fractional-order fuzzy neural networks with uncertain parameters, providing a useful perspective on adaptive control of uncertain nonlinear systems. Vamvoudakis and Lewis then developed an online solution for nonlinear two-player zero-sum games using synchronous policy iteration [
20]. More recently, Wang et al. combined safe reinforcement learning, fixed-time stability, and zero-sum games to address nonlinear systems with external disturbances and obstacle-avoidance awareness [
21]. A novel ESO-enhanced actor–critic RL method for robust marine-vessel tracking was developed in [
22]. However, fixed-time bounds can still be conservative, and the designer may not be able to set the actual convergence time directly as a mission parameter.
To directly set the convergence time, Yuan et al. developed neural adaptive fixed-time control for nonlinear systems with full-state constraints, illustrating the value of neural approximation in time-constrained control [
23]. An efficient fixed-time event-triggered control was proposed for distributed underwater vehicles in [
24]. Wang, Cao, and Liu then proposed adaptive fuzzy control with both predefined time and predefined accuracy, emphasizing that convergence time and convergence accuracy can be assigned in advance for engineering systems [
25]. A new standard was developed in [
26] to measure the lowest convergence rate for leader–follower consensus under communication delays. Efficient cooperative consensus tracking control was investigated in [
27] for multi-agent systems on cooperation–competition networks with asynchronous information exchange. Ni and Shi developed predefined-time adaptive neural control with output constraints, and Liu et al. studied predefined-time backstepping for strict-feedback nonlinear systems [
28,
29]. Related predefined-time stability results further support the use of user-specified time parameters in nonlinear systems [
30]. However, a zero-sum game-based predefined-time reinforcement learning method has not been designed for highly flexible aircraft.
Motivated by the above discussions, this work develops a zero-sum game-based practical predefined-time reinforcement learning method for robust tracking control of highly flexible aircraft. The main contributions are summarized as follows:
(1) Compared with existing RL-based control methods for highly flexible aircraft [
2], where external disturbances are treated as exogenous signals, a zero-sum differential-game architecture is constructed to explicitly characterize the worst-case interaction between the controller and disturbance, thereby converting the robust tracking problem into a min–max optimal control problem.
(2) Compared with conventional zero-sum ADP/RL methods [
18,
20] and the recent fixed-time zero-sum RL method [
21], the critic weight-estimation error is guaranteed to converge to a bounded neighborhood within the predefined time.
(3) In contrast with predefined-time adaptive neural/fuzzy and backstepping control methods [
25,
28,
29], the proposed predefined-time mechanism is incorporated into the critic learning process of the zero-sum game. The practical predefined-time convergence of the critic weight-estimation error and uniform ultimate boundedness of the overall closed-loop tracking system are rigorously established.
4. Stability Analysis
The stability analysis is developed in two stages. First, the learning dynamics of the critic network are analyzed to establish practical predefined-time convergence of the critic weight-estimation error. This result further provides boundedness of the learned value-function approximation and the corresponding policy reconstruction errors. Second, the critic-learning dynamics and the physical tracking dynamics are incorporated into a composite Lyapunov analysis to establish uniform ultimate boundedness of the overall closed-loop system. Therefore, Theorem 1 provides the learning-layer guarantee required for the closed-loop stability analysis in Theorem 2.
Theorem 1. Under Assumption 1, if the current-and-history fractional-power network update law (46) is applied, the practical predefined-time stability of the critic approximation error can be guaranteed. Proof. Construct the Lyapunov functional reflecting the critic error energy:
Taking its derivative and substituting the current-and-history update law (
46) gives
Using (
38) and (
39), Equation (
51) can be rewritten as
For the ideal approximation case—namely,
and
,
—one obtains
Since the current term is non-positive, it can be retained in the learning law but discarded in the upper-bound estimate. Hence,
According to Lemmas 2 and 3, one has
According to the definition of
and Assumption 1,
Moreover, since
one obtains
Substituting (
55)–(
59) into (
54) gives
According to the definitions of
and
in (
49),
According to Lemma 1, the ideal critic weight-estimation error converges to zero within the predefined upper bound ().
For the nonideal case, let
and
Outside the residual-dominated region, i.e., when
,
, it holds that
Therefore, Equation (
52) yields
Using the absolute-value inequalities, i.e.,
Equation (
65) is upper-bounded by
Because
and
are bounded and the normalized regressors are bounded, there exists a positive constant (
) satisfying
Then, the residual-induced terms in (
68) satisfy
Dropping the non-positive current term and using the same historical-data excitation argument as in the ideal case gives
The compact set is defined as
where
For
, inequality (
72) implies
Using (
59) further yields
According to (
49), one has
Thus, according to Lemma 1, converges to the compact set () within the predefined upper bound () and remains therein thereafter. This completes the proof. □
Remark 2. Theorem 1 provides a practical predefined-time learning guarantee for the critic network. In the ideal approximation case, the critic weight-estimation error converges to zero within the predefined upper bound (). In the nonideal case, critic error converges to the compact set () within the prescribed learning time. Therefore, characterizes the learning-time requirement, whereas the size of reflects the effects of approximation accuracy.
Based on Theorem 1, the critic weight-estimation error converges to a bounded region within the predefined time, which further ensures bounded errors in the learned value-function gradient and the reconstructed control policy. This learning-layer result provides the basis for analysis of the overall closed-loop stability in Theorem 2.
Theorem 2. For the disturbed physical system (12), apply the RL control policy in (28) and the current-and-history network parameter update law (46). If Assumption 1 holds, then the closed-loop system is bounded. Proof. The Lyapunov function is defined as
Taking the derivative yields
By further mathematical manipulation,
According to the optimality condition,
Moreover, according to the extremum conditions of the optimal control and disturbance laws, i.e.,
one has
Applying Young’s inequality to the cross terms gives
Substituting the bounds into
and eliminating the
term yields
Combining the critic neural-network approximation property, the errors between the practical and optimal policies are expressed by the parameter reconstruction error as
Theorem 1 has already shown that the network weight-estimation error (
) converges to a bounded region within the predefined time. On a compact set,
,
, and
are bounded. Therefore, we define a positive constant (
) such that the residual terms induced by the network approximation error and external bounded disturbance satisfy
The quadratic form (
) satisfies
Substituting this result into the derivative of the composite Lyapunov functional and using Theorem 1 gives, for
,
When
, the critic error is already bounded. Therefore, when
one has
. According to Lyapunov stability theory, the coupled closed-loop system is uniformly ultimately bounded. This completes the proof. □
5. Simulation
To comprehensively validate the efficacy of the proposed algorithm, this article conducts a series of experiments designed to evaluate its generality, performance, and robustness. Specifically, the experimental validation consists of three parts. Firstly, different trim points are selected to verify the generality and broad applicability of the proposed algorithm. Secondly, the proposed algorithm is compared with traditional Integral Reinforcement Learning (IRL) algorithms to demonstrate its performance superiority. Finally, model uncertainty and measurement noise are introduced into the experimental setting to verify the robustness of the proposed algorithm.
5.1. Different Trim Points
The applicable range of this linear model is when the dihedral angle is no more than 15 degrees [
31]. In order to verify that this method has a good generalization ability for this model, the paper selects two equilibrium points with dihedral angles of 15 degrees and 5 degrees, which are relatively large and small, respectively. The HFA is trimmed under the following conditions:
V = 68 ft/s,
h = 40,000 ft,
,
,
or
, and
.
The state-space matrix and input matrix (
and
) are shown as follows:
The bounded disturbance used for plant propagation is
The zero-sum game weights and critic-learning parameters are
in (
46) is set as
, and the predefined learning time is
. The fractional parameter is
, the number of basis functions is
, the critic gain matrix is
, and the initial critic weight is
. The simulation step is
s, the final time is 20 s, and the initial state is
. The reference command and its derivative are
Figure 3,
Figure 4 and
Figure 5 compare the tracking performance of the two parameter sets for velocity, angle of attack, and dihedral angle. Both cases show a short initial adjustment followed by bounded tracking. The second parameter set contains stronger pitch-rate and dihedral angle coupling, especially in the fourth row of
and in the fourth row of
, but the controller still keeps the main rigid and elastic responses close to their commands. This indicates that the zero-sum RL policy addsa degree of robustness to the considered model-parameter variation.
Figure 6 shows that the five control channels remain smooth.
Figure 7,
Figure 8,
Figure 9 and
Figure 10 describe the critic-learning process. The critic weights change rapidly during the early learning interval, then remain bounded. The weight increments decrease after the main adaptation phase, showing that the learning process is concentrated around the prescribed learning window. The Bellman residual is not forced exactly to zero because the basis functions and disturbance estimate are approximate, but it stays bounded in both cases. The weight-update norm also decreases after the initial transient, which supports the practical predefined-time critic convergence stated in Theorem 1.
Overall, the two-case simulations verify three points. First, the zero-sum controller suppresses bounded disturbances. Second, the predefined-time critic update concentrates learning activity within the prescribed time window and keeps the Bellman residual bounded afterward. Third, the same controller parameters work for both HFA parameter sets, which indicates robustness to the considered model-parameter variation.
5.2. Comparison with Traditional Methods
In this section, the proposed algorithm is compared with traditional Integral Reinforcement Learning (IRL) algorithms introduced by Ma [
32]. The linearized state matrix and input matrix is
, which is introduced in the previous section.
is
, and other simulation conditions are the same as the previous section.
Based on the simulation results presented in the figures, a comparative analysis between the proposed method and the conventional nominal Integral Reinforcement Learning (IRL) approach reveals significant performance advantages of the former. The superiority of the proposed method is primarily evidenced by its rapid convergence characteristics and high-precision tracking capabilities.
As illustrated in the state trajectory plots in
Figure 11, the proposed controller demonstrates a markedly faster transient response compared to the nominal IRL method. Specifically, for states
and
, the proposed method achieves convergence to the desired reference values within approximately
. In contrast, the nominal IRL method exhibits a sluggish response, failing to reach the steady-state target, even after
of operation. This indicates that the prescribed-time framework effectively overcomes the asymptotic convergence limitations often associated with traditional adaptive or learning-based controllers, guaranteeing system stabilization within a user-defined timeframe independent of initial conditions.
The tracking-error dynamics further corroborate the robustness of the proposed scheme in
Figure 12. While the nominal IRL method maintains non-negligible steady-state errors, the proposed method drives tracking errors
,
, and
to zero with high precision. The error trajectories for the proposed method remain virtually flat at zero after the settling time, whereas the nominal IRL shows continuous drift or offset, highlighting the superior disturbance rejection and parameter-estimation accuracy of the proposed algorithm.
Regarding the control input activity in
Figure 13, although the proposed method initially utilizes higher control authority to enforce the fast convergence, it settles to a stable equilibrium rapidly within 5 s. Conversely, the nominal IRL method applies significantly smaller control magnitudes but fails to stabilize the system effectively, resulting in prolonged transient periods. This trade-off suggests that the proposed method efficiently utilizes available control energy to achieve superior dynamic performance, whereas the nominal method’s conservative input leads to inadequate system regulation.
In summary, the comparative results validate that the proposed control method offers substantial improvements over the nominal IRL method, particularly in terms of convergence rate, steady-state accuracy, and transient response quality.
5.3. Model Uncertainty
In this section, model uncertainty is introduced to assess the robustness of the algorithm proposed in this paper. The linearized state matrix and input matrix is
, which is introduced in the previous section. The linearized system has eigenvalues with
,
,
, and
[
32]. The system-matched uncertainties are caused by flexible effects of dihedral angle [
33].
In the simulation environment,
replaces
in the calculation of environmental information feedback, when the proposed algorithm has no information about it and still uses
in the calculation of controller, where
Different values are settled as follows:
Case I: no model uncertainty, ;
Case II: low model uncertainty, ;
Case III: high model uncertainty,
Case IV: negative model uncertainty, .
Other simulation conditions are the same as in the previous section.
Figure 14,
Figure 15 and
Figure 16 demonstrate that the proposed controller preserves essentially invariant tracking performance across all tested uncertainty levels. The commanded and measured trajectories overlap closely after a short transient of approximately 4–5 s, with negligible steady-state error and nearly identical rise and settling times, irrespective of the sign or magnitude of
.
Under the same uncertainty cases, the control inputs remain bounded, smooth, and free of saturation or sustained oscillations in
Figure 17. The transient profiles and peak amplitudes are nearly indistinguishable across
, and no chattering is observed; input rates stay within acceptable limits, indicating benign actuation demands and strong robustness of the control law.
Figure 18,
Figure 19,
Figure 20 and
Figure 21 characterize the learning process. The critic weights converge to bounded constants, the weight-update increments rapidly decay to small values, and the Bellman residual drops by several orders of magnitude and stays near zero. The norm of the weight-update vector exhibits an early peak, then quickly vanishes. Closed-loop performance is preserved, even after learning is frozen at
, evidencing stable adaptation and insensitivity to modeling errors up to every tested
.
5.4. Measurement Noise
The measurement noiseis represented as
where
denotes the measurement-noise vector. The resulting tracking errors and critic-weight responses are given below.
Figure 22 and
Figure 23 present the results in the presence of measurement noise. The tracking errors remain bounded and decrease to a small neighborhood of zero after a short transient, while the critic weights exhibit bounded learning responses and gradually approach steady values before
. Therefore, the introduced measurement noise does not significantly degrade the tracking performance across tested measurement-noise levels.