Next Article in Journal
Analysis and Enhancement of Steady Climb Performance with Control Input Redundancy for a Dual-Propulsion VTOL UAV
Next Article in Special Issue
Adaptive Autopilot Design and Implementation for Cessna Citation X
Previous Article in Journal
DDA-SIM-ATT: A Synergistic Multi-Module Fusion Model for High-Precision Prediction of Departure Flight Taxi-Out Time
Previous Article in Special Issue
Adaptive Incremental Nonlinear Dynamic Inversion Control with Guaranteed Stability for Aerial Manipulators
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Incremental Nonlinear Reinforcement Learning Control for a Civil Aircraft with Model Uncertainties and Actuator Faults

The School of Aeronautics and Astronautics, Shanghai Jiao Tong University, Shanghai 200240, China
*
Author to whom correspondence should be addressed.
Aerospace 2026, 13(4), 315; https://doi.org/10.3390/aerospace13040315
Submission received: 11 February 2026 / Revised: 24 March 2026 / Accepted: 25 March 2026 / Published: 27 March 2026
(This article belongs to the Special Issue Challenges and Innovations in Aircraft Flight Control (2nd Edition))

Abstract

The problem of fault-tolerant attitude tracking control for the civil aircraft with model uncertainties and actuator faults is studied. A robust multiple inversion-based incremental nonlinear dynamic inversion (RMI-INDI) fault-tolerant control method is proposed for the problem. Firstly, considering that the higher-order term is neglected in the INDI method, an RMI method is proposed to deal with the higher-order term and model uncertainties of the INDI control. Secondly, to achieve the optimal control parameters for the INDI controller, a reinforcement learning (RL) method is suggested, where a Deep Deterministic Policy Gradient (DDPG) algorithm with a smooth reward function is designed. Finally, performances of the proposed RL-RMI-INDI fault-tolerant controller are demonstrated by using two scenario simulations. Compared with the SMC control, RMI-NDI control and INDI control without RL, tracking errors and overshoots are greatly reduced by the proposed RL-RMI-INDI controller for attitude tracking missions, even under model uncertainties and actuator faults.

1. Introduction

Flight safety is a hot topic for civil aircraft recently [1]. For instance, many flight accidents result from faults of flight control system, including actuator stuck and sensor failure, etc. These faults not only cause large model uncertainties but also threaten the stability of the system [2]. Therefore, designing a novel fault-tolerant flight control system that maintain aircraft stability under fault conditions and uncertainties becomes a significant challenge for aviation safety.
Due to the strong nonlinear characteristics of the aircraft under actuator faults and model uncertainties, the traditional linear control design at the balance points can hardly meet the control requirements. Feedback linearization (FL) is a kind of important nonlinear control method [3], but this method makes it vulnerable in the presence of model mismatch or unmodeled dynamics. Therefore, researchers have proposed various other methods to enhance robustness, such as H∞ control [4,5,6] and a robust multiple inversion method [7]. To achieve optimal control gain for the robust stabilization system, a differential evolution (DE) algorithms combined with linear matrix inequalities (LMIs) was proposed [8], but the parameters that resulted from DE easily made the system unstable. The Robust Granular Feedback Linearization (RGFL) controller was proposed to combine with Evolutionary Participatory Learning (EPL), which could estimate and compensate model uncertainties online [9,10], but there are no systematic comparisons for measurement noises, model mismatch, and actuator faults. Nonlinear dynamic inversion (NDI) control was achieved by effectively compensating for the nonlinear effects of the system inversion [11,12,13,14]. However, the performances of the NDI controllers were highly dependent on the accuracy of the model [15]. To enhance robustness for model uncertainties, an incremental nonlinear dynamic inversion (INDI) method has been proposed [16,17]. The INDI applies angular acceleration as feedback, which reduces reliance on precise models and overcomes the impact of model uncertainties. Due to its low sensitivity to model parameters, INDI has been successfully applied to many kinds of the aircraft [18,19].
Although INDI reduces the reliance on accurate models and improves robustness by an incremental form, it is a critical issue to maintain performances and to ensure stability of the system with actuator faults. In the existing studies, there are very few studies on fault-tolerant control and advanced intelligent algorithms for the civil aircraft. Furthermore, parameters optimization of INDI control is critical but challenging, especially for fault-tolerant control systems. Currently the incremental nonlinear dynamic inversion parameters are optimization mainly was depended on engineering experience [20]. When sudden disturbances or actuator faults occur, effectiveness of the controller may be significantly reduced. Therefore, it is imperative to explore novel adaptive parameter optimization methods.
Several intelligent algorithms have been applied to parameter optimization. Han [21] used a frequency domain analysis algorithm and neural network to adaptively optimize the parameters of the closed-loop control system. Wang [22] proposed an improved fuzzy particle swarm optimization method to adjust the control parameters. Jiang [23] applied the improved fuzzy particle swarm optimization algorithm to identify and optimize parameters of the aircraft, which showed good robustness.
With the development of artificial intelligence technology, reinforcement learning is widely applied in the field of flight control. Yu and Choong [24] used an RL method to optimize the parameters of a fixed-wing UAV, and time consumption is significantly reduced. For the two-degree-of-freedom flight simulator, Mo [25] discussed intelligent control methods integrating fuzzy logic and neural networks and also explored in-depth various aspects of each control technique concerning stability, disturbance rejection, response time, asymptotic, exponential convergence, and finite-time convergence. For the actuator fault problem in highly flexible aircraft, Ma [26] proposed a model-free adaptive fault-tolerant control method based on integral reinforcement learning. By introducing an augmented system with tracking errors and a reference model-based observer, the proposed method can effectively compensate for uncertainties and external disturbances without system model information. For strictly feedback nonlinear systems affected by intermittent actuator faults, Yang [27] designed an adaptive fault-tolerant control scheme based on RL. This scheme automatically identified and switched the control through a switching function and achieved optimal tracking control through approximating the Hamilton–Jacobi–Bellman (HJB) equation by using neural networks. For actuator effectiveness loss and bias faults in nonlinear servo systems, Xie [28] proposed a fault-tolerant control strategy that combined the prescribed performance with reinforcement learning. An identification–critic–actor architecture and a fuzzy logic system were employed, and this strategy assured the tracking error remained within prescribed bounds and performance optimization simultaneously. For the problem of uncertain parameters and time-varying state constraints in multi-actuator aerospace systems, Liu [29] constructed a state constraint observer based on collision detection, which combined a barrier function with an actor–critic architecture, and it could implement cooperative fault-tolerant control and collision avoidance.
Although the RL algorithm has shown potential value, research on reinforcement learning in fault-tolerant control systems of the large civil aircraft, especially for parameter optimization under actuator faults, is still greatly limited. To overcome the effect of faults and uncertainties on the flight control of a civil aircraft, a fault-tolerant control method based on RL-RMI-INDI is proposed. The main contributions can be summarized as follows:
  • For the uncertainties problem of higher-order terms in INDI control, an estimation method based on robust multiple inversion is proposed. The RMI method is employed to estimate and compensate for the high-order truncation terms that are neglected in the INDI, specifically under actuator faults in aircraft control. By constructing an augmented system and designing an integral loop to calculate the RMI factor, effective estimation of high-order uncertain terms is achieved while feedback linearization effect of the system remains unaffected.
  • For the parameter optimization problem of the RMI-INDI controller in the case of actuator faults, a reinforcement learning approach based on DDPG is proposed. A smooth reward function based on hyperbolic tangent is designed that incorporates system tracking error with control input and state constraints, thereby controller parameter optimization is achieved.
  • A RL-RMI-INDI controller is designed for fault-tolerant control of the aircraft with actuator faults. The tracking control of attitude is achieved by the outer loop NDI, while the RL-RMI-INDI is designed for fault-tolerant control in the angular rate inner loop, and the stability of the closed-loop system under varying gain condition has been proven through Lyapunov theory.
The rest of the paper is organized as follows: the dynamics modeling, uncertainties analysis and problem formulation are presented in Section 2; the design process of fault-tolerant controller based on RMI-INDI and a RL method based on the DDPG are presented for parameter optimization in Section 3; simulation and analysis are presented in Section 4; and the conclusion is presented in Section 5.

2. Dynamics Modeling and Problem Formulation

2.1. Dynamics Modeling of the Civil Aircraft

The attitude dynamics model of the aircraft is as follows [30]
x ˙ 1 = cos α cos β 0 sin α sin β 1 0 sin α cos β 0 cos α 1 T B V χ ˙ sin γ γ ˙ χ ˙ cos γ + p d q d r d
where x 1 = μ , α , β T is the attitude angles vector, μ is the kinematic bank angle. α and β are the angle of attack and angle of sideslip, respectively. χ is the kinematic azimuth angle and γ is the flight path angle. u 1 = x 2 d = ω d = p d , q d , r d T is the desired angular rate vector. T B V is the coordinate transformation matrix from the velocity frame to the body frame, that is,
T B V = cos α cos β cos α sin β cos μ + sin α sin μ cos α sin β sin μ sin α cos μ sin β cos β cos μ cos β sin μ sin α cos β sin α sin β cos μ cos α sin μ sin α sin β sin μ + cos α cos μ
The dynamics model of the angular rate is expressed as follows [31]
x ˙ 2 = ω ˙ = J 1 M ω × J ω
where x 2 = ω = p , q , r T ; the inertia matrix J is expressed as follows
J = I x x I x y I x z I x y I y y I y z I x z I y z I z z
where I x x I y y I z z T is the moment of inertia vector, and I x y I x z I y z T is the inertia product vector.
The moment M in Equation (3) can be presented as follows
M = M 0 + M u = M 0 + C M u u 2
where
M 0 = q ¯ S b C l β , p , r , V T c ¯ C m α , q , V T b C n β , p , r , V T
M u = C M u u = q ¯ S b C l δ a 0 b C l δ r 0 c ¯ C m δ e 0 b C n δ a 0 b C n δ r δ a δ e δ r
where b is the wing span, S is the wing area, q d p is the dynamic pressure, and c ¯ is the mean aerodynamic chord, and V T is the flight velocity, u 2 = δ a , δ e , δ r T denotes deflection of aileron, elevator and rudder, respectively. C M u is the coefficient related to M u , which can be expressed as follows
C M u = b C l δ a 0 b C l δ r 0 c ¯ C m δ e 0 b C n δ a 0 b C n δ r
Equation (3) can be rewritten in an affine form as follows
ω ˙ = f 2 ( ω ) + G 2 ( ω ) u 2
where f 2 ( ω ) = J 1 M 0 ω × J ω , G 2 ( ω ) = J 1 C M u .

2.2. Uncertainties Analysis Due to the INDI Design

Consider a MIMO affine nonlinear control system as follows
x ˙ = f ( x ) + G ( x ) u
y = x
where x is the state vector x R n , u is the control input vector u R n , and y is the output vector y R n . f ( x ) is the state function, G ( x ) is the control input matrix function, and G ( x ) is assumed to be non-singular.
According to the INDI principle [32], Equation (10a) can be expanded by Taylor expansion as follows
x ˙ = x ˙ 0 + f ( x ) + G ( x ) u x x = x 0 u = u 0 ( x x 0 ) + f ( x ) + G ( x ) u u x = x 0 u = u 0 ( u u 0 ) + Δ H . O . T
where x 0 and u 0 represent x and u at the previous sampling time, and Δ H . O . T denotes the higher-order truncation term.
It is assumed that the sampling time Δ t is small enough and u will change faster than x . Therefore, state x can be approximated by x 0 , but control variable u cannot be approximated by u 0 .
Remark 1. 
The higher-order terms  Δ H . O . T can be neglected if the sampling time  Δ t  is sufficiently small, and compared with the control input, the system state changes slowly. This is a standard assumption in the INDI design; that is, if  x x 0 is of order O ( Δ t 2 )  and the dynamics are Lipschitz continuous, then the residual is bounded through a constant multiplied by Δ t 2 , and for small values of Δ t , this residual can be neglected.
So, Equation (11) can be simplified as
x ˙ = x ˙ 0 + G ( x 0 ) ( u u 0 )
where x ˙ is replaced by the designed pseudo-inverse control input v , and the INDI control law can be obtained from Equation (12).
u = u 0 + G 1 ( x 0 ) ( v x ˙ 0 )
Substituting Equation (13) into Equation (12) yields
x ˙ = x ˙ 0 + G ( x 0 ) ( u 0 + G 1 ( x 0 ) ( v x ˙ 0 ) u 0 ) = x ˙ 0 + G ( x 0 ) G 1 ( x 0 ) ( v x ˙ 0 ) = v
In the INDI control law design for a practical aircraft, assuming that x ˙ 0 is an angular acceleration, x ˙ 0 cannot be measured directly because the aircraft is not equipped with angular acceleration sensors. Therefore, x ˙ 0 is obtained by using a differential combined with a second-order low-pass filter. The second-order low-pass filter Z ( s ) is designed as follows
Z ( s ) = s ω n 2 s 2 + 2 ξ n ω n s + ω n 2
where ω n is the angular frequency of the filter, ξ n is the damping ratio of the filter, and s denotes the Laplace factor.
It can be seen that model uncertainties Δ G ( x 0 ) of the INDI are considered for G ( x 0 ) , and f ( x ) in Equation (10a) is replaced by x ˙ 0 , so Equation (12) can be rewritten as follows:
x ˙ = x ˙ 0 + ( G ( x 0 ) + Δ G ( x 0 ) ) ( u u 0 )
Substituting Equation (13) into Equation (16) yields
x ˙ = x ˙ 0 + ( G ( x 0 ) + Δ G ( x 0 ) ) ( G 1 ( x 0 ) ( v x ˙ 0 ) ) = v Δ G ( x 0 ) G 1 ( x 0 ) x ˙ 0 + Δ G ( x 0 ) G 1 ( x 0 ) v = ( I + Δ G ( x 0 ) G 1 ( x 0 ) ) v Δ G ( x 0 ) G 1 ( x 0 ) x ˙ 0
Define Δ G = Δ G ( x 0 ) G 1 ( x 0 ) , then Equation (17) can be transformed as follows:
x ˙ = Δ G x ˙ 0 + ( I + Δ G ) v
According to Equation (18), the control scheme is designed as Figure 1, where Δ G denotes the model uncertainties, PID denotes Proportional Integral Derivative (PID) algorithm, and s , I / s denote integration or differentiation respectively.
The pseudo-inverse control input v ( t ) = ( K P + K I s + K D s ) e ( t ) is obtained by PID design and the transfer function of the system (18) is defined as G f ( s ) . The transfer function G f ( s ) is presented as follows:
G f ( s ) = v ( s ) ( I + Δ G ) I / s I + s Δ G / s 1 + ( v ( s ) ( I + Δ G ) I / s I + s Δ G / s ) = v ( s ) I + Δ G s ( I + Δ G ) 1 + v ( s ) I + Δ G s ( I + Δ G ) = v ( s ) s + v ( s )
It can be seen that model uncertainties Δ G = Δ G ( x 0 ) G 1 ( x 0 ) in Equation (19) have been eliminated. However, the INDI method is based on Taylor’s expansion principle and the higher-order term of Δ H . O . T in Equation (11) is omit. To improve the control precision of the aircraft, the effect of higher-order terms must be considered. So, Equation (12) can be rewritten as follows:
x ˙ = x ˙ 0 + G ( x 0 ) ( u u 0 ) + Δ H . O . T
Therefore, it is necessary to find a method to deal with the model uncertainties due to the higher-order term Δ H . O . T of the INDI.

2.3. Problem Formulation

To realize precision control for the civil aircraft, the main objective of this study is to design a reinforcement learning-based INDI fault-tolerant controller, which meets the following requirements:
(1)
The closed-loop system is stable with model uncertainties including Δ H . O . T and actuator faults;
(2)
The system state x is driven towards the desired x d , and the tracking error e = x d x , meets lim e 2 < ε , where the ε is a prescribed constant. e 2 = e T e represents the 2-norm of the vector e .

3. RL-RMI-INDI-Based Controller Design

To improve attitude tracking control of a civil aircraft under actuator faults and model uncertainties, an RL-RMI-INDI-based fault-tolerant control method is proposed in this section.
The control structure based on RL-RMI-INDI is shown in Figure 2, which includes a DDPG-based reinforcement learning loop, an update of the gain K ω , an attitude control loop, an angular rate control loop, and a RMI estimator.

3.1. Attitude Control Loop

The objective of the attitude control loop is to track the reference command vector x 1 d = μ r e f , α r e f , β r e f T , where β r e f is assumed to be zero, because it is small during flight or coordinated turn. The control input is the angular rate, u 1 = x 2 d = ω d . Since there are no model uncertainties in the attitude control loop, the NDI method is adopted [33].
x ˙ 1 = f 1 + G 1 u 1 = f 1 + G 1 ω d
where x 1 = μ , α , β T , u 1 = x 2 d = ω d = p d , q d , r d T .
f 1 = cos α cos β 0 sin α sin β 1 0 sin α cos β 0 cos α 1 ( T B V χ ˙ sin γ γ ˙ χ ˙ cos γ )
G 1 = cos α cos β 0 sin α sin β 1 0 sin α cos β 0 cos α 1
where T B V is the coordinate transformation matrix, as shown in Equation (2).
By using the dynamic inverse method, the desired angular rate can be obtained by Equation (21):
u 1 = ω d = G 1 1 ( x ˙ 1 f 1 ) = G 1 1 ( v 1 f 1 )
The virtual control v 1 is designed by the PID control.
v 1 = K P 1 e 1 ( t ) + K I 1 0 t e 1 ( τ ) d τ + K D 1 d e 1 ( t ) d t
where the attitude error e 1 ( t ) = x 1 d x 1 , K P 1 , K I 1 , K D 1 are the control parameters to be designed.

3.2. Angular Rate Control Loop

The angular rate inner loop controller structure based on the RMI-INDI is shown in Figure 3, which includes the RMI-INDI controller, a robust multiple inversion-based estimator, actuator model and aircraft model.

3.2.1. The Robust Multiple Inversion Estimator Design

Consider the angular rate system of the aircraft with model uncertainties as follows:
ω ˙ m = f 2 m ( ω ) + G 2 m ( ω ) u 2
where f 2 m ( ω ) = f 2 ( ω ) + Δ f 2 m ( ω ) , G 2 m ( ω ) = G 2 ( ω ) + Δ G 2 m ( ω ) , Δ f 2 m ( ω ) and Δ G 2 m ( ω ) represent the model uncertainties associated with the functions of f 2 ( ) and G 2 ( ) , respectively. By using the NDI control, the control input u 2 is obtained as follows:
u 2 = G 2 m 1 ( ω ) ( v 2 f 2 m ( ω ) )
where v 2 denotes the pseudo-control input. Equation (26) can be expressed as follows:
ω ˙ m = f 2 ( ω ) + G 2 ( ω ) u 2 + Δ m
where Δ m = Δ f 2 m ( ω ) + Δ G 2 m ( ω ) u 2 .
To reduce the effects of model uncertainties, a novel RMI estimator system ω ^ ˙ R M I is designed. Let ω ˙ a u g = ω ˙ ω ˙ m ω ^ ˙ R M I , the augmentation system is presented as follows:
ω ˙ a u g = ω ˙ ω ˙ m ω ^ ˙ R M I = f 2 ( ω ) + G 2 ( ω ) u 2 f 2 ( ω ) + G 2 ( ω ) u 2 + Δ m f 2 ( ω ) + G 2 ( ω ) u 2 + Δ m Δ R M I
The robust multiple inversion factor Δ R M I in Equation (29) is designed by an error integral as follows:
Δ ^ R M I = d i a g ( 1 τ k ) 1 k p ( ω ^ ˙ R M I ω ˙ ) d t = d i a g ( 1 τ k ) 1 k p ( Δ m Δ ^ R M I ) d t
where d i a g ( 1 τ k ) 1 k p = d i a g ( 1 / τ 1 , 1 / τ 2 , , 1 / τ p ) , and τ k > 0 , k is the relative degree of the system, p is the maximum relative degree of the system, and Δ ^ R M I is the estimated value of the robust multiple inversion factor Δ R M I .
Since the static error of the system can be eliminated by the integral control, the convergence process can be obtained by Equations (29) and (30) as follows:
lim t Δ ^ R M I ( t ) = Δ m
Substituting Equation (31) into Equation (30) yields
( Δ m Δ ^ R M I ) d t = 0
According to Equations (30) and (32), Δ ^ R M I will keep un-changeable Therefore, Δ m = Δ ^ R M I = Δ R M I can be obtained, this illustrates that Δ m can be estimated by the RMI-based estimator.
Substituting Δ m = Δ ^ R M I = Δ R M I into Equation (29) yields
ω ^ ˙ R M I = f 2 ( ω ) + G 2 ( ω ) u 2 + Δ m Δ R M I = f 2 ( ω ) + G 2 ( ω ) u 2 = ω ˙
Therefore, the RMI system ω ^ ˙ R M I is reversible.
According to the NDI control principle, the NDI control law u 2 of the RMI system ω ^ ˙ R M I is as follows:
u 2 = G 2 m 1 ( v 2 f 2 m ( ω ) + Δ ^ R M I )
where v 2 is the pseudo-input. Substituting Δ m = Δ ^ R M I = Δ R M I and Equation (34) into the RMI system (33) yields
ω ^ ˙ R M I = f 2 m ( ω ) + G 2 m ( ω ) [ G 2 m 1 ( v 2 f 2 m ( ω ) + Δ ^ R M I ) ] Δ R M I = f 2 m ( ω ) + v 2 f 2 m ( ω ) + Δ ^ R M I Δ R M I = v 2
By using Equations (33) and (35), the following result can be obtained:
ω ^ ˙ R M I = ω ˙ = v 2
In conclusion, not only Δ m can be well-estimated by the RMI-based estimator, but also the feedback linearization of the system will not be affected.

3.2.2. RMI-INDI-Based the Angular Rate Fault-Tolerant Control Law Design

A. Fault-Tolerant Control of INDI
The schematic diagram of the actuator fault is shown in Figure 4.
In Figure 4, u d e s = ( δ a d e s , δ e d e s , δ r d e s ) is the expected control variable of the control surface calculated by the controller, u r e a l = ( δ a r e a l , δ e r e a l , δ r r e a l ) is the control surface deflection with the influence of actuator fault, and u exp = ( δ a exp , δ e exp , δ r exp ) is the control surface deflection without actuator fault.
Taking the fault of the aileron as an example, that is:
δ a r e a l = λ f δ a exp + δ ¯ a s t u c k
where δ ¯ a s t u c k is the aileron stuck angle and the proportional coefficient λ f is a positive coefficient with λ f < 1 . It can be divided into four fault cases:
(1)
λ f = 1 , δ ¯ a s t u c k = 0 : it indicates no fault on the control surface.
(2)
0 < λ f < 1 , δ ¯ a s t u c k = 0 : it indicates partial failure of the control surface.
(3)
λ f = 0 , δ ¯ a s t u c k = 0 : it indicates that the control surface has completely failed.
(4)
λ f = 0 , δ ¯ a s t u c k 0 : it indicates that the control surface is stuck.
For Case (4), an INDI-based fault tolerance control is proposed. Now consider the attitude dynamics system with actuator faults:
ω ˙ = f 2 ( ω ) + G 2 ( ω ) u r e a l
where the practical control input can be described as follows by Equation (11):
u r e a l = u exp + u f a u l t
Substituting Equation (39) into Equation (38) yields
ω ˙ = f 2 ( ω ) + G 2 ( ω ) ( u exp + u f a u l t ) = f 2 ( ω ) + G 2 ( ω ) u exp + G 2 ( ω ) u f a u l t
where u f a u l t is the control input generated by a single control surface with stuck fault. According to the INDI principle, and by using the Taylor expansion principle and ignoring the error of higher-order terms, Equation (40) can be rewritten as
ω ˙ = ω ˙ 0 + ω f 2 ( ω ) + G 2 ( ω ) u exp + G 2 ( ω ) u f a u l t ω = ω 0 , u exp = u exp 0 , u f a u l t = u 0 f a u l t ω ω 0   + u exp f 2 ( ω ) + G 2 ( ω ) u exp + G 2 ( ω ) u f a u l t ω = ω 0 , u exp = u exp 0 , u f a u l t = u 0 f a u l t ( u exp u exp 0 )   + u f a u l t f 2 ( ω ) + G 2 ( ω ) u exp + G 2 ( ω ) u f a u l t ω = ω 0 , u exp = u exp 0 , u f a u l t = u 0 f a u l t ( u f a u l t u 0 f a u l t )
Assumption 1. 
The actuator fault type is stuck, and the fault mode is the step jump form, that is,
u t f a u l t = 0 0 < t < t f a u l t u ¯ s t u c k t f a u l t t
where  u t f a u l t  denotes the stuck fault of the control surface, and  t f a u l t  is the time when the stuck fault occurs.
Define the sampling time as Δ t , then u f a u l t u 0 f a u l t in Equation (41) can be rewritten as
u f a u l t u 0 f a u l t = u t f a u l t + ( i + 1 ) Δ t f a u l t u t f a u l t + i Δ t f a u l t ( i = 1 , 0 , 1 , 2 , 3 )
Define the actuator fault impact factor:
Δ f a u l t = u f a u l t f 2 ( ω ) + G 2 ( ω ) u exp + G 2 ( ω ) u f a u l t ω = ω 0 , u exp = u exp 0 , u f a u l t = u 0 f a u l t ( u f a u l t u 0 f a u l t )
When t = t f a u l t , u f a u l t u 0 f a u l t = u ¯ s t u c k 0 ; so Equation (44) can be written as follows:
Δ f a u l t = u f a u l t f 2 ( ω ) + G 2 ( ω ) u exp + G 2 ( ω ) u f a u l t ω = ω 0 , u exp = u exp 0 , u f a u l t = u 0 f a u l t u ¯ s t u c k 0
Equation (45) indicates that the sudden actuator stuck fault impacts on the controller only at t f a u l t time.
When t t f a u l t , u f a u l t u 0 f a u l t = 0 ; so Equation (44) can be written as follows:
Δ f a u l t = u f a u l t f 2 ( ω ) + G 2 ( ω ) u exp + G 2 ( ω ) u f a u l t ω = ω 0 , u exp = u exp 0 , u f a u l t = u 0 f a u l t ( u f a u l t u 0 f a u l t ) = 0
Thus, INDI design can implement fault-tolerant control.
B. RMI-INDI-Based Angular Rate Fault-Tolerant Control Law Design
Considering the model uncertainties in the angular rate inner loop, the RMI-INDI control method is designed to achieve accurate control for angular rate. The dynamics model in the angular rate loop with model uncertainties is presented as follows:
ω ˙ m = f 2 m ( ω ) + G 2 m ( ω ) u 2 = f 2 ( ω ) + G 2 ( ω ) u 2 + Δ m = J 1 M 0 ω × J ω + J 1 C M u u 2 + Δ m
where ω = p , q , r T , u 2 = δ a , δ e , δ r T , f 2 ( ω ) J 1 M 0 ω × J ω , G 2 ( ω ) J 1 C M u . The first term f 2 ( ω ) in Equation (47) indicates the uncertainties of nonlinear cross-coupling existing in the model of the angular rate model. The second term G 2 ( ω ) represents the control effectiveness matrix. Δ m = Δ f 2 m ( ω ) + Δ G 2 m ( ω ) u 2 represents model uncertainties.
It can be found that in the INDI control law of Equation (13), the complex and variable aerodynamic moment M 0 in f 2 ( ω ) is absent, which means that the INDI controller is not sensitive to the aerodynamic moment of the aircraft. Compared with NDI, the robustness of the system has been improved.
According to the Taylor series expansion, Equation (47) can be simplified as
ω ˙ m I N D I = ω ˙ 0 + G 2 m ( ω 0 ) ( u 2 u 2 0 ) + Δ H . O . T = ω ˙ 0 + G 2 ( ω 0 ) ( u 2 u 2 0 ) + Δ m + Δ H . O . T
Now the RMI method can be applied to estimate and compensate for the higher-order term Δ H . O . T .
Using Equation (29), the RMI-INDI system ω ^ ˙ R M I I N D I in the angular rate loop is described as follows:
ω ^ ˙ R M I I N D I = ω ˙ 0 + G 2 m ( ω 0 ) ( u 2 u 2 0 ) + Δ ^ R M I I N D I
Let ω ˙ a u g I N D I = ω ˙ I N D I ω ˙ m I N D I ω ^ ˙ R M I I N D I ; the augmentation system is presented as follows:
ω ˙ a u g I N D I = ω ˙ I N D I ω ˙ m I N D I ω ^ ˙ R M I I N D I = ω ˙ 0 + G 2 m ( ω 0 ) ( u 2 u 2 0 ) ω ˙ 0 + G 2 m ( ω 0 ) ( u 2 u 2 0 ) + Δ H . O . T ω ˙ 0 + G 2 m ( ω 0 ) ( u 2 u 2 0 ) + Δ R M I I N D I
According to Equation (30), the RMI factor Δ R M I I N D I in Equation (50) can be designed by the error integration as follows:
Δ ^ R M I I N D I = d i a g ( 1 τ k ) 1 k p ( ω ˙ m I N D I ω ^ ˙ R M I I N D I ) d t = d i a g ( 1 τ k ) 1 k p ( Δ H . O . T Δ ^ R M I I N D I ) d t
where d i a g ( 1 τ k ) 1 k p = d i a g ( 1 / τ 1 , 1 / τ 2 , , 1 / τ p ) , τ k > 0 , Δ ^ R M I I N D I is the estimated value of the robust multiple inversion factor Δ R M I I N D I , and the detailed RMI loop structure is shown in Figure 3.
As the above, the RMI-based estimator can effectively estimate the higher-order term Δ H . O . T in Equation (48), and by using Equation (50), the RMI-INDI-based angular rate control law u 2 is designed as follows:
u 2 = u 2 0 + Δ u 2 = u 2 0 + G 2 m 1 ( ω 0 ) ( v 2 ω ˙ 0 Δ ^ R M I I N D I )
where Δ u 2 = Δ δ a , Δ δ e , Δ δ r T , and the pseudo-inverse control input v 2 is designed as
v 2 = K P 2 e 2 ( t ) + K I 2 0 t e 2 ( τ ) d τ + K D 2 d e 2 ( t ) d t + ω ˙ d
where ω ˙ d is the derivative of the angular rate command ω d , e 2 ( t ) = ω d ω ^ R M I I N D I , ω ^ R M I I N D I is the system feedback angular rate, K ω = ( K P 2 , K I 2 , K D 2 ) are control parameters to be designed.

3.3. Parameters Optimization of the RL-Based RMI-INDI Controller

To achieve the optimization of control parameters of the civil aircraft with actuator faults, a fault-tolerant control parameter optimization method is proposed based on the DDPG algorithm. DDPG is chosen because it is simple and effective in continuous control tasks that employ deterministic strategies. This method can realize adaptive parameters optimization offline for the control system and significantly improves flight control performances under actuator faults.
Since the function approximation is used by the neural network, it cannot guarantee obtaining the global optimal solution of the Hamilton–Jacobi–Bellman (HJB) equation under the strong nonlinear characteristics of aircraft dynamics. However, the proposed RL framework is to find a set of high-performance and near-optimal control gains that minimize the cumulative cost of the reward function. The objective is to significantly improve the efficiency of tuning parameters and the performance of baseline controllers.
The DDPG algorithm can be regarded as a combination of deterministic policy gradients and deep neural network; it overcomes the defects that traditional RL can only output discrete actions. Therefore, the DDPG algorithm is applied to optimize the control parameters K ω = ( K P 2 , K I 2 , K D 2 ) in Equation (53).
Reinforcement learning takes the Markov decision process as its basic framework and makes the optimal decision through the continuous interaction between the controller system and the environment [34]. The basic model can be represented by s , a , p , R , where s represents the total set of states, a represents the set of all actions, p represents the probability of state transition, which is the probability that the state changes from s t to the next state s t + 1 under action a t , and can be defined as p ( s t + 1 s t , a t ) . R represents the total reward function. The accumulated reward value R t during the training process is
R t = r t + λ r t + 1 + λ 2 r t + 2 + = k = 1 λ k r t + k + 1
where λ is the discount factor, satisfying λ 0 1 , r t is the reward at the time t .
In the DDPG design, an actor is a deterministic policy function that can be expressed as π . The parameters of the actor network to be learned are represented as θ Q , and the parameters of the critic network to be learned are represented as θ μ . Each action of the aircraft control system is directly calculated by a t = μ ( s t ) and does not need to be sampled from the random policy [35]. In the state s t , the action a t is performed through policy π to obtain the next state s t + 1 and the reward value r t .
The actor network comprises two hidden layers by using the ReLU activation function, which includes 256 neurons and 128 ones, respectively. The tanh activation function is used by the output layer, and the output is scaled according to the action range. The critic network also comprises two hidden layers, which includes 256 neurons and 128 ones, respectively, and the state and action are concatenated before they are feedback to the second hidden layer.
The DDPG produce is as follows:
Step 1: Select the action according to the current policy: a t = μ ( s t θ μ ) .
Step 2: Perform the action a t , obtain the reward r t , and the environment state changes from s t to s t + 1 . Store transition s t , a t , r t , s t + 1 in D.
Step 3: Sample a random minibatch of N transitions { ( s i , a i , r i , s i + 1 ) } i = 1 , , N from D.
Step 4: Calculate using the target network: y i = r i + λ Q ( s i + 1 , μ ( s i + 1 θ μ ) θ Q )
Step 5: Update the current critic network by minimizing the target loss function:
L = 1 N i ( y i Q ( s i , a i θ Q ) ) 2
Step 6: Update the current network by calculating the sampled policy gradient:
θ μ μ s i 1 N i a Q ( s , a θ Q ) s = s i , a = μ ( s i ) θ μ μ ( s θ μ ) | s i
Step 7: Update the target network by using exponential smoothing:
θ Q ρ θ Q + ( 1 ρ ) θ Q θ μ ρ θ μ + ( 1 ρ ) θ μ
where ρ ( 0 , 1 ) represents the learning rate.
The Algorithm 1 is as follows:
Algorithm 1 DDPG
1: Set hyperparameters: soft update factor ρ , reward discount factor λ .
2: Randomly initialize weight parameter θ Q of the critic network Q ( s t θ Q ) and weight parameter θ μ of the actor network μ ( s t θ μ ) .
3: Initialize target networks Q and μ with weight parameters θ Q θ Q , θ μ θ μ .
4: Initialize replay buffer D.
5: for episode = 1, …, M do
6: Receive initial observation s 1 .
7: for t = 1, …, T do
8: Select the action according to the current policy: a t = μ ( s t θ μ ) .
9: Perform the action a t , obtain the reward r t , and the environment state changes from s t to s t + 1 .
10: Store transition s t , a t , r t , s t + 1 in D.
11: Sample a random minibatch of N transitions { ( s i , a i , r i , s i + 1 ) } i = 1 , , N from D.
12: Calculate using the target network: y i = r i + λ Q ( s i + 1 , μ ( s i + 1 θ μ ) θ Q ) .
13: Update the current critic network by minimizing the target loss function: L = 1 N i ( y i Q ( s i , a i θ Q ) ) 2
14: Update the current network by calculating the sampled policy gradient: θ μ μ s i 1 N i a Q ( s , a θ Q ) s = s i , a = μ ( s i ) θ μ μ ( s θ μ ) | s i
15: Update the target network by using exponential smoothing: θ Q ρ θ Q + ( 1 ρ ) θ Q θ μ ρ θ μ + ( 1 ρ ) θ μ
16: end for
17: end for
To ensure that the control parameters can achieve the desired flight control performances in cases of faults and uncertainties, thus ensuring flight safety, the DDPG method is used to optimize the control parameter K ω = ( K P 2 , K I 2 , K D 2 ) in Equation (53), The algorithm application process is as follows:
(1)
State space design
The state space represents aircraft state information and is the basis for control parameters evaluation. First, define the angular rate error:
e 2 ( t ) = e p e q e r = p d p q d q r d r
where ( e p , e q , e r ) T is the angular rate error, ( p d , q d , r d ) T is the angular rate, and ( p , q , r ) T is the angular rate real-time output.
The state space based on the angular rate real-time output and angular rate error can be represented as
s = e p , e q , e r , p , q , r
(2)
Action space design
Considering the parameter optimization of the aircraft control system in scenarios with faults and uncertainties, the action space can be defined as follows:
a = K ω = ( K P 2 , K I 2 , K D 2 )
The parameter K ω is dynamically offline optimized in each period of 0 T ; K ω is related to the action a directly.
(3)
Reward function design
In this part, the current state of the aircraft is effectively evaluated by designing a composite reward function, which is designed according to the state error, the control input and state constraints. To achieve the best result of the DDPG algorithm, the novel reward function is designed as follows:
r t = w e e 2 ( t ) 2 w u u 2 2 w ω tanh ( w α ω ( t ) 2 )
where the tracking error term w e e 2 ( t ) 2 is used to penalize the deviation from the target angular acceleration and directly reflects the control accuracy; the control input term w u u 2 2 is used to constrain excessive deflection of the control surface and to prevent actuator saturation; the state term w ω tanh ( w α ω ( t ) 2 ) is used to penalize excessive angular rates by a smooth saturation function, which can assure that the aircraft flies within the safety envelope while preventing instability in the learning process due to hard constraints; w e , w u , w ω , w α are configuration parameters of the reward function.
Remark 2. 
The hyperbolic tangent function t a n h ( · ) is chosen because it is smooth and bounded; these properties can effectively suppress excessive angular acceleration while discontinuous phenomenon is avoided.

3.4. Stability Analysis

By using RMI systems (50) and the control law (52), the derivative of the tracking error e ˙ 2 ( t ) can be obtained as follows:
e ˙ 2 ( t ) = ω ˙ d ω ^ ˙ R M I I N D I = ω ˙ d ( ω ˙ 0 + G 2 m ( ω 0 ) Δ u 2 + Δ R M I I N D I ) = ω ˙ d ω ˙ 0 + G 2 m ( ω 0 ) G 2 m 1 ( ω 0 ) ( v 2 ω ˙ 0 Δ ^ R M I I N D I ) + Δ R M I I N D I = ω ˙ d ω ˙ 0 + v 2 ω ˙ 0 + Δ R M I I N D I Δ ^ R M I I N D I = ω ˙ d v 2 + Δ ^ R M I I N D I Δ R M I I N D I
Substituting Equation (53) into Equation (62) yields
e ˙ 2 ( t ) = ω ˙ d ( K P 2 e 2 ( t ) + K I 2 0 t e 2 ( τ ) d τ + K D 2 d e 2 ( t ) d t ) ω ˙ d + Δ ^ R M I I N D I Δ R M I I N D I e ˙ 2 ( t ) = ( K P 2 e 2 ( t ) + K I 2 0 t e 2 ( τ ) d τ + K D 2 d e 2 ( t ) d t ) + Δ ^ R M I I N D I Δ R M I I N D I e ˙ 2 ( t ) = ( K P 2 e 2 ( t ) + K I 2 0 t e 2 ( τ ) d τ + K D 2 e ˙ 2 ( t ) ) + Δ ^ R M I I N D I Δ R M I I N D I ( I + K D 2 ) e ˙ 2 ( t ) = K P 2 e 2 ( t ) K I 2 0 t e 2 ( τ ) d τ + Δ ^ R M I I N D I Δ R M I I N D I
and then define η = 0 t e 2 ( τ ) d τ , P ( t ) = ( I + K D 2 ) , Δ e = Δ ^ R M I I N D I Δ R M I I N D I , where η ˙ = e 2 ( t ) , Δ e represents all uncertainties of the RMI system, and Equation (63) can be rewritten as follows:
e ˙ 2 ( t ) = P 1 ( t ) [ ( K P 2 e 2 ( t ) + K I 2 η Δ e ]
To ensure the feasibility of stability analysis, the following assumptions are introduced:
Assumption 2. 
There exists a constant d 0 such that Δ e d , and if the RMI estimator converges asymptotically, then d can be set arbitrarily small.
Since K P 2 , K I 2 , K D 2 is obtained through the RL dynamic training, it can be assumed that the K P 2 , K I 2 , K D 2 are time-varying bounded gain matrices, and its rate of change is also bounded during the training process.
Assumption 3. 
There exist constants and 0 < λ m λ M < such that for all t 0 ,
λ m I K P 2 , K I 2 , P ( t ) λ M I
where K P 2 , K I 2 , P ( t )  are symmetric gain matrices.
Assumption 4. 
There exists a constant ε 0 , and ε is sufficiently small such that
P ˙ ( t ) ε , K ˙ P 2 ( t ) ε , K ˙ I 2 ( t ) ε
Remark 3. 
In the actual implementation of RL, these assumptions can be ensured as follows: (1) DDPG gains can be designed in the prescribed range by projection clipping of its output; (2) low-pass filtering can constrain the change rate of the DDPG updated gains; (3) the parameters of the RMI estimator can be designed reasonably.
Based on the above assumptions, the stability of the closed-loop system is proven by the Lyapunov theory. The detailed proof process is shown in the Appendix A.

4. Simulation and Analysis

In this section, the civil aircraft is used as a research model with the structural parameters shown in Table 1, where x c g is the center of gravity position, and x c g r e f is the reference center of gravity position.
Actuator dynamics of the civil aircraft modeled as a first-order inertial transfer function [32], G a ( s ) = 1 / ( τ a s + 1 ) , where the constant value is τ a , is shown in Table 2.
The filter parameters of ω n 1 = ω n 2 = 25   rad / s and ξ n 1 = ξ n 2 = 0.8 are shown in Figure 3. The simulation experiments are carried out under the following initial flight conditions: V T = 90   m / s , H 0 = 2000   m , throttle δ t = 0.38 , initial attitude angle ( μ , α , β ) T = ( 0 , 3.2 , 0 ) T ( deg ) .

4.1. Simulation of Parameter Optimization

Learning parameters of the DDPG are configured as follows: the total number of rounds M = 1000; the period of a single round T = 60 s ; the discount factor is λ = 0.99 ; the replacement parameter ρ = 0.02 of the target network, the learning rates for the actor and critic are 10−3 and 10−4, respectively; the replay buffer size is 105 and the batch size is 64; and the reward function configuration parameters are set as w e = 10 , w u = 1 , w ω = 5 , w α = 0.01 .
The RL training is performed offline in a simulated environment. The DDPG is used to optimize the control parameters; the training process is shown in Figure 5.
To evaluate the sensitivity of the hyperparameter of the DDPG algorithm, a series of experiments are carried out by one hyperparameter changed at a time while others are remained. The final average episode reward is recorded for each training (where more than five seeds are used). The results are summarized in Table 3.
The test values of the learning rate including {10−4, 5 × 10−4, 10−3, 5 × 10−3}, and 10−3 (actor) and 10−4 (critic) are selected based on the convergence speed; the test range of the discount factor λ is [0.9, 0.99]; three network structures are tested: [128, 64], [256, 128], [512, 256]; [256, 128] is selected to balance capacity and efficiency.
From Table 3, it can be seen that the algorithm remains stable within the reasonable range of these parameters, where the discount factor λ has the greatest influence, and it remains stable within the range of λ [ 0.95 , 0.99 ] .
To verify the consistency of the RL training, a Monte Carlo simulation is carried out and the DDPG optimization process is tested by five different random seeds. The results in Table 3 show that all responses converge to a similar reward level after about 190 episodes, where a final standard deviation of the average reward is less than 8%. This demonstrates that the proposed RL algorithm is feasible and stable.
Scenario 1. 
Angular rate control loop with actuator faults.
In this scenario, stuck faults of the actuators are studied, the configuration is shown in Table 4.
As in Figure 6 and Table 5, when the actuator is stuck, the tracking performances of the angular rate by the RL-INDI with fault-tolerant control are significantly better than those by the LQR design. Therefore, the RL-INDI design has better control performances for sudden faults.

4.2. Simulation of Attitude Tracking Control with Model Uncertainties and Actuator Faults

Scenario 2. 
Attitude tracking control with model uncertainties
In this Scenario, the controller parameters and the RMI time constant matrix τ k of the RL-RMI-INDI controller are designed as shown in Table 6.
In this Scenario, an attitude tracking mission is carried out with uncertainties, the model uncertainties Δ m are set as 20% reduction in the values of model parameters m , I x x , I y y , and I x z as shown in Table 1 and the higher-order term Δ H . O . T = 0.5 sin ( t ) 0.5 sin ( t ) 0.5 sin ( t ) T . Using the proposed algorithm, the simulation results are shown in Figure 7 and Figure 8.
As in Figure 7 and Figure 8, the attitude tracking responses fluctuate when there are model uncertainties in the aircraft system all the time. Among which as Table 7, the root mean square errors of the RL-RMI-INDI are the smallest, that is E = [0.059, 0.026, 0.057] (deg), while those of the RMI-NDI are the largest, and the RMI-NDI and INDI controls are less robust in this case. Compared with the SMC, RMI-NDI, and INDI methods, the attitude tracking responses have lower overshoots by the RL-RMI-INDI, which shows it has better robust performance than the RMI-NDI and INDI designs.
Scenario 3. 
Attitude tracking control with model uncertainties and actuator faults
In this Scenario, the controller parameters and the RMI time constant matrix τ k of the RL-RMI-INDI controller are designed as Table 8.
The model uncertainties Δ m are set as a 20% reduction in the values of model parameters m , I x x , I y y , and I x z as shown in Table 1, the higher-order term Δ H . O . T = 0.5 sin ( t ) 0.5 sin ( t ) 0.5 sin ( t ) T and the specific time of fault occurrence and the angle of jamming are shown in Table 9. In this scenario, the simulation results by using the proposed algorithm are shown in Figure 9 and Figure 10.
As in Figure 9 and Figure 10, the attitude tracking response initially fluctuates when the faults in elevator, aileron, and rudder occur at times of 50 s; however, the closed-loop tracking system eventually converge. Among which as Table 10, the root mean square errors of the RL-RMI-INDI are the smallest, that is E = [0.078, 0.028, 0.070] (deg), while those of the RMI-NDI are the largest; the RMI-NDI and INDI controls are less robust in this case. Compared with the SMC, RMI-NDI, and INDI methods, RL-RMI-INDI has lower overshoots, which shows that it has more robust performance than the RMI-NDI and INDI designs.
The results show that RMI effectively estimates and compensates the high-order truncation terms through its robust multi-inverse integral structure, but they are ignored in the INDI method, and thus, the RMI design can significantly improve the tracking accuracy of the system. RL uses the DDPG algorithm to adaptively optimize the controller gains and dynamically adjusts the control parameters under actuator faults; thus, the tracking accuracy of the system is effectively improved. From above, the RL-RMI-INDI architecture is proposed, which combines the estimation ability of RMI, the optimization ability of RL, and the fault-tolerant ability of INDI. The overall architecture of the RL-RMI-INDI achieves the best performance, and the effectiveness of the three parts in the proposed framework is demonstrated by the above simulations.

4.3. Flight Simulator Test

In order to simulate the practical flight environment as realistically as possible, this study conducted tests on the flight simulator in the flight control laboratory in Shanghai Jiao Tong University; the results are shown in Figure 11.
The experiments are carried out under the following initial flight conditions: V T = 90   m / s , H 0 = 2000   m , and initial attitude angle of ( μ , α , β ) T = ( 0 , 0 , 0 ) T ( rad ) . The actuator (left aileron) fault stuck angle of 0.12 rad ( t > 50   s ) and the angular rate sensor noise/fault are set as shown in Table 11.
As in Figure 12, Figure 13 and Figure 14 and Table 12, the results indicate that the proposed RL-RMI-INDI controller, combined with the signal reconstruction method based on the Kalman filter, can still maintain satisfactory performance under sensor faults. This shows that the proposed method can realize not only fault-tolerant control of actuator faults but also reconstruct the sensor faults.

5. Conclusions

A RL-RMI-INDI method is proposed to achieve fault-tolerant control of a civil aircraft. The precise attitude tracking controller is designed through two cascaded loops of attitude angle and angular rate control. The RL-RMI-INDI controller not only stabilizes the attitude and angular rate control system of the aircraft in the presence of the uncertainties but also accurately tracks the ideal attitude angular of the civil aircraft under actuator faults. In addition, robust multi-inversion is designed to estimate the uncertainties. Stability analysis indicates that the closed-loop system is stable, by the Lyapunov theory. Two scenarios of uncertainties and faults have been simulated, demonstrating the robustness of the proposed RL-RMI-INDI control. Compared with RMI-NDI and INDI controllers, the RL-RMI-INDI control can achieve more robust tracking performances under uncertainties and faults. Therefore, the effectiveness and capability of the RL-RMI-INDI design have been demonstrated. Although the effectiveness of DDPG is demonstrated by this work, future research can involve a comprehensive comparison with more advanced reinforcement learning algorithms, such as the TD3 and the SAC, and a Hardware-In-the-Loop (HIL) semi-physical simulation platform test and eventually conduct actual flight experiments in the future.

Author Contributions

Conceptualization, Q.Z., W.L., C.Y. and J.C.; Methodology, Q.Z., W.L., C.Y. and J.C.; Software, Q.Z. and J.C.; Validation, Q.Z.; Investigation, Q.Z., W.L. and C.Y.; Writing – original draft, Q.Z.; Supervision, S.L.; Funding acquisition, S.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research is funded by the National Natural Science Foundation of China (No. 52272400).

Data Availability Statement

The original contributions presented in the study are included in the article, further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflict of interest.

Appendix A

Now construct a candidate composite Lyapunov function as follows:
V ( t , e 2 , η ) = V 1 ( t , e 2 ) + V 2 ( η ) + V 3 ( e 2 , η )
where V 1 ( t , e 2 ) = 1 2 e 2 T P ( t ) e 2 , V 2 ( η ) = 1 2 η 2 T Q η , V 3 ( e 2 , η ) = σ 2 e 2 + ϑ η 2 = σ 2 ( e 2 + ϑ η ) T ( e 2 + ϑ η ) , Q = q m I is a positive definite constant matrix and q m > 2 1 + λ M . q m , σ , and ϑ are constants to be undetermined.
From Assumption 3, according to the vector norm inequality, there exist positive constants c 1 , c 2 such that
c min ( e 2 2 + η 2 ) V c max ( e 2 2 + η 2 )
where c min = min ( λ m / 2 , q m / 2 ) , c max = max ( λ M / 2 + γ , q m / 2 + σ ϑ 2 ) .
The derivative of the Lyapunov function Equation (A1) is as follows:
V ˙ = V ˙ 1 + V ˙ 2 + V ˙ 3
where
V ˙ 1 = 1 2 e 2 T P ˙ ( t ) e 2 + e 2 T P ( t ) e ˙ 2
Substituting Equation (64) into Equation (A4) yields
V ˙ 1 = 1 2 e 2 T P ˙ ( t ) e 2 e 2 T K P 2 ( t ) e 2 e 2 T K I 2 ( t ) η + e 2 T Δ e ( t )
where
V ˙ 2 = η T Q η ˙ = η T Q e 2
and
V ˙ 3 = σ ( e 2 + ϑ η ) T ( e ˙ 2 + ϑ η ˙ )
Substituting Equation (64) and η ˙ = e 2 ( t ) into Equation (A7) yields
V ˙ 3 = σ ( e 2 + ϑ η ) T P 1 ( t ) K P 2 ( t ) e 2   σ ( e 2 + ϑ η ) T P 1 ( t ) K I 2 ( t ) η   + σ ϑ ( e 2 + ϑ η ) T e 2   + σ ( e 2 + ϑ η ) T P 1 ( t ) Δ e ( t )
Substituting Equations (A5), (A6), and (A8) into Equation (A3) yields
V ˙ = V ˙ 1 + V ˙ 2 + V ˙ 3 = 1 2 e 2 T P ˙ ( t ) e 2 e 2 T K P 2 ( t ) e 2 e 2 T K I 2 ( t ) η + e 2 T Δ e ( t )   + η T Q e 2   σ ( e 2 + ϑ η ) T P 1 ( t ) K P 2 ( t ) e 2   σ ( e 2 + ϑ η ) T P 1 ( t ) K I 2 ( t ) η   + σ ϑ ( e 2 + ϑ η ) T e 2   + σ ( e 2 + ϑ η ) T P 1 ( t ) Δ e ( t )
Let σ = q m 2 λ M , ϑ = q m 4 λ M , and then by Young’s inequality, it holds that the quadratic term bound of e 2 is
e 2 T K P 2 ( t ) e 2 σ e 2 T P 1 ( t ) K P 2 ( t ) e 2 + 1 2 e 2 T P ˙ ( t ) e 2 + σ ϑ e 2 T e 2 ( q m σ q m λ M + σ ϑ + ε 2 ) e 2 2
the quadratic term bound of η is
σ ϑ η T P 1 ( t ) K I 2 ( t ) η σ ϑ q m λ M η 2
For the boundary of the coupling term between e 2 and η , there exists a constant m such that
e 2 T K I 2 ( t ) η + η T Q e 2 σ ϑ e 2 T P 1 ( t ) K P 2 ( t ) η σ e 2 T P 1 ( t ) K I 2 ( t ) η + σ ϑ 2 η T e 2 q m 4 e 2 2 + m 2 q m η 2
and the bound of the disturbance term is
e 2 T Δ e + σ ( e 2 + ϑ η ) T P 1 ( t ) Δ e ( 1 2 + σ 2 λ m ) e 2 2 + σ ϑ 2 λ m η 2 + ( 1 2 + σ 2 λ m + σ ϑ 2 λ m ) Δ e 2
According to Equations (A10)–(A13), there exist positive constants c 1 , c 2 , c 3 such that
V ˙ c 1 e 2 2 c 2 η 2 + c 3 Δ e 2
Let c = min ( c 1 , c 2 ) , and Assumption 2 shows that
V ˙ c ( e 2 2 + η 2 ) + c 3 d 2
Combined with Equation (A2), the following can be obtained:
V c 2 ( e 2 2 + η 2 )
Substituting Equation (A16) into Equation (A15) yields
V ˙ c c 2 V + c 3 d 2
Following LaSalle’s theorem [36], all states of the closed-loop system are semi-global uniformly bounded, that is, V V ( t 0 ) e c c 2 t + c 3 d 2 c 2 c . Therefore, the closed-loop system of Equation (50) is stable.

References

  1. Wang, X.; Wang, S.; Yang, Z. Active fault-tolerant control strategy of large civil aircraft under elevator failures. Chin. J. Aeronaut. 2015, 28, 1658–1666. [Google Scholar] [CrossRef] [Scilit]
  2. Yu, X.; Jiang, J. A survey of fault-tolerant controllers based on safety-related issues. Annu. Rev. Control 2015, 39, 46–57. [Google Scholar] [CrossRef] [Scilit]
  3. Oliveira, L.; Leite, V.; Bento, A. Robust granular feedback linearization. In IEEE International Conference on Fuzzy Systems (FUZZ-IEEE); IEEE: New York, NY, USA, 2019; pp. 1–6. [Google Scholar]
  4. Franco, A.L.D.; Bourlès, H. Robust Nonlinear Control Associating Robust Feedback Linearization and H∞ Control. IEEE Trans. Autom. Control 2006, 51, 1200–1207. [Google Scholar] [CrossRef]
  5. Farlane, D.M.; Glover, K. A loopshaping design procedure using H∞ synthesis. IEEE Trans. Autom. Control 1992, 37, 759–769. [Google Scholar] [CrossRef] [Scilit]
  6. Guillard, H.; Bourles, H. Robust feedback linearization. In Proceedings of the 14th International Symposium on Mathematical Theory of Networks and Systems, Perpignan, France, 19–23 June 2000. [Google Scholar]
  7. Lavergne, F.; Villaume, F.; Jeanneau, M. Nonlinear robust Autoland. In Proceedings of the AIAA Guidance, Navigation, and Control Conference and Exhibit, San Francisco, CA, USA, 15–18 August 2005; p. AIAA 2005-5848. [Google Scholar]
  8. Chakraborty, U.K. Advances in Differential Evolution; Springer Science & Business Media: New York, NY, USA, 2008. [Google Scholar]
  9. Oliveira, L.; Bento, A.; Leite, V. Robust evolving granular feedback linearization. In International Fuzzy Systems Association World Congress; Springer International Publishing: Cham, Switzerland, 2019; pp. 442–452. [Google Scholar]
  10. Lima, E.; Gomide, F.; Ballini, R. Participatory evolving fuzzy modeling. In International Symposium on Evolving Fuzzy Systems; IEEE: New York, NY, USA, 2006; pp. 36–41. [Google Scholar]
  11. Li, Y.; Sun, L.; Qu, X. Acceleration measurement-based incremental nonlinear flight control for air-breathing hypersonic vehicles. Aerosp. Sci. Technol. 2016, 58, 235–247. [Google Scholar] [CrossRef] [Scilit]
  12. Lungu, M. Backstepping and dynamic inversion combined controller for auto-landing of fixed wing UAVs. Aerosp. Sci. Technol. 2020, 96, 105526. [Google Scholar] [CrossRef] [Scilit]
  13. Mullen, J.; Bailey, S.C.C.; Hoagg, J.B. Filtered dynamic inversion for altitude control of fixed-wing unmanned air vehicles. Aerosp. Sci. Technol. 2016, 54, 241–252. [Google Scholar] [CrossRef] [Scilit]
  14. Wu, G.; Meng, X.; Wang, F. Improved nonlinear dynamic inversion control for a flexible air-breathing hypersonic vehicle. Aerosp. Sci. Technol. 2018, 78, 734–743. [Google Scholar] [CrossRef] [Scilit]
  15. Lungu, M. Auto-landing of UAVs with variable centre of mass using the backstepping and dynamic inversion control. Aerosp. Sci. Technol. 2020, 103, 105912. [Google Scholar] [CrossRef] [Scilit]
  16. Cheng, H.H.; Liu, S.Q.; Ma, X.J. Neural network incremental dynamic inversion-based trajectory tracking control of an aircraft with pilot mishandling. J. Aeronaut. Astronaut. Aviat. 2021, 53, 43–56. [Google Scholar]
  17. Liu, S.Q.; Whidborne, J.F.; Lyv, W.Z.; Zhang, Q. Observer based incremental backstepping terminal sliding-mode control with learning rate for a multi-vectored propeller airship. Aerosp. Sci. Technol. 2023, 140, 108490. [Google Scholar] [CrossRef] [Scilit]
  18. Rudin, K.; Ducard, G.J.J.; Siegwart, R.Y. Active fault-tolerant control with imperfect fault detection information: Applications to UAVs. IEEE Trans. Aerosp. Electron. Syst. 2019, 56, 2792–2805. [Google Scholar] [CrossRef] [Scilit]
  19. Fourlas, G.K.; Karras, G.C. A survey on fault diagnosis and fault-tolerant control methods for unmanned aerial vehicles. Machines 2021, 9, 197. [Google Scholar] [CrossRef] [Scilit]
  20. Solihin, M.I.; Tack, L.F.; Kean, M.L. Tuning of PID controller using particle swarm optimization (PSO). Int. J. Adv. Sci. Eng. Inf. Technol. 2011, 1, 458–461. [Google Scholar] [CrossRef] [Scilit]
  21. Han, J.; Zhu, Z.; Jiang, Z. Simple PID parameter tuning method based on outputs of the closed loop system. Chin. J. Mech. Eng. 2016, 29, 465–474. [Google Scholar] [CrossRef] [Scilit]
  22. Wang, Z.; Wang, Q.; He, D. An improved particle swarm optimization algorithm based on fuzzy PID control. In Proceedings of the 4th International Conference on Information Science and Control Engineering (ICISCE), Changsha, China, 21–23 July 2017; IEEE: New York, NY, USA, 2017; pp. 835–839. [Google Scholar]
  23. Jiang, T.; Li, J.; Huang, K. Longitudinal parameter identification of a small unmanned aerial vehicle based on modified particle swarm optimization. Chin. J. Aeronaut. 2015, 28, 865–873. [Google Scholar] [CrossRef] [Scilit]
  24. Yu, M.P.; Chong, S.H. Deploying UAV based on Reinforcement Learning for Throughput Maximization in UAV Environments. J. Kiise 2020, 47, 700–706. [Google Scholar]
  25. Mo, H.; Ghulam, F. Nonlinear and adaptive intelligent control techniques for quadrotor UAV—A Survey. Asian J. Control 2019, 21, 989–1008. [Google Scholar] [CrossRef] [Scilit]
  26. Ma, J.; Peng, C. Adaptive model-free fault-tolerant control based on integral reinforcement learning for a highly flexible aircraft with actuator faults. Aerosp. Sci. Technol. 2021, 119, 107204. [Google Scholar] [CrossRef] [Scilit]
  27. Yang, Q.; Li, H.; Ruan, Z.; Fan, B.; Sam Ge, S. Reinforcement Learning-Based Fault-Tolerant Control of Uncertain Strict-Feedback Nonlinear Systems with Intermittent Actuator Faults. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 10464–10478. [Google Scholar] [CrossRef] [Scilit]
  28. Xie, N.; Lv, S.; Huang, Z.; Shen, H. Reinforcement learning-based fault-tolerant control of nonlinear servo systems with performance guarantees. Nonlinear Dyn. 2026, 114, 120. [Google Scholar] [CrossRef] [Scilit]
  29. Liu, B.; Fu, Z.; Hua, Z.; Fang, Z.; Jing, F.; Yao, J. Reinforcement-Learning-Based Cooperative Fault-Tolerant Control for Multiactuator System with Uncertain Parameters and State Constraints. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 17084–17095. [Google Scholar] [CrossRef] [Scilit]
  30. Bacon, B.; Gregory, I. General equations of motion for a damaged asymmetric aircraft. In Proceedings of the AIAA Atmospheric Flight Mechanics Conference and Exhibit, Hilton Head, SC, USA, 20–23 August 2007; p. AIAA 2007-6306. [Google Scholar]
  31. Li, Y.; Liu, X.; Lu, P. Angular acceleration estimation-based incremental nonlinear dynamic inversion for robust flight control. Control Eng. Pract. 2021, 117, 104938. [Google Scholar] [CrossRef] [Scilit]
  32. Lu, P.; Van, K.E.J.; De, V.C. Aircraft fault-tolerant trajectory control using incremental nonlinear dynamic inversion. Control Eng. Pract. 2016, 57, 126–141. [Google Scholar] [CrossRef] [Scilit]
  33. Liu, S.Q.; Lyu, W.Z.; Zhang, Q. Neural-Network-Based Incremental Backstepping Sliding Mode Control for Flying-Wing Aircraft. J. Guid. Control Dyn. 2025, 48, 600–614. [Google Scholar] [CrossRef] [Scilit]
  34. Zhang, Q.; Liu, S.Q.; Lyu, W.Z. Reinforcement Learning-Incremental Nonlinear Dynamic Inversion Based Intelligent Fault-Tolerant Control. In International Conference on Guidance, Navigation and Control; Springer Nature: Singapore, 2024; pp. 534–544. [Google Scholar]
  35. Bouhamed, O.; Ghazzai, H.; Besbes, H. Autonomous UAV navigation: A DDPG-based deep reinforcement learning approach. In Proceedings of the IEEE International Symposium on circuits and systems (ISCAS), Seville, Spain, 12–14 October 2020; IEEE: New York, NY, USA, 2020; pp. 1–5. [Google Scholar]
  36. Lyu, W.; Zhang, Q.; Yang, C.; Chen, J.; Liu, S.; Whidborne, J.F. LPV-Based Adaptive Robust-Augmented Control for a Blended-Wing-Body Aircraft Under Icing, Faults, Disturbances, and Uncertainties. IEEE Trans. Autom. Sci. Eng. 2026, 23, 612–630. [Google Scholar] [CrossRef] [Scilit]
Figure 1. The closed-loop control system with uncertainties.
Figure 1. The closed-loop control system with uncertainties.
Aerospace 13 00315 g001
Figure 2. The RL-RMI-INDI-based control structure of the civil aircraft.
Figure 2. The RL-RMI-INDI-based control structure of the civil aircraft.
Aerospace 13 00315 g002
Figure 3. Block of the RMI-INDI controller.
Figure 3. Block of the RMI-INDI controller.
Aerospace 13 00315 g003
Figure 4. The schematic diagram of the actuator faults.
Figure 4. The schematic diagram of the actuator faults.
Aerospace 13 00315 g004
Figure 5. DDPG training process for angle rate control loop.
Figure 5. DDPG training process for angle rate control loop.
Aerospace 13 00315 g005
Figure 6. Angle rate tracking responses under actuator faults: (a) the output tracking responses by LQR Control; (b) the output tracking responses by the RL-INDI Control.
Figure 6. Angle rate tracking responses under actuator faults: (a) the output tracking responses by LQR Control; (b) the output tracking responses by the RL-INDI Control.
Aerospace 13 00315 g006
Figure 7. Attitude angle tracking responses under model uncertainties: (a) kinematic roll angle tracking responses; (b) angle of attack tracking responses; (c) sideslip angle tracking responses.
Figure 7. Attitude angle tracking responses under model uncertainties: (a) kinematic roll angle tracking responses; (b) angle of attack tracking responses; (c) sideslip angle tracking responses.
Aerospace 13 00315 g007aAerospace 13 00315 g007b
Figure 8. Angle rate tracking responses under model uncertainties: (a) pitch angle rate tracking responses; (b) roll angle rate tracking responses; (c) yaw angle rate tracking responses.
Figure 8. Angle rate tracking responses under model uncertainties: (a) pitch angle rate tracking responses; (b) roll angle rate tracking responses; (c) yaw angle rate tracking responses.
Aerospace 13 00315 g008aAerospace 13 00315 g008b
Figure 9. Attitude angle tracking responses under model uncertainties and faults: (a) kinematic roll angle tracking responses; (b) angle of attack tracking responses; (c) sideslip angle tracking responses.
Figure 9. Attitude angle tracking responses under model uncertainties and faults: (a) kinematic roll angle tracking responses; (b) angle of attack tracking responses; (c) sideslip angle tracking responses.
Aerospace 13 00315 g009aAerospace 13 00315 g009b
Figure 10. Angle rate tracking responses under model uncertainties and faults: (a) pitch angle rate tracking responses; (b) roll angle rate tracking responses; (c) yaw angle rate tracking responses.
Figure 10. Angle rate tracking responses under model uncertainties and faults: (a) pitch angle rate tracking responses; (b) roll angle rate tracking responses; (c) yaw angle rate tracking responses.
Aerospace 13 00315 g010aAerospace 13 00315 g010b
Figure 11. The civil aircraft simulator.
Figure 11. The civil aircraft simulator.
Aerospace 13 00315 g011
Figure 12. Flight process diagram of the simulator: (a) the aircraft in a rightward flight state; (b) the aircraft in a level flight state; (c) the aircraft in a leftward flight state.
Figure 12. Flight process diagram of the simulator: (a) the aircraft in a rightward flight state; (b) the aircraft in a level flight state; (c) the aircraft in a leftward flight state.
Aerospace 13 00315 g012
Figure 13. Attitude angle tracking responses.
Figure 13. Attitude angle tracking responses.
Aerospace 13 00315 g013
Figure 14. Angle rate tracking responses.
Figure 14. Angle rate tracking responses.
Aerospace 13 00315 g014
Table 1. Parameters for the civil aircraft.
Table 1. Parameters for the civil aircraft.
ParameterValuesUnit
m 4500 kg
S 24.99 m 2
c ¯ 1.9910 m
b 13.32500 m
x c g 6.340 m
x c g r e f 6.500 m
I x x 1.12 × 104 kg m 2
I y y 2.29 × 104 kg m 2
I z z 3.20 × 104 kg m 2
I x z 1.93 × 103 kg m 2
I x y , I y z 0 kg m 2
Table 2. Actuator parameters.
Table 2. Actuator parameters.
ActuatorMinimum Position
Values (rad)
Maximum Position
Values (rad)
τ a ( s )
Elevator−0.360.3614
Aileron−0.660.66
Rudder−0.400.40
Table 3. Hyperparameter sensitivity results.
Table 3. Hyperparameter sensitivity results.
HyperparameterTested ValuesFinal Reward (mean ± std)
Actor LR5 × 10−4, 1 × 10−3, 5 × 10−3−125.3 ± 4.2, −120.1 ± 3.8, −122.5 ± 5.1
Critic LR5 × 10−5, 1 × 10−4, 5 × 10−4−122.8 ± 4.5, −120.1 ± 3.8, −124.2 ± 4.9
Discount λ 0.95, 0.99, 0.995−124.5 ± 4.7, −120.1 ± 3.8, −123.8 ± 5.2
Hidden neurons[128, 64], [256, 128], [512, 256]−121.8 ± 4.1, −120.1 ± 3.8, −119.5 ± 4.3
Table 4. Actuator faults.
Table 4. Actuator faults.
Time IntervalActuatorStuck Angle (rad)
t > 50   s Aileron (left)0.34
t > 50   s Elevator (left)0.52
t > 50   s Rudder (upper)0.2
Table 5. Performance comparative data under actuator faults.
Table 5. Performance comparative data under actuator faults.
ControllerVariableOvershoot (deg/s)Settling Time (s)
LQRp−13.21.2
q−0.90.23
r−1.91.3
RL-INDIp−30.3
q−0.750.2
r−0.50.16
Table 6. Controller parameters and time constant matrix.
Table 6. Controller parameters and time constant matrix.
( K P 1 , K I 1 , K D 1 ) T ( K P 2 , K I 2 , K D 2 ) T τ k
K P 1 ( p , q , r ) K I 1 ( p , q , r ) K D 1 ( p , q , r ) = 10 10 10 0.25 0.25 0.25 0.5 0.5 0.5 K P 2 ( p , q , r ) K I 2 ( p , q , r ) K D 2 ( p , q , r ) = 25 25 50 10 5 15 1 0.5 2.5 20 0 0 0 20 0 0 0 6
Table 7. Performance comparative data under uncertainties.
Table 7. Performance comparative data under uncertainties.
ControllerVariableOvershoot (deg)RMSE (deg)
SMC μ −1.9230.094
α 0.3940.071
β 0.7200.093
RMI-NDI μ −1.0530.205
α 0.2340.029
β 0.7350.065
INDI μ −2.0120.100
α 0.3520.075
β 0.6200.085
RL-RMI-INDI μ 0.7510.059
α 0.0770.026
β 0.6020.057
Table 8. Controller parameters and time constant matrix.
Table 8. Controller parameters and time constant matrix.
( K P 1 , K I 1 , K D 1 ) T ( K P 2 , K I 2 , K D 2 ) T τ k
K P 1 ( p , q , r ) K I 1 ( p , q , r ) K D 1 ( p , q , r ) = 10 10 10 0.25 0.25 0.25 0.5 0.5 0.5 K P 2 ( p , q , r ) K I 2 ( p , q , r ) K D 2 ( p , q , r ) = 25 25 25 15 15 15 1 0.5 1 20 0 0 0 20 0 0 0 6
Table 9. Actuator faults.
Table 9. Actuator faults.
Time IntervalActuatorStuck Angle (rad)
t > 50   s Aileron (left)0.28
t > 50   s Elevator (left)0.13
t > 50   s Rudder (upper)0.19
Table 10. Performance comparative data under uncertainties and actuator faults.
Table 10. Performance comparative data under uncertainties and actuator faults.
ControllerVariableOvershoot (deg)RMSE (deg)
SMC μ −1.1020.151
α −0.1320.043
β 0.6360.936
RMI-NDI μ −1.0860.233
α 0.2320.062
β 0.7360.094
INDI μ 2.0010.107
α 0.3020.053
β 0.6980.071
RL-RMI-INDI μ 0.9310.078
α 0.0720.028
β 0.6070.070
Table 11. The angular rate sensor noise/fault.
Table 11. The angular rate sensor noise/fault.
The Angular Rate Sensor Noise   σ s e n s o r 2 Bias Fault (rad/s)
p10−60.02 ( 45   s > t > 35   s )
q10−60.02 ( 30   s > t > 20   s )
r10−60.02 ( 15   s > t > 5   s )
Table 12. Simulator test data of RMSE for p, q, r.
Table 12. Simulator test data of RMSE for p, q, r.
p (rad/s)q (rad/s)r (rad/s)
7.6 × 10−47.4 × 10−47.3 × 10−4
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhang, Q.; Lyu, W.; Yang, C.; Chen, J.; Liu, S. Incremental Nonlinear Reinforcement Learning Control for a Civil Aircraft with Model Uncertainties and Actuator Faults. Aerospace 2026, 13, 315. https://doi.org/10.3390/aerospace13040315

AMA Style

Zhang Q, Lyu W, Yang C, Chen J, Liu S. Incremental Nonlinear Reinforcement Learning Control for a Civil Aircraft with Model Uncertainties and Actuator Faults. Aerospace. 2026; 13(4):315. https://doi.org/10.3390/aerospace13040315

Chicago/Turabian Style

Zhang, Qian, Weizhi Lyu, Congjie Yang, Jiaxin Chen, and Shiqian Liu. 2026. "Incremental Nonlinear Reinforcement Learning Control for a Civil Aircraft with Model Uncertainties and Actuator Faults" Aerospace 13, no. 4: 315. https://doi.org/10.3390/aerospace13040315

APA Style

Zhang, Q., Lyu, W., Yang, C., Chen, J., & Liu, S. (2026). Incremental Nonlinear Reinforcement Learning Control for a Civil Aircraft with Model Uncertainties and Actuator Faults. Aerospace, 13(4), 315. https://doi.org/10.3390/aerospace13040315

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop