Next Article in Journal
Optimization of Land-Based Impact Zones for Spent Rocket Stages Launched from the Baikonur Cosmodrome
Next Article in Special Issue
Online Trajectory Optimization Based on Pseudospectra Convex Optimization for Morphing Gliding Reentry Vehicles
Previous Article in Journal
Distant Retrograde Orbit and near Rectilinear Halo Orbit Determination and Time Synchronization Based on BeiDou Signals
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Fault-Tolerant Attitude Control of Flexible Spacecraft via Reinforcement Learning

School of Aeronautics and Astronautics, Shanghai Jiao Tong University, 800 Dongchuan Road, Shanghai 200240, China
*
Author to whom correspondence should be addressed.
Aerospace 2026, 13(7), 571; https://doi.org/10.3390/aerospace13070571
Submission received: 7 June 2026 / Revised: 19 June 2026 / Accepted: 19 June 2026 / Published: 24 June 2026

Abstract

This paper proposes an integrated attitude control framework for flexible spacecraft subject to external disturbances, rigid–flexible dynamic coupling, and actuator faults. The control framework combines the Twin Delayed Deep Deterministic Policy Gradient (TD3) reinforcement learning algorithm with an adaptive fault-tolerant (AFT) compensator. First, a rigid–flexible coupling dynamic model is formulated using Modified Rodrigues Parameters. Second, an observer-based TD3 attitude controller is designed, where a hierarchical reward function incorporating the observer-estimated flexible modal displacement η ^ is constructed to train the agent for simultaneous attitude convergence and vibration suppression. Third, a composite fault-tolerant control structure is developed by integrating the trained TD3 policy with an adaptive sliding mode compensator that handles both partial loss-of-effectiveness faults and time-varying additive faults. The proposed framework is evaluated under a progressive five-scenario uncertainty evaluation framework encompassing measurement noise, parameter mismatch, external disturbances, and actuator faults. Simulation results demonstrate that (i) the η ^ -augmented reward enables substantial improvements in vibration suppression over the baseline reward, achieving a better balance between pointing accuracy and vibration attenuation; (ii) under the most demanding fault scenario, the AFT compensator proves essential for precise convergence, and the composite TD3+AFT architecture achieves the best overall performance among the four compared control schemes.

1. Introduction

Flexible spacecraft attitude control represents a critical yet challenging problem in modern space missions, characterized by rigid–flexible coupling effects, multi-source uncertainties, and stringent performance requirements [1,2]. Unlike rigid spacecraft, flexible spacecraft feature large lightweight appendages (such as solar arrays and antenna reflectors) that introduce low-frequency elastic vibrations intrinsically coupled with rigid-body attitude motion, degrading pointing accuracy and risking persistent oscillations during maneuvers [2].
Recent advances have addressed flexible vibration suppression, disturbance rejection, actuator fault tolerance, and agile maneuvering through disturbance observers, sliding mode control (SMC), input shaping, and prescribed performance control. Observer-based and SMC approaches include Golestani et al.’s fixed-time disturbance observer with prescribed performance control [1], Shahid et al.’s composite anti-unwinding finite-time adaptive sliding mode with a modal vibration observer [3], and Dian et al.’s scheduled input shaping with sinusoidal path planning and a nonlinear disturbance observer [4]. For synchronous maneuver and vibration suppression, He and Cao designed cooperative linear quadratic regulator (LQR) and proportional-derivative (PD) with input shaping [2], Zhong et al. employed the internal model principle to asymptotically reject sinusoidal disturbances without modal measurements [5], and Song and Agrawal experimentally validated pulse width pulse frequency (PWPF) modulation with positive position feedback to attenuate thruster-induced vibrations [6]. Robust and adaptive approaches addressing model uncertainties include Yu et al.’s gain-scheduled H control for large-angle maneuvers [7], Angeletti et al.’s end-to-end robust design via linear fractional transformation [8], and Khoroshylov et al.’s adaptive notch-filter strategies for space-based synthetic aperture radar (SAR) platforms [9].
Deep reinforcement learning (DRL) has emerged as a broader trend in spacecraft guidance, navigation, and control, as surveyed in recent reviews [10,11,12]. Among DRL algorithms, the Twin Delayed Deep Deterministic policy gradient (TD3) [13] is selected in this work for its deterministic policy formulation well-suited for continuous torque commands, off-policy sample efficiency critical for computationally expensive flexible-body simulations, and twin-critic architecture that mitigates value overestimation bias. TD3 has been extensively applied across aerospace scenarios including unmanned aerial vehicle (UAV) operations [14,15,16,17,18], spacecraft attitude control [19], hypersonic vehicle maneuvering [20], and aero-engine regulation [21], with algorithmic modifications such as proportional-integral-derivative (PID)-guided exploration [19], Steffensen value iteration (SVI) for faster convergence [22], frame stacking for energy-efficient UAV path planning [14], and 3D environment encoding for dynamic obstacle avoidance [16].
More recently, several works have explored the intersection of TD3/DRL with flexible spacecraft dynamics and fault-tolerant control. Shi et al. [23] applied TD3 to cooperative control of ultra-close flexible spacecraft formation flying. Li et al. [24] integrated a DRL-based disturbance observer with the fully-actuated system approach for dumbbell-shaped flexible spacecraft. In the direction of fault tolerance, El-Dalahmeh et al. [25] developed TD3-HD with Hindsight Experience Replay for attitude fault recovery, and Henna et al. [26] introduced expert-guided exploration for satellite fault-tolerant attitude control. Nevertheless, two gaps remain to be addressed. First, Shi et al. [23] and Li et al. [24] address flexible attitude control without considering actuator faults. Second, El-Dalahmeh et al. [25] and Henna et al. [26] address fault-tolerant control but treat the spacecraft as rigid, neglecting rigid–flexible coupling dynamics. To the best of the authors’ knowledge, no existing work has simultaneously addressed flexible vibration suppression and actuator fault tolerance within a unified TD3-based framework.
Motivated by this gap, this paper develops an integrated attitude control framework for flexible spacecraft stabilization, capable of simultaneously addressing external disturbances, rigid–flexible coupling dynamics, and actuator faults. The framework employs the TD3 algorithm integrated with adaptive fault-tolerant control.The main contributions of this work are: (1) A hierarchical phased reward function for TD3-based flexible spacecraft attitude control. Unlike prior reward designs that rely solely on attitude error [27] or lack vibration-aware shaping [23,24], the proposed reward comprises four components (base, convergence, smoothness, and terminal) implementing a phased convergence strategy, with the convergence term augmented by observer-estimated flexible modal displacement η ^ to explicitly guide oscillation suppression. (2) A composite fault-tolerant control structure integrating the trained TD3 agent with an adaptive sliding mode fault-tolerant (AFT) compensator, extending the integral-type sliding mode framework of [28] from rigid to flexible spacecraft. The sliding manifold and AFT compensator are reformulated with rigid–flexible coupling terms ( δ T η ¨ and δ T η ˙ ) using observer-estimated modal coordinates.
The source code and simulation scripts accompanying this work are openly available at https://github.com/EllencePeng/Aerospace-RL-Flexible-Spacecraft (accessed on 18 June 2026).
The organization of this paper is as follows: Section 1 presents the research background and motivation. Section 2 presents the flexible spacecraft dynamics, fault model, flexible state observer, and problem statement. Section 3 establishes the RL training framework. Section 4 designs the adaptive fault-tolerant controller. Section 5 verifies the proposed schemes via simulations. Section 6 concludes this paper by summarizing its main contributions and putting forward prospects for future research.

2. Preliminaries

2.1. Flexible Spacecraft Attitude Dynamics and Kinematics

Two reference frames are used throughout this paper: an inertial reference frame I and a body-fixed frame B attached to the spacecraft’s main body. The angular velocity vector ω R 3 represents the inertial angular velocity of the spacecraft with respect to I and expressed in B . The attitude orientation of B with respect to I is described by the Modified Rodrigues Parameters (MRPs) [29], denoted by p R 3 . The corresponding attitude kinematic model, derived by referring to the modeling method in [30], is given by
p ˙ = 1 4 1 p T p I 3 + 2 p × + 2 p p T ω = T ( p ) ω ,
where p R 3 represents the MRP vector, and I 3 denotes the 3 × 3 identity matrix. The compact form T ( p ) ω is adopted for brevity in subsequent equations. The symbol p × denotes the skew-symmetric matrix of p , defined as
p × = 0 p 3 p 2 p 3 0 p 1 p 2 p 1 0 .
Compared with conventional attitude parameterizations, MRPs provide a computationally efficient and numerically robust attitude description [29], well-suited for precise spacecraft attitude control simulation. MRPs are globally valid for eigenaxis rotations up to 360°, but encounter a singularity where p when the rotation angle reaches exactly 360°. This is readily handled by shadow-set switching. Since all maneuvers considered in this work involve rotation angles far below 360°, p remains well within the unit sphere and shadow-set switching is unnecessary here.
The attitude dynamics of flexible spacecraft is adopted from [31] in low Earth orbit (LEO), formulated as
J ω ˙ + δ T η ¨ = ω × J ω + δ T η ˙ + τ η ¨ + C η ˙ + K η = δ ω ˙ ,
where J R 3 × 3 denotes the inertial matrix of the spacecraft, τ R 3 is the control torque input, δ represents the coupling matrix between the rigid main body and flexible appendages, and η R m denotes the flexible modal coordinates vector, treated as dimensionless. m is the total number of flexible vibration modes, C = diag { 2 ξ i ω n i , i = 1 , , m } and K = diag { ω n i 2 , i = 1 , , m } are the modal damping and stiffness matrices respectively, ω n i is the natural frequency of the i-th flexible mode, and ξ i is the corresponding damping ratio.

2.2. Fault Model

Assume that the dynamic response characteristics of spacecraft actuators can be ignored. An actuator is defined as fault-free when its output signal is completely consistent with the input command. The typical actuator faults considered in this study are modeled based on the existing research framework as follows [28]:
τ = I 3 E ( t ) u + u ¯ ,
where E ( t ) = diag { e 1 ( t ) , e 2 ( t ) , e 3 ( t ) } R 3 × 3 represents the actuator efficiency loss matrix of the spacecraft, satisfying 0 e i ( t ) < 1 ( i = 1 , 2 , 3 ) . Specifically, the condition e i ( t ) = 0 means the i th actuator operates in a normal state, while 0 < e i ( t ) < 1 indicates the i th actuator suffers from partial efficiency loss without complete functional failure. Additionally, u ¯ R 3 denotes the bounded time-varying additive fault occurring in the actuator. On this basis, the nonlinear attitude dynamic model of the spacecraft, which integrates the actuator fault model (4) and external disturbance d , can be reconstructed as the following expression:
J ω ˙ + δ T η ¨ = ω × J ω + δ T η ˙ + I 3 E ( t ) u + u ¯ + d .
To lay a solid foundation for the design of the composite fault-tolerant controller, three reasonable assumptions are proposed for the actuator faults and external disturbances, as detailed below.
Assumption 1.
The external disturbance d is bounded, satisfying d d m , where d m is a positive scalar and · represents the Euclidean norm.
Assumption 2.
For the actuator efficiency loss matrix E ( t ) = diag { e 1 ( t ) , e 2 ( t ) , e 3 ( t ) } , the inequality 0 max { e 1 , e 2 , e 3 } e m < 1 holds, where e m is a positive scalar.
Assumption 3.
The time-varying additive actuator fault u ¯ is bounded, satisfying u ¯ u ¯ m , with u ¯ m being a positive constant.
Remark 1.
Assumptions 1–3 are consistent with the actual operating environment of most spacecraft systems. Assumption 1 states that the external disturbance d acting on the flexible spacecraft is norm-bounded, which is consistent with the fact that environmental perturbations (atmospheric drag, gravity gradient, and solar radiation pressure) are always limited in actual space missions. Assumption 2 describes the actuator partial loss of effectiveness fault, where the upper bound 0 e m < 1 ensures that the actuator still retains a certain control ability rather than complete failure. Assumption 3 restricts the additive actuator fault to be bounded, reflecting that the deviation or bias caused by component aging, signal drift or mechanical wear is always finite in real physical systems. These physically meaningful assumptions provide a practical prerequisite for the design and analysis of the fault-tolerant control scheme.

2.3. Flexible State Observer

In practice, the flexible modal variables η and η ˙ are not directly measurable by onboard sensors. To address this, the nonlinear observer developed by Zou et al. [30] is adopted. The observer is of the high-gain type: it exploits the rigid–flexible coupling through the matrix δ to infer the unmeasurable modal states from the available attitude measurement p . The correction terms θ θ 1 p ˜ and θ 2 θ 2 p ˜ , driven by the attitude estimation error p ˜ = p p ^ , propagate through the coupled dynamics to update the estimates of v ^ , η ^ , and ψ ^ . A single design parameter θ > 0 governs the convergence rate; by choosing a sufficiently large θ , the estimation errors ( p ˜ , v ˜ , η ˜ , ψ ˜ ) are guaranteed to converge to zero exponentially. The complete observer dynamics are given by
p ^ ˙ = v ^ + θ θ 1 p ˜ , v ^ ˙ = f ( p , v ^ , η ^ , ψ ^ ) + g τ + θ 2 θ 2 p ˜ , η ^ ˙ = ψ ^ δ P v ^ , ψ ^ ˙ = C ψ ^ K η ^ + C δ P v ^ ,
where v = p ˙ , ψ = η ˙ + δ ω , P = T 1 , g = T J ¯ 1 , J ¯ = J δ T δ and f ( p , v , η , ψ ) = T P ˙ v T J ¯ 1 ( P v ) × J ¯ P v + δ T ψ + T J ¯ 1 δ T C ψ + K η C δ P v . θ , θ 1 and θ 2 are positive constants, p ^ with p ^ ( 0 ) = p ( 0 ) and v ^ with v ^ ( 0 ) = 0 are estimates of p and v , respectively, p ˜ = p p ^ and ψ ^ , η ^ are estimates of ψ , η . The observer gains in this work are set as θ = 5.0 , θ 1 = 2.0 , and θ 2 = 5.0 .

2.4. Problem Statement

Problem 1.
For a flexible spacecraft system modeled in as (1) and (3), under concurrent external disturbances and actuator faults described by (4), design a composite control scheme combining reinforcement learning and adaptive fault-tolerant control to achieve attitude stabilization and vibration suppression.

3. TD3-Based Controller Design

This section formulates a baseline controller utilizing the TD3 algorithm for flexible spacecraft attitude stabilization under fault-free conditions. The TD3 algorithm is first briefly reviewed, followed by three design aspects that instantiate it for the problem at hand: the state and action representation, the reward function, and the network architectures.

3.1. TD3 Algorithm

TD3 is an off-policy actor–critic algorithm for continuous control, proposed by Fujimoto et al. [13] as an improvement over DDPG. It employs twin critic networks and delayed policy updates to mitigate value overestimation, and a deterministic policy that is well-suited for generating smooth continuous torque commands. The training procedure follows the standard TD3 protocol established in [13].

3.2. State and Action Representation

It is assumed that only the attitude and angular velocity of the flexible spacecraft are directly measurable, while the flexible modal displacement η is not directly measurable but can be estimated by the observer introduced in Section 2.3. Consequently, flexible modal coordinates are excluded from the state vector provided to the TD3 policy.
The state vector provided to the TD3 policy is defined as
s R L = p T ω T T R 6 ,
The action generated by the policy corresponds to the continuous control torque applied to the spacecraft body. The action vector is defined as
a = τ R 3 ,
where τ represents the control torque along the three principal body axes.
To ensure actuator feasibility, the control torque is subjected to saturation constraints,
τ u max ,
where u max = 1 N·m denotes the maximum allowable actuator torque, a value consistent with the torque capacity of reaction wheels typically employed on small-to-medium-class spacecraft.

3.3. Reward Function Design

The reward function implements a phased convergence strategy with four components:
R = R base + R convergence + R smoothness + R terminal
where the 1-norm of p is defined as p 1 = | p 1 | + | p 2 | + | p 3 | .
The base reward provides dense per-step signals during the large-error phase ( p 1 0.5 ):
R base = 1 τ 2 , if p 1 0.5 and ( p 1 < p prev 1 ) 1 τ 2 , if p 1 0.5 and ( ω 1 < ω prev 1 ) p 1 10 ω 1 τ 2 , otherwise
The piecewise monotonic-improvement condition rewards the agent with 1 τ 2 when the attitude or angular velocity error decreases, and penalizes all error sources otherwise. The coefficient 10 on ω 1 normalizes it to the scale of p 1 .
When p 1 < 0.5 , inverse-proportional rewards are activated to accelerate fine convergence:
R convergence , 1 = 1 p 1 + 0.01 + 0.05 ω 1 + 0.01 , p 1 < 0.5 0 , otherwise
The form 1 / ( p 1 + 0.01 ) is chosen because its gradient steepens as the error shrinks, naturally accelerating fine convergence. The + 0.01 offset prevents division by zero and caps the reward, avoiding destabilizing Q-value targets during critic training.
To explicitly suppress flexible vibrations, the observer-estimated modal displacement η ^ (Section 2.3) is incorporated into the convergence reward:
R convergence , 2 = 1 p 1 + 0.01 + 0.05 ω 1 + 0.01 + 0.05 η ^ 2 + 0.01 , p 1 < 0.5 0 , otherwise
The additional term uses the same inverse-proportional form. The weight 0.05 (vs. 1.0 for attitude) encodes a 20:1 priority ratio, reflecting that vibration suppression is subordinate to attitude convergence. The agents trained with these two formulations—agent A 0 (without η ^ ) and agent A (with η ^ )—are compared in Section 5.3.
Smoothness penalties discourage abrupt actuation:
R smoothness = 10 · τ τ prev 1 + 10 · τ 2 , p 1 < 0.01 0 , otherwise
where τ prev denotes the RL output at the previous timestep. The L1-norm penalizes torque chattering linearly without being dominated by occasional large transients; the quadratic penalty τ 2 activates only near convergence to prevent wasteful energy expenditure.
Upon termination, a graduated reward is assigned based on final convergence precision:
R terminal = 100 , p 1 or ω 1 rad / s 1000 , p 1 < 0.01 200 , 0.01 p 1 < 0.05 100 , 0.05 p 1 < 0.1 0 , otherwise
where p = max ( | p 1 | , | p 2 | , | p 3 | ) . The graduated tiers ( + 1000 , + 200 , + 100 ) provide a curriculum-like signal: partial credit during early training helps the agent discover the basin of attraction, while progressively tighter thresholds refine convergence as the policy improves. The 100 divergence penalty supplements the opportunity cost of forfeited future rewards.
All weights and thresholds are calibrated so that each component’s magnitude remains comparable at its intended phase.

3.4. Network Structures

Both the actor and critic networks follow the standard TD3 architecture of Fujimoto et al. [13], employing two hidden layers with 400 and 300 units, respectively. The network structures are summarized in Table 1.
The actor output is passed through a tanh activation scaled by u max to respect the saturation limits defined in Section 3.2. The critic concatenates the state and action at its input layer and outputs a scalar Q-value. All hidden layers use ReLU activations; no batch normalization or dropout is applied.

4. Fault-Tolerant Attitude Controller Design

While the reinforcement learning-based controller offers significant advantages in adaptability and reduces manual tuning, its limitations in handling specific actuator fault modes necessitate the development of a dedicated fault-tolerant control scheme. Building upon the integral-type sliding mode framework of [28], an adaptive fault-tolerant (AFT) controller is formulated in this section. The proposed composite structure extends the framework of [28] in two key aspects: (i) the nominal controller is replaced by the trained TD3 agent rather than the saturated PD law in [28], which was designed for rigid-body dynamics only; and (ii) the sliding manifold and AFT compensator are reformulated to accommodate rigid–flexible coupling terms ( δ T η ¨ and δ T η ˙ ) with observer-estimated modal coordinates, thereby enabling fault tolerance in the presence of flexible dynamics that are not addressed in [28].

4.1. Adaptive Fault-Tolerant Controller

The controller is structured as
u = u R L + u a N ,
where u R L is the nominal controller generated by the trained TD3 agent, and u a N is the fault-tolerant controller used for compensating actuator faults. The pretrained TD3 agent is integrated into the composite control framework, as shown in Figure 1.
Based on the standard sliding mode manifold design framework, the sliding manifold is constructed as follows:
s = Γ ω ( t ) ω ( t 0 ) t 0 t J 1 ω ( σ ) × ( J ω + δ T η ˙ ) δ T η ¨ + u RL d σ ,
where Γ R 3 × 3 denotes a constant matrix. The matrix Γ is selected such that the product Γ J 1 is non-singular (invertible). The unmeasurable flexible modal states η is calculated by their observer estimates η ^ from Section 2.3. It should be noted that at time instant t = t 0 , the sliding mode surface satisfies s ( ω ( t 0 ) , t 0 ) = 0 , thereby eliminating the sliding mode reaching phase entirely.
The fault-tolerant compensator is designed as
u a N = α ^ Γ J 1 T s Γ J 1 T s , if α ^ Γ J 1 T s ϵ α ^ 2 Γ J 1 T s ϵ , if α ^ Γ J 1 T s < ϵ
where ϵ is a small positive scalar alleviating undesirable chattering, and α ^ 0 is obtained by the following adaptive law:
α ^ ˙ = γ Γ J 1 T s λ α ^ , with α ^ ( 0 ) 0
where γ and λ denote positive scalar constants, and the second term λ α ^ serves to enhance robustness against external disturbances and unmodeled dynamic characteristics, while mitigating the excessive growth of the adaptive gain parameter. In this work, these parameters are set to γ = 10.0 and λ = 2.5 × 10 4 , with the initial adaptive gain α ^ ( 0 ) = 1.0 and the boundary layer thickness ϵ = 0.01 .
The fault-tolerant mechanism of the compensator (18) can be understood through the integral-type sliding mode (ISM) principle. By construction of the sliding manifold (17), the nominal rigid–flexible dynamics are embedded in the integral term; differentiating s cancels these nominal terms and yields (23), where only the fault, disturbance, and the deviation u u RL remain. When the compensator maintains s = 0 , the condition s ˙ = 0 enforces ( I 3 E ) u + u ¯ + d = u RL , i.e., the equivalent control exactly counterbalances the fault and disturbance effects. Substituting this relation back into (3) recovers the nominal fault-free dynamics driven solely by u RL . Hence, the actuator effectiveness loss E ( t ) , additive fault u ¯ , and disturbance d are all algebraically rejected on the sliding manifold—without requiring explicit fault identification. The compensator (18) enforces this sliding motion: its direction ( Γ J 1 ) T s / ( Γ J 1 ) T s is the steepest-descent direction for the Lyapunov function 1 2 s T s , and the adaptive gain α ^ in (19) self-tunes to dominate the unknown combined fault–disturbance magnitude, thereby guaranteeing reachability without prior knowledge of the fault bounds.

4.2. Stability Analysis

Lemma 1.
Based on the switching gain update rule in (19), α ^ is upper bounded. In other words, for all t > 0 , there exists a positive scalar α ¯ satisfying
α ^ α ¯ , α α ¯ , t > 0 ,
where α is expressed as α = 3 e m u max + u ¯ m + d m + ε 1 e m , with u max denoting the upper bound of the RL output, e m , d m , and u ¯ m defined in Assumptions 1–3, and ϵ being a bounded positive constant.
Proof. 
The proof follows the same lines as in the appendix of [28]. Therefore, it is omitted here. The essential idea is as follows. Consider the Lyapunov candidate V s = 1 2 s T s , whose derivative along the closed-loop trajectories satisfies V ˙ s ( 1 e m ) ( α ^ α ) ( Γ J 1 ) T s , where α is the finite constant defined above. To establish boundedness of α ^ , four mutually exclusive cases are examined based on whether α ^ exceeds α and whether ( Γ J 1 ) T s exceeds a small threshold related to the σ -modification parameter λ . In each case, either V ˙ s 0 or the comparison principle is invoked, and in all cases a uniform upper bound for α ^ is established. □
Theorem 1.
Consider the flexible spacecraft attitude kinematics and dynamics described by (1) and (3) in the presence of fault described by (4). Suppose that Assumptions 1–3 hold and only the boundedness of actuator faults/disturbances is known (their upper bounds are unknown). Under the action of the adaptive controller in (16) and the adaptive law in (19), the system trajectories are able to enter a bounded region around the sliding manifold s = 0 in finite time.
Proof. 
First, construct the candidate Lyapunov function that fuses the sliding manifold error and the adaptive gain estimation error:
V = 1 2 s T s + 1 e m 2 γ α ^ α ¯ 2 ,
where α ¯ is the upper bound of α ^ defined in Lemma 1, and γ > 0 is the adaptive gain. We analyze the time derivative V ˙ in two cases based on the boundary layer of the smoothed control u a N , as defined in Equation (18).
Case I: α ^ ( Γ J 1 ) T s ϵ
Taking the time derivative of Equation (17), the first-order derivative of the sliding surface is obtained as
s ˙ = Γ ω ˙ ( t ) J 1 ω ( t ) × ( J ω ( t ) + δ T η ˙ ( t ) ) δ T η ¨ ( t ) + u RL ( t ) .
Multiplying both sides of Equation (3) by J 1 and rearranging terms, the angular velocity derivative ω ˙ is solved as
ω ˙ = J 1 ω × J ω + δ T η ˙ δ T η ¨ + I 3 E ( t ) u + u ¯ + d .
Substituting Equation (22) into Equation (21) yields
s ˙ = Γ J 1 I 3 E ( t ) u + u ¯ + d u RL .
Taking the time derivative of V, and substituting Equation (23), we obtain
V ˙ = s T s ˙ + 1 e m γ α ^ α ¯ α ^ ˙ 1 e m α ¯ 3 e m u max + u ¯ m + d m 1 e m ( Γ J 1 ) T s λ ( 1 e m ) α ^ α ¯ α ^ .
From Lemma 1, we have the following key inequalities:
α ¯ 3 e m u max + u ¯ m + d m 1 e m ε 1 e m , | α ^ α ¯ | α ¯ .
Substituting these into Equation (24), we further simplify V ˙ to
V ˙ λ ( 1 e m ) | α ^ α ¯ | ε ( Γ J 1 ) T s + λ ( 1 e m ) | α ^ α ¯ | + 1 4 α ¯ 2 = δ 1 | α ^ α ¯ | δ 2 s + η 1 ,
where δ 1 = λ ( 1 e m ) , δ 2 = ε Γ J 1 , and η 1 = λ ( 1 e m ) α ¯ + 1 4 α ¯ 2 are positive scalars. Rewriting Equation (25) in terms of V (the square root of the Lyapunov function) for finite-time stability analysis:
V ˙ 2 γ δ 1 2 1 e m 1 e m 2 γ ( α ^ α ¯ ) 2 2 δ 2 1 2 s T s + η 1 δ m V 1 / 2 + η 1 ,
where δ m = 2 min γ δ 1 2 1 e m , δ 2 is a positive constant.
Case II: α ^ ( Γ J 1 ) T s < ϵ
For the case inside the boundary layer, the time derivative of V (Equation (20)) is derived as
V ˙ = 1 e m ϵ α ^ 2 ( Γ J 1 ) T s 2 + ( 1 e m ) α ^ ( Γ J 1 ) T s ( 1 e m ) α ^ 3 e m u max + u ¯ m + d m 1 e m ( Γ J 1 ) T s + ( 1 e m ) α ^ α ¯ ( Γ J 1 ) T s λ ( 1 e m ) α ^ 2 α ^ α ¯ .
The quadratic term 1 e m ϵ α ^ 2 ( Γ J 1 ) T s 2 + ( 1 e m ) α ^ ( Γ J 1 ) T s attains its maximum value ( 1 e m ) ϵ 4 at α ^ ( Γ J 1 ) T s = ϵ 2 . Following the same derivation as Case I and using Lemma 1, we obtain
V ˙ δ m V 1 / 2 + η 2 ,
where η 2 = 1 e m 4 λ α ¯ 2 + 4 λ α ¯ + ϵ is a positive scalar.
Combining Case I and Case II, we obtain a unified inequality for V ˙ that holds for all t > 0 :
V ˙ δ m V 1 / 2 + η m ,
where η m = max { η 1 , η 2 } = 1 e m 4 λ α ¯ 2 + 4 λ α ¯ + ϵ . From the practical finite-time stability theory for nonlinear systems, the decrease of V drives the system trajectories into the set V 1 / 2 η m ( 1 θ 0 ) δ m in finite time, where 0 < θ 0 < 1 is a positive scalar. For the sliding manifold s, this implies the ultimate boundedness (the system converges to a small neighborhood of s = 0 ):
s 2 η m ( 1 θ 0 ) δ m = ( 1 e m ) λ α ¯ 2 + 4 λ α ¯ + ϵ 4 ( 1 θ 0 ) · min λ γ ( 1 e m ) , ε Γ J 1 .
Hence, the trajectory of the system remains bounded within a small neighborhood surrounding the sliding manifold s = 0 over a finite time horizon. This concludes the proof. □

5. Simulation and Discussion

In this section, the training environment settings are introduced and hyperparameter configurations are summarized in Table 2. To systematically assess robustness, five evaluation scenarios (S0–S4) form a progressive uncertainty evaluation framework, as summarized in Table 3. Starting from a noise-free baseline (S0), each subsequent scenario introduces an additional source of non-ideality: measurement noise (S1), non-zero initial angular velocity (S2), parameter mismatch and external disturbance (S3), and finally concurrent actuator faults (S4). Detailed scenario configurations are provided in the notes accompanying Table 3.
This section validates the proposed control framework through numerical simulations, organized into three parts. First, the training setup, spacecraft physical parameters, and hyperparameter configurations are documented. Second, the trained TD3 agents are evaluated under fault-free conditions (Scenarios S0–S3), comparing the η ^ -augmented agent A against the baseline agent A 0 to assess the effectiveness of vibration-aware reward shaping. Third, the composite TD3+AFT controller is tested under concurrent actuator faults (Scenario S4) and benchmarked against PD, PD+AFT, and the standalone TD3 agent.

5.1. Training Settings

During the training, each component of the initial attitude p 0 is independently sampled from the uniform distribution over the interval [ 1 , 1 ] . No measurement noise, parameter perturbation or fault condition is introduced in the training phase. The environment is treated as ideal except for a sinusoidal disturbance d train ( t ) , as follows:
d train , i ( t ) = A i sin 2 π 40 t + ϕ i + 10 7 , i = 1 , 2 , 3 A i U ( 0 , 0.05 ) [ N · m ] , i = 1 , 2 , 3 ϕ i U ( 0 , 2 π ) [ rad ] , i = 1 , 2 , 3
The frequency 2 π 40 rad/s is chosen to be representative and to maintain the same order of magnitude as Equation (34). The amplitude A i is intentionally set substantially larger than the actual disturbance magnitude. The exaggerated disturbance magnitude and fully randomized phases are designed to expose the agent to a diverse set of disturbance realizations during training, thereby encouraging the learning of robust control policies that generalize beyond the specific disturbance patterns encountered in evaluation.
The inertia matrix J is configured as
J = 120 3 4 3 100 10 4 10 120 kg · m 2 .
Different from [32], where the off-diagonal components are neglected, this paper takes the inertial coupling into account to better match the practical engineering situation. The flexible parameters are selected with reference to [31]. The number of flexible modes m is set as 4, and the coupling matrix δ is set as
δ = 6.45637 1.27814 2.15629 1.25819 0.91756 1.67264 1.11687 2.48901 0.83674 1.23637 2.65810 1.12530 .
The flexible modal parameters are given by
ω n = 0.7681 1.1038 1.8733 2.5496 , ξ = 0.0056 0.0086 0.0128 0.0252 ,
where ω n denotes natural frequencies (rad/s) and ξ represents corresponding damping ratios.
The training procedure follows the standard TD3 protocol [13]. At each timestep, the agent observes the state s R L , t = [ p T , ω T ] T , selects an action a t = π ϕ ( s R L , t ) + ϵ t with exploration noise ϵ t N ( 0 , σ explore 2 I ) , and receives a reward computed by the phased reward function (13). Transitions are stored in an experience replay buffer of capacity 10 6 and sampled uniformly in minibatches for network updates. Training begins after a warm-up phase of 1000 steps to populate the buffer. The complete hyperparameter configuration is listed in Table 2.
It is worth noting that the later TD3 policy network itself only receives the 6-dimensional state vector [ p T , ω T ] T as defined in Section 3.2; the observer-estimated flexible states are used internally by the environment for reward computation and by the AFT control layer, and are never passed to the RL policy network.
The training is implemented in Python 3.10 using PyTorch 1.12.0. The flexible spacecraft environment is simulated with a fixed-step integration using SciPy’s solve_ivp function with the RK45 method at a time step of Δ t = 1 s. All simulations are performed on a workstation with an Intel Core i7-13700 CPU and an NVIDIA RTX 3060 GPU; typical training duration is approximately 12 h for 20,000 episodes.
The effectiveness of incorporating the observer-estimated flexible state η ^ into the reward function is verified through an ablation experiment that compares the control performance of two agents trained by different reward functions (12) and (13). For clear distinction, the agent trained using the reward function (12) is designated as agent A 0 , and the agent trained with the reward function (13) is denoted as agent A. The reward accumulation of the two agents is visualized in Figure 2. Both agents exhibit a rapid ascent after approximately 5000 episodes, indicating that the exploration phase transitions into policy exploitation around this point. Agent A 0 produces an overall smoother reward curve, whereas agent A exhibits slightly larger fluctuations but converges to a higher steady-state reward value. These differences stem from distinct reward structures: Agent A 0 optimizes solely with respect to attitude and angular velocity errors, yielding a simpler and more stationary reward signal; agent A additionally incorporates the observer-estimated flexible modal displacement η ^ , which enriches the learning signal with vibration-suppression information but simultaneously introduces greater stochasticity due to the coupling between the flexible dynamics and the rigid-body motion. The richer feedback enables agent A to attain a higher asymptotic reward at the cost of moderately increased training variance.

5.2. Test Scenario Settings

To systematically evaluate agent robustness under progressively challenging conditions, five evaluation scenarios (S0–S4) are constructed as a progressive uncertainty evaluation framework. The complete scenario configurations are summarized in Table 3, and each scenario is described in detail below.
All five scenarios share a common initial attitude p 0 = [ 0.3 , 0.1 , 0.2 ] , with a simulation duration of 200 s and a time step of 1 s. The target attitude is zero ( p = 0 , ω = 0 ) in all cases.
Scenario S0 (pure baseline) represents the simplest evaluation condition: no measurement noise, no parameter mismatch, no external disturbance, and no actuator faults. The spacecraft starts from rest ( ω 0 = 0 ) and the controller operates with perfect state information and a perfectly known plant model. This scenario establishes the nominal performance reference against which all subsequent, more challenging scenarios are compared.
Scenario S1 (measurement noise) introduces Gaussian white noise on both the MRP attitude and angular velocity measurements. The noise standard deviations are set to σ p = 8.0 × 10 5 [33] and σ ω = 5 × 10 5 rad/s [34]. All other conditions remain identical to S0. This scenario evaluates the agents’ tolerance to realistic sensor imperfections.
Scenario S2 (non-zero initial angular velocity) retains the measurement noise of S1 and additionally assigns a non-zero initial angular velocity ω 0 = [ 0.02 , 0.03 , 0.04 ] rad/s. The spacecraft thus possesses initial kinetic energy at the start of the maneuver, making the attitude reorientation more demanding than the rest-to-rest cases of S0 and S1.
Scenario S3 (parameter mismatch and external disturbance) layers model-plant mismatch and persistent external disturbance on top of the measurement noise of S1. The parameter mismatch perturbs three off-diagonal elements of the moment of inertia matrix (from nominal values J 12 = J 21 = 3 , J 13 = J 31 = 4 to J 12 = J 21 = 6 , J 13 = J 31 = 7 kg m2), while the flexible coupling matrix δ , natural frequencies ω n , and damping ratios ξ are all increased by + 10 % relative to their nominal values. The external disturbance is a multi-sinusoidal torque adopted from [35]:
d ( t ) = 0.2 × 10 3 · 3 cos ( 10 ω d t ) + 4 sin ( 3 ω d t ) 10 1.5 sin ( 2 ω d t ) + 3 cos ( 5 ω d t ) + 15 3 sin ( 10 ω d t ) 8 sin ( 4 ω d t ) + 10 Nm .
where ω d = 0.1 rad/s in this work. The initial angular velocity is zero. Scenario S3 represents the most demanding fault-free condition in the evaluation suite.
Scenario S4 (actuator faults) incorporates all the non-idealities of S3—measurement noise, parameter mismatch, and external disturbance—and adds a time-varying actuator fault profile. Each actuator experiences a partial loss-of-effectiveness fault at t = 10 s. Subsequently, at t = 50 s, these actuators are also subjected to an additive time-varying fault, which is injected into the spacecraft dynamics in an additive manner. The parameters in (4) are selected as follows:
e i ( t ) = 0 , t < 10 0.5 , t 10 u ¯ ( t ) = 0 , t < 50 0.3 + 0.05 sin ( t ) , t 50
The efficiency loss e i = 0.5 and the additive bias u ¯ = 0.3 together represent a severe fault condition relative to the maximum actuator torque u max = 1 N·m: the actuators lose half of their control effectiveness while being simultaneously subjected to a constant offset of 0.3 N·m, corresponding to static actuator anomalies such as sudden installation misalignment, mechanical wear, circuit zero-point drift, or long-term aging-induced output bias. The sinusoidal term 0.05 sin ( t ) , introduces a time-varying component into the additive fault term, capturing the fact that actuator faults may not remain constant but can evolve over time—for instance, due to mechanical/electrical periodic perturbations such as reaction wheel bearing periodic friction, thruster jet pulsation, or electromagnetic circuit ripple.
This scenario represents the most challenging condition in the evaluation suite, testing the agents’ resilience to concurrent sensor noise, modeling errors, environmental disturbances, and actuator degradation.
Scenarios S0–S3 are designed to compare the performance of agent A and agent A 0 , demonstrating the benefit of incorporating the observer-estimated flexible modal state η ^ into the reward function. Scenario S4 serves as the validation platform for the adaptive fault-tolerant (AFT) control framework, where four control schemes are evaluated and compared under concurrent actuator faults.

5.3. Fault-Free Scenarios

The performance of agent A and agent A 0 under S0–S3 is visualized in Figure 3, Figure 4, Figure 5 and Figure 6. Each figure presents the time histories of the MRP attitude, angular velocity, control torque output, and flexible modal displacement for both agents in a side-by-side comparison.
Across all four scenarios, agent A consistently exhibits noticeably cleaner trajectories with reduced oscillations and smaller vibration peaks, whereas agent A 0 often produces persistent back-and-forth chattering in the control output accompanied by sustained flexible-mode excitation. This difference is a direct consequence of the η ^ -augmented reward function (13): by rewarding the suppression of the observer-estimated modal displacement during training, agent A learns a policy that inherently avoids control actions that would excite the flexible modes. The resulting closed-loop behavior achieves simultaneous attitude regulation and vibration damping, yielding smoother state trajectories and a cleaner control effort. In contrast, agent A 0 , trained without any vibration-related feedback (12), optimizes only for attitude errors and may inadvertently induce and sustain flexible oscillations, which manifest as noisier signals and persistent torque chattering.
To provide a quantitative assessment, steady-state performance metrics (computed over t > 160 s) of both agents across S0–S3 are summarized in Table 4. The metrics include the steady-state infinity norms of the MRP attitude ( p ) and angular velocity ( ω ), the root mean square (RMS) of the flexible modal displacement vector norm ( η RMS) as a measure of residual vibration energy, the settling time of the flexible modal displacement ( η 0.01 ), and the RMS of the applied control torque (u RMS) as a measure of control effort. Here η 0.01 is defined as the earliest time after which the peak absolute modal displacement across all four flexible modes ( max ( | η 1 | , | η 2 | , | η 3 | , | η 4 | ) ) remains permanently below 0.01 , i.e., the vibration settling time. “—” denotes failure to converge within the 200 s evaluation horizon.
The quantitative results in Table 4 reveal several important findings. First, in terms of vibration suppression, agent A consistently and substantially outperforms agent A 0 across all four scenarios. Compared with agent A 0 , agent A reduces the η RMS by approximately 92.7 % (S0), 59.3 % (S1), 67.2 % (S2), and 88.0 % (S3). Correspondingly, compared with agent A 0 , agent A reduces η 0.01 (vibration settling time) by 68.4 % (S0); in all other scenarios, agent A 0 fails to converge within the 200 s evaluation horizon. This demonstrates that the η ^ -augmented reward function (13) provides a robust and generalizable vibration suppression capability that persists even as non-idealities accumulate.
Second, a consistent trade-off between attitude pointing accuracy and vibration suppression is observed. Although agent A achieves better vibration damping across all scenarios, agent A 0 attains marginally better pointing accuracy in S0–S2, at the cost of severely degraded vibration damping. However, in S3, the most demanding fault-free scenario, agent A surpasses agent A 0 across all metrics including attitude pointing ( p = 6.55 × 10 4 vs. 8.91 × 10 4 ), indicating that the η ^ -guided policy learns a more robust control strategy that does not sacrifice pointing accuracy under compound uncertainties.
Third, the torque expenditure of agent A is moderately higher than that of agent A 0 across S1–S2 (e.g., 5.80 × 10 2 vs. 4.89 × 10 2 Nm in S1), reflecting the additional control effort required for active vibration suppression. In S3, however, agent A achieves substantially lower torque RMS ( 6.33 × 10 2 vs. 0.194 Nm), demonstrating that under compound uncertainties the η ^ -guided policy simultaneously delivers superior vibration attenuation and higher control efficiency. In S0, agent A achieves significantly lower torque RMS ( 6.16 × 10 4 vs. 4.91 × 10 3 Nm), indicating that in the absence of perturbations, the learned policy achieves both superior vibration attenuation and higher control efficiency.
Overall, the ablation study across S0–S3 confirms that incorporating the observer-estimated flexible modal displacement η ^ into the reward function enables the agent to strike a better balance between attitude pointing accuracy and vibration suppression. The resulting policy generalizes effectively across diverse non-ideal conditions at a modest increase in control effort under perturbed scenarios.

5.4. Fault Scenario

Given that agent A demonstrated superior vibration suppression and robust performance across S0–S3 in the ablation study, its trained policy weights are adopted as the TD3 nominal controller in this section. Four control architectures are compared under Scenario S4: (i) pure PD ( u p d = k p p k d ω with k p = 30 , k d = 60 ), (ii) pure TD3, (iii) PD augmented with the AFT compensator (PD+AFT), and (iv) TD3 augmented with the AFT compensator (TD3+AFT).
To evaluate the adaptability of the proposed control framework under actuator fault conditions, Scenario S4 is employed, whose configuration is detailed in Section 5.2. This simulation validates the efficacy of the composite control framework proposed in Section 4.
The simulation outcomes of the four control architectures under S4 are illustrated in Figure 7. To further provide a quantitative assessment, key performance metrics are summarized in Table 5.
In Table 5, all steady-state metrics are computed over t > 160 s. Integrated Torque RMS is computed over the full trajectory. Peak η and peak ω capture the maximum values over the entire 200 s trajectory. Integrated η RMS is computed over the full trajectory as a proxy for total vibration energy. MRPs converging to 0.005 indicates the earliest time after which p remains below 0.005 (“—” denotes failure to converge within 200 s).
Several observations can be drawn from the quantitative comparison in Table 5. First, AFT augmentation is essential for precise convergence under actuator faults: pure PD and pure TD3 both fail to bring the MRP error below 0.005 within the 200 s evaluation horizon, whereas TD3+AFT achieves this threshold at 74 s and PD+AFT at 94 s. The proposed TD3+AFT architecture delivers the best overall pointing accuracy, reducing p by a factor of approximately 43.3 relative to pure TD3 ( 1.75 × 10 2 vs. 4.04 × 10 4 ) and by a factor of approximately 50.0 relative to pure PD ( 2.02 × 10 2 vs. 4.04 × 10 4 ).
Second, TD3 consistently outperforms PD across all transient and steady-state vibration metrics. In steady state, TD3 achieves η RMS of 8.43 × 10 3 compared to 1.37 × 10 2 for PD, representing a 38.5% reduction. Among AFT-augmented architectures, TD3+AFT achieves lower η RMS ( 1.04 × 10 2 vs. 1.17 × 10 2 for PD+AFT). In the transient phase, TD3 and TD3+AFT tie for the lowest peak η (both 0.117 ) and integrated η RMS (both 2.91 × 10 2 ), confirming that the learned TD3 policy inherently produces smoother maneuvers with reduced flexible mode excitation.
Third, a comparison between the two AFT-augmented methods shows that TD3+AFT achieves better pointing accuracy, convergence speed, and vibration suppression: 30.5% lower steady-state p ( 4.04 × 10 4 vs. 5.81 × 10 4 ), 21.3% faster MRP convergence (74 s vs. 94 s), and 11.1% lower steady-state η RMS ( 1.04 × 10 2 vs. 1.17 × 10 2 ). PD+AFT retains marginal advantages in peak angular velocity ( ω : 8.12 × 10 2 vs. 8.22 × 10 2 rad/s) and integrated torque RMS ( 1.112 vs. 1.118 N·m), though the differences are small. Overall, the learned TD3 nominal policy coordinates with the AFT compensator more effectively than the fixed-gain PD law, as the RL agent learns to generate control actions that complement the fault-tolerant layer rather than compete with it.
Regarding control effort, all four controllers operate within the u max = 1 N·m actuator saturation limit stated in Section 3.2. The steady-state torque RMS values are tightly clustered ( 1.042 1.045 N·m), and the integrated torque RMS also remains within a narrow band ( 1.082 1.118 N·m). The fact that TD3+AFT achieves substantially superior pointing and vibration performance while expending comparable control effort indicates that both the RL policy and the AFT layer contribute to control efficiency in complementary ways: the TD3 agent provides smooth nominal commands that minimize unnecessary actuation, while the AFT compensator injects targeted corrections only when faults are active.

6. Conclusions

This study focuses on the attitude maneuver problem of flexible spacecraft under external disturbances, flexible dynamics, and actuator faults. An integrated control framework combining TD3-based RL agent with an adaptive fault-tolerant compensator is proposed. Simulation results demonstrate that incorporating the observer-estimated flexible modal displacement into the reward function enables the RL agent to learn active vibration suppression while maintaining competitive attitude accuracy. Under actuator faults, the AFT compensator significantly improves control precision, and the composite TD3+AFT architecture achieves the best pointing accuracy, convergence speed, and vibration suppression among all configurations, while maintaining comparable control effort. The RL agent and AFT compensator operate in a complementary manner: the RL agent provides smooth nominal commands that inherently suppress flexible mode excitation, while the AFT layer injects targeted fault-tolerant corrections when needed.
Future work includes experimental validation on physical testbeds to assess performance under real sensing and actuation imperfections beyond the Gaussian noise models used in the present simulations. Additionally, extending the RL policy to support on-orbit incremental learning would enable real-time adaptation to evolving spacecraft parameters and fault conditions, further improving the controller’s generalization capability in complex operational environments.

Author Contributions

Conceptualization, Q.S.; methodology, Q.S.; software, Z.P.; validation, Z.P.; formal analysis, Z.P.; investigation, Z.P.; resources, Q.S.; data curation, Z.P.; writing—original draft preparation, Z.P.; writing—review and editing, Q.S.; visualization, Z.P.; supervision, Q.S.; project administration, Q.S.; funding acquisition, Q.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National Natural Science Foundation of China (No. U24B20157) and the Science and Technology Commission of Shanghai Municipality Key Technology R&D Plan (No. 25DZ3100703).

Data Availability Statement

The source code presented in this study is openly available at https://github.com/EllencePeng/Aerospace-RL-Flexible-Spacecraft (accessed on 18 June 2026).

Conflicts of Interest

The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

References

  1. Golestani, M.; Zhang, W.; Yang, Y.; Xuan-Mung, N. Disturbance observer-based constrained attitude control for flexible spacecraft. IEEE Trans. Aerosp. Electron. Syst. 2022, 59, 963–972. [Google Scholar] [CrossRef]
  2. He, G.; Cao, D. Dynamic Modeling and Attitude–Vibration Cooperative Control for a Large-Scale Flexible Spacecraft. Actuators 2023, 12, 167. [Google Scholar] [CrossRef]
  3. Shahid, F.; Luo, H.; Jiang, Y.; Noman Hasan, M. Finite-Time Adaptive Sliding Mode Fault-Tolerant Attitude Control for Flexible Spacecraft. IEEE Trans. Control Syst. Technol. 2025, 33, 1700–1711. [Google Scholar] [CrossRef]
  4. Dian, W.; Yunhua, W.; Hongyi, X.; Franco, B.Z.; Chen, X. Scheduled input shaping based attitude agile control for flexible spacecraft with vibration suppression. Aerosp. Sci. Technol. 2025, 164, 110355. [Google Scholar] [CrossRef]
  5. Zhong, C.; Chen, Z.; Guo, Y. Attitude Control for Flexible Spacecraft with Disturbance Rejection. IEEE Trans. Aerosp. Electron. Syst. 2017, 53, 101–110. [Google Scholar] [CrossRef]
  6. Song, G.; Agrawal, B.N. Vibration suppression of flexible spacecraft during attitude control. Acta Astronaut. 2001, 49, 73–83. [Google Scholar] [CrossRef]
  7. Yu, Y.; Meng, X.; Li, K.; Xiong, F. Robust control of flexible spacecraft during large-angle attitude maneuver. J. Guid. Control. Dyn. 2014, 37, 1027–1033. [Google Scholar] [CrossRef]
  8. Angeletti, F.; Iannelli, P.; Gasbarri, P.; Sabatini, M. End-to-end design of a robust attitude control and vibration suppression system for large space smart structures. Acta Astronaut. 2021, 187, 416–428. [Google Scholar]
  9. Khoroshylov, S.; Martyniuk, S.; Sushko, O.; Vasyliev, V.; Medzmariashvili, E.; Woods, W. Dynamics and attitude control of space-based synthetic aperture radar. Nonlinear Eng. 2023, 12, 20220277. [Google Scholar] [CrossRef]
  10. Izzo, D.; Märtens, M.; Pan, B. A survey on artificial intelligence trends in spacecraft guidance dynamics and control. Astrodynamics 2019, 3, 287–299. [Google Scholar] [CrossRef]
  11. Khoroshylov, S.V.; Redka, M.O. Deep learning for space guidance, navigation, and control. Space Sci. Technol. 2021, 27, 38–52. [Google Scholar]
  12. Silvestrini, S.; Lavagna, M. Deep learning and artificial neural networks for spacecraft dynamics, navigation and control. Drones 2022, 6, 270. [Google Scholar] [CrossRef]
  13. Fujimoto, S.; Hoof, H.; Meger, D. Addressing function approximation error in actor-critic methods. In Proceedings of the International Conference on Machine Learning, PMLR, Stockholm, Sweden, 10–15 July 2018; pp. 1587–1596. [Google Scholar]
  14. Hong, D.; Lee, S.; Cho, Y.H.; Baek, D.; Kim, J.; Chang, N. Energy-Efficient Online Path Planning of Multiple Drones Using Reinforcement Learning. IEEE Trans. Veh. Technol. 2021, 70, 9725–9740. [Google Scholar] [CrossRef]
  15. Abo Mosali, N.; Shamsudin, S.S.; Alfandi, O.; Omar, R.; Al-Fadhali, N. Twin Delayed Deep Deterministic Policy Gradient-Based Target Tracking for Unmanned Aerial Vehicle with Achievement Rewarding and Multistage Training. IEEE Access 2022, 10, 23545–23559. [Google Scholar] [CrossRef]
  16. Zhang, D.; Li, X.; Ren, G.; Yao, J.; Chen, K.; Li, X. Three-Dimensional Path Planning of UAVs in a Complex Dynamic Environment Based on Environment Exploration Twin Delayed Deep Deterministic Policy Gradient. Symmetry 2023, 15, 1371. [Google Scholar] [CrossRef]
  17. Yan, T.; Liu, C.; Gao, M.; Jiang, Z.; Li, T. A Deep Reinforcement Learning-Based Intelligent Maneuvering Strategy for the High-Speed UAV Pursuit-Evasion Game. Drones 2024, 8, 309. [Google Scholar] [CrossRef]
  18. Zhou, T.; Liu, Z.; Jin, W.; Han, Z. Intelligent maneuver decision-making for UAVs using the TD3–LSTM reinforcement learning algorithm under uncertain information. Front. Robot. AI 2025, 12, 1645927. [Google Scholar] [CrossRef] [PubMed]
  19. Zhang, Z.; Li, X.; An, J.; Man, W.; Zhang, G. Model-Free Attitude Control of Spacecraft Based on PID-Guide TD3 Algorithm. Int. J. Aerosp. Eng. 2020, 2020, 8874619. [Google Scholar] [CrossRef]
  20. Gao, M.; Yan, T.; Li, Q.; Fu, W.; Zhang, J. Intelligent Pursuit–Evasion Game Based on Deep Reinforcement Learning for Hypersonic Vehicles. Aerospace 2023, 10, 86. [Google Scholar] [CrossRef]
  21. Zhu, J.; Tang, W.; Dong, J. Design of Intelligent Controller for Aero-engine Based on TD3 Algorithm. Inf. Technol. Control 2023, 52, 1010–1024. [Google Scholar] [CrossRef]
  22. Cheng, Y.; Chen, L.; Chen, C.L.P.; Wang, X. Off-Policy Deep Reinforcement Learning Based on Steffensen Value Iteration. IEEE Trans. Cogn. Dev. Syst. 2021, 13, 1023–1032. [Google Scholar] [CrossRef]
  23. Shi, J.; Liu, X.; Cai, G.; Liu, F.; Sun, J.; Zhu, D. Reinforcement Learning for Cooperative Control of the Ultra-Close Flexible Spacecraft Formation Flying. IEEE Aerosp. Electron. Syst. Mag. 2026, 41, 54–68. [Google Scholar] [CrossRef]
  24. Li, Y.; Zhang, F.; Gao, Z. Attitude control for dumbbell-shaped spacecraft subject to parameter uncertainty via fully-actuated system approach and deep reinforcement learning. Int. J. Syst. Sci. 2026, 1–18. [Google Scholar] [CrossRef]
  25. El-Dalahmeh, G.; Jabbarpour, M.R.; Vo, B.Q.; Kowalczyk, R. Intelligent Spacecraft Attitude Fault Recovery Using Deep Reinforcement Learning. IEEE Access 2026, 14, 6238–6260. [Google Scholar]
  26. Henna, H.; Toubakh, H.; Kafi, M.R.; Gürsoy, Ö.; Sayed-Mouchaweh, M.; Djemai, M. Satellite fault tolerant attitude control based on expert guided exploration of reinforcement learning agent. J. Exp. Theor. Artif. Intell. 2025, 37, 987–1011. [Google Scholar]
  27. Mahfouz, A.; Valiullin, A.; Lukashevichus, A.; Pritykin, D. Reinforcement learning for attitude control of A spacecraft with flexible appendages. In Proceedings of the 73rd International Astronautical Congress. International Astronautical Federation, Paris, France, 18–22 September 2022. [Google Scholar]
  28. Shen, Q.; Wang, D.; Zhu, S.; Poh, E.K. Integral-Type Sliding Mode Fault-Tolerant Control for Attitude Stabilization of Spacecraft. IEEE Trans. Control Syst. Technol. 2015, 23, 1131–1138. [Google Scholar] [CrossRef]
  29. Terzakis, G.; Lourakis, M.; Ait-Boudaoud, D. Modified Rodrigues Parameters: An Efficient Representation of Orientation in 3D Vision and Graphics. J. Math. Imaging Vis. 2018, 60, 422–442. [Google Scholar] [CrossRef]
  30. Zou, A.M.; De Ruiter, A.H.J.; Dev Kumar, K. Distributed attitude synchronization control for a group of flexible spacecraft using only attitude measurements. Inf. Sci. 2016, 343–344, 66–78. [Google Scholar] [CrossRef]
  31. Di Gennaro, S. Output stabilization of flexible spacecraft with active vibration suppression. IEEE Trans. Aerosp. Electron. Syst. 2003, 39, 747–759. [Google Scholar] [CrossRef]
  32. Peng, Z.; Shen, Q.; Song, C. Attitude Control of Spacecraft Based on Deep Reinforcement Learning TD3 Algorithm. In Proceedings of the International Conference on Guidance, Navigation and Control; Springer: Singapore, 2024; pp. 318–327. [Google Scholar] [CrossRef]
  33. Zou, A.M.; Fan, Z. Fixed-time attitude tracking control for rigid spacecraft without angular velocity measurements. IEEE Trans. Ind. Electron. 2019, 67, 6795–6805. [Google Scholar] [CrossRef]
  34. Javaid, U.; Basin, M.V.; Ijaz, S. Spacecraft Attitude Stabilization Control Under Actuator Faults and Input Saturation. In Proceedings of the 2024 IEEE 18th International Conference on Advanced Motion Control (AMC), Kyoto, Japan; IEEE: Piscataway, NJ, USA, 2024; pp. 1–6. [Google Scholar] [CrossRef]
  35. Hu, Q.; Li, B.; Qi, J. Disturbance observer based finite-time attitude control for rigid spacecraft under input saturation. Aerosp. Sci. Technol. 2014, 39, 13–21. [Google Scholar] [CrossRef]
Figure 1. Composite controller.
Figure 1. Composite controller.
Aerospace 13 00571 g001
Figure 2. Reward growth of agents trained by different reward functions.
Figure 2. Reward growth of agents trained by different reward functions.
Aerospace 13 00571 g002
Figure 3. Agent performance comparison in Scenario S0 (pure baseline).
Figure 3. Agent performance comparison in Scenario S0 (pure baseline).
Aerospace 13 00571 g003
Figure 4. Agent performance comparison in Scenario S1 (measurement noise).
Figure 4. Agent performance comparison in Scenario S1 (measurement noise).
Aerospace 13 00571 g004
Figure 5. Agent performance comparison in Scenario S2 (measurement noise + non-zero ω 0 ).
Figure 5. Agent performance comparison in Scenario S2 (measurement noise + non-zero ω 0 ).
Aerospace 13 00571 g005
Figure 6. Agent performance comparison in Scenario S3 (measurement noise + parameter mismatch + external disturbance).
Figure 6. Agent performance comparison in Scenario S3 (measurement noise + parameter mismatch + external disturbance).
Aerospace 13 00571 g006
Figure 7. Performance comparison of four control architectures under Scenario S4 (actuator faults).
Figure 7. Performance comparison of four control architectures under Scenario S4 (actuator faults).
Aerospace 13 00571 g007
Table 1. Network architectures of the actor and critic.
Table 1. Network architectures of the actor and critic.
ComponentArchitectureActivation
Actor 6 400 300 3 ReLU (hidden), tanh (output)
Critic 9 400 300 1 ReLU (hidden), linear (output)
Table 2. Hyperparameter settings.
Table 2. Hyperparameter settings.
ParameterValue
Learning Rates
    Actor learning rate ( α actor ) 3 × 10 4
    Critic learning rate ( α critic ) 3 × 10 4
Network Architecture
    Hidden layer dimensions400, 300
    Number of hidden layers2
RL Algorithm Parameters
    Discount factor ( γ R L )0.99
    Soft update coefficient ( τ )0.01
    Experience replay buffer size1,000,000
    Minimum buffer size for training1000
    Batch size250
    Policy update delay (d)3
    Exploration noise std ( σ explore )0.1
    Target policy noise std ( σ )0.2
    Target noise clip bound (c)0.5
Optimizer and Reproducibility
    OptimizerAdam
    Random seeds (Python, NumPy, PyTorch)2
Training Schedule
    Total training episodes ( N episodes )20,000
    Episode length ( T episode )200 steps
    Time step duration ( Δ t )1 s
Table 3. Configurations of the five evaluation scenarios.
Table 3. Configurations of the five evaluation scenarios.
SettingS0S1S2S3S4
Measurement noise
Non-zero initial ω 0
Parameter mismatch
External disturbance
Actuator faults
Table 4. Quantitative comparison of agent A vs A0 across Scenarios S0–S3 (steady-state, t > 160 s).
Table 4. Quantitative comparison of agent A vs A0 across Scenarios S0–S3 (steady-state, t > 160 s).
ScenarioAgent p ω (rad/s) η RMS η 0.01 (s)u RMS (Nm)
S0A 4.57 × 10 4 7.00 × 10 6 5.00 × 10 5 30 6.16 × 10 4
A0 2.87 × 10 4 1.17 × 10 4 6.87 × 10 4 95 4.91 × 10 3
S1A 5.44 × 10 4 9.97 × 10 4 2.24 × 10 3 30 5.80 × 10 2
A0 3.63 × 10 4 1.72 × 10 3 5.50 × 10 3 4.89 × 10 2
S2A 5.44 × 10 4 9.95 × 10 4 2.25 × 10 3 56 5.79 × 10 2
A0 4.02 × 10 4 1.83 × 10 3 6.85 × 10 3 4.95 × 10 2
S3A 6.55 × 10 4 1.12 × 10 3 2.32 × 10 3 42 6.33 × 10 2
A0 8.91 × 10 4 6.54 × 10 3 1.94 × 10 2 0.194
Table 5. Quantitative comparison of four control methods under Scenario 4 (actuator faults).
Table 5. Quantitative comparison of four control methods under Scenario 4 (actuator faults).
PDTD3PD+AFTTD3+AFT
Steady-state ( t > 160  s)
MRP p 2.02 × 10 2 1.75 × 10 2 5.81 × 10 4 4.04 × 10 4
ω (rad/s) 8.43 × 10 4 8.54 × 10 4 8.34 × 10 4 7.54 × 10 4
η RMS 1.37 × 10 2 8.43 × 10 3 1.17 × 10 2 1.04 × 10 2
Torque RMS (Nm) 1.042 1.0451.0441.044
Transient/full-trajectory
Peak η 0.132 0.117 0.132 0.117
Peak ω (rad/s) 8.05 × 10 2 8.12 × 10 2 8.12 × 10 2 8.22 × 10 2
Integrated η RMS 2.98 × 10 2 2.91 × 10 2 3.03 × 10 2 2.91 × 10 2
MRPs converge to 0.005 (s)9474
Integrated Torque RMS (Nm) 1.082 1.086 1.112 1.118
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Peng, Z.; Shen, Q. Fault-Tolerant Attitude Control of Flexible Spacecraft via Reinforcement Learning. Aerospace 2026, 13, 571. https://doi.org/10.3390/aerospace13070571

AMA Style

Peng Z, Shen Q. Fault-Tolerant Attitude Control of Flexible Spacecraft via Reinforcement Learning. Aerospace. 2026; 13(7):571. https://doi.org/10.3390/aerospace13070571

Chicago/Turabian Style

Peng, Zhuoyue, and Qiang Shen. 2026. "Fault-Tolerant Attitude Control of Flexible Spacecraft via Reinforcement Learning" Aerospace 13, no. 7: 571. https://doi.org/10.3390/aerospace13070571

APA Style

Peng, Z., & Shen, Q. (2026). Fault-Tolerant Attitude Control of Flexible Spacecraft via Reinforcement Learning. Aerospace, 13(7), 571. https://doi.org/10.3390/aerospace13070571

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop