A Transformer-Based Deep Reinforcement Learning Method for Controller Parameter Modulation in Fault-Tolerant Control
Abstract
1. Introduction
- (1)
- Transformer-guided parameter modulation framework. A deep reinforcement learning framework is proposed for adaptive controller parameter modulation, enabling online fault recovery without replacing the baseline controller structure.
- (2)
- Parameter tokenization and attention-based coupling modeling. Controller parameters are represented as token embeddings and processed using a Transformer policy network, allowing self-attention mechanisms to capture cross-parameter dependencies and coordinate parameter adaptation.
- (3)
- Fault-aware cross-attention mechanism for adaptive recovery. A cross-attention mechanism is introduced to incorporate fault type and severity information into the policy network, enabling context-aware parameter adjustment under diverse fault conditions.
- (4)
- Temporal belief representation for fault evolution modeling. A sequential state encoder is designed to extract temporal patterns from historical system observations, allowing the controller to anticipate fault evolution and perform proactive parameter adaptation.
2. Problem Formulations
2.1. System Modeling and Adaptive Parameter Modulation
2.2. Optimization Model
3. Methodology
3.1. Sequential State Encoder
3.2. Parameter Token Embedding
3.3. Transformer Policy Network
3.4. Training via Proximal Policy Optimization
| Algorithm 1. Transformer-DRL Parameter Modulation training |
| Input: Environment E, nominal params θ0, bound ε, hyperparams a1, a2, a3, a4 |
| 1: Initialize policy πϕ (Transformer), value network Vψ |
| 2: Initialize replay buffer B ← ∅ |
| 3: for episode = 1, 2, …, Nep do |
| 4: Sample fault scenario ϕa ∼ p(ϕ) |
| 5: Reset environment; observe s0 |
| 6: for t = 0, 1, …, T do |
| 7: Encode history: hc = SequentialEncoder(O{t – M + 1:t}) ▷ Equations (11)–(15) |
| 8: Tokenize params: Z = ParamEmbedding(θ0) ▷ Equations (16) and (17) |
| 9: Compute (μϕ, σϕ) = TransformerPolicy(Z, hC) ▷ Equations (18)–(22) |
| 10: Sample at ∼ N() ▷ Equation (23) |
| 11: Modulate: Δθt = ϵ · tanh(at) ▷ Equation (24) |
| 12: Apply: θ′_t = θ0+ Δθ_t) ▷ Equation (4) |
| 13: Execute ut = C(xt, xref, θ′t); observe rt, st+1 ▷ Equation (5) |
| ▷ Equation (27) |
| 15: Store (st, at, rt, st+1) in B |
| 16: end for |
| 17: Compute advantages Ât via GAE |
| 18: for k = 1, …, K do |
| 19: Compute ratio rt(ϕ) = πϕ(at|st)/πϕold(at|st) |
| 20: LCLIP(ϕ) = Et[min(rt(ϕ)Ât, clip(rt(ϕ), 1 − δ, 1 + δ)Ât)] ▷ Equation (29) |
| 21: Update ϕ by maximizing LCLIP(ϕ) |
| 22: Update ψ by minimizing value loss |
| 23: end for |
| 24: end for |
| 25: return πϕ* |
3.5. Stability Analysis
4. Experimental Results
4.1. Comparison Study with Genetic Algorithm-Based Neural Network
4.2. Comparison Study with PSO-Based Optimal Gain Scheduling Backstepping Controller
4.2.1. Scenario 1: Circular Trajectory Tracking Results
4.2.2. Scenario 2: Waypoint Trajectory Tracking Results
4.3. Comparison Study with TD3 Reinforcement Learning for PI Controller Gain Scheduling
4.3.1. Scenario 1: Bias Fault Experiment
4.3.2. Scenario 2: Sinusoidal Fault Experiment
4.4. Case Study
5. Discussion
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Yazdjerdi, P.; Meskin, N. Actuator fault detection and isolation of differential drive mobile robots using multiple model algorithm. In Proceedings of the 2017 4th International Conference on Control, Decision and Information Technologies (CoDIT), Barcelona, Spain, 5–7 April 2017; IEEE: New York, NY, USA, 2017; pp. 439–443. [Google Scholar]
- Yazdjerdi, P.; Meskin, N. Design and real-time implementation of actuator fault-tolerant control for differential-drive mobile robots based on multiple-model approach. Proc. Inst. Mech. Eng. Part I J. Syst. Control Eng. 2018, 232, 652–661. [Google Scholar] [CrossRef] [Scilit]
- Sun, H.; Zhao, S. Fault diagnosis for bearing based on 1DCNN and LSTM. Shock Vib. 2021, 2021, 1221462. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Wang, G.; Hu, P.; Duan, L.Y.; Kot, A.C. Global context-aware attention LSTM networks for 3D action recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1647–1656. [Google Scholar]
- Lundgren, A.; Jung, D. Data-driven fault diagnosis analysis and open-set classification of time-series data. Control Eng. Pract. 2022, 121, 105006. [Google Scholar] [CrossRef] [Scilit]
- Ma, Z.L.; Li, X.J. A data-driven fault isolation and estimation approach for unknown linear systems. J. Process Control 2023, 124, 118–128. [Google Scholar] [CrossRef] [Scilit]
- Parzinger, M.; Hanfstaengl, L.; Sigg, F.; Spindler, U.; Wellisch, U.; Wirnsberger, M. Residual analysis of predictive modelling data for automated fault detection in building’s heating, ventilation and air conditioning systems. Sustainability 2020, 12, 6758. [Google Scholar] [CrossRef] [Scilit]
- Zhu, F.; Shan, Y.; Tang, Y. Actuator and sensor fault detection and isolation for uncertain switched nonlinear system based on sliding mode observers. Int. J. Control Autom. Syst. 2021, 19, 3075–3086. [Google Scholar] [CrossRef] [Scilit]
- Tran, A.M.D.; Vu, T.V. Robust MIMO LQR control with integral action for differential drive robots: A Lyapunov-cost function approach. Eng. Technol. Appl. Sci. Res. 2025, 15, 24775–24781. [Google Scholar] [CrossRef] [Scilit]
- Karaboga, D.; Kaya, E. Adaptive network based fuzzy inference system (ANFIS) training approaches: A comprehensive survey. Artif. Intell. Rev. 2019, 52, 2263–2293. [Google Scholar] [CrossRef] [Scilit]
- Shi, P.; Wang, X.; Meng, X.; He, M.; Mao, Y.; Wang, Z. Adaptive fault-tolerant control for open-circuit faults in dual three-phase PMSM drives. IEEE Trans. Power Electron. 2022, 38, 3676–3688. [Google Scholar] [CrossRef] [Scilit]
- Stavrinidis, S.; Zacharia, P. An ANFIS-based strategy for autonomous robot collision-free navigation in dynamic environments. Robotics 2024, 13, 124. [Google Scholar] [CrossRef] [Scilit]
- Zhong, K.; Yang, Z.; Yu, S.; Li, K. Deep reinforcement learning-based multi-layer cascaded resilient recovery for cyber-physical systems. IEEE Trans. Serv. Comput. 2024. [Google Scholar]
- Nikanjam, A.; Morovati, M.M.; Khomh, F.; Ben Braiek, H. Faults in deep reinforcement learning programs: A taxonomy and a detection approach. Autom. Softw. Eng. 2022, 29, 8. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Abtahi, S.M.; Chahari, M.; Zhao, T. An adaptive neuro-fuzzy model for attitude estimation and control of a 3 DOF system. Mathematics 2022, 10, 976. [Google Scholar] [CrossRef] [Scilit]
- Lawrence, N.P.; Forbes, M.G.; Loewen, P.D.; McClement, D.G.; Backström, J.U.; Gopaluni, R.B. Deep reinforcement learning with shallow controllers: An experimental application to PID tuning. Control Eng. Pract. 2022, 121, 105046. [Google Scholar] [CrossRef] [Scilit]
- Joseph, S.B.; Dada, E.G.; Abidemi, A.; Oyewola, D.O.; Khammas, B.M. Metaheuristic algorithms for PID controller parameters tuning: Review, approaches and open problems. Heliyon 2022, 8, e09399. [Google Scholar] [CrossRef] [Scilit]
- Süpürtülü, M.; Hatipoğlu, A.; Yılmaz, E. An analytical benchmark of feature selection techniques for industrial fault classification leveraging time-domain features. Appl. Sci. 2025, 15, 1457. [Google Scholar] [CrossRef] [Scilit]
- Wang, L.; Zhao, W.; Liu, Z.; Dang, Q.; Zou, X.; Wang, K. Incipient fault detection and reconstruction using an adaptive sliding-mode observer for the actuators of fixed-wing aircraft. Aerospace 2023, 10, 422. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Surur, K.; Kabir, I.; Ahmad, G.; Abido, M.A. Optimal gain scheduling for fault-tolerant control of quadrotor UAV using genetic algorithm-based neural network. Arab. J. Sci. Eng. 2025, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Derrouaoui, S.H.; Bouzid, Y.; Guiatni, M. PSO based optimal gain scheduling backstepping flight controller design for a transformable quadrotor. J. Intell. Robot. Syst. 2021, 102, 67. [Google Scholar] [CrossRef] [Scilit]
- Kuryło, P. Modeling of two-wheeled self-balancing robot driven by DC gearmotors. Int. J. Appl. Mech. Eng. 2017, 22, 739–747. [Google Scholar]
- Moritz, J.; Musa, M.; Wejinya, U. Design and modeling of a two-wheeled differential drive robot. Int. J. Control Autom. Syst. 2024, 22, 2273–2282. [Google Scholar] [CrossRef] [Scilit]
- Sardashti, A.; Nazari, J. A learning-based approach to fault detection and fault-tolerant control of permanent magnet DC motors. J. Eng. Appl. Sci. 2023, 70, 109. [Google Scholar] [CrossRef] [Scilit]
- Ye, S.; Jiang, J.; Li, J.; Liu, Y.; Zhou, Z.; Liu, C. Fault diagnosis and tolerance control of five-level nested NPP converter using wavelet packet and LSTM. IEEE Trans. Power Electron. 2019, 35, 1907–1921. [Google Scholar] [CrossRef] [Scilit]
- Shen, Q.; Shi, P.; Lim, C.P. Fuzzy adaptive fault-tolerant stability control against novel actuator faults and its application to mechanical systems. IEEE Trans. Fuzzy Syst. 2024, 32, 2331–2340. [Google Scholar] [CrossRef] [Scilit]
- Kong, N.J.; Li, C.; Council, G.; Johnson, A.M. Hybrid iLQR model predictive control for contact implicit stabilization on legged robots. IEEE Trans. Robot. 2023, 39, 4712–4727. [Google Scholar] [CrossRef] [Scilit]
- Pan, J.; Qu, L.; Peng, K. Deep residual neural-network-based robot joint fault diagnosis method. Sci. Rep. 2022, 12, 17158. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lakhani, A.I.; Chowdhury, M.A.; Lu, Q. Stability-preserving automatic tuning of PID control with reinforcement learning. arXiv 2021, arXiv:2112.15187. [Google Scholar] [CrossRef] [Scilit]
- Mehta, N.S.; Bhaiya, V.; Patel, K.A.; Farsangi, E.N. Predictive active control of building structures using LQR and artificial intelligence. Earthq. Eng. Eng. Vib. 2024, 23, 489–502. [Google Scholar] [CrossRef] [Scilit]








| Method | No Fault | = 0.7 Fault |
|---|---|---|
| Manual | 2.8357 | 22.1196 |
| GA (normal tuning) | 0.8047 | 1.1897 |
| GA (fault tuning) | 1.0533 | 0.9820 |
| DRL + Transformer | 0.7853 | 0.9644 |
| Performance Metric | TD3 | Transformer + PPO |
|---|---|---|
| Overshoot at fault | 0.12 | 0.08 |
| Recovery time (s) | 0.45 | 0.32 |
| Steady-state error | 0.023 | 0.015 |
| RMSE (5–10 s) | 0.089 | 0.064 |
| IAE (5–10 s) | 0.156 | 0.137 |
| Performance Metric | TD3 | Transformer + PPO |
|---|---|---|
| Overshoot at fault | 0.38 | 0.19 |
| Recovery time (s) | 0.62 | 0.33 |
| Steady-state error | 0.044 | 0.019 |
| RMSE (5–10 s) | 0.165 | 0.073 |
| IAE (5–10 s) | 0.412 | 0.188 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, C.; Li, X. A Transformer-Based Deep Reinforcement Learning Method for Controller Parameter Modulation in Fault-Tolerant Control. Mathematics 2026, 14, 1409. https://doi.org/10.3390/math14091409
Zhang C, Li X. A Transformer-Based Deep Reinforcement Learning Method for Controller Parameter Modulation in Fault-Tolerant Control. Mathematics. 2026; 14(9):1409. https://doi.org/10.3390/math14091409
Chicago/Turabian StyleZhang, Chenfei, and Xiangning Li. 2026. "A Transformer-Based Deep Reinforcement Learning Method for Controller Parameter Modulation in Fault-Tolerant Control" Mathematics 14, no. 9: 1409. https://doi.org/10.3390/math14091409
APA StyleZhang, C., & Li, X. (2026). A Transformer-Based Deep Reinforcement Learning Method for Controller Parameter Modulation in Fault-Tolerant Control. Mathematics, 14(9), 1409. https://doi.org/10.3390/math14091409

