Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation
Abstract
1. Introduction
2. Materials and Methods
2.1. Dynamic Environment of RL
2.1.1. Aircraft Equations of Motion
2.1.2. Micro Downburst Model
2.2. RL Algorithm with Active Exploration Reward Function
2.2.1. Design of the Active Reward Function
2.2.2. Design of the Manual Reward Function
2.3. Multivariable Decoupling and Signal Reconstruction via W-MSSA
2.3.1. Denoising of Control Signals via Discrete Wavelet Transform
2.3.2. Multivariable Singular Spectrum Analysis
2.4. Rule Extraction and Standard Operating Procedure Generation
| Algorithm 1 Explainable Reinforcement Learning Framework for Autonomous Windshear Escape |
Phase 1: AR-PPO Training (Actor-Critic Architecture)
Phase 2: Multivariable Decoupling and Signal Reconstruction (W-MSSA)
Phase 3: Rule Extraction and SOP Generation
|
3. Results
3.1. Simulation Setup and Parameter Settings
3.2. Experimental Results
3.2.1. Comparison in Nominal Windshear Environment
3.2.2. Robustness Analysis
3.3. Design for Physical SOP Extraction
4. Discussion
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AEDRL | Active exploration deep reinforcement learning |
| AR-PPO | Active-reward proximal policy optimization |
| CART | Classification and regression tree |
| DRL | Deep reinforcement learning |
| IL | Imitation learning |
| IRL | Inverse reinforcement learning |
| PIRL | Programmatically interpretable RL |
| PPO | Proximal policy optimization |
| RL | Reinforcement learning |
| SILVER | Shapley value-based interpretable policy via explanation verification |
| XRL | Explainable reinforcement learning |
| DWT | Discrete wavelet transform |
| IDWT | Inverse discrete wavelet transform |
| MRA | Multi-resolution analysis |
| MSSA | Multivariate singular spectrum analysis |
| SNR | Signal-to-noise ratio |
| SVD | Singular value decomposition |
| W-MSSA | Wavelet analysis combined with multivariate singular spectrum analysis |
| 6-DOF | Six-degree-of-freedom |
| CBTA | Competency-based training and assessment |
| CFIT | Controlled flight into terrain |
| EBT | Evidence-based training |
| FDR | Flight data recorder |
| PIO | Pilot-induced oscillation |
| SOP | Standard operating procedure |
| TOGA | Take-off/go-around |
References
- Razzaghi, P.; Tabrizian, A.; Guo, W.; Chen, S.; Taye, A.; Thompson, E.; Bregeon, A.; Baheri, A.; Wei, P. A survey on reinforcement learning in aviation applications. Eng. Appl. Artif. Intell. 2024, 136, 108911. [Google Scholar] [CrossRef] [Scilit]
- Lombaerts, T.; Looye, G.; Ellerbroek, J.; Martin, M.R.Y. Design and piloted simulator evaluation of adaptive safe flight envelope protection algorithm. J. Guid. Control Dyn. 2017, 40, 1902–1924. [Google Scholar] [CrossRef] [Scilit]
- Harms, M.; Lim, J.; Rohr, D.; Rockenbauer, F.; Lawrance, N.; Siegwart, R. Towards robust optimization-based autonomous dynamic soaring with a fixed-wing UAV. IEEE Robot. Autom. Lett. 2026, 7620–7627. [Google Scholar] [CrossRef] [Scilit]
- Liu, F.; Dai, S.; Zhao, Y. Learning to Have a Civil Aircraft Take Off under Crosswind Conditions by Reinforcement Learning with Multimodal Data and Preprocessing Data. Sensors 2021, 21, 1386. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Park, S.; Fanjoy, A.; Golubev, V.V. Application of reinforcement learning for autonomous dynamic soaring. In Proceedings of the AIAA Scitech 2025 Forum, Orlando, FL, USA, 6–10 January 2025. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Lu, J.; Yin, Y.; Huang, J.; Xiang, Y.; Liu, H. Learning step-level dynamic soaring in shear flow. arXiv 2026, arXiv:2604.12413. [Google Scholar] [CrossRef] [Scilit]
- Ma, Q.; Wu, Y.; Shoukat, M.U.; Yan, Y.; Wang, J.; Yang, L.; Yan, F.; Yan, L. Deep reinforcement learning-based wind disturbance rejection control strategy for UAV. Drones 2024, 8, 632. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Chai, S.; Zhang, Y.; Huang, D.; Ge, Q. Cascaded TD3-PID hybrid controller for quadrotor trajectory tracking. arXiv 2026, arXiv:2604.13505. [Google Scholar] [CrossRef] [Scilit]
- Ren, Y.; Zhu, F.; Sui, S.; Yi, Z.; Chen, K. Enhancing quadrotor control robustness with multi-proportional–integral–derivative self-attention-guided deep reinforcement learning. Drones 2024, 8, 315. [Google Scholar] [CrossRef] [Scilit]
- Nigam, R.; Choi, J.; Parikh, N.; Li, M.Z.; Tran, H.T. Survey of inverse reinforcement learning in aviation and future outlooks. J. Aerosp. Inf. Syst. 2026, 23, 305–321. [Google Scholar] [CrossRef] [Scilit]
- Song, L.; Guo, Q.; Channa, I.A.; Wang, Z. A survey of maximum entropy-based inverse reinforcement learning: Methods and applications. Drones 2025, 17, 1632. [Google Scholar] [CrossRef] [Scilit]
- Zare, M.; Kebria, P.M.; Khosravi, A.; Nahavandi, S. A survey of imitation learning: Algorithms, recent developments, and challenges. IEEE Trans. Cybern. 2024, 54, 7173–7186. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kayal, A.; Pignatelli, E.; Toni, L. The impact of intrinsic rewards on exploration in reinforcement learning. Neural Comput. Appl. 2025, 37, 16269–16303. [Google Scholar] [CrossRef] [Scilit]
- Pathak, D.; Agrawal, P.; Efros, A.A.; Darrell, T. Deep reinforcement learning-based wind disturbance rejection control strategy for UAV. In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, 6–11 August 2017; Available online: https://proceedings.mlr.press/v70/pathak17a.html (accessed on 4 August 2026).
- Burda, Y.; Edwards, H.; Storkey, A.; Klimov, O. Exploration by random network distillation. arXiv 2017, arXiv:1810.12894. [Google Scholar] [CrossRef] [Scilit]
- Saglam, B.; Kozat, S.S. Deep intrinsically motivated exploration in continuous control. Mach. Learn. 2023, 112, 4959–4993. [Google Scholar] [CrossRef] [Scilit]
- Oh, J.; Farquhar, G.; Kemaev, I.; Calian, D.A.; Hessel, M.; Zintgraf, L.; Singh, S.; van Hasselt, H.; Silver, D. Discovering state-of-the-art reinforcement learning algorithms. Nature 2025, 648, 312–319. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lu, R.; Shao, Z.; Ding, Y.; Chen, R.; Wu, D.; Su, H.; Yang, T.; Zhang, F.; Wang, J.; Shi, Y.; et al. Discovery of the reward function for embodied reinforcement learning agents. Nat. Commun. 2025, 16, 11064. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhao, D.; Xu, H.; Zhang, X. Active exploration deep reinforcement learning for continuous action space with forward prediction. Int. J. Comput. Intell. Syst. 2024, 17, 6. [Google Scholar] [CrossRef] [Scilit]
- Russo, A.; Proutiere, A. Model-free active exploration in reinforcement learning. In Proceedings of the 37th Conference on Neural Information Processing Systems (NeurIPS), New Orleans, LA, USA, 10–16 December 2023. [Google Scholar] [CrossRef] [Scilit]
- Eysenbach, B.; Gupta, A.; Ibarz, J.; Levine, S. Diversity is all you need: Learning skills without a reward function. arXiv 2019, arXiv:1802.06070. [Google Scholar] [CrossRef] [Scilit]
- Zhao, R.; Gao, Y.; Abbeel, P.; Tresp, V.; Xu, W. Mutual information state intrinsic control. arXiv 2021, arXiv:2103.08107. [Google Scholar] [CrossRef] [Scilit]
- As, Y.; Sukhija, B.; Treven, L.; Sferrazza, C.; Coros, S.; Krause, A. ActSafe: Active exploration with safety constraints for reinforcement learning. arXiv 2025, arXiv:2410.09486. [Google Scholar] [CrossRef] [Scilit]
- Bush, T.; Chung, S.; Anwar, U.; Garriga-Alonso, A.; Krueger, D. Interpreting emergent planning in model-free reinforcement learning. arXiv 2025, arXiv:2504.01871. [Google Scholar] [CrossRef] [Scilit]
- Qian, Y.; Nguyen, S.; Chen, C.; Zhou, Q.; Zhao, L. Interpret policies in deep reinforcement learning using SILVER with RL-guided labeling. arXiv 2025, arXiv:2510.19244. [Google Scholar] [CrossRef] [Scilit]
- Gokhale, G.; Karimi Madahi, S.S.; Claessens, B.; Develder, C. Distill2Explain: Differentiable decision trees for explainable reinforcement learning in energy application controllers. In Proceedings of the 15th ACM International Conference on Future and Sustainable Energy Systems, Singapore, 4–7 June 2024. [Google Scholar] [CrossRef] [Scilit]
- Wen, Y.; Li, S.; Zuo, R.; Yuan, L.; Mao, H.; Liu, P. SkillTree: Explainable skill-based deep reinforcement learning for long-horizon control tasks. In Proceedings of the AAAI Conference on Artificial Intelligence, Philadelphia, PA, USA, 25 February–4 March 2025. [Google Scholar] [CrossRef] [Scilit]
- Zhang, M.; Miao, Z.; Nan, X.; Ma, N.; Liu, R. Explainable reinforcement learning for the initial design optimization of compressors. Biomimetics 2025, 10, 497. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Engelhardt, R.C.; Lange, M.; Wiskott, L.; Konen, W. Sample-based rule extraction for explainable reinforcement learning. In Machine Learning, Optimization, and Data Science; Nicosia, G., Ojha, V., La Malfa, E., La Malfa, G., Pardalos, P., Di Fatta, G., Giuffrida, G., Umeton, R., Eds.; Publishing House: Cham, Switzerland, 2023; pp. 330–345. [Google Scholar] [CrossRef] [Scilit]
- Shawon, R.U.; Liu, S.; Siddiqua, A. Explainable reinforcement learning for multi-agent systems. In Proceedings of the 2025 IEEE 37th International Conference on Tools with Artificial Intelligence (ICTAI), Athens, Greece, 3–5 November 2025. [Google Scholar] [CrossRef] [Scilit]
- Verma, A.; Murali, V.; Singh, R.; Kohli, P.; Chaudhuri, S. Programmatically interpretable reinforcement learning. In Proceedings of the 35th International Conference on Machine Learning (PMLR), New Orleans, LA, USA, 10–15 July 2018; Available online: https://proceedings.mlr.press/v80/verma18a.html (accessed on 4 August 2026).
- Champion, K.; Lusch, B.; Kutz, J.N.; Brunton, S.L. Data-driven discovery of coordinates and governing equations. Proc. Natl. Acad. Sci. USA 2019, 116, 22445–22451. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Stevens, B.L.; Lewis, F.L.; Johnson, E.N. Aircraft Control and Simulation: Dynamics, Controls Design, and Autonomous Systems, 3rd ed.; John Wiley & Sons: Hoboken, NJ, USA, 2015; Available online: https://pure.psu.edu/en/publications/aircraft-control-and-simulation-dynamics-controls-design-and-auto/ (accessed on 4 August 2026).
- Schultz, T.A. Multiple vortex ring model of the DFW microburst. J. Aircr. 1990, 27, 163–168. Available online: https://arc.aiaa.org/doi/pdf/10.2514/3.45913?casa_token=IgxKdQQfL0gAAAAA:A-ajzBNu0EgZTBXTKFGz8nHssQ-n34gwqiob3WCPYnf-2dVXMK8B9yYb9fr5wLbQU-lbPUMi (accessed on 4 August 2026). [CrossRef] [Scilit] [PubMed]
- Akbar, H.; Arshad, I.; Rahman, S.; Aamir, F.; Fatima, M.; Khan, H.U.R. TRUST-X: Real-time explainable decision-making framework for reinforcement learning-based UAV control. In Proceedings of the 2025 20th International Conference on Emerging Technologies (ICET), Peshawar, Pakistan, 18–19 November 2025. [Google Scholar] [CrossRef] [Scilit]











| Category | Parameter | Value |
|---|---|---|
| Aircraft and environment | Initial airspeed (m/s) | 75.0 |
| Altitude (m) | 500 | |
| Micro downburst max downdraft (m/s) | 15.0 (Train)/22.0 (Test) | |
| Throttle | [0.1, 1.0] | |
| Elevator (°) | [−15.0, 15.0]/15.0 | |
| Rudder/Aileron (°) | [−20.0, 20.0]/20.0 | |
| Active reward function parameters | Temporal discount factor | 0.99 |
| Meta-objective terminal penalty | −100 | |
| Meta-Objective: Terminal Reward | 50 | |
| Manual reward function parameters | 1.0 | |
| 0.5 | ||
| 0.3 | ||
| 0.1 | ||
| 0.1 | ||
| 0.1 | ||
| 1000 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Y.; Zhang, Y.; Gao, Z. Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation. Aerospace 2026, 13, 721. https://doi.org/10.3390/aerospace13080721
Wang Y, Zhang Y, Gao Z. Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation. Aerospace. 2026; 13(8):721. https://doi.org/10.3390/aerospace13080721
Chicago/Turabian StyleWang, Yitan, Yangyang Zhang, and Zhenxing Gao. 2026. "Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation" Aerospace 13, no. 8: 721. https://doi.org/10.3390/aerospace13080721
APA StyleWang, Y., Zhang, Y., & Gao, Z. (2026). Explainable Reinforcement Learning Framework for Autonomous Windshear Escape with Policy Distillation. Aerospace, 13(8), 721. https://doi.org/10.3390/aerospace13080721

