Safe UAV Control Against Wind Disturbances via Demonstration-Guided Reinforcement Learning
Highlights
- A hybrid controller combining PPO (RL), a PID expert (DGRL), and a real-time CBF safety filter achieves safer and more accurate hovering under gusts in both AirSim and indoor CoDrone tests.
- The hybrid design converges faster and more stably while incurring fewer safety violations during training and deployment.
- Demonstration-guided safe RL enables deployable, safety-aware UAV control in disturbed, uncertain environments without sacrificing steady-state precision.
- The framework provides a practical Sim-to-Real pathway for small UAVs, supporting precision tasks where robustness and formal safety constraints are critical.
Abstract
1. Introduction
- Unified hybrid-DGRL framework: We propose a three-layer control architecture that integrates PPO-based adaptive control, a calibrated PID expert for DGRL, and a CBF module for formally grounded, real-time safety enforcement.
- Fundamentally novel DGRL and accelerated safe training: We introduce a DGRL mechanism that converts expert behavior into an intrinsic reward signal for PPO optimization. This design fundamentally differs from conventional imitation learning by avoiding approaches that rely on an explicit imitation loss, ensuring safe policy initialization and significantly accelerating convergence.
- Sim-to-real validation on a resource-constrained UAV platform: We deploy and validate the proposed hybrid-DGRL policy on a lightweight, resource-constrained CoDrone platform. Experiments under controlled airflow disturbance demonstrate that our method achieves superior hover stability, training efficiency, and safety guarantees compared with PID and PPO baseline controllers.
- Computationally feasible safety and non-interference: We detail the implementation of a simplified CBF-inspired Analytical Clipping Filter which, unlike resource-intensive QP solvers, provides effective safety bounds in real-time. This guarantees the UAV state remains within a predefined safe set during exploration while minimizing processing time and policy distortion.
2. Methods
2.1. PPO-Based RL
2.1.1. State Space and Action Space
2.1.2. Reward Function Design
2.1.3. Policy Update Mechanism
2.2. CBF Safety Filter
2.3. DGRL
2.3.1. Expert Demonstrator and PID Tunings
2.3.2. Demonstration-Guided Reward Shaping
3. Experimental Setup
3.1. Simulator and Real-World Platforms
3.2. Baselines and Implementation Details
3.2.1. PPO Training Parameters
3.2.2. PID Parameters
3.3. Performance Benchmarks and Strategy Nomenclature
4. Results and Discussion
4.1. Primary Validation: Safety-Aware Hovering Under Gusts
4.1.1. Flight Evaluation and Discussion in the Simulation Environment
4.1.2. Indoor Real-World Flight Evaluation and Results Analysis
4.2. Training Efficiency and Stability (Convergence Curves)
4.3. Ablation Study and Steady-State Performance
4.3.1. Ablation Study on Simulation Results
4.3.2. Ablation Study on Real-World UAV Results
4.3.3. Cross-Environment Synthesis
4.4. Discussion
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Raparelli, E.; Bajocco, S. A bibliometric analysis on the use of unmanned aerial vehicles in agricultural and forestry studies. Int. J. Remote Sens. 2019, 40, 9070–9083. [Google Scholar] [CrossRef] [Scilit]
- Calamoneri, T.; Corò, F.; Mancini, S. A realistic model to support rescue operations after an earthquake via uavs. IEEE Access 2022, 10, 6109–6125. [Google Scholar] [CrossRef] [Scilit]
- Koch, W.; Mancuso, R.; West, R.; Bestavros, A. Reinforcement learning for UAV attitude control. ACM Trans. Cyber-Phys. Syst. 2019, 3, 1–21. [Google Scholar] [CrossRef] [Scilit]
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Ma, B.; Liu, Z.; Dang, Q.; Zhao, W.; Wang, J.; Cheng, Y.; Yuan, Z. Deep reinforcement learning of UAV tracking control under wind disturbances environments. IEEE Trans. Instrum. Meas. 2023, 72, 2510913. [Google Scholar] [CrossRef] [Scilit]
- Wu, J.; Yang, Z.; Zhuo, H.; Xu, C.; Zhang, C.; He, N.; Liao, L.; Wang, Z. A Supervised Reinforcement Learning Algorithm for Controlling Drone Hovering. Drones 2024, 8, 69. [Google Scholar] [CrossRef] [Scilit]
- Yang, H.; Yu, C.; Chen, S. Hybrid policy optimization from imperfect demonstrations. Adv. Neural Inf. Process. Syst. 2023, 36, 4653–4663. [Google Scholar]
- Kim, G.; Chang, K.; Byun, Y.; Baek, J.-G. Autonomous PID tuning: Two-phase reinforcement learning through adversarial imitation learning under imperfect demonstrations. IEEE Trans. Autom. Sci. Eng. 2025, 22, 20280–20295. [Google Scholar] [CrossRef] [Scilit]
- Joshi, B.; Kapur, D.; Kandath, H. Sim-to-real deep reinforcement learning based obstacle avoidance for UAVs under measurement uncertainty. In Proceedings of the 2024 10th International Conference on Automation, Robotics and Applications (ICARA), Athens, Greece, 22–24 February 2024; pp. 278–284. [Google Scholar] [CrossRef] [Scilit]
- Ajani, O.S.; Hur, S.-H.; Mallipeddi, R. Evaluating domain randomization in deep reinforcement learning locomotion tasks. Mathematics 2023, 11, 4744. [Google Scholar] [CrossRef] [Scilit]
- Nakamoto, M.; Zhai, S.; Singh, A.; Sobol Mark, M.; Ma, Y.; Finn, C.; Kumar, A.; Levine, S. Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning. Adv. Neural Inf. Process. Syst. 2023, 36, 62244–62269. [Google Scholar]
- Kalaria, D.; Lin, Q.; Dolan, J.M. Disturbance Observer-based Control Barrier Functions with Residual Model Learning for Safe Reinforcement Learning. arXiv 2024, arXiv:2410.06570. [Google Scholar] [CrossRef] [Scilit]
- Ugurlu, H.I.; Pham, X.H.; Kayacan, E. Sim-to-real deep reinforcement learning for safe end-to-end planning of aerial robots. Robotics 2022, 11, 109. [Google Scholar] [CrossRef] [Scilit]
- Ames, A.D.; Coogan, S.; Egerstedt, M.; Notomista, G.; Sreenath, K.; Tabuada, P. Control barrier functions: Theory and applications. In Proceedings of the 2019 18th European Control Conference (ECC), Naples, Italy, 25–28 June 2019; pp. 3420–3431. [Google Scholar] [CrossRef] [Scilit]
- Du, D.; Han, S.; Qi, N.; Ammar, H.B.; Wang, J.; Pan, W. Reinforcement learning for safe robot control using control lyapunov barrier functions. arXiv 2023, arXiv:2305.09793. [Google Scholar] [CrossRef] [Scilit]
- Chen, H.; Huang, D.; Wang, C.; Ding, L.; Song, L.; Liu, H. Collision-Free Path Planning for Multiple Drones Based on Safe Reinforcement Learning. Drones 2024, 8, 481. [Google Scholar] [CrossRef] [Scilit]
- Jembre, Y.Z.; Nugroho, Y.W.; Khan, M.T.R.; Attique, M.; Paul, R.; Shah, S.H.A.; Kim, B. Evaluation of reinforcement and deep learning algorithms in controlling unmanned aerial vehicles. Appl. Sci. 2021, 11, 7240. [Google Scholar] [CrossRef] [Scilit]
- Zandavi, S.M.; Chung, V.; Anaissi, A. Accelerated control using stochastic dual simplex algorithm and genetic filter for drone application. IEEE Trans. Aerosp. Electron. Syst. 2021, 58, 2180–2191. [Google Scholar] [CrossRef] [Scilit]
- Ross, S.; Gordon, G.; Bagnell, D. A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, JMLR Workshop and Conference Proceedings, Fort Lauderdale, FL, USA, 11–13 April 2011; pp. 627–635. [Google Scholar]
- Smith, C. A Drone’s-Eye View: Small drones for easy lending and programs. Am. Libr. 2020, 51, 52–54. [Google Scholar]
- Nahrendra, I.M.A.; Tirtawardhana, C.; Yu, B.; Lee, E.M.; Myung, H. Retro-RL: Reinforcing nominal controller with deep reinforcement learning for tilting-rotor drones. IEEE Robot. Autom. Lett. 2022, 7, 9004–9011. [Google Scholar] [CrossRef] [Scilit]
- Kalidas, A.P.; Joshua, C.J.; Md, A.Q.; Basheer, S.; Mohan, S.; Sakri, S. Deep reinforcement learning for vision-based navigation of UAVs in avoiding stationary and mobile obstacles. Drones 2023, 7, 245. [Google Scholar] [CrossRef] [Scilit]
- Naeem, M.; Rizvi, S.T.H.; Coronato, A. A gentle introduction to reinforcement learning and its application in different fields. IEEE Access 2020, 8, 209320–209344. [Google Scholar] [CrossRef] [Scilit]
- Joseph, S.B.; Dada, E.G.; Abidemi, A.; Oyewola, D.O.; Khammas, B.M. Metaheuristic algorithms for PID controller parameters tuning: Review, approaches and open problems. Heliyon 2022, 8, e09399. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Rivera, C.E.O.; Tyni, K.; Nguyen, S. AirPilot: Interpretable PPO-based DRL Auto-Tuned Nonlinear PID Drone Controller for Robust Autonomous Flights. arXiv 2024, arXiv:2404.00204. [Google Scholar] [CrossRef] [Scilit]
- Dogru, O.; Velswamy, K.; Ibrahim, F.; Wu, Y.; Sundaramoorthy, A.S.; Huang, B.; Xu, S.; Nixon, M.; Bell, N. Reinforcement learning approach to autonomous PID tuning. Comput. Chem. Eng. 2022, 161, 107760. [Google Scholar] [CrossRef] [Scilit]
- Gao, Y. PID-based search algorithm: A novel metaheuristic algorithm based on PID algorithm. Expert Syst. Appl. 2023, 232, 120886. [Google Scholar] [CrossRef] [Scilit]
- Mosali, N.A.; Shamsudin, S.S.; Alfandi, O.; Omar, R.; Al-Fadhali, N. Twin delayed deep deterministic policy gradient-based target tracking for unmanned aerial vehicle with achievement rewarding and multistage training. IEEE Access 2022, 10, 23545–23559. [Google Scholar] [CrossRef] [Scilit]
- Gu, Y.; Cheng, Y.; Chen, C.P.; Wang, X. Proximal policy optimization with policy feedback. IEEE Trans. Syst. Man Cybern. Syst. 2021, 52, 4600–4610. [Google Scholar] [CrossRef] [Scilit]
- Mohamadi, N.; Niaki, S.T.A.; Taher, M.; Shavandi, A. An application of deep reinforcement learning and vendor-managed inventory in perishable supply chain management. Eng. Appl. Artif. Intell. 2024, 127, 107403. [Google Scholar] [CrossRef] [Scilit]
- Meng, W.; Zheng, Q.; Shi, Y.; Pan, G. An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning. IEEE Trans. Neural Netw. Learn. Syst. 2021, 33, 2223–2235. [Google Scholar] [CrossRef] [Scilit]
- Yang, Y.; Liu, J.; Tan, S. A constrained multi-objective evolutionary algorithm based on decomposition and dynamic constraint-handling mechanism. Appl. Soft Comput. 2020, 89, 106104. [Google Scholar] [CrossRef] [Scilit]
- Chia, K.S. Ziegler-nichols based proportional-integral-derivative controller for a line tracking robot. Indones. J. Electr. Eng. Comput. Sci. 2018, 9, 221–226. [Google Scholar] [CrossRef] [Scilit]
- Shah, S.; Dey, D.; Lovett, C.; Kapoor, A. Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and Service Robotics: Results of the 11th International Conference; Springer: Cham, Switzerland, 2017; pp. 621–635. [Google Scholar] [CrossRef] [Scilit]
- Raffin, A.; Hill, A.; Gleave, A.; Kanervisto, A.; Ernestus, M.; Dormann, N. Stable-baselines3: Reliable reinforcement learning implementations. J. Mach. Learn. Res. 2021, 22, 1–8. [Google Scholar]
- Dong, D.; Thacker, T.; Burgos, R.; Wang, F.; Boroyevich, D. On Zero Steady-State Error Voltage Control of Single-Phase PWM Inverters With Different Load Types. IEEE Trans. Power Electron. 2011, 26, 3285–3297. [Google Scholar] [CrossRef] [Scilit]












| Symbol | Value | |
|---|---|---|
| Rollout Buffe Size | ||
| Learning Rate | ||
| Clip Epsilon | 0.2 | |
| Discount Factor | 0.99 | |
| GAE | 0.95 | |
| Policy/Value Network Architecture | - | MLP: 64,64 |
| Batch Size/Updates | Batch size | 64 |
| Simplified Policy Naming | Full Configuration | Role in Evaluation |
|---|---|---|
| Hybrid-DGRL | PPO, PID, CBF | The framework proposed in this study integrates RL performance optimization, PID-based demonstrations, and formal safety guarantees. |
| PPO-Safety | PPO, CBF | The standard PPO policy, where actions are executed through a CBF safety filter. |
| PPO | PPO | The standard PPO baseline policy. |
| PID | PID (ZN-Tuned) | The classical control baseline. |
| PPO | PPO-Safety | PID | Hybrid-DGRL | |
|---|---|---|---|---|
| RMSE (cm) | 5.3707 ± 0.3689 | 5.7902 ± 0.9679 | 3.2774 ± 0.4187 | 4.7660 ± 1.2634 |
| SD of Position Jitter (cm) | 4.3671 ± 0.6893 | 5.5565 ± 1.1896 | 3.2687 ± 0.3956 | 4.3713 ± 1.0213 |
| TTR(s) | 4.1749 ± 3.5683 | 4.0892 ± 3.8723 | 6.2767 ± 3.9148 | 6.0031 ± 4.7395 |
| Number of Safety Boundary Violations | N/A | |||
| CVR (%) | 0.50501 | 0.0935 | N/A | 0.0675 |
| Convergence steps | N/A |
| PPO | PID | Hybrid-DGRL | |
|---|---|---|---|
| RMSE (cm) | 7.7982 | 5.0104 | |
| SD of Position Jitter (cm) | 4.1062 | 3.4011 | |
| TTR(s) | 4.5 | N/A | 2.34 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Huang, Y.-H.; Liu, E.-J.; Wu, B.-C.; Ning, Y.-J. Safe UAV Control Against Wind Disturbances via Demonstration-Guided Reinforcement Learning. Drones 2026, 10, 2. https://doi.org/10.3390/drones10010002
Huang Y-H, Liu E-J, Wu B-C, Ning Y-J. Safe UAV Control Against Wind Disturbances via Demonstration-Guided Reinforcement Learning. Drones. 2026; 10(1):2. https://doi.org/10.3390/drones10010002
Chicago/Turabian StyleHuang, Yan-Hao, En-Jui Liu, Bo-Cing Wu, and Yong-Jie Ning. 2026. "Safe UAV Control Against Wind Disturbances via Demonstration-Guided Reinforcement Learning" Drones 10, no. 1: 2. https://doi.org/10.3390/drones10010002
APA StyleHuang, Y.-H., Liu, E.-J., Wu, B.-C., & Ning, Y.-J. (2026). Safe UAV Control Against Wind Disturbances via Demonstration-Guided Reinforcement Learning. Drones, 10(1), 2. https://doi.org/10.3390/drones10010002

