EdgeTwin-DRL: Real-Time Counter-UAS Detection and Response Optimization Using Edge-Assisted Digital Twins and Multi-Agent Deep Reinforcement Learning
Abstract
1. Introduction
- (1)
- This paper proposed a four-layer architecture, which synchronously optimizes both detection confidence and compute allocation to a multi-agent DRL controller, as well as response selection, with a distributed edge processing layer coupled with a digital twin of protected airspace.
- (2)
- The digital twin is used in two ways: it is a synchronized operational model for multi-modal sensor fusion and it is a simulation environment for the training and pre-deployment of policies and checking the countermeasures.
- (3)
- This paper represents the decision problem as a cooperative multi-agent Markov decision process (MAMDP) and it is solved with PPO with decentralized actors and a centralized critic adapted to the distributed edge nodes.
- (4)
- An assessment of the framework was performed in a simulated testbed which emulates the edge tier and integrates radar, RF, EO/IR and acoustic observations. In the simulated environment, EdgeTwin-DRL reduces detection-to-response latency by up to 72% when compared to a cloud-centralized baseline with a 1.4% false-positive rate across three scenarios—urban, suburban, and open field.
2. Related Work
2.1. Counter-Drone Detection Technologies
2.2. Digital Twins for Security and Intrusion Detection
2.3. Deep Reinforcement Learning for Edge Computing and Resource Management
2.4. Summary and Research Gap
3. System Model
3.1. Deployment Scenario and Threat Model
3.2. Heterogeneous Sensor Architecture
Multimodal Observation Alignment and Fusion
3.3. Edge Computing Infrastructure
3.4. Digital Twin Design
3.4.1. Online State Synchronization
3.4.2. Simulation and Policy Training
3.4.3. Candidate-Response Evaluation and Decision Feedback
4. Proposed Methodology
4.1. Multi-Agent Markov Decision Process Formulation
4.2. Observation Space
4.3. Action Space Design
4.4. Reward Function Design
- Detection Reward (): The higher the reward that can be gained for an intrusion when it is detected earlier following geofence entry, the better.
- 2.
- Response Effectiveness Reward (): The response term combines the digital twin’s estimated probability of response success with the modeled collateral-cost penalty:
- 3.
- Resource Efficiency Penalty (): To discourage the controller from allocating excessive computational resources to every sensing stream, a resource-utilization penalty is defined as follows:where denotes the normalized computational-resource utilization of agent at time , is the number of agents, and is the resource-utilization penalty coefficient. This term penalizes unnecessary computational-resource consumption while allowing additional resources to be allocated when they contribute to the primary detection and response objectives.
- 4.
- Cooperation Bonus (): If two modalities are able to detect the same event within two seconds, a +3 bonus is added for each confirmation. This is not about giving incentives for more and more confidence in the same stream, but for the confirmation.
4.5. Cooperative Multi-Agent Proximal Policy Optimization
| Algorithm 1. EdgeTwin-DRL Training Procedure |
|
4.6. Digital Twin Integration in Training and Deployment
4.7. Computational and Communication Complexity
5. Experimental Setup
5.1. Computing Platform
5.2. Real-World Datasets for Sensor Model Calibration
5.3. Digital Twin Simulation Environment
5.4. Baseline Methods
5.4.1. Cloud-Centralized DRL (Cloud-DRL)
5.4.2. Independent PPO (IPPO)
5.4.3. Multi-Agent Deep Deterministic Policy Gradient (MADDPG)
5.4.4. Rule-Based Expert System (RB-Expert)
5.4.5. MAPPO Without Digital Twin (MAPPO-NoCalib)
5.5. Evaluation Metrics
5.6. Training Configuration and Hyperparameters
6. Results and Discussion
6.1. Overall Comparative Performance
6.2. Training Convergence Analysis
6.3. Scenario-Specific Performance
6.4. Per-Threat-Category Analysis
6.5. Ablation Studies
6.5.1. Reward Weight Sensitivity (Figure 6, Left)
6.5.2. Agent Scalability (Figure 6, Right)
6.6. Discussion
Safety and Deployment Considerations
7. Conclusions and Future Works
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Chamola, V.; Kotesh, P.; Agarwal, A.; Gupta, N.; Guizani, M. A comprehensive review of unmanned aerial vehicle attacks and neutralization techniques. Ad Hoc Netw. 2021, 111, 102324. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kang, H.; Joung, J.; Kim, J.; Kang, J.; Cho, Y.S. Protect your sky: A survey of counter unmanned aerial vehicle systems. IEEE Access 2020, 8, 168671–168710. [Google Scholar] [CrossRef] [Scilit]
- Kratky, M.; Farlik, J. Countering UAVs—The mover of research in military technology. Def. Sci. J. 2018, 68, 460–466. [Google Scholar] [CrossRef] [Scilit]
- Çetin, E.; Barrado, C.; Salami, E.; Pastor, E. Analyzing deep reinforcement learning model decisions with Shapley additive explanations for counter-drone operations. Appl. Intell. 2024, 54, 12095–12111. [Google Scholar] [CrossRef] [Scilit]
- Brown, A.D. Radar challenges, current solutions, and future advancements for the counter-UAS mission. IEEE Aerosp. Electron. Syst. Mag. 2023, 38, 34–50. [Google Scholar] [CrossRef] [Scilit]
- Mao, B.; Liu, J.; Wu, Y.; Kato, N. Security and privacy on 6G network edge: A survey. IEEE Commun. Surv. Tutor. 2023, 25, 1095–1127. [Google Scholar] [CrossRef] [Scilit]
- Shi, W.; Pallis, G.; Xu, Z. Edge computing. Proc. IEEE 2019, 107, 1474–1481. [Google Scholar]
- Mach, P.; Becvar, Z. Mobile edge computing: A survey on architecture and computation offloading. IEEE Commun. Surv. Tutor. 2017, 19, 1628–1656. [Google Scholar] [CrossRef] [Scilit]
- Fuller, A.; Fan, Z.; Day, C.; Barlow, C. Digital twin: Enabling technologies, challenges and open research. IEEE Access 2020, 8, 108952–108971. [Google Scholar] [CrossRef] [Scilit]
- Lu, Y.; Huang, X.; Dai, Y.; Maharjan, S.; Zhang, Y. Low-latency federated learning and blockchain for edge association in digital twin empowered 6G networks. IEEE Trans. Ind. Inform. 2021, 17, 5098–5107. [Google Scholar] [CrossRef] [Scilit]
- Kaufmann, E.; Bauersfeld, L.; Loquercio, A.; Müller, M.; Koltun, V.; Scaramuzza, D. Champion-level drone racing using deep reinforcement learning. Nature 2023, 620, 982–987. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, J.; Cao, X.; Yang, P.; Xiao, M.; Ren, S.; Zhao, Z.; Wu, D.O. Deep reinforcement learning-based resource allocation in multi-UAV-aided MEC networks. IEEE Trans. Commun. 2023, 71, 296–309. [Google Scholar] [CrossRef] [Scilit]
- Çetin, E.; Barrado, C.; Pastor, E. Countering a drone in a 3D space: Analyzing deep reinforcement learning methods. Sensors 2022, 22, 8863. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Han, S.-K.; Lee, J.-H.; Jung, Y.-H. Convolutional Neural Network-Based Drone Detection and Classification Using Overlaid Frequency-Modulated Continuous-Wave (FMCW) Range–Doppler Images. Sensors 2024, 24, 5805. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Al-Sa’d, M.F.; Al-Ali, A.; Mohamed, A.; Erbad, A.; Guizani, M. RF-based drone detection and identification using deep learning approaches: An initiative towards a large open source drone database. Future Gener. Comput. Syst. 2019, 100, 86–97. [Google Scholar] [CrossRef] [Scilit]
- Ding, S.; Guo, X.; Peng, T.; Huang, X.; Hong, X. Drone detection and tracking system based on fused acoustical and optical approaches. Adv. Intell. Syst. 2023, 5, 2300251. [Google Scholar] [CrossRef] [Scilit]
- Frid, A.; Ben-Shimol, Y.; Manor, E.; Greenberg, S. Drone detection using a fusion of RF and acoustic features and deep neural networks. Sensors 2024, 24, 2427. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, H.; Han, S.; Byeon, J.I.; Han, S.; Myung, R.; Joung, J.; Choi, J. CNN-based UAV detection and classification using sensor fusion. IEEE Access 2023, 11, 68791–68808. [Google Scholar] [CrossRef] [Scilit]
- Jeremiah, S.R.; El Azzaoui, A.; Xiong, N.N.; Park, J.H. A comprehensive survey of digital twins: Applications, technologies and security challenges. J. Syst. Archit. 2024, 151, 103120. [Google Scholar] [CrossRef] [Scilit]
- Balta, E.C.; Pease, M.; Moyne, J.; Barton, K.; Tilbury, D.M. Digital twin-based cyber-attack detection framework for cyber-physical manufacturing systems. IEEE Trans. Autom. Sci. Eng. 2023, 21, 1695–1712. [Google Scholar] [CrossRef] [Scilit]
- El-Hajj, M. Leveraging digital twins and intrusion detection systems for enhanced security in IoT-based smart city infrastructures. Electronics 2024, 13, 3941. [Google Scholar] [CrossRef] [Scilit]
- Krishnaveni, S.; Sivamohan, S.; Jothi, B.; Chen, T.M.; Sathiyanarayanan, M. TwinSec-IDS: An enhanced intrusion detection system in SDN digital-twin-based industrial cyber-physical systems. Concurr. Comput. Pract. Exp. 2025, 37, e8334. [Google Scholar] [CrossRef] [Scilit]
- Yigit, Y.; Nguyen, L.D.; Ozdem, M.; Kinaci, O.K.; Hoang, T.; Canberk, B.; Duong, T.Q. TwinPort: 5G drone-assisted data collection with digital twin for smart seaports. Sci. Rep. 2023, 13, 12310. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yigit, Y.; Bal, B.; Karameseoglu, A.; Duong, T.Q.; Canberk, B. Digital twin-enabled intelligent DDoS detection mechanism for autonomous core networks. IEEE Commun. Stand. Mag. 2022, 6, 38–44. [Google Scholar] [CrossRef] [Scilit]
- Iqbal, D.; Buhnova, B. Digital twin design for autonomous drones. In Proceedings of the 19th International Conference on Computer Science and Information Systems (FedCSIS), Belgrade, Serbia, 8–11 September 2024; pp. 119–130. [Google Scholar] [CrossRef] [Scilit]
- Xiong, Z.; Kang, J.; Niyato, D.; Ye, H.; Kim, D.I.; Poor, H.V. Deep reinforcement learning for mobile 5G and beyond. IEEE J. Sel. Areas Commun. 2019, 37, 2239–2253. [Google Scholar] [CrossRef] [Scilit]
- Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The surprising effectiveness of PPO in cooperative multi-agent games. In Advances in Neural Information Processing Systems; NeurIPS: Sydney, Australia, 2022; Volume 35, pp. 24611–24624. [Google Scholar] [CrossRef] [Scilit]
- Seid, A.M.; Boateng, G.O.; Mareri, B.; Sun, G.; Jiang, W. Multi-agent deep reinforcement learning for task offloading and resource allocation in multi-UAV-enabled IoT edge networks. IEEE Trans. Netw. Serv. Manag. 2021, 18, 4531–4547. [Google Scholar] [CrossRef] [Scilit]
- Suzuki, A.; Kobayashi, M.; Oki, E. Multi-agent deep reinforcement learning for cooperative computing offloading and route optimization in multi-cloud edge networks. IEEE Trans. Netw. Serv. Manag. 2023, 20, 4416–4434. [Google Scholar] [CrossRef] [Scilit]
- Zhao, L.; Zhao, Z.; Zhang, E.; Hawbani, A.; Al-Dubai, A.Y.; Tan, Z.; Hussain, A. A digital twin-assisted intelligent partial offloading approach for vehicular edge computing. IEEE J. Sel. Areas Commun. 2023, 41, 3386–3400. [Google Scholar] [CrossRef] [Scilit]
- Prevot, T.; Rios, J.; Kopardekar, P.; Robinson, J.E., III; Johnson, M.; Jung, J. UAS traffic management (UTM) concept of operations to safely enable low-altitude flight operations. In Proceedings of the AIAA Aviation Forum, Washington, DC, USA, 13–17 June 2016; pp. 2016–3292. [Google Scholar] [CrossRef] [Scilit]
- NVIDIA Corporation. NVIDIA Jetson AGX Orin Technical Brief; NVIDIA Corporation: Santa Clara, CA, USA, 2022; Available online: https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/ (accessed on 12 January 2026).
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems; NeurIPS: Sydney, Australia, 2017; Volume 30. [Google Scholar]
- Schulman, J.; Moritz, P.; Levine, S.; Jordan, M.; Abbeel, P. High-dimensional continuous control using generalized advantage estimation. In Proceedings of the International Conference on Learning Representations (ICLR), San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar]
- Tobin, J.; Fong, R.; Ray, A.; Schneider, J.; Zaremba, W.; Abbeel, P. Domain randomization for transferring deep neural networks from simulation to the real world. In Proceedings of the IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS), Vancouver, BC, Canada, 24–28 September 2017; pp. 23–30. [Google Scholar] [CrossRef] [Scilit]
- Filali, A.; Abouaomar, A.; Cherkaoui, S.; Kobbane, A.; Guizani, M. Multi-access edge computing: A survey. IEEE Access 2020, 8, 197017–197046. [Google Scholar] [CrossRef] [Scilit]
- Allahham, M.; Al-Sa’d, M.F.; Al-Ali, A.; Mohamed, A.; Khattab, T.; Erbad, A. DroneRF dataset: A dataset of drones for RF-based detection, classification, and identification. Data Brief 2019, 26, 104313. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yu, N.; Mao, S.; Zhou, C.; Sun, G.; Shi, Z.; Chen, J. DroneRFa: A large-scale dataset of drone radio frequency signals for detecting low-altitude drones. J. Electron. Inf. Technol. 2024, 46, 1147–1156. [Google Scholar]
- Brewczyński, K.D.; Życzkowski, M.; Cichulski, K.; Kamiński, K.A.; Petsioti, P.; De Cubber, G. Methods for assessing the effectiveness of modern counter-unmanned aircraft systems. Remote Sens. 2024, 16, 3714. [Google Scholar] [CrossRef] [Scilit]
- Zhao, W.; Queralta, J.P.; Westerlund, T. Sim-to-real transfer in deep reinforcement learning for robotics: A survey. arXiv 2020, arXiv:2009.13303. [Google Scholar]






| Ref. | Approach | DT | Edge | DRL | Multi-Agent | Joint Detection and Response |
|---|---|---|---|---|---|---|
| [5] | Radar counter-UAS survey | ✗ | ✗ | ✗ | ✗ | ✗ |
| [17] | RF + acoustic fusion DNN | ✗ | ✗ | ✗ | ✗ | ✗ |
| [18] | CNN sensor fusion (radar + camera) | ✗ | ✗ | ✗ | ✗ | ✗ |
| [13] | DRL drone interception (3D) | ✗ | ✗ | ✓ | ✗ | ✗ |
| [4] | Explainable DRL counter-drone | ✗ | ✗ | ✓ | ✗ | ✗ |
| [20] | DT cyber-attack detection (CPS) | ✓ | ✗ | ✗ | ✗ | ✗ |
| [22] | TwinSec-IDS (SDN + DT) | ✓ | ✗ | ✗ | ✗ | ✗ |
| [23] | TwinPort (DT + drone data) | ✓ | ✗ | ✗ | ✗ | ✗ |
| [28] | MADRL task offloading (UAV-MEC) | ✗ | ✓ | ✓ | ✓ | ✗ |
| [30] | DT-assisted offloading (vehicular) | ✓ | ✓ | ✓ | ✗ | ✗ |
| Ours | EdgeTwin-DRL | ✓ | ✓ | ✓ | ✓ | ✓ |
| Parameter | Value/Range | Randomization |
|---|---|---|
| Airspace dimensions | 5 km × 5 km × 0.5 km | Fixed |
| Decision step frequency | 10 Hz | Fixed |
| Episode duration | 300 s (3000 steps) | Fixed |
| Number of drones (Cat. I) | 1–2 | Uniform |
| Number of drones (Cat. II) | 1–3 | Uniform |
| Number of drones (Cat. III) | 3–5 | Uniform |
| Drone speed | 10–80 km/h | ±15% |
| Drone RCS | 0.001–0.1 m2 | ±25% |
| Radar detection range | 200–5000 m | ±15% |
| EO/IR detection range | 100–2000 m | ±20% |
| RF detection range | 50–3000 m | ±15% |
| Acoustic detection range | 10–300 m | ±20% |
| Sensor noise levels | SNR 5–30 dB | ±20% |
| Wind speed | 0–40 km/h | Uniform |
| Visibility | 500 m–10 km | Uniform |
| Precipitation | None/Light/Heavy | Categorical |
| Time of day | Day/Dusk/Night | Categorical |
| Clutter density (radar) | Low/Medium/High | ±30% |
| Edge inter-node latency | 0.3 ± 0.1 ms (log-normal) | Stochastic |
| Cloud round-trip latency | 50–200 ms (uniform) | Stochastic |
| Hyperparameter | Value |
|---|---|
| Total training steps | 5 × 106 |
| Discount factor (γ) | 0.99 |
| GAE parameter (λ) | 0.95 |
| PPO clipping parameter (ε) | 0.2 |
| Learning rate (actor and critic) | 3 × 10−4 (Adam) |
| Batch size | 4096 transitions |
| PPO epochs per update | 15 |
| Mini-batch size | 512 transitions |
| Actor hidden layers | [256, 128, 64] FC + ReLU + LayerNorm |
| Critic hidden layers | [256, 128, 64] FC + ReLU + LayerNorm |
| Entropy coefficient (c2) | 0.01 |
| Value loss coefficient (c1) | 0.5 |
| Max gradient norm | 0.5 |
| Number of agents (n) | 4 |
| Episode length | 3000 steps (300 s simulated) |
| Parallel environments | 64 |
| Reward weights (w1, w2, w3, w4) | 1.0, 0.8, 0.1, 0.3 |
| Engagement threshold (τ_eng) | 0.7 |
| Random seeds | 5 (0, 42, 123, 456, 789) |
| Method | DA (%) | F1 (%) | FPR (%) | RSR (%) | Lat. (ms) | GPU (%) | |
|---|---|---|---|---|---|---|---|
| RB-Expert | 87.2 ± 1.3 | 81.5 ± 1.8 | 6.8 ± 0.9 | 72.4 ± 2.1 | 1417 ± 95 | 25.0 ± 0.0 | 124 ± 12 |
| Cloud-DRL | 91.8 ± 1.5 | 87.3 ± 1.9 | 4.2 ± 0.7 | 78.1 ± 2.4 | 1910 ± 234 | 22.0 ± 1.2 | 138 ± 18 |
| IPPO | 93.1 ± 1.2 | 89.6 ± 1.5 | 3.5 ± 0.6 | 81.3 ± 2.0 | 970 ± 115 | 58.3 ± 3.1 | 148 ± 15 |
| MADDPG | 94.5 ± 1.0 | 91.2 ± 1.3 | 2.8 ± 0.5 | 84.7 ± 1.8 | 880 ± 102 | 52.1 ± 2.8 | 176 ± 14 |
| MAPPO-NoCalib | 95.8 ± 0.8 | 93.1 ± 1.1 | 2.1 ± 0.4 | 87.5 ± 1.5 | 713 ± 75 | 45.6 ± 2.4 | 221 ± 16 |
| EdgeTwin-DRL | 97.3 ± 0.6 | 95.8 ± 0.8 | 1.4 ± 0.3 | 92.1 ± 1.2 | 527 ± 52 | 41.2 ± 2.0 | 283 ± 13 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Alnaim, A.K.; Alwakeel, A.M. EdgeTwin-DRL: Real-Time Counter-UAS Detection and Response Optimization Using Edge-Assisted Digital Twins and Multi-Agent Deep Reinforcement Learning. Sensors 2026, 26, 5632. https://doi.org/10.3390/s26175632
Alnaim AK, Alwakeel AM. EdgeTwin-DRL: Real-Time Counter-UAS Detection and Response Optimization Using Edge-Assisted Digital Twins and Multi-Agent Deep Reinforcement Learning. Sensors. 2026; 26(17):5632. https://doi.org/10.3390/s26175632
Chicago/Turabian StyleAlnaim, Abdulrahman K., and Ahmed M. Alwakeel. 2026. "EdgeTwin-DRL: Real-Time Counter-UAS Detection and Response Optimization Using Edge-Assisted Digital Twins and Multi-Agent Deep Reinforcement Learning" Sensors 26, no. 17: 5632. https://doi.org/10.3390/s26175632
APA StyleAlnaim, A. K., & Alwakeel, A. M. (2026). EdgeTwin-DRL: Real-Time Counter-UAS Detection and Response Optimization Using Edge-Assisted Digital Twins and Multi-Agent Deep Reinforcement Learning. Sensors, 26(17), 5632. https://doi.org/10.3390/s26175632

