A Review of AI-Enabled UAV-Based Systems for Defense Applications
Highlights
- This paper provides an up-to-date review of Artificial Intelligence (AI)-enabled Unmanned Aerial Vehicle (UAV)-based defense systems, covering autonomous air combat, cooperative UAV operations, path planning, target tracking, and cybersecurity.
- An integrated system architecture and a functional classification framework are presented together with an analysis of major AI paradigms, recent advances, open challenges, and future research directions.
- This paper provides a unified reference by consolidating recent developments across AI-enabled UAV-based defense technologies within a common analytical framework.
- The identified challenges and future directions can guide the development of more autonomous, resilient, secure, and scalable next-generation UAV-based defense systems.
Abstract
1. Introduction
1.1. Contributions
- Overview of AI-Enabled UAV-Based Defense Systems: This paper provides a foundational overview of UAV platforms, sensing and perception technologies, communication infrastructures, and AI techniques that underpin intelligent UAV operation in defense environments.
- Integrated Architecture and Functional Classification Framework: This paper presents an integrated system architecture and a functional classification framework for AI-enabled UAV-based defense systems. The proposed framework provides a unified system-level perspective by organizing the key functional layers, including UAV platforms, sensing and perception, communication and networking, AI-enabled intelligence and computing, autonomous control and decision-making, cybersecurity and resilience, and the operational environment. Furthermore, the interactions among these functional layers are analyzed to provide a holistic perspective on the AI-enabled UAV-based defense ecosystem.
- Review of Major Defense Application Domains: Complementing existing review papers which primarily organize the literature around individual technologies or specific application areas, this paper classifies recent state-of-the-art research into four complementary and interconnected defense operational domains to provide a unified perspective on AI-enabled UAV-based defense systems: (i) autonomous air combat, cooperative multi-UAV operations, and swarm combat intelligence; (ii) autonomous navigation and path planning; (iii) target tracking, detection, and classification; and (iv) cybersecurity, electronic warfare protection, and resilient UAV operation. These domains were identified through a search of major scientific databases, including ACM Digital Library, IEEE Xplore, PubMed, Web of Science, and others using keywords related to AI, UAVs, autonomous drones, autonomous air combat, swarm intelligence, multi-UAV systems, path planning, autonomous navigation, target tracking, cybersecurity, electronic warfare, resilient communications, and intelligent defense systems. The primary body of literature reviewed in this paper consists of peer-reviewed journal publications written in English and published between January 2023 and June 2026. Earlier publications were included where necessary to cite seminal contributions or provide technical background, while conference papers, publicly available datasets, software frameworks, standards, and online resources were referenced when directly relevant to the surveyed technologies. Furthermore, studies involving civilian or dual-use UAV applications were considered when they introduced AI methodologies or communication technologies that are equally applicable to defense-oriented UAV systems. Inclusion criteria considered studies directly investigating AI-enabled UAV technologies and defense-oriented applications, while exclusion criteria omitted works focusing solely on civilian UAV use cases, conventional non-AI methodologies, or systems not involving UAVs. Following the literature search, duplicate records were removed where applicable, then the remaining publications were screened based on their titles and abstracts to assess their relevance to the scope of this review, after which the full text of the shortlisted studies was examined to determine their eligibility according to the predefined inclusion and exclusion criteria.
- Analysis of AI Techniques and Their Operational Roles: This paper investigates the application of major AI paradigms, including Machine Learning (ML), Deep Learning (DL), Reinforcement Learning (RL), Deep RL (DRL), Multi-Agent RL (MARL), transformer-based architectures, Federated Learning (FL), and Explainable AI (XAI). Their effectiveness, advantages, limitations, scalability characteristics, and practical tradeoffs are analyzed across diverse defense-oriented UAV applications.
- Lessons Learned and Future Research Directions: This paper synthesizes key findings from recent studies, identifies current limitations and open research challenges, and outlines promising future research directions.
1.2. Structure
2. Previous Review Papers
3. Background and System Architecture of AI-Enabled UAV-Based Defense Systems
3.1. UAV-Based Systems in Defense Applications
3.2. AI Techniques for Intelligent UAV-Based Defense Systems
3.3. Integrated System Architecture of AI-Enabled UAV-Based Defense Systems
4. Autonomous Air Combat and Cooperative UAV Operations
4.1. Autonomous Air Combat Maneuver Decision-Making
4.2. Cooperative Multi-UAV Air Combat
| Reference | Objective | AI/ML Models | Operational Scenario | UAV Configuration | Main Components | Key Results |
|---|---|---|---|---|---|---|
| Wang, B. et al., 2024 [40] | Cooperative multi-UAV air combat maneuver decision-making | E-MATD3, evolutionary MARL, attention-based learning | 3D within-visual-range air combat | 5 vs. 5 cooperative combat UAV teams | CTDE, evolutionary population training, curriculum learning, attention-based policies, death masking, crossover and mutation operators, TrueSkill-based selection | Outperformed MADDPG, MATD3, DYAN-SUM, and ATT-MATD3 while enabling coordinated target pursuit and cooperative combat maneuvers |
| Xu et al., 2025 [44] | Hierarchical cooperative multi-UAV air combat decision-making | Hierarchical MARL, PPO, attention-based value decomposition | High-fidelity multi-UAV air combat | 3 vs. 3–6 vs. 6 UAV operational scenarios | HRL, CTDE, PPO, GAE, virtual-opponent modeling, attention-based value decomposition, pretrained low-level maneuver policies | Outperformed MAPPO and related MARL baselines while enabling coordinated target selection and collaborative target engagement |
| Yang et al., 2024 [47] | Cooperative UAV swarm operational scenario and tactical coordination | DP-MADDPG, PER-enhanced MADDPG | POMG-based multi-UAV battlefield | Eight combat UAVs and two reconnaissance UAVs per team | CTDE, dual-critic learning, PER, local/global reward optimization, cooperative coordination | Achieved a convergent reward of 695.5 vs. 592.9 (MADDPG) and 344.8 (ILDDPG), with 96% and 80.5% win rates against rule-based and ILDDPG opponents, respectively |
| Ren et al., 2023 [48] | Cooperative maneuver generation under uncertainty | MADDPG, Bayesian inference, dynamic game theory | Incomplete-information multi-UAV air combat | 2 vs. 2, 2 vs. 3, and 3 vs. 2 combat scenarios | PBE analysis, DBN-based intention inference, CTDE, cooperative situation assessment | Improved convergence and exploitability characteristics compared with TD3 while generating adaptive cooperative combat tactics |
| Ding et al., 2025 [50] | Collaborative multi-UAV air combat decision-making | MAPPO-LDC, PPO-based MARL | Adversarial 3D air combat | 2 vs. 2, 4 vs. 4, and 6 vs. 6 UAV teams | CTDE, dual-center critics, delayed policy updates, curriculum-style training | Achieved a 93% win rate in 2 vs. 2 combat, improving over MAPPO by 89% while reducing draw rates by 84% |
| Tang et al., 2026 [52] | Risk-aware cooperative multi-UAV air combat decision-making | Hierarchical RL, IQN-based distributional RL, Stackelberg game | A2/AD cooperative air combat | Multi-UAV cooperative combat swarm | Stackelberg-game task allocation, hierarchical RL, IQN, CVaR-based risk-aware decision-making, resilient predictive control | Outperformed MAPPO and CBBA while improving mission survivability, tactical coordination, and robustness, with over 60% lower computational load |
| Wang, H. et al., 2024 [53] | Hierarchical cooperative multi-UAV combat decision-making | HRL, QMIX-based MARL | JSBSim-based adversarial air combat | 4 vs. 4 and 8 vs. 8 UAV teams | DEC-POMDP modeling, hierarchical maneuver and attack decision layers, experience decomposition, QMIX value decomposition | Achieved win rates of 64.7% and 82.9% in 4 vs. 4 and 8 vs. 8 combat, respectively, outperforming VDN, COMA, and QMIX |
| Luo et al., 2024 [55] | Cooperative pursuit-evasion maneuver decision-making | CommNet-A2C, communication-aware MADRL | 3D pursuit-evasion air combat | Two pursuit UAVs versus one evasive UAV | Communication interaction layers, GRU memory modules, shared rewards, communication protocol learning, NASA-inspired maneuver primitives | Converged after approximately 700 episodes and achieved a pursuit success rate of 92.7%, compared with 68.5% for CommNet-REINFORCE |
| Jianhong et al., 2025 [56] | Cooperative combat planning and UAV task scheduling | MDRL-DQN, multimodal DRL | Electronic warfare attack-defense environments | Multi-UAV attack and defense teams | Multimodal fusion, CNN-based image processing, RNN-based sensor processing, adaptive rewards, self-attention task prioritization, actor–critic evaluation | Achieved task success rates of 89.6% and 94.8% in dispersed and concentrated defense scenarios, respectively, outperforming PPO, A3C, and DDPG |
| Han et al., 2025 [58] | Communication-efficient cooperative UAV air combat | SIIS-MARL, ToM-based MARL | Communication-constrained multi-UAV combat | 2 vs. 2–16 vs. 16 UAV swarms | ToM-based intention inference, sparse communication, LSTM temporal modeling, CTDE | Reduced communication overhead and achieved winning rates of 99.3% and 76.8% in 2 vs. 2 and 16 vs. 16 combat, respectively, surpassing MADDPG, COMA, DAACMP, and MASIA |
4.3. Large-Scale Swarm Air Combat and Scalable Combat Intelligence
5. Path Planning and Autonomous Navigation
6. Target Tracking, Detection, and Classification
| Reference | Objective | AI/ML Models | Operational Scenario | UAV Configuration | Main Components | Key Results |
|---|---|---|---|---|---|---|
| Lee et al., 2024 [76] | Autonomous military reconnaissance and concealed target detection | YOLOv4-tiny, PPO, ResNet encoder | Battlefield ISR and concealed tank reconnaissance | Reconnaissance UAV with camera, LRF, and GNSS | Unity-based simulator, StyleGAN3 data augmentation, PID flight control, RL-guided viewpoint optimization, confidence enhancement | After 5M training steps, image + vector observations achieved reward 230 and episode length 987, compared with 137 and 734 for vector-only observations |
| Salameh et al., 2024 [77] | Autonomous target detection under uncertain mobility | Q-learning, SARSA | Dynamic surveillance environments | Single UAV with charging station | MDP formulation, Markov-chain target mobility, energy-aware navigation, adaptive surveillance optimization | Detection rates of 29– vs. 14– for baseline methods; detection time reduced from 7 to units and energy consumption reduced by up to |
| Ud Din et al., 2025 [78] | Robust UAV detection under adversarial attacks | CNNs, Q-learning, adversarial ML | Urban and rural UAV surveillance | AI-enabled UAV detection systems | VisDrone dataset, adaptive image segmentation, adversarial training, defensive distillation, feature-space regularization | Detection accuracy reached in rural daytime environments and in urban nighttime scenarios; adversarial defenses improved accuracy and precision to nearly |
| Sayed et al., 2024 [79] | Radar-based UAV detection and classification | DopplerNet-based CNN | Counter-UAV radar surveillance | Seven UAV types using SISO and MIMO radar sensing | 79 GHz FMCW radar, HFSS electromagnetic simulations, range-Doppler processing, 2Tx/4Rx MIMO architecture | Classification accuracy of (synthetic) and >98% (real data); detection probability near at −17.5 dB SNR and false-alarm probability |
| Han et al., 2025 [81] | GNSS-denied autonomous visual target tracking | SSD-MobileNetV2, KCF, Kalman Filter, Re-ID | Reconnaissance and autonomous tracking missions | Quadrotor UAV with Raspberry Pi 5 | IBVS, hierarchical PID control, Kalman prediction, Re-ID, SIL/HIL validation | Achieved IoU values of – and processing speeds of 11–30 FPS while maintaining robust target tracking under GNSS-denied conditions |
| Kant et al., 2023 [88] | Resilient UAV trajectory prediction during communication failures | LSTM-AE | Communication-loss tracking and localization | Fixed-wing UAV and GDT infrastructure | Trajectory preprocessing, LSTM autoencoder, ECEF/NED transformations, radar look-angle estimation | Approximately prediction accuracy versus for Vanilla LSTM; loss reduced to with MAE and RMSE |
| Zuo et al., 2025 [91] | UAV-to-UAV small-target detection for counter-UAV operations | UAV-STD, attention-based DL | Dynamic air-to-air environments | UAV surveillance and counter-UAV systems | AMSTD module, SSP Head, FPN feature fusion, NWD-CIoU loss, Coordinate Attention | AP50 = and APs = ; AP50 = for targets below 10 × 10 pixels while using only 8M parameters and GFLOPs |
| Kim et al., 2025 [95] | Detection, interception, and inspection of balloon-borne aerial threats | YOLOv8, CNNs | Hazardous-balloon interception and aerial defense | Cooperative four-drone swarm | YOLOv8 detection, swarm-net interception, CNN-based X-ray inspection, decentralized coordination, HITL supervision | mAP@0.5, recall, false-positive rate, overall detection accuracy, capture success rate, and response time <7.8 s |
7. Cybersecurity, Electronic Warfare Protection, and Resilient UAV Operation
| Reference | Objective | AI/ML Models | Operational Scenario | UAV Configuration | Main Components | Key Results |
|---|---|---|---|---|---|---|
| Masadeh et al., 2024 [96] | RL-based surveillance and intrusion detection | Q-learning, SARSA | Stochastic target-mobility environments | Single energy-constrained UAV with charging station | MDP formulation, Markov-chain mobility, adaptive exploration, energy-aware navigation | Detection rates of 29–32% vs. 14–21%; detection time reduced from 7 to units and energy per detection from 36–41 to 18 units |
| Ma et al., 2025 [98] | Covert UAV navigation spoofing and deceptive trajectory manipulation | SIE-SAC, SAC, MERL | Counter-UAV GPS/INS spoofing | Single UAV with GPS/INS-integrated navigation | SIE, NIS-based stealth constraints, GPS/INS-integrated KF, MDP-based spoofing optimization | Return convergence approaching 500 vs. 300 for SAC; smoother deceptive trajectories while maintaining NIS below detection thresholds |
| Zhao et al., 2026 [101] | AI-driven anti-jamming communication | DT, DRL | Electronic warfare and adaptive jamming | Multiple UAV relays and base station | Offline trajectory learning, cloud-edge collaboration, adaptive reconfiguration, jammer detection | BER , SNR = 25.5 dB, PDR = 97.1%, detection time s, latency ms |
| Hickling et al., 2023 [102] | Adversarial attack detection for DRL-based navigation | DDPG, PER, SHAP, CNN-AD, LSTM-AD | Autonomous navigation under FGSM/BIM attacks | LiDAR-equipped UAV | APF guidance, XAI, SHAP-value analysis, adversarial detectors | Navigation success reduced from to under BIM attacks; CNN and LSTM detectors achieved approximately and accuracy |
| Mynuddin et al., 2024 [106] | Trojan attack detection and trustworthy UAV navigation | MobileNetV3, DeepFool-UAP | Backdoor attacks against autonomous navigation | DroNet (ResNet-8) UAV navigation system | Trigger-pattern injection, UAP generation, feature extraction, Trojan classifier | Reduced Trojan attack success rate from to |
| Hadi et al., 2024 [110] | Collaborative IDS for UAV communication networks | FFCNN | Cyberattack-prone UAV networks | Multiple UAVs and distributed monitoring nodes | UAVIDS dataset, ReLU-TaLU activation, distributed event correlation, trust-aware management, incident response | Detection accuracy of with high precision and sensitivity and low false alarm rates |
| Agnew et al., 2024 [112] | Cybersecurity enhancement of SD-UAV relay networks | Predictive Queuing Analysis | Urban battlefield communication environments | 40 SD-UAV relay nodes | SDN, Jackson Open Networks, predictive queuing analysis, OpenFlow routing, ONOS controller | Theoretical estimates of inter-arrival time, transmission delay, and packet count closely matched simulation results, validating the predictive model |
| Jadav et al., 2023 [114] | Secure UAV communication over 6G networks | RF, blockchain | Military reconnaissance and battlefield communications | UAV swarms with U2U and satellite links | Block-USB, IPFS storage, 6G communications, ML-based intrusion detection | RF achieved near- classification accuracy; processing delay reduced from ms to ms and scalability improved from 45 to 63 units |
| Martínez Beltrán et al., 2025 [116] | Resilient decentralized FL for military reconnaissance | DFL, VGG16 | Adversarial SAR reconnaissance and mosaic warfare | Four military aircraft with SAR sensing | Flighter, adaptive differential privacy, asynchronous communications, situational awareness defense, decentralized aggregation | F1-scores of (MSTAR) and (SAMPLE); maintained >89–90% F1 under geopositioning and trajectory manipulation attacks |
| Nguyen et al., 2026 [118] | Resilient autonomous UAV swarm coordination | Agentic AI, LLMs, TinyLLaMA, GPT-4.1 | Mission-critical edge-enabled swarm environments | Standalone, edge-enabled, and hybrid UAV swarms | YOLO detection, U-Net segmentation, mesh networking, multimodal sensing, distributed reasoning, edge-assisted coordination | Coverage rates of 98–100%; 8-UAV missions completed in <25 min and 12-UAV missions in <17 min |
8. Lessons Learned and Research Insights
- Autonomous Air Combat and Cooperative UAV Operations
- –
- Prominent Role of DRL and MARL: The reviewed studies indicate that DRL and MARL have emerged as among the most widely investigated AI paradigms for autonomous air combat, maneuver decision-making, cooperative engagement, and swarm-level tactical coordination. PPO, DQN, DDPG, MATD3, MAPPO, QMIX, hierarchical RL, and transformer-assisted MARL frameworks were widely employed to enable UAVs to learn adaptive maneuvering, target engagement, pursuit, evasion, and cooperative attack strategies in adversarial environments.
- –
- Emergent Cooperative Combat Behaviours: Several MARL-based studies demonstrated that cooperative UAV teams can autonomously develop advanced tactical behaviors such as pincer attacks, coordinated flanking, decoy maneuvers, cross-attacks, target allocation, cooperative encirclement, fire-attraction tactics, synchronized target engagement, and collaborative target elimination. These behaviors emerged through reward-driven learning rather than explicitly programmed combat rules.
- –
- Increasing Importance of Opponent Modeling: Recent studies increasingly consider opponent capabilities, tactical behaviors, and threat levels as part of the decision-making process. Explicit modeling of adversarial agents enables more informed maneuver planning, adaptive target prioritization, and improved tactical coordination in dynamic air-combat scenarios.
- –
- Scalability of Swarm Combat Intelligence: Large-scale swarm combat studies increasingly employed attention-assisted, transformer-based, communication-aware, and value-decomposition architectures to improve coordination across formations ranging from small teams to large-scale engagements such as 16 vs. 16 and 20 vs. 20 operational scenarios.
- –
- Growing Adoption of Transformer and Attention Mechanisms: Recent studies increasingly incorporated transformer architectures and attention mechanisms for battlefield representation, scalable swarm coordination, target prioritization, temporal feature extraction, and adaptive tactical reasoning. Attention-based learning consistently improved situational awareness and enabled UAVs to focus on tactically relevant allies, opponents, and battlefield events.
- –
- Importance of Hierarchical Learning: Hierarchical RL frameworks were repeatedly employed to decompose complex combat tasks into multiple decision layers. More recently, some studies have combined hierarchical RL with game-theoretic modeling to support cooperative task allocation, adversarial decision-making, and risk-aware tactical planning under battlefield uncertainty. High-level policies typically generated tactical objectives, virtual targets, or risk-aware tactical decisions, while low-level policies executed maneuver commands. This decomposition reduced exploration difficulty, improved training efficiency, enhanced scalability in large combat environments, and supported resilient decision-making under battlefield uncertainty.
- –
- Communication-Efficient Cooperation: Communication-aware MARL frameworks increasingly adopted sparse communication, intention inference, CommNet-based interaction, and attention-assisted information exchange to improve cooperative coordination while reducing communication overhead.
- –
- Role of Realistic Simulation Environments: The reviewed air combat studies increasingly relied on high-fidelity simulators and digital environments such as Unity3D, JSBSim, PyAirCombat, MaCA, and digital-twin battlefield platforms. These environments enabled the evaluation of realistic flight dynamics, radar sensing, missile engagements, probabilistic damage models, and cooperative combat tactics.
- –
- Growing Interest in Explainability: Explainable RL approaches employed reward decomposition, tactical-region analysis, and visual explanation mechanisms to improve the interpretability of autonomous combat policies and facilitate operator understanding.
- Autonomous Navigation and Path Planning
- –
- Broad Adoption of DRL and Hybrid Navigation Frameworks: The reviewed navigation studies show that frameworks based on PPO, DQN, actor–critic, and Q-learning can substantially improve autonomous UAV navigation, obstacle and collision avoidance, target visitation, and tactical path planning. Several studies further combined DRL with optimization-based planners, semantic terrain analysis, and sampling-based path planning techniques to improve robustness and mission effectiveness.
- –
- Navigation Beyond Shortest-Path Optimization: Several studies emphasized that defense-oriented UAV navigation must consider survivability, concealment, threat avoidance, semantic terrain information, situational awareness, and adversarial exposure rather than merely minimizing path length. AI-enabled navigation frameworks are increasingly able to exploit terrain cover, avoid regions that are under enemy threat, and adapt their trajectories according to evolving battlefield conditions.
- –
- GNSS-Denied and Communication-Constrained Operation: Many reviewed studies investigated GNSS-denied, communication-constrained, and interference-prone environments using telemetry prediction, DRL-based autonomous control, and sensor-based navigation to improve operational resilience.
- –
- Energy-Aware Mission Execution: Energy consumption emerged as an important design consideration in target visitation, surveillance, swarm operation, and path planning studies. Several frameworks incorporated residual energy, flight duration, mission time, and energy-aware reward mechanisms to improve mission efficiency and operational endurance.
- –
- Hybrid AI and Traditional Planning Synergy: Multiple studies demonstrated that combining AI-driven decision-making with traditional path planning algorithms such as A*, RRT, and genetic optimization methods can improve convergence, navigation reliability, computational efficiency, and path quality in complex operational environments.
- –
- Simulation-Based Validation and Transfer Challenges: Many navigation studies relied on Gazebo, Parrot Sphinx, AirSim, Unreal Engine, Unity ML-Agents, and battlefield simulation environments. Although several studies incorporated real UAV experiments and telemetry-based validation, simulation remained the main evaluation methodology.
- Target Tracking, Detection, and Classification
- –
- Advances in AI-Driven Battlefield Perception: The reviewed studies demonstrate that CNNs, YOLO-based detectors, SSD-MobileNetV2, DopplerNet, attention-assisted architectures, Kalman filtering, KCF tracking, and RL-based perception frameworks can improve UAV-based target detection, classification, tracking, reconnaissance, and situational awareness.
- –
- Persistent Small-Target Detection Challenges: Detecting small UAVs and distant aerial targets remains particularly challenging due to sparse visual features, long observation distances, cluttered backgrounds, rapid motion, viewpoint changes, and adverse environmental conditions. Attention mechanisms, multiscale feature fusion, specialized detection heads, and enhanced loss functions were commonly employed to address these limitations.
- –
- Trend Toward Multimodal Sensing: The reviewed studies increasingly combined multiple sensing modalities, including electro-optical imagery, radar sensing, MIMO FMCW radar, range-Doppler processing, telemetry information, LiDAR measurements, onboard sensors, and X-ray inspection. Multimodal sensing consistently improved robustness, detection reliability, classification accuracy, and operational awareness.
- –
- Embedded and Real-Time AI Operation: Several studies demonstrated real-time or near-real-time operation using lightweight detectors, embedded AI hardware, Raspberry Pi platforms, onboard processors, and computationally efficient tracking pipelines. These studies highlight the growing emphasis on computationally efficient AI for onboard deployment.
- –
- Adversarial Robustness of Perception Systems: The reviewed adversarial ML-based studies demonstrated that AI-enabled UAV perception systems remain vulnerable to adversarial perturbations, spoofing attacks, deceptive inputs, and cyber–physical manipulation. The reviewed literature explored adversarial training, defensive distillation, feature-space regularization, and explainability-based detection to improve perception robustness.
- –
- HITL and Human-Supervised Decision-Making: Although AI has the potential to improve autonomous target detection, tracking, and classification, safety-critical defense applications continue to benefit from HITL and human-supervised operation. The reviewed studies indicate that combining AI-enabled perception with operator oversight enhances decision reliability, supports target-engagement verification, and increases trust in mission-critical environments.
- Cybersecurity, Electronic Warfare Protection, and Resilient UAV Operation
- –
- Cyber–Physical Threats to UAV-Based Systems: The reviewed cybersecurity studies addressed a broad spectrum of cyber–physical threats, including GPS spoofing, GPS jamming, communication attacks, actuator faults, adversarial attacks, intrusion attempts, intelligent electromagnetic jamming, and attacks targeting AI-driven decision-making mechanisms. These threats can significantly affect UAV navigation, sensing, communication, control, and mission execution.
- –
- AI-Enabled Intrusion and Attack Detection: DL, CNNs, LSTMs, transformer-based architectures, FL, and RL were widely employed for intrusion detection, anomaly identification, spoofing detection, anti-jamming control, adversarial defense, and fault diagnosis. The reviewed studies consistently reported strong attack-detection and mitigation performance across diverse UAV security scenarios.
- –
- Distributed and Decentralized Resilience: Collaborative intrusion detection, decentralized FL, trust-aware node management, distributed event validation, and SDN were investigated to reduce single points of failure and improve resilience in distributed UAV networks and swarms.
- –
- Anti-Jamming and Secure Communication: The reviewed works showed that AI-driven anti-jamming methods, DRL-based communication adaptation, frequency hopping, spread-spectrum techniques, communication reconfiguration, and hybrid mitigation strategies can improve communication robustness under intelligent jamming and adverse wireless conditions.
- –
- Explainability for Cyber Defense: Explainability-driven mechanisms, including SHAP-based analysis and feature-activation interpretation, demonstrated the ability to identify abnormal behaviors and adversarial manipulations. These approaches improve the interpretability of AI-enabled cyber defense mechanisms and support operator understanding of security decisions.
- Simulation-Based Evaluation and Practical Validation
- –
- Central Role of Simulation-Based Evaluation: Across the reviewed application domains, simulation environments, digital twins, synthetic datasets, physics-based simulators, battlefield simulators, and RL training platforms constituted the primary means of evaluating AI-enabled UAV autonomy, air combat, navigation, perception, and cybersecurity mechanisms.
- –
- Increasing Emphasis on Practical Validation: Several reviewed studies complemented simulation-based evaluation with real UAV experiments, telemetry datasets, embedded AI hardware, onboard implementation, SIL, and HIL validation. These efforts improve practical relevance and increase confidence in the applicability of the proposed AI frameworks.
- Key Lessons for Future AI-Enabled UAV-Based Defense Systems
- –
- AI as an Enabler of Intelligent Defense Operations: The reviewed literature demonstrates that many defense problems, including autonomous air combat, cooperative multi-UAV coordination, navigation in dynamic and contested environments, multi-target tracking, and cyberattack mitigation, are difficult to address using traditional rule-based approaches due to the complexity, uncertainty, and rapidly changing nature of battlefield conditions. AI-based methods, particularly DRL, MARL, DL, and hybrid learning frameworks, enable adaptive decision-making, continuous learning, and improved generalization in scenarios where manually designed rules are difficult to develop, maintain, or scale.
- –
- Robust Autonomous Operation Across Defense Applications: Across the reviewed domains, robust UAV operation increasingly relies on resilient perception, secure communications, reliable autonomous decision-making, and effective operation under uncertain battlefield conditions.
- –
- Importance of Trustworthy and Interpretable AI: The reviewed literature highlights the increasing role of trustworthy and interpretable AI in supporting transparency, operator confidence, and responsible deployment of autonomous UAV-based defense systems.
- –
- Toward Integrated UAV-Based Defense Ecosystems: An emerging trend is the convergence of perception, communications, autonomy, cybersecurity, and distributed intelligence into integrated UAV-based defense ecosystems capable of supporting increasingly complex missions.
9. Challenges, Open Issues, and Future Research Directions
- Autonomous Air Combat Intelligence and Cooperative UAV Operations: According to the reviewed literature, several important challenges remain, including partial observability, the design of effective reward functions, training instability, limited interpretability of learned policies, communication overhead in cooperative UAV networks, scalability to large-scale combat formations, and the limited operational validation of policies trained primarily in simulation. Moreover, increasing battlefield complexity and communication constraints further complicate cooperative decision-making in large-scale aerial engagements. Future research should focus on communication-aware and scalable MARL frameworks, hierarchical and risk-aware cooperative decision-making, distributional RL for uncertainty-aware planning, improved learning efficiency through curriculum learning and self-play strategies, explainable tactical reasoning, and more realistic battlefield validation under communication-constrained and adversarial operational environments.
- Autonomous Navigation and Mission Planning: Despite the advances of recent AI-driven navigation frameworks, maintaining reliable navigation under localization uncertainty, communication degradation, adversarial interference, dynamic obstacles, uncertain environmental conditions, and long-duration missions remains challenging. In addition, robust autonomous decision-making under incomplete situational awareness continues to be an open problem. Future research should investigate resilient multi-sensor navigation, adaptive mission re-planning, learning-assisted navigation under degraded sensing and communication conditions, and validation in operationally representative battlefield environments.
- Target Tracking, Detection, and Battlefield Perception: Based on the reviewed literature, robust perception remains challenging in cluttered environments, adverse weather, camouflage, low-visibility conditions, long-range observations, and low-SNR scenarios. In addition, limited defense-oriented datasets, computational constraints, and insufficient operational validation continue to hinder practical deployment. Future research should emphasize robust multi-modal perception, computationally efficient detection algorithms suitable for onboard deployment, improved radar-based recognition, more representative defense-oriented datasets, and evaluation under realistic battlefield conditions.
- Cybersecurity, Electronic Warfare, and Resilient UAV Operations: The reviewed literature demonstrates increasing attention toward protecting UAV-based defense systems against spoofing, jamming, communication attacks, malicious data manipulation, adversarial AI attacks, and distributed cyber threats. Although AI-assisted intrusion detection, FL, and XAI have demonstrated promising capabilities, ensuring trustworthy and resilient autonomous operation under contested battlefield environments remains an important research challenge. Future work should focus on adversarially robust AI, secure collaborative learning, lightweight distributed intrusion detection, communication-aware cyber defense mechanisms, and XAI techniques that improve operator confidence and mission reliability.
- Resource-Efficient Embedded AI: Many reviewed studies highlight the need to execute increasingly sophisticated AI algorithms directly onboard UAV platforms while operating under stringent computational, memory, energy, and communication constraints. Maintaining reliable perception, navigation, and autonomous decision-making within these resource limitations remains challenging, particularly for long-duration and communication-constrained missions. Future research should investigate lightweight neural architectures, computationally efficient onboard inference, communication-efficient distributed AI, energy-aware mission execution, and integrated hardware–software optimization for embedded UAV intelligence.
- Simulation, Validation, and Operational Deployment: A recurring observation throughout the reviewed literature is the extensive reliance on simulation environments, synthetic datasets, digital twins, and relatively limited real-world experimentation. Although these approaches have accelerated algorithm development, they often fail to fully capture the complexity, uncertainty, and communication conditions encountered during real defense operations. Therefore, bridging the gap between laboratory-scale evaluation and operational deployment remains one of the most important challenges identified in this review. Future research should prioritize realistic defense-oriented datasets, digital twins, SIL and HIL validation, standardized benchmarking methodologies, and large-scale real-flight experimentation.
- Ethical, Legal, Safety, and Governance Considerations: As AI-enabled UAV-based defense systems become increasingly autonomous, ethical, legal, safety, and governance considerations are becoming increasingly important alongside the technical challenges discussed in this review. While recent advances in AI have significantly enhanced battlefield perception, autonomous navigation, cooperative operations, and decision support, the growing autonomy of defense platforms raises broader concerns regarding meaningful human oversight, accountability, transparency, explainability, and operator trust, particularly when AI supports mission-critical functions. In addition, future operational deployment will require appropriate governance frameworks together with rigorous verification, validation, certification, and assurance processes to improve the reliability and trustworthiness of AI-enabled autonomous behaviors. Future research should investigate trustworthy AI, robust assurance and certification methodologies, governance frameworks, human–AI collaboration, and interdisciplinary approaches that support the responsible and safe deployment of AI-enabled UAV-based defense systems.
10. Conclusions
Author Contributions
Funding
Data Availability Statement
DURC Statement
Conflicts of Interest
Abbreviations
| AA | Aspect Angle |
| A2/AD | Anti-Access/Area-Denial |
| ACMDM | Autonomous Air Combat Maneuver Decision-Making |
| AI | Artificial Intelligence |
| AMSTD | Attention Mechanism-Based Small Target Detection |
| AP50 | Average Precision at IoU threshold 50% |
| APF | Artificial Potential Field |
| APs | Average Precision for small objects |
| ARIS | Aerial Reconfigurable Intelligent Surface |
| ATA | Antenna Train Angle |
| BER | Bit Error Rate |
| BFM | Basic Fighter Maneuvering |
| BIM | Basic Iterative Method |
| BIT | Batch Informed Trees |
| BVLOS | Beyond-Visual Line-of-Sight |
| BVR | Beyond Visual Range |
| C2 | Command-and-Control |
| CBS | Convolution-Batch Normalization-SiLU |
| CIDS | Collaborative Intrusion Detection System |
| CNN | Convolutional Neural Network |
| CSP | Cross-Stage Partial |
| CTDE | Centralized Training and Decentralized Execution |
| CVaR | Conditional Value-at-Risk |
| D3QN | Dueling Double DQN |
| DBN | Dynamic Bayesian Network |
| DCPA | Distance at Closest Point of Approach |
| DDPG | Deep Deterministic Policy Gradient |
| DEC-POMDP | Decentralized Partially Observable Markov Decision Process |
| DFL | Decentralized Federated Learning |
| DL | Deep Learning |
| DNN | Deep Neural Network |
| DQN | Deep Q-Network |
| DRL | Deep Reinforcement Learning |
| DT | Decision Transformer |
| DTDE | Decentralized Training and Decentralized Execution |
| DTPA | Dynamic Threat Prioritization Assessment |
| ECEF | Earth-Centered Earth-Fixed |
| EO/IR | Electro-Optical/Infrared |
| FFCNN | Feedforward CNN |
| FGSM | Fast Gradient Sign Method |
| FHP | Frequency of Hazardous Proximity |
| FL | Federated Learning |
| FMCW | Frequency-Modulated Continuous Wave |
| FPS | Frames Per Second |
| FPN | Feature Pyramid Network |
| FSM | Finite State Machine |
| GA | Genetic Algorithm |
| GAE | Generalized Advantage Estimation |
| GDT | Ground Data Terminal |
| GFLOPs | Giga Floating-Point Operations per Second |
| GIS | Geographic Information System |
| GNSS | Global Navigation Satellite System |
| GPS | Global Positioning System |
| GRU | Gated Recurrent Unit |
| HALE | High-Altitude Long-Endurance |
| HCA | Horizontal Crossing Angle |
| HIL | Hardware-in-the-Loop |
| HITL | Human-in-the-Loop |
| HRL | Hierarchical Reinforcement Learning |
| IBVS | Image-Based Visual Servoing |
| ICEM | Improved Cross-Entropy Method |
| ICM | Intrinsic Curiosity Module |
| IDS | Intrusion Detection Systems |
| IG | Information Gain |
| IMU | Inertial Measurement Unit |
| INAV | INertial NAVigation |
| INS | Inertial Navigation System |
| IoD | Internet of Drones |
| IoU | Intersection over Union |
| IPFS | InterPlanetary File System |
| IPPO | Independent PPO |
| IQN | Implicit Quantile Networks |
| ISR | Intelligence, Surveillance, and Reconnaissance |
| ISTAR | Intelligence, Surveillance, Target Acquisition, and Reconnaissance |
| I2C-MATD3 | Improved Cross-Entropy Method with Intrinsic Curiosity-enhanced Multi-Agent |
| Twin Delayed DDPG | |
| KCF | Kernelized Correlation Filter |
| KF | Kalman Filter |
| LDA | Linear Discriminant Analysis |
| LiDAR | Light Detection and Ranging |
| LLM | Large Language Model |
| LoS | Line-of-Sight |
| LRF | Laser Range Finder |
| LSTM | Long Short-Term Memory |
| LSTM-AE | LSTM Autoencoder |
| MALE | Medium-Altitude Long-Endurance |
| MAE | Mean Absolute Error |
| MAPE | Mean Absolute Percentage Error |
| MAPPO | Multi-Agent Proximal Policy Optimization |
| MARL | Multi-Agent Reinforcement Learning |
| MaCA | Multi-agent Combat Arena |
| MCLDPPO | Motivational Curriculum Learning Distributed Proximal Policy Optimization |
| MEC | Mobile Edge Computing |
| MDP | Markov Decision Process |
| MDRL-DQN | Multimodal DRL DQN |
| MERL | Maximum Entropy Reinforcement Learning |
| MHA | Multihead Attention |
| MIMO | Multiple-Input Multiple-Output |
| ML | Machine Learning |
| MLP | Multi-Layer Perceptron |
| mIoU | mean Intersection over Union |
| mAP@0.5 | mean Average Precision at IoU threshold 0.5 |
| MTVO | Multi-Agent Transformer introducing Virtual Objects |
| NB | Naive Bayes |
| NED | North-East-Down |
| NIS | Normalized Innovation Square |
| NWD-CIoU | normalized Wasserstein distance and Complete IoU |
| ONOS | Open Network Operating System |
| PBE | Perfect Bayes–Nash Equilibrium |
| PCSP | Population Curriculum Self-Play |
| PDR | Packet Delivery Ratio |
| PER | Prioritized Experience Replay |
| PID | Proportional-Integral-Differential |
| POMG | Partially Observable Markov Game |
| PPO | Proximal Policy Optimization |
| QR | Quick Response |
| QDA | Quadratic Discriminant Analysis |
| Q-SAPF | Q-learning-based Strategic Artificial Potential Field |
| R2 | Coefficient of Determination (R-squared) |
| Re-ID | Re-Identification |
| RF | Radio-Frequency |
| RGB | Red–Green–Blue |
| RL | Reinforcement Learning |
| RMSE | Root Mean Square Error |
| RNN | Recurrent Neural Network |
| ROC | Receiver Operating Characteristic |
| RRT | Rapidly-exploring Random Trees |
| SAPF | Strategic Artificial Potential Field |
| SAR | Synthetic Aperture Radar |
| SARSA | State-Action-Reward-State-Action |
| SDN | Software-Defined Networking |
| SD-UAV | Software-Defined UAV |
| SHAP | SHapley Additive exPlanations |
| SIE | Spatial Information Entropy |
| SIIS-MARL | Sparse Inferred Intention Sharing MARL |
| SIL | Software-in-the-Loop |
| SISO | Single-Input Single-Output |
| SLAM | Simultaneous Localization and Mapping |
| SMA | Scalable Mixing Networks based on Attention |
| SNR | Signal-to-Noise Ratio |
| SPPF | Spatial Pyramid Pooling Fast |
| SSP | Spatial-aware and Scale-aware Prediction |
| TaLU | Tanh Linear Unit |
| TCPA | Time to Closest Point of Approach |
| ToM | Theory of Mind |
| U2G | UAV-to-Ground |
| U2S | UAV-to-Satellite |
| U2U | UAV-to-UAV |
| UAV | Unmanned Aerial Vehicle |
| UAV-STD | UAV-to-UAV Small Target Detection |
| UCAV | Unmanned Combat Aerial Vehicle |
| UAP | Universal Adversarial Perturbation |
| VTOL | Vertical Takeoff and Landing |
| WEZ | Weapon Engagement Zone |
| XAI | Explainable AI |
References
- Rashid, A.B.; Kausik, A.K.; Sunny, A.H.; Ahamed, B.; Hassan, M. Artificial Intelligence in the Military: An Overview of the Capabilities, Applications, and Challenges. Int. J. Intell. Syst. 2023, 8676366. [Google Scholar] [CrossRef]
- Bistron, M.; Piotrowski, Z. Artificial Intelligence Applications in Military Systems and Their Influence on Sense of Security of Citizens. Electronics 2021, 10, 871. [Google Scholar] [CrossRef]
- Alcántara Suárez, E.J.; Monzon Baeza, V. Evaluating the Role of Machine Learning in Defense Applications and Industry. Mach. Learn. Knowl. Extr. 2023, 5, 1557–1569. [Google Scholar] [CrossRef]
- Hadlington, L.; Binder, J.; Gardner, S.; Karanika-Murray, M.; Knight, S. The Use of Artificial Intelligence in a Military Context: Development of the Attitudes Toward AI in Defense (AAID) Scale. Front. Psychol. 2023, 14, 1164810. [Google Scholar] [CrossRef] [PubMed]
- Gargalakos, M. The Role of Unmanned Aerial Vehicles in Military Communications: Application Scenarios, Current Trends, and Beyond. J. Def. Model. Simul. Appl. Methodol. Technol. 2024, 21, 313–321. [Google Scholar] [CrossRef]
- Yu, A.; Kolotylo, I.; Hashim, H.A.; Eltoukhy, A.E.E.E. Electronic Warfare Cyberattacks, Countermeasures, and Modern Defensive Strategies of UAV Avionics: A Survey. IEEE Access 2025, 13, 68660–68681. [Google Scholar] [CrossRef]
- Michailidis, E.T.; Potirakis, S.M.; Kanatas, A.G. AI-Inspired Non-Terrestrial Networks for IIoT: Review on Enabling Technologies and Applications. IoT 2020, 1, 21–48. [Google Scholar] [CrossRef]
- Michailidis, E.T.; Maliatsos, K.; Skoutas, D.N.; Vouyioukas, D.; Skianis, C. Secure UAV-Aided Mobile Edge Computing for IoT: A Review. IEEE Access 2022, 10, 86353–86383. [Google Scholar] [CrossRef]
- Michailidis, E.T.; Vouyioukas, D. A Review on Software-Based and Hardware-Based Authentication Mechanisms for the Internet of Drones. Drones 2022, 6, 41. [Google Scholar] [CrossRef]
- Hadi, H.J.; Cao, Y.; Un Nisa, K.; Jamil, A.M.; Ni, Q. A Comprehensive Survey on Security, Privacy Issues and Emerging Defence Technologies for UAVs. J. Netw. Comput. Appl. 2023, 213, 103607. [Google Scholar] [CrossRef]
- Alsadie, D. Cybersecurity and Artificial Intelligence in Unmanned Aerial Vehicles: Emerging Challenges and Advanced Countermeasures. IET Inf. Secur. 2025, 2046868. [Google Scholar] [CrossRef]
- Oli, A.; Mahalal, E. UAV Security: Attacks, Defenses, and Open Challenges. IEEE Access 2025, 13, 215606–215635. [Google Scholar] [CrossRef]
- Ogab, M.; Zaidi, S.; Bourouis, A.; Calafate, C.T. Machine Learning-Based Intrusion Detection Systems for the Internet of Drones: A Systematic Literature Review. IEEE Access 2025, 13, 96681–96714. [Google Scholar] [CrossRef]
- AL-Syouf, R.; Bani-Hani, R.; AL-Jarrah, O.Y. Machine Learning Approaches to Intrusion Detection in Unmanned Aerial Vehicles (UAVs). Neural Comput. Appl. 2024, 36, 18009–18041. [Google Scholar] [CrossRef]
- Islam, M.S.; Mahmoud, A.S.; Sheltami, T.R. AI-Enhanced Intrusion Detection for UAV Systems: A Taxonomy and Comparative Review. Drones 2025, 9, 682. [Google Scholar] [CrossRef]
- Papathanasiou, D.; Zacharakis, E.; Liaperdos, J.; Kotsilieris, T.; Livieris, I.E.; Ioannou, K. Secure Communication Protocols and AI-Based Anomaly Detection in UAV-GCS. Appl. Sci. 2026, 16, 3339. [Google Scholar] [CrossRef]
- Kacem, T.; Benjamin, K. Artificial Intelligence Methods for Unmanned Aerial Vehicles Cybersecurity: A Comprehensive Survey. Drones 2026, 10, 400. [Google Scholar] [CrossRef]
- Skarka, W.; Ashfaq, R. Hybrid Machine Learning and Reinforcement Learning Framework for Adaptive UAV Obstacle Avoidance. Aerospace 2024, 11, 870. [Google Scholar] [CrossRef]
- Fagundes-Junior, L.A.; de Carvalho, K.B.; Ferreira, R.S.; Brandão, A.S. Machine Learning for Unmanned Aerial Vehicles Navigation: An Overview. SN Comput. Sci. 2024, 5, 256. [Google Scholar] [CrossRef]
- Meng, W.; Zhang, X.; Zhou, L.; Guo, H.; Hu, X. Advances in UAV Path Planning: A Comprehensive Review of Methods, Challenges, and Future Directions. Drones 2025, 9, 376. [Google Scholar] [CrossRef]
- Zhai, L.; Wu, H.; Lai, L.; Gao, Z. Intelligent Optimization Algorithms for Multi-UAV Path Planning: A Comprehensive Review. IEEE Access 2025, 13, 101106–101130. [Google Scholar] [CrossRef]
- Teixeira, K.; Miguel, G.; Silva, H.S.; Madeiro, F. A Survey on Applications of Unmanned Aerial Vehicles Using Machine Learning. IEEE Access 2023, 11, 117582–117621. [Google Scholar] [CrossRef]
- Caballero-Martin, D.; Lopez-Guede, J.M.; Estevez, J.; Graña, M. Artificial Intelligence Applied to Drone Control: A State of the Art. Drones 2024, 8, 296. [Google Scholar] [CrossRef]
- Yang, Z.; Zhang, Y.; Zeng, J.; Yang, Y.; Jia, Y.; Song, H.; Lv, T.; Sun, Q.; An, J. AI-Driven Safety and Security for UAVs: From Machine Learning to Large Language Models. Drones 2025, 9, 392. [Google Scholar] [CrossRef]
- Wu, P.; Li, Y.; Xue, D. UAV Target Tracking: A Survey. Artif. Intell. Rev. 2025, 58, 358. [Google Scholar] [CrossRef]
- Costa, A.N.; Dantas, J.P.A.; Scukins, E.; Medeiros, F.L.L.; Ögren, P. Simulation and Machine Learning in Beyond Visual Range Air Combat: A Survey. IEEE Access 2025, 13, 76755–76774. [Google Scholar] [CrossRef]
- Al-Kamali, F.; Chan, F.; Ammar, H.A.; Bayes, J.H.; D’Amours, C. UAV-Mounted Aerial Relays in Military Communications: A Comprehensive Survey. IEEE Open J. Commun. Soc. 2026, 7, 1096–1136. [Google Scholar] [CrossRef]
- Khawaja, W.; Ezuma, M.; Semkin, V.; Erden, F.; Ozdemir, O.; Guvenc, I. A Survey on Detection, Classification, and Tracking of AAVs Using Radar and Communications Systems. IEEE Commun. Surv. Tutor. 2026, 28, 3272–3310. [Google Scholar] [CrossRef]
- Michailidis, E.T.; Maliatsos, K.; Vouyioukas, D. Software-Defined Radio Deployments in UAV-Driven Applications: A Comprehensive Review. IEEE Open J. Veh. Technol. 2024, 5, 1545–1586. [Google Scholar] [CrossRef]
- Javed, S.; Hassan, A.; Ahmad, R.; Ahmed, W.; Ahmed, R.; Saadat, A.; Guizani, M. State-of-the-Art and Future Research Challenges in UAV Swarms. IEEE Internet Things J. 2024, 11, 19023–19045. [Google Scholar] [CrossRef]
- Ekechi, C.C.; Elfouly, T.; Alouani, A.; Khattab, T. A Survey on UAV Control with Multi-Agent Reinforcement Learning. Drones 2025, 9, 484. [Google Scholar] [CrossRef]
- Phadke, A.; Medrano, F.A. Towards Resilient UAV Swarms—A Breakdown of Resiliency Requirements in UAV Swarms. Drones 2022, 6, 340. [Google Scholar] [CrossRef]
- Chen, B.; Lin, B.; Li, M.; Li, Z.; Zhang, X.; Shi, M.; Qin, K. Event-Triggered-Based Neuroadaptive Bipartite Containment Tracking for Networked Unmanned Aerial Vehicles. Drones 2025, 9, 317. [Google Scholar] [CrossRef]
- McEnroe, P.; Wang, S.; Liyanage, M. A Survey on the Convergence of Edge Computing and AI for UAVs: Opportunities and Challenges. IEEE Internet Things J. 2022, 9, 15435–15459. [Google Scholar] [CrossRef]
- Ding, Y.; Yang, Z.; Pham, Q.-V.; Hu, Y.; Zhang, Z.; Shikh-Bahaei, M. Distributed Machine Learning for UAV Swarms: Computing, Sensing, and Semantics. IEEE Internet Things J. 2024, 11, 7447–7473. [Google Scholar] [CrossRef]
- Zhu, J.; Kuang, M.; Zhou, W.; Shi, H.; Zhu, J.; Han, X. Mastering Air Combat Game with Deep Reinforcement Learning. Def. Technol. 2024, 34, 295–312. [Google Scholar] [CrossRef]
- Zheng, Z.; Duan, H. UAV Maneuver Decision-Making via Deep Reinforcement Learning for Short-Range Air Combat. Intell. Robot. 2023, 3, 76–94. [Google Scholar] [CrossRef]
- Saldiran, E.; Hasanzade, M.; Inalhan, G.; Tsourdos, A. Towards Global Explainability of Artificial Intelligence Agent Tactics in Close Air Combat. Aerospace 2024, 11, 415. [Google Scholar] [CrossRef]
- Wang, X.; Wang, Y.; Su, X.; Wang, L.; Lu, C.; Peng, H.; Liu, J. Deep Reinforcement Learning-Based Air Combat Maneuver Decision-Making: Literature Review, Implementation Tutorial and Future Direction. Artif. Intell. Rev. 2024, 57, 1. [Google Scholar] [CrossRef]
- Wang, B.; Gao, X.; Xie, T. An Evolutionary Multi-Agent Reinforcement Learning Algorithm for Multi-UAV Air Combat. Knowl.-Based Syst. 2024, 299, 112000. [Google Scholar] [CrossRef]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-Agent Actor–Critic for Mixed Cooperative-Competitive Environments. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS 2017); Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 6382–6393. [Google Scholar]
- Fujimoto, S.; van Hoof, H.; Meger, D. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning (ICML 2018); Proceedings of Machine Learning Research; PMLR: New York, NY, USA, 2018; Volume 80, pp. 1587–1596. [Google Scholar]
- Wang, W.; Yang, T.; Liu, Y.; Hao, J.; Hao, X.; Hu, Y.; Chen, Y.; Fan, C.; Gao, Y. From Few to More: Large-Scale Dynamic Multiagent Curriculum Learning. Proc. Aaai Conf. Artif. Intell. 2020, 34, 7293–7300. [Google Scholar] [CrossRef]
- Xu, X.; Wang, Y.; Guo, X.; Huang, K.; Zhang, X. Multi-UAV Air Combat Cooperative Game Based on Virtual Opponent and Value Attention Decomposition Policy Gradient. Expert Syst. Appl. 2025, 267, 126069. [Google Scholar] [CrossRef]
- Barto, A.G.; Mahadevan, S. Recent Advances in Hierarchical Reinforcement Learning. Discret. Event Dyn. Syst. 2003, 13, 341–379. [Google Scholar] [CrossRef]
- Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2022; pp. 24611–24624. [Google Scholar]
- Yang, J.; Yang, X.; Yu, T. Multi-Unmanned Aerial Vehicle Confrontation in Intelligent Air Combat: A Multi-Agent Deep Reinforcement Learning Approach. Drones 2024, 8, 382. [Google Scholar] [CrossRef]
- Ren, Z.; Zhang, D.; Tang, S.; Xiong, W.; Yang, S.-H. Cooperative Maneuver Decision Making for Multi-UAV Air Combat Based on Incomplete Information Dynamic Game. Def. Technol. 2023, 27, 308–317. [Google Scholar] [CrossRef]
- Dankwa, S.; Zheng, W. Twin-Delayed DDPG: A Deep Reinforcement Learning Technique to Model a Continuous Movement of an Intelligent Robot Agent. In Proceedings of the 3rd International Conference on Vision, Image and Signal Processing, Vancouver, BC, Canada, 26–28 August 2019; pp. 1–5. [Google Scholar]
- Ding, Z.; Wang, X.; Cai, C.; Jia, L.; Xu, Z. Multi-UAV Intelligent Decision-Making Method with Layer Delay Dual-Center MAPPO for Air Combat. Appl. Intell. 2025, 55, 811. [Google Scholar] [CrossRef]
- Song, Y.; Chen, H. Active Interception for Multi-Target Encirclement by Heterogeneous UAVs: An LSTM-Enhanced Independent PPO Algorithm. Designs 2026, 10, 26. [Google Scholar] [CrossRef]
- Tang, M.; Chen, R.; Zhu, J.; Jiang, Y.; Zhang, R.; Cui, Y.; Zhang, B. Hierarchical Game-Theoretic and Risk-Aware Predictive Control Framework for Resilient Multi-UAV Cooperative Combat. IEEE Access 2026, 14, 35678–35704. [Google Scholar] [CrossRef]
- Wang, H.; Wang, J. Enhancing Multi-UAV Air Combat Decision Making via Hierarchical Reinforcement Learning. Sci. Rep. 2024, 14, 4458. [Google Scholar] [CrossRef] [PubMed]
- Wu, Q.; Chen, L.; Liu, K.; Lü, J. A Simulation Platform for MARL Training and Evaluation in Swarm Confrontation. IEEE Robot. Autom. Lett. 2026, 11, 6050–6057. [Google Scholar] [CrossRef]
- Luo, D.; Fan, Z.; Yang, Z.; Xu, Y. Multi-UAV Cooperative Maneuver Decision-Making for Pursuit-Evasion Using Improved MADRL. Def. Technol. 2024, 35, 187–197. [Google Scholar] [CrossRef]
- Xie, J.; Li, G. Enhanced Q Learning and Deep Reinforcement Learning for Unmanned Combat Intelligence Planning in Adversarial Environments. Sci. Rep. 2025, 15, 28364. [Google Scholar] [CrossRef] [PubMed]
- Zhao, M.; Wang, G.; Fu, Q.; Quan, W.; Wen, Q.; Wang, X.; Li, T.; Chen, Y.; Xue, S.; Han, J. Intelligent Decision-Making System of Air Defense Resource Allocation via Hierarchical Reinforcement Learning. Int. J. Intell. Syst. 2024, 2024, 7777050. [Google Scholar] [CrossRef]
- Han, J.; Yan, Y.; Zhang, B. Towards Efficient Multi-UAV Air Combat: An Intention Inference and Sparse Transmission Based Multiagent Reinforcement Learning Algorithm. IEEE Trans. Artif. Intell. 2025, 6, 3441–3452. [Google Scholar] [CrossRef]
- Rabinowitz, N.; Perbet, F.; Song, F.; Zhang, C.; Eslami, S.M.A.; Botvinick, M. Machine Theory of Mind. In Proceedings of the International Conference on Machine Learning (ICML), Stockholm, Sweden, 10–15 July 2018; pp. 4218–4227. [Google Scholar]
- Mao, H.; Zhang, Z.; Xiao, Z.; Gong, Z.; Ni, Y. Learning Multi-Agent Communication with Double Attentional Deep Reinforcement Learning. Auton. Agents Multi-Agent Syst. 2020, 34, 32. [Google Scholar] [CrossRef]
- Guan, C.; Chen, F.; Yuan, L.; Wang, C.; Yin, H. Efficient Multiagent Communication via Self-Supervised Information Aggregation. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 1020–1033. [Google Scholar]
- Jiang, F.; Xu, M.; Li, Y.; Cui, H.; Wang, R. Short-Range Air Combat Maneuver Decision of UAV Swarm Based on Multi-Agent Transformer Introducing Virtual Objects. Eng. Appl. Artif. Intell. 2023, 123, 106358. [Google Scholar] [CrossRef]
- Kong, W.; Zhou, D.; Du, Y.; Zhou, Y.; Zhao, Y. Reinforcement Learning for Multiaircraft Autonomous Air Combat in Multisensor UCAV Platform. IEEE Sens. J. 2023, 23, 20596–20606. [Google Scholar] [CrossRef]
- Kim, G.S.; Lee, S.; Woo, T.; Park, S. Cooperative Reinforcement Learning for Military Drones over Large-Scale Battlefields. IEEE Trans. Intell. Veh. 2024, 1–11. [Google Scholar] [CrossRef]
- Li, B.; Huang, J.; Zhang, H.; Huai, L.; Neretin, E. Enhanced Exploration for Multi-UAV Cooperative Roundup: An I2C-MATD3 Reinforcement Learning Framework. Def. Technol. 2026, 58, 374–389. [Google Scholar] [CrossRef]
- Borja-Jaimes, V.; Valdez-Martínez, J.S.; Beltrán-Escobar, M.; Ramírez-Zúñiga, G.; Reyes-Mayer, A.; Calixto-Rodríguez, M. Robust Backstepping-Sliding Control of a Quadrotor UAV with Disturbance Compensation. Computation 2026, 14, 51. [Google Scholar] [CrossRef]
- Soliman, A.; Al-Ali, A.; Mohamed, A.; Gedawy, H.; Izham, D.; Bahri, M.; Erbad, A.; Guizani, M. AI-Based UAV Navigation Framework with Digital Twin Technology for Mobile Target Visitation. Eng. Appl. Artif. Intell. 2023, 123, 106318. [Google Scholar] [CrossRef]
- Shen, Y.; Zhang, X.; Li, Y.; Zhang, W. Deep Reinforcement Learning-Based Adaptive Collision Avoidance Method for UAV in Joint Operational Airspace. Def. Technol. 2026, 56, 142–159. [Google Scholar] [CrossRef]
- Zhang, J.; Xian, Y.; Zhu, X.; Deng, H. A Hybrid Deep Learning Model for UAV Path Planning in Dynamic Environments. IEEE Access 2025, 13, 67459–67475. [Google Scholar] [CrossRef]
- García-Gascón, C.; Castelló-Pedrero, P.; Chinesta, F.; García-Manrique, J.A. Artificial Intelligence-Driven Aircraft Systems to Emulate Autopilot and GPS Functionality in GPS-Denied Scenarios Through Deep Learning. Drones 2025, 9, 250. [Google Scholar] [CrossRef]
- Sujecki, P.; Frąszczak, D. Artificial Intelligence-Based Decision Support System for UAV Control in a Simulated Environment. Sensors 2026, 26, 2436. [Google Scholar] [CrossRef] [PubMed]
- Chronis, C.; Anagnostopoulos, G.; Politi, E.; Dimitrakopoulos, G.; Varlamis, I. Dynamic Navigation in Unconstrained Environments Using Reinforcement Learning Algorithms. IEEE Access 2023, 11, 117984–118001. [Google Scholar] [CrossRef]
- Lee, J.; Seo, Y. Q-Learning Based on Strategic Artificial Potential Field for Path Planning Enabling Concealment and Cover in Ground Battlefield Environments. Appl. Intell. 2024, 54, 7170–7200. [Google Scholar] [CrossRef]
- Kim, S.M.; Lee, J.; Kwon, H. Tactical Path Planning in Battlefield Environments Using Semantic Segmentation and Deep Reinforcement Learning. IEEE Access 2026, 14, 42838–42849. [Google Scholar] [CrossRef]
- Zhan, H.; Zhang, Y.; Huang, J.; Song, Y.; Xing, L.; Wu, J.; Gao, Z. A Reinforcement Learning-Based Evolutionary Algorithm for the Unmanned Aerial Vehicles Maritime Search and Rescue Path Planning Problem Considering Multiple Rescue Centers. Memetic Comput. 2024, 16, 373–386. [Google Scholar] [CrossRef]
- Lee, M.; Choi, M.; Yang, T.; Kim, J.; Kim, J.; Kwon, O.; Cho, N. A Study on the Advancement of Intelligent Military Drones: Focusing on Reconnaissance Operations. IEEE Access 2024, 12, 55964–55975. [Google Scholar] [CrossRef]
- Salameh, H.B.; Hussienat, A.; Alhafnawi, M.; Al-Ajlouni, A. Autonomous UAV-Based Surveillance System for Multi-Target Detection Using Reinforcement Learning. Clust. Comput. 2024, 27, 9381–9394. [Google Scholar] [CrossRef]
- Ud Din, I.; Almogren, A.; Rodrigues, J.J.P.C. Adversarial Machine Learning for Robust and Secure UAV Detection in Consumer Applications. IEEE Trans. Consum. Electron. 2025, 71, 3360–3367. [Google Scholar] [CrossRef]
- Sayed, A.N.; Abedi, H.; Ramahi, O.M.; Shaker, G. Enhanced UAV Detection and Classification Using Machine Learning and MIMO Radars. IEEE Trans. Microw. Theory Tech. 2024, 72, 6716–6727. [Google Scholar] [CrossRef]
- Ansys Inc. Ansys HFSS. 2022. Available online: https://www.ansys.com/products/electronics/ansys-hfss (accessed on 22 July 2026).
- Han, T.-T.; Duy, A.L.; Van, T.N.; Tan, H.D.; Trung, A.D. High Performance Autonomous Target Tracking and Control Method for Quadrotors Using Artificial Intelligence. IEEE Access 2025, 13, 214236–214252. [Google Scholar] [CrossRef]
- Yang, H.; Gao, S.; Wu, X.; Zhang, Y. Online Multi-Object Tracking Using KCF-Based Single-Object Tracker with Occlusion Analysis. Multimed. Syst. 2020, 26, 655–669. [Google Scholar] [CrossRef]
- Han, K. Image Object Tracking Based on Temporal Context and MOSSE. Clust. Comput. 2017, 20, 1259–1269. [Google Scholar] [CrossRef]
- Farkhodov, K.; Lee, S.; Kwon, K. Object Tracking Using CSRT Tracker and RCNN. In Proceedings of the 13th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC), Valletta, Malt, 24–26 February 2020; pp. 209–212. [Google Scholar]
- Bhat, G.; Danelljan, M.; Van Gool, L.; Timofte, R. Learning Discriminative Model Prediction for Tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6181–6190. [Google Scholar]
- Chen, X.; Yan, B.; Zhu, J.; Wang, D.; Yang, X.; Lu, H. Transformer Tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 8122–8131. [Google Scholar]
- Ye, B.; Chang, H.; Ma, B.; Shan, S. Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework. In Computer Vision—ECCV 2022; Springer: Cham, Switzerland, 2022; pp. 341–357. [Google Scholar]
- Kant, R.; Saini, P.; Kumari, J. Long Short-Term Memory Auto-Encoder-Based Position Prediction Model for Fixed-Wing UAV During Communication Failure. IEEE Trans. Artif. Intell. 2023, 4, 173–181. [Google Scholar] [CrossRef]
- Liu, Y.; Su, Z.; Li, H.; Zhang, Y. An LSTM Based Classification Method for Time Series Trend Forecasting. In Proceedings of the 14th IEEE Conference on Industrial Electronics and Applications (ICIEA), Xi’an, China, 19–21 June 2019; pp. 402–406. [Google Scholar]
- Li, C.; Wei, F.; Dong, W.; Wang, X.; Liu, Q.; Zhang, X. Dynamic Structure Embedded Online Multiple-Output Regression for Streaming Data. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 323–336. [Google Scholar] [CrossRef] [PubMed]
- Zuo, G.; Zhou, K.; Wang, Q. UAV-to-UAV Small Target Detection Method Based on Deep Learning in Complex Scenes. IEEE Sens. J. 2025, 25, 3806–3820. [Google Scholar] [CrossRef]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Computer Vision—ECCV 2016; Springer: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [PubMed]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-Time Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar]
- Kim, J.; Joe, I. Deep Learning-Based Drone Defense System for Autonomous Detection and Mitigation of Balloon-Borne Threats. Electronics 2025, 14, 1553. [Google Scholar] [CrossRef]
- Masadeh, A.; Alhafnawi, M.; Salameh, H.A.B.; Musa, A.; Jararweh, Y. Reinforcement Learning-Based Security/Safety UAV System for Intrusion Detection Under Dynamic and Uncertain Target Movement. IEEE Trans. Eng. Manag. 2024, 71, 12498–12508. [Google Scholar] [CrossRef]
- Mignon, A.d.S.; Rocha, R.L.d.A.d. An Adaptive Implementation of ϵ-Greedy in Reinforcement Learning. Procedia Comput. Sci. 2017, 109, 1146–1151. [Google Scholar] [CrossRef]
- Ma, X.; Sun, T.; Gao, M. A Reinforcement-Learning-Enhanced Spoofing Algorithm for UAV With GPS/INS-Integrated Navigation. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 8659–8673. [Google Scholar] [CrossRef]
- Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor–Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning (ICML 2018), Proceedings of Machine Learning Research, Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1861–1870. [Google Scholar]
- Ma, X.; Gao, M.; Zhao, Y.; Yu, M. A Novel Navigation Spoofing Algorithm for UAV Based on GPS/INS-Integrated Navigation. IEEE Trans. Veh. Technol. 2024, 73, 15424–15439. [Google Scholar] [CrossRef]
- Zhao, R.; Wen, H.; Hou, W.; Jiang, L.; Chen, Z.; Tang, T.; Feng, X. Anti-Jamming Communication Framework for UAV Based Defense Operations to Control AI-Enabled Attacks. IEEE Trans. Consum. Electron. 2026; early access. [CrossRef]
- Hickling, T.; Aouf, N.; Spencer, P. Robust Adversarial Attacks Detection Based on Explainable Deep Reinforcement Learning for UAV Guidance and Planning. IEEE Trans. Intell. Veh. 2023, 8, 4381–4394. [Google Scholar] [CrossRef]
- Lundberg, M.S.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 4768–4777. [Google Scholar]
- Musa, A.; Vishi, K.; Rexha, B. Attack Analysis of Face Recognition Authentication Systems Using Fast Gradient Sign Method. Appl. Artif. Intell. 2021, 35, 1346–1360. [Google Scholar] [CrossRef]
- Kurakin, A.; Goodfellow, I.J.; Bengio, S. Adversarial Examples in the Physical World. In Artificial Intelligence Safety and Security; Chapman and Hall/CRC: Boca Raton, FL, USA, 2018; pp. 99–112. [Google Scholar]
- Mynuddin, M.; Khan, S.U.; Ahmari, R.; Landivar, L.; Mahmoud, M.N.; Homaifar, A. Trojan Attack and Defense for Deep Learning-Based Navigation Systems of Unmanned Aerial Vehicles. IEEE Access 2024, 12, 89887–89907. [Google Scholar] [CrossRef]
- Loquercio, A.; Maqueda, A.I.; del-Blanco, C.R.; Scaramuzza, D. DroNet: Learning to Fly by Driving. IEEE Robot. Autom. Lett. 2018, 3, 1088–1095. [Google Scholar] [CrossRef]
- Moosavi-Dezfooli, S.-M.; Fawzi, A.; Fawzi, O.; Frossard, P. Universal Adversarial Perturbations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1765–1773. [Google Scholar]
- Howard, A.; Sandler, M.; Chen, B.; Wang, W.; Chen, L.-C.; Tan, M.; Chu, G.; Vasudevan, V.; Zhu, Y.; Pang, R.; et al. Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1314–1324. [Google Scholar]
- Hadi, H.J.; Cao, Y.; Li, S.; Hu, Y.; Wang, J.; Wang, S. Real-Time Collaborative Intrusion Detection System in UAV Networks Using Deep Learning. IEEE Internet Things J. 2024, 11, 33371–33391. [Google Scholar] [CrossRef]
- Unmanned Aerial Vehicle (UAV) Intrusion Detection [Dataset]. UCI Machine Learning Repository. 2020. Available online: https://archive.ics.uci.edu/dataset/564/unmanned+aerial+vehicle+uav+intrusion+detection (accessed on 22 July 2026).
- Agnew, D.; Aguila, A.D.; McNair, J. Enhanced Network Metric Prediction for Machine Learning-Based Cyber Security of a Software-Defined UAV Relay Network. IEEE Access 2024, 12, 54202–54219. [Google Scholar] [CrossRef]
- del Aguila, A.; Mendoza, J.V.; Mandavilli, S.B.; McNair, J. Remote and Rural Connectivity via Multi-Tier Systems Through SDN-Managed Drone Networks. In Proceedings of the IEEE Military Communications Conference (MILCOM), Rockville, MD, USA, 28 November–2 December 2022; pp. 956–961. [Google Scholar]
- Jadav, N.K.; Rathod, T.; Gupta, R.; Tanwar, S.; Kumar, N.; Iqbal, R.; Atalla, S.; Mohammad, H.; Al-Rubaye, S. Blockchain-Based Secure and Intelligent Data Dissemination Framework for UAVs in Battlefield Applications. IEEE Commun. Stand. Mag. 2023, 7, 16–23. [Google Scholar] [CrossRef]
- Gupta, R.; Tanwar, S.; Kumar, N. B-IoMV: Blockchain-Based Onion Routing Protocol for D2D Communication in an IoMV Environment Beyond 5G. Veh. Commun. 2021, 32, 100401. [Google Scholar] [CrossRef]
- Beltrán, E.T.M.; Sánchez, P.M.S.; Bovet, G.; Stiller, B.; Pérez, G.M.; Celdrán, A.H. Flighter: Decentralized Federated Learning and Situational Awareness for Secure Military Aerial Reconnaissance. IEEE Commun. Mag. 2025, 63, 136–142. [Google Scholar] [CrossRef]
- Beltrán, E.T.M. CyberDataLab/Flighter. GitHub Repository. 2024. Available online: https://github.com/CyberDataLab/flighter (accessed on 22 July 2026).
- Nguyen, T.M.; Truong, V.T.; Le, L.B. Agentic AI Meets Edge Computing in Autonomous UAV Swarms. IEEE Internet Things Mag. 2026, 9, 87–95. [Google Scholar] [CrossRef]



| Reference | UAVs | AI Paradigms | Defense Applications | Air Combat | Path Planning | Target Tracking | Cybersecurity | Electronic Warfare |
|---|---|---|---|---|---|---|---|---|
| Rashid et al., 2023 [1] | Partially | ✓ | ✓ | Partially | ✗ | ✗ | ✓ | Partially |
| Yu et al., 2025 [6] | ✓ | Partially | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ |
| Hadi et al., 2023 [10] | ✓ | Partially | Partially | ✗ | ✗ | ✗ | ✓ | ✓ |
| Alsadie et al., 2025 [11] | ✓ | ✓ | Partially | ✗ | ✗ | ✗ | ✓ | ✓ |
| Oli et al., 2025 [12] | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ |
| Ogab et al., 2025 [13] | ✓ | ✓ | Partially | ✗ | ✗ | ✗ | ✓ | Partially |
| Al-Syouf et al., 2024 [14] | ✓ | ✓ | Partially | ✗ | ✗ | ✗ | ✓ | ✓ |
| Islam et al., 2025 [15] | ✓ | ✓ | Partially | ✗ | ✗ | ✗ | ✓ | ✓ |
| Papathanasiou et al., 2026 [16] | ✓ | ✓ | Partially | ✗ | ✗ | ✗ | ✓ | ✓ |
| Kacem et al., 2026 [17] | ✓ | ✓ | Partially | ✗ | ✗ | ✗ | ✓ | ✓ |
| Skarka et al., 2024 [18] | ✓ | ✓ | Partially | ✗ | ✓ | ✗ | ✗ | ✗ |
| Fagundes-Junior et al., 2024 [19] | ✓ | ✓ | Partially | ✗ | ✓ | ✗ | Partially | ✗ |
| Meng et al., 2025 [20] | ✓ | ✓ | Partially | ✗ | ✓ | ✗ | ✗ | ✗ |
| Zhai et al., 2025 [21] | ✓ | ✓ | ✓ | Partially | ✓ | ✗ | ✗ | ✗ |
| Teixeira et al., 2023 [22] | ✓ | ✓ | Partially | ✗ | Partially | Partially | Partially | ✗ |
| Caballero-Martin et al., 2024 [23] | ✓ | ✓ | Partially | ✗ | Partially | Partially | Partially | ✗ |
| Yang et al., 2025 [24] | ✓ | ✓ | Partially | ✗ | Partially | ✗ | ✓ | ✓ |
| Wu et al., 2025 [25] | ✓ | ✓ | Partially | ✗ | Partially | ✓ | ✗ | ✗ |
| Costa et al., 2025 [26] | Partially | ✓ | ✓ | ✓ | Partially | ✗ | ✗ | ✗ |
| Al-Kamali et al., 2026 [27] | ✓ | Partially | ✓ | ✗ | ✗ | ✗ | Partially | Partially |
| This Paper | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ |
| Reference | Objective | AI/ML Models | Operational Scenario | UAV Configuration | Main Components | Key Results |
|---|---|---|---|---|---|---|
| Zhu et al., 2024 [36] | Autonomous air combat maneuver and engagement decision-making | MCLDPPO, distributed PPO, LSTM-based actor–critic DRL | Digital twin-enabled air combat with BVR missile engagements | Single combat UAV | POMDP formulation, motivational curriculum learning, hierarchical maneuver actions, missile threat interruption mechanism, distributed battlefield simulation | Achieved approximately 66% win rate against expert-level FSM opponents, outperforming predictive game trees, SAC, DDPG, and DQN while demonstrating advanced tactical combat behaviors |
| Zheng et al., 2023 [37] | Autonomous maneuver decision-making for short-range UAV air combat | GRU-enhanced PPO, actor–critic DRL | One-versus-one short-range aerial combat | Single combat UAV | GRU-based temporal feature extraction, phased curriculum training, dense/event/terminal rewards, adversarial opponent policy, probabilistic damage modeling | Improved convergence, policy stability, temporal situational awareness, and combat effectiveness enabling adaptive offensive and evasive maneuver generation |
| Saldiran et al., 2024 [38] | Explainable decision-making for RL-based air combat agents | DQN, Double DQN, XAI-enabled RL | Two-dimensional (2D) close-range aerial combat | Single combat UAV | Reward decomposition, ATA/AA/LoS-based analysis, tactical-region explainability, visual explanation maps | Identified tactical strengths, weaknesses, asymmetries, and hidden behavioral patterns while improving transparency, interpretability, and debugging capability |
| Wang et al., 2024 [39] | Implementation and evaluation of DRL-based air combat maneuver decision-making | DQN-based DRL framework | Three-dimensional (3D) 1 vs. 1 UAV dogfight | Single combat UAV | Six-Degrees-of-Freedom (DOF) UAV kinematics, NASA-inspired BFM actions, prioritized experience replay, -greedy exploration, curriculum learning, reward shaping | Demonstrated autonomous target pursuit, attack positioning, maneuver adaptation, and engagement decision-making while providing a practical DRL implementation framework for intelligent air combat systems |
| Reference | Objective | AI/ML Models | Operational Scenario | UAV Configuration | Main Components | Key Results |
|---|---|---|---|---|---|---|
| Jiang et al., 2023 [62] | Large-scale UAV swarm combat maneuver decision-making | MTVO, transformer-based PPO | Short-range 3D UAV swarm air combat | 1 vs. 1–20 vs. 20 UAV swarm operational scenarios | Transformer self-attention, virtual object global representation, mask mechanisms, PPO-based CTDE, attention aggregation, reward shaping, situation assessment, actor–critic RL | Achieved win rates exceeding 70% in 5 vs. 5 combat while maintaining stable convergence; demonstrated robustness under numerical disadvantage and verified the importance of self-attention, attention aggregation, and virtual object modules through ablation studies |
| Kong et al., 2023 [63] | Scalable multi-aircraft autonomous close-range air combat | QMIX-based MARL, MHA, SMA, PCSP | Partially observable multi-sensor 3D air combat | Variable-size cooperative UCAV formations | CTDE, MHA-based information embedding, SMA, PCSP, GRU-based recurrent Q-networks, multi-sensor information fusion, competitive self-play | Achieved winning rates exceeding 70% and the highest Elo scores among self-play methods; outperformed QMIX, VDN, and MeanPool while improving convergence stability, scalability, and combat effectiveness |
| Kim et al., 2024 [64] | Cooperative RL for large-scale military drone operations | CommNet, MADRL, CTDE | Large-scale dynamic battlefield environments | Cooperative military drone swarms based on Shahed-136 UAVs | Communication-aware coordination, hidden state information sharing, DEC-POMDP modeling, swarm flight maintenance, energy-aware bombing, realistic flight dynamics and energy consumption modeling | Outperformed independent DQN, PPO, and MADDPG baselines in swarm coordination, bombing success rate, and residual energy preservation; maintained stable formations and improved energy efficiency over a battlefield of approximately |
| Li et al., 2026 [65] | Autonomous cooperative roundup and UAV interception | I2C-MATD3, ICEM, ICM | Obstacle-constrained cooperative interception environment | Three cooperative UAVs versus one adversarial UAV | MATD3, ICEM, ICM, CTDE, replay buffer sharing, global-elite evolutionary optimization, LiDAR sensing, curiosity-driven exploration | Achieved near-100% roundup success after approximately 1200 training episodes; outperformed MATD3 and IMTD3 while maintaining success rates above 71% under wind disturbances and approximately 70% under severe communication degradation |
| Reference | Objective | AI/ML Models | Operational Scenario | UAV Configuration | Main Components | Key Results |
|---|---|---|---|---|---|---|
| Soliman et al., 2023 [67] | Path planning for autonomous mobile-target search and reconnaissance | PPO-based DRL, MLP Actor-Critic | Reconnaissance and surveillance with mobile targets | Parrot ANAFI UAV | MDP formulation, digital twin simulation, QR-code target identification, energy-aware navigation, mobility-aware target tracking | Performance close to an ideal benchmark policy and superior to area-scanning baselines for up to 50 mobile targets and 64 grid cells |
| Shen et al., 2026 [68] | Adaptive collision avoidance under uncertainty | HPER-D3QN | Joint manned-unmanned operational airspace | Up to 25 heterogeneous aircraft | DTPA, TCPA/DCPA threat assessment, wind-aware navigation, hierarchical PER, dual-layer safety zones, partial observation | Achieved final reward of approximately 2.36 and collision-avoidance success rate of 96.3% while reducing FHP and task-completion time |
| Zhang et al., 2025 [69] | Hybrid intelligent UAV path planning in adversarial environments | CNNs, Bi-LSTM, attention-assisted Informed RRT | Dynamic obstacle-rich battlefield environments | Single UAV navigation system | GIS/radar/camera fusion, attention mechanism, adaptive informed sampling, DL-enhanced Informed RRT, Bézier smoothing | Achieved approximately 91% success rate in narrow channels and 88% in trap-obstacle scenarios while reducing path length by 7.82% versus DDPG |
| Garcia-Gascon et al., 2025 [70] | Autonomous navigation in GPS-denied environments | LSTM, MLP-based DL | GPS-denied waypoint-following missions | Fixed-wing UAV | Telemetry-based GPS prediction, INAV autopilot emulation, onboard sensor fusion, 49 input features | Approximately 99% prediction accuracy with –0.999 and MAPE of 0.14–0.16% using 432,000 telemetry measurements |
| Sujecki et al., 2026 [71] | Autonomous UAV navigation in contested environments | PPO, REINFORCE | GPS-denied and communication-constrained environments | Quadrotor UAV | Unity ML-Agents, PyTorch, continuous-control learning, reward shaping, 146-dimensional observation space | PPO achieved cumulative reward of 3.8658 vs. 1.5712 for REINFORCE and reduced average episode length from 200 to 110 steps |
| Chronis et al., 2023 [72] | Autonomous navigation in unknown dynamic environments | PPO-based DRL | Urban and rural BVLOS missions | Single UAV with low-cost sensors | AirSim/Unreal Engine, actor–critic navigation, obstacle avoidance, lightweight sensing | Approximately 99% success rate with near-zero collision probability, outperforming A2C and A* |
| Lee et al., 2024 [73] | Concealment-aware battlefield navigation | Q-learning, Q-SAPF | Battlefield terrain environments | Autonomous military UAV | 50 × 50 battlefield maps, concealment-aware rewards, terrain-aware scoring, strategic potential fields | Mission success rates of 99.5–99.7%; path-length reduction of 1.96–8.77 units and execution-time reduction of 14.57–69.29 s |
| Kim et al., 2026 [74] | Threat-aware tactical battlefield navigation | SwiftFormer, DQN-based DRL | Battlefield mobility and adversarial terrain environments | Autonomous military UAV | Semantic segmentation, mobility-aware terrain mapping, threat zone modeling, adaptive DQN decision-making, A* routing | 93 FPS segmentation, 93.0% pixel accuracy, 72.8% mIoU, zero enemy-zone exposure, and over 50-fold execution-time reduction |
| Zhan et al., 2024 [75] | Path planning for large-scale maritime search-and-rescue missions | Q-learning-enhanced genetic algorithm | Maritime SAR and emergency-response operations | Multiple UAVs and rescue centers | Adaptive population management, heuristic initialization, adaptive crossover, elite repositories, population perturbation | Improved convergence, workload balancing, and reduced mission duration for 100–1000 tasks, 2–10 rescue centers, and 80 km2 search regions |
| UAV-Based Defense Domain | DRL | MARL | CNNs | Attention/ Transformers | Radar-Based Learning | FL | XAI | Digital Twins | Multi-Modal Sensing | Comm.-Aware AI |
|---|---|---|---|---|---|---|---|---|---|---|
| Autonomous Air Combat and Cooperative UAV Operations | • | • | • | |||||||
| Path Planning and Autonomous Navigation | • | • | • | • | • | • | ||||
| Target Tracking, Detection, and Classification | • | • | • | • | • | • | ||||
| Cybersecurity, Electronic Warfare Protection, and Resilient UAV Operation | • | • | • | |||||||
| Overall/ Cross-Cutting |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Michailidis, E.T.; Karanasiou, I.S. A Review of AI-Enabled UAV-Based Systems for Defense Applications. Drones 2026, 10, 602. https://doi.org/10.3390/drones10080602
Michailidis ET, Karanasiou IS. A Review of AI-Enabled UAV-Based Systems for Defense Applications. Drones. 2026; 10(8):602. https://doi.org/10.3390/drones10080602
Chicago/Turabian StyleMichailidis, Emmanouel T., and Irene S. Karanasiou. 2026. "A Review of AI-Enabled UAV-Based Systems for Defense Applications" Drones 10, no. 8: 602. https://doi.org/10.3390/drones10080602
APA StyleMichailidis, E. T., & Karanasiou, I. S. (2026). A Review of AI-Enabled UAV-Based Systems for Defense Applications. Drones, 10(8), 602. https://doi.org/10.3390/drones10080602

