Next Article in Journal
GeoSOT-H-Enabled Risk-Aware Hierarchical Path Planning and Emergency Replanning for Urban Low-Altitude UAV Missions
Previous Article in Journal
Design and Wind Tunnel Test of Control Laws for High Angle of Attack Flight of Low-Aspect-Ratio Flying-Wing UAVs Based on NDI
Previous Article in Special Issue
Certification-Oriented Requirements and Model Verification Methodology for UAV Systems
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

A Review of AI-Enabled UAV-Based Systems for Defense Applications

by
Emmanouel T. Michailidis
* and
Irene S. Karanasiou
Division of Mathematics and Engineering Sciences, Department of Military Science, Hellenic Army Academy, 16673 Vari, Greece
*
Author to whom correspondence should be addressed.
Drones 2026, 10(8), 602; https://doi.org/10.3390/drones10080602
Submission received: 1 July 2026 / Revised: 2 August 2026 / Accepted: 3 August 2026 / Published: 5 August 2026

Highlights

What are the main findings?
  • This paper provides an up-to-date review of Artificial Intelligence (AI)-enabled Unmanned Aerial Vehicle (UAV)-based defense systems, covering autonomous air combat, cooperative UAV operations, path planning, target tracking, and cybersecurity.
  • An integrated system architecture and a functional classification framework are presented together with an analysis of major AI paradigms, recent advances, open challenges, and future research directions.
What are the implications of the main findings?
  • This paper provides a unified reference by consolidating recent developments across AI-enabled UAV-based defense technologies within a common analytical framework.
  • The identified challenges and future directions can guide the development of more autonomous, resilient, secure, and scalable next-generation UAV-based defense systems.

Abstract

Unmanned aerial vehicles have become indispensable components of modern defense systems by conducting Intelligence, Surveillance, and Reconnaissance (ISR) missions to collect critical operational information through onboard sensing technologies, supporting secure communication and information sharing among distributed military assets, and enhancing battlefield situational awareness through real-time sensing and data fusion. In addition, UAVs are increasingly capable of executing a wide range of defense missions, including target search and tracking, electronic warfare, search-and-rescue, and combat support. The integration of AI into UAV-based systems has the potential to enhance these operational capabilities by enabling intelligent perception, autonomous decision-making, adaptive mission planning, autonomous navigation, resilient communications, and cooperative multi-UAV coordination, thereby enabling the autonomous and collaborative execution of complex defense missions. This paper presents an up-to-date review of AI-enabled UAV-based defense systems, focusing on major operational domains including autonomous air combat and cooperative UAV operations, path planning and autonomous navigation, target tracking/detection/classification, cybersecurity, electronic warfare protection, and resilient UAV operation. In addition to surveying the recent literature, this paper provides an integrated system architecture, a functional classification framework, and an analysis of the AI paradigms enabling next-generation UAV-based defense systems. Furthermore, this review synthesizes the key technological trends, lessons learned, and cross-domain research challenges identified across the reviewed studies, providing a unified perspective on the current state of the field. Finally, this paper highlights promising future research directions for resilient, scalable, secure, and intelligent next-generation AI-enabled UAV-based defense systems.

1. Introduction

Artificial intelligence has emerged as a transformative technology in modern defense systems, enabling autonomous decision-making, intelligent sensing, adaptive mission execution, and enhanced situational awareness in increasingly complex and contested operational environments [1,2,3]. Driven by challenges such as rapidly evolving tactical conditions, cyber–physical threats, electronic warfare, communication disruptions, and the growing volume of heterogeneous battlefield data, AI techniques can be adopted to improve the autonomy, adaptability, resilience, and operational effectiveness of military platforms [1,2,3]. AI technologies have found applications across a broad spectrum of defense functions—including surveillance and reconnaissance, target detection and recognition, cyber defense, logistics, predictive analytics, mission planning, and decision support—while simultaneously raising important concerns regarding transparency, reliability, accountability, trustworthiness, and ethical deployment [2,3,4].
Among the various defense platforms benefiting from AI integration, UAVs have become critical assets due to their mobility, operational flexibility, relatively low deployment cost, and ability to perform missions in hazardous environments without directly exposing personnel to risk [5,6]. UAVs can be employed in Intelligence, Surveillance, Target Acquisition, and Reconnaissance (ISTAR) missions, autonomous air combat, electronic warfare, communication relaying, border monitoring, and search-and-rescue missions [1,5]. In addition, they can operate collaboratively in swarm configurations to improve mission effectiveness, scalability, and resilience. Advances in sensing technologies [7], wireless communications, radar systems, embedded computing, Mobile Edge Computing (MEC) [8], edge intelligence, and miniaturized hardware have significantly expanded the capabilities of UAV-enabled defense ecosystems and Internet of Drones (IoD) architectures [9]. In this context, the integration of AI with UAV technologies is driving the evolution of intelligent aerial systems capable of autonomous navigation, threat assessment, and adaptive mission planning in future defense environments [1].
Despite the significant progress achieved in recent years, existing research remains dispersed across multiple disciplines and application domains, often focusing on specific technologies or operational scenarios in isolation. Consequently, there is value in providing a broader cross-domain synthesis that integrates recent advances in AI-enabled UAV technologies, defense-oriented applications, enabling architectures, and emerging research directions within a unified analytical framework.

1.1. Contributions

Motivated by the rapid convergence of AI and UAV technologies in modern defense applications, this paper presents a structured review of recent advances in AI-enabled UAV-based defense systems. In the context of this review, such systems are defined as defense systems in which one or more UAVs constitute the primary operational platform, while AI techniques are integrated to enhance perception, autonomous operation, decision-making, communication, cybersecurity, and overall mission effectiveness. The main contributions of this paper are summarized as follows:
  • Overview of AI-Enabled UAV-Based Defense Systems: This paper provides a foundational overview of UAV platforms, sensing and perception technologies, communication infrastructures, and AI techniques that underpin intelligent UAV operation in defense environments.
  • Integrated Architecture and Functional Classification Framework: This paper presents an integrated system architecture and a functional classification framework for AI-enabled UAV-based defense systems. The proposed framework provides a unified system-level perspective by organizing the key functional layers, including UAV platforms, sensing and perception, communication and networking, AI-enabled intelligence and computing, autonomous control and decision-making, cybersecurity and resilience, and the operational environment. Furthermore, the interactions among these functional layers are analyzed to provide a holistic perspective on the AI-enabled UAV-based defense ecosystem.
  • Review of Major Defense Application Domains: Complementing existing review papers which primarily organize the literature around individual technologies or specific application areas, this paper classifies recent state-of-the-art research into four complementary and interconnected defense operational domains to provide a unified perspective on AI-enabled UAV-based defense systems: (i) autonomous air combat, cooperative multi-UAV operations, and swarm combat intelligence; (ii) autonomous navigation and path planning; (iii) target tracking, detection, and classification; and (iv) cybersecurity, electronic warfare protection, and resilient UAV operation. These domains were identified through a search of major scientific databases, including ACM Digital Library, IEEE Xplore, PubMed, Web of Science, and others using keywords related to AI, UAVs, autonomous drones, autonomous air combat, swarm intelligence, multi-UAV systems, path planning, autonomous navigation, target tracking, cybersecurity, electronic warfare, resilient communications, and intelligent defense systems. The primary body of literature reviewed in this paper consists of peer-reviewed journal publications written in English and published between January 2023 and June 2026. Earlier publications were included where necessary to cite seminal contributions or provide technical background, while conference papers, publicly available datasets, software frameworks, standards, and online resources were referenced when directly relevant to the surveyed technologies. Furthermore, studies involving civilian or dual-use UAV applications were considered when they introduced AI methodologies or communication technologies that are equally applicable to defense-oriented UAV systems. Inclusion criteria considered studies directly investigating AI-enabled UAV technologies and defense-oriented applications, while exclusion criteria omitted works focusing solely on civilian UAV use cases, conventional non-AI methodologies, or systems not involving UAVs. Following the literature search, duplicate records were removed where applicable, then the remaining publications were screened based on their titles and abstracts to assess their relevance to the scope of this review, after which the full text of the shortlisted studies was examined to determine their eligibility according to the predefined inclusion and exclusion criteria.
  • Analysis of AI Techniques and Their Operational Roles: This paper investigates the application of major AI paradigms, including Machine Learning (ML), Deep Learning (DL), Reinforcement Learning (RL), Deep RL (DRL), Multi-Agent RL (MARL), transformer-based architectures, Federated Learning (FL), and Explainable AI (XAI). Their effectiveness, advantages, limitations, scalability characteristics, and practical tradeoffs are analyzed across diverse defense-oriented UAV applications.
  • Lessons Learned and Future Research Directions: This paper synthesizes key findings from recent studies, identifies current limitations and open research challenges, and outlines promising future research directions.
Figure 1 presents an overview of the major operational domains and enabling capabilities of AI-enabled UAV-based defense systems. At the core of the ecosystem, AI integrates five key functional pillars: (i) sensing and perception, (ii) communication and networking, (iii) intelligence and computing, (iv) autonomy and decision-making, and (v) control and actuation. Cybersecurity, secure communications, and resilience constitute cross-cutting system capabilities that span and support all five functional pillars to ensure secure, reliable, and resilient operation across the entire UAV-based system. Collectively, these interconnected capabilities provide the technological foundation for situational awareness, intelligent information processing, autonomous control, and effective mission execution. The ecosystem established by this foundation enables the major operational domains considered in this paper.

1.2. Structure

The remainder of this paper is organized as follows. Section 2 examines prior survey papers and positions the present work within the existing literature. Section 3 provides background on UAV platforms and AI techniques, then introduces an integrated system architecture for AI-enabled UAV-based defense systems. Section 4 analyzes AI-enabled approaches for autonomous air combat, cooperative multi-UAV air combat, and large-scale swarm combat intelligence. Section 5 discusses AI-enabled path planning and autonomous navigation methods. Section 6 examines AI-enabled target tracking, detection, and classification. Section 7 describes AI-enabled approaches for cybersecurity, electronic warfare protection, and resilient UAV operation. Section 8 summarizes the key lessons learned from the surveyed literature. Section 9 identifies the remaining challenges, highlights open issues, and suggests future research directions. Finally, Section 10 concludes the paper.

2. Previous Review Papers

The recent literature contains a diverse collection of review papers examining the integration of AI into UAV-based systems, autonomous aerial platforms, and defense-related applications. Several existing papers concentrate primarily on cybersecurity, electronic warfare, resilient communications, and cyber–physical protection mechanisms. The work in [10] reviewed security vulnerabilities, privacy threats, secure communication mechanisms, Intrusion Detection Systems (IDS), and blockchain-assisted protection frameworks for UAV-assisted networks. Similarly, the work in [6] investigated electronic warfare cyberattacks, avionics vulnerabilities, Global Positioning System (GPS) spoofing, jamming, malware propagation, and anti-jamming defensive mechanisms. In [11], AI-driven cybersecurity threats, adversarial AI attacks, blockchain-based protection, post-quantum cryptography, XAI, and FL were analyzed for secure UAV operation. The study in [12] provided a comprehensive cross-layer taxonomy of security threats, attack–defense risk evaluation frameworks, security countermeasures, datasets, regulatory frameworks, and post-quantum protection mechanisms. In parallel, the work in [13,14,15,16] explored ML-based IDS, anomaly detection frameworks, federated IDS architectures, and AI-assisted communication security mechanisms. More recently, a comprehensive survey of AI-driven cybersecurity was presented in [17], covering ML, DL, FL, RL, Graph Neural Network (GNN), and generative AI approaches for securing UAV-based systems. Nevertheless, the primary focus of these works remains on cyber–physical protection, communication security, IDS, and attack mitigation rather than broader AI-enabled operations.
Another important category of existing review papers focuses on AI-enabled autonomous navigation, trajectory optimization, and path planning for UAV-based systems. The work in [18] reviewed ML and RL techniques for UAV navigation, obstacle avoidance, collision prevention, and autonomous flight control in dynamic environments. The work in [19] investigated UAV navigation frameworks based on ML, DL, RL, and DRL, including those involving simultaneous localization and mapping (SLAM), visual navigation, localization, and adaptive flight control. Moreover, the study in [20] categorized graph-based, optimization-based, bio-inspired, and AI-driven UAV path planning approaches while analyzing cooperative multi-UAV navigation and energy-efficient routing. On the other hand, the work in [21] focused specifically on intelligent optimization algorithms for battlefield-oriented multi-UAV trajectory planning, collaborative swarm coordination, and dynamic battlefield path optimization. Although these studies provide extensive analyses of AI-driven navigation and trajectory optimization techniques, they mainly address autonomous navigation and cooperative path planning rather than integrated AI-enabled UAV-based defense systems.
Several existing review papers have also examined AI-enabled UAV technologies from broader civilian and industrial perspectives rather than focusing specifically on defense applications. The work in [22] reviewed ML-enabled applications in smart cities, communications, disaster management, surveillance, infrastructure inspection, and intelligent automation, while the work in [23] examined AI-assisted drone control, RL, swarm coordination, logistics, infrastructure inspection, and autonomous operational optimization for civilian UAV-based systems. In [24], the authors analyzed AI-driven safety and security technologies including anomaly detection, multimodal sensing, Large Language Model (LLM)-assisted UAV reasoning, and resilient communication mechanisms. While these studies cover a broad range of AI-enabled technologies, their primary focus remains on civilian drone applications, intelligent automation, operational optimization, and safety management rather than defense-oriented applications.
Other existing review papers have focused on specific operational functionalities, including target tracking and air combat simulation. The work in [25] reviewed target tracking technologies, active and passive tracking paradigms, swarm-based tracking coordination, multi-sensor fusion, and RL-assisted tracking frameworks. Moreover, the work in [26] surveyed simulation environments, ML techniques, RL approaches, and decision-support frameworks for Beyond-Visual-Range (BVR) air combat environments. Although these works provide important contributions in their respective areas, they focus on isolated operational functionalities rather than presenting a unified review of AI-enabled UAV-based defense systems across multiple interconnected operational domains.
Another emerging research direction concerns military UAV-based communications and airborne networking architectures designed to support resilient connectivity in contested environments. The work in [27] presented a survey of UAV-mounted aerial relays in military communications, examining the role of active aerial relays and Aerial Reconfigurable Intelligent Surface (ARIS) technologies for resilient battlefield connectivity. However, the primary focus of this work was on communication infrastructure and networking aspects of military UAV-based systems, and it did not address broader AI-enabled UAV-based defense domains. In parallel, a broader defense-oriented AI survey in [1] considered UAVs as just one component within the wider military and defense domains.
Consequently, as summarized in Table 1, the existing review literature primarily emphasizes specific technologies, operational functionalities, or application domains such as cybersecurity, autonomous navigation, target tracking, or communications. In addition, many existing reviews focus on civilian UAV applications, generic AI methodologies, or communication-layer security mechanisms, with fewer studies adopting a broader perspective that jointly examines autonomous decision-making, battlefield intelligence, resilient communications, swarm coordination, cybersecurity, electronic warfare, and mission-oriented UAV operations in contested and adversarial environments. Building upon these complementary contributions, this paper provides a broader cross-domain synthesis by integrating multiple defense application domains, enabling AI paradigms, system architectures, emerging research trends, open challenges, and future research directions within a unified analytical framework.

3. Background and System Architecture of AI-Enabled UAV-Based Defense Systems

3.1. UAV-Based Systems in Defense Applications

UAV-based systems have emerged as indispensable components of modern defense operations thanks to their operational flexibility, mobility, relatively low deployment cost, and ability to perform missions in hazardous and contested environments without directly exposing human personnel to risk. Their versatility enables deployment across a broad spectrum of military applications, including Intelligence, Surveillance, and Reconnaissance (ISR) missions, target acquisition, battlefield assessment, electronic warfare, communication relaying, logistics support, search-and-rescue operations, and autonomous combat missions [27]. Modern military forces increasingly rely on UAVs to provide persistent situational awareness, extend operational reach, support command-and-control (C2) functions, and enhance mission effectiveness across complex multi-domain battlefields.
Depending on their aerodynamic configuration and operational characteristics, UAVs are generally classified into fixed-wing, rotary-wing, and hybrid platforms [27]. Fixed-wing UAVs offer long flight endurance, high-altitude operation, and extensive area coverage, making them particularly suitable for persistent ISR, strategic reconnaissance, and communications relay missions; representative platforms include the RQ-4 Global Hawk, MQ-9 Reaper, and MQ-1 Predator. In contrast, rotary-wing UAVs provide Vertical Take-Off and Landing (VTOL), hovering capability, and superior maneuverability, allowing for their effective operation in urban, densely populated, and geographically constrained environments; typical examples include the MQ-8 Fire Scout and RQ-16 T-Hawk. Hybrid UAVs combine the endurance advantages of fixed-wing systems with the maneuverability of rotary-wing platforms to support diverse mission requirements ranging from tactical communications and electronic warfare to contested resupply operations and autonomous reconnaissance; representative hybrid platforms include the V-BAT and Arcturus Jump 20.
Military UAVs are further categorized according to their size, range, altitude, and endurance capabilities [27]. Micro and mini UAVs such as the Black Hornet Nano, RQ-11 Raven, and Wasp AE are primarily employed for close-range reconnaissance, force protection, and urban operations. Tactical UAVs such as the RQ-7 Shadow and Bayraktar TB2 provide persistent battlefield intelligence, target acquisition, and combat support. At the strategic level, Medium-Altitude Long-Endurance (MALE) and High-Altitude Long-Endurance (HALE) platforms such as the MQ-9 Reaper and RQ-4 Global Hawk enable wide-area surveillance, long-range reconnaissance, communications relaying, and persistent monitoring over extended operational regions. These categories involve distinct tradeoffs among endurance, payload capacity, survivability, and operational flexibility, leading modern defense forces to employ heterogeneous UAV fleets that combine strategic situational awareness with adaptable tactical support capabilities.
As defense missions continue to increase in complexity, modern UAV platforms are being equipped with advanced sensing, communication, computing, and control technologies that significantly expand their operational capabilities. These systems commonly integrate Electro-Optical/Infrared (EO/IR) cameras, radar systems, Light Detection and Ranging (LiDAR) sensors, Radio-Frequency (RF) sensing modules, Inertial Measurement Units (IMUs), GPS receivers, and onboard computing resources to support environmental perception, target detection and tracking, autonomous navigation, and situational awareness [28]. Furthermore, UAVs increasingly operate as components of interconnected and distributed communication architectures supported by UAV-to-UAV (U2U), UAV-to-Ground (U2G), and UAV-to-Satellite (U2S) links, enabling cooperative sensing, information sharing, coordinated mission execution, and swarm-level operation [29].
Beyond individual platforms, contemporary military operations increasingly rely on multi-UAV systems and autonomous swarms capable of cooperative sensing, distributed intelligence, adaptive networking, and decentralized decision-making [30]. Compared with single-UAV deployments, swarm-based architectures offer enhanced scalability, fault tolerance, operational resilience, and mission continuity by mitigating single points of failure and dynamically reallocating tasks among cooperating agents. Such capabilities are particularly valuable in contested environments characterized by electronic warfare, GPS denial, communication disruption, cyber–physical attacks, and Anti-Access/Area-Denial (A2/AD) strategies [27]. In these scenarios, UAVs not only support surveillance, reconnaissance, and strike missions but also contribute to electronic attack, spectrum monitoring, anti-jamming communications, and resilient battlefield networking.

3.2. AI Techniques for Intelligent UAV-Based Defense Systems

AI has emerged as a fundamental enabler of intelligent UAV-based defense systems by providing capabilities for perception, learning, reasoning, adaptation, and autonomous decision-making. Unlike conventional rule-based approaches, AI techniques can extract knowledge from large volumes of heterogeneous data, continuously improve their performance through experience, and adapt to dynamic and uncertain operational environments; consequently, AI has become increasingly important for enhancing the autonomy, resilience, scalability, and operational effectiveness of next-generation UAV platforms operating in complex and contested defense scenarios [1,2]. Recent defense-oriented studies have further highlighted that AI enables intelligent systems to process large volumes of information efficiently through advanced data analysis, pattern recognition, and predictive analytics. These capabilities support self-control, self-regulation, self-actuation, decision support, and autonomous operation, thereby enhancing the performance of defense-oriented UAV-based systems [1,2].
ML forms the foundation of many intelligent systems by enabling models to learn patterns and relationships directly from data without requiring explicitly programmed rules. Among ML approaches, DL has attracted significant attention due to its ability to automatically learn hierarchical feature representations from large-scale and high-dimensional datasets. Through multi-layer neural network architectures, DL can efficiently process heterogeneous information originating from sensors, communication systems, radar platforms, and autonomous vehicles while reducing the need for manual feature engineering. Moreover, ML and DL techniques can improve automation, prediction accuracy, anomaly detection, and data-driven decision-making, making them key enabling technologies for modern intelligent defense systems [1,3]. Several neural network architectures have emerged as important building blocks of modern AI systems: Convolutional Neural Networks (CNNs) are designed to extract spatial features through hierarchical convolution operations, whereas Recurrent Neural Networks (RNNs) and Long Short-Term Memory (LSTM) networks extend learning capabilities to sequential and temporal information by capturing long-term dependencies and historical observations [15,22]; more recently, transformer-based architectures have emerged as powerful alternatives to conventional recurrent models. By employing self-attention, transformers can efficiently model long-range dependencies and contextual relationships while supporting highly parallel computation [17].
Reinforcement Learning (RL) represents a fundamentally different learning paradigm in which an autonomous agent learns optimal behavior through continuous interaction with its environment. Rather than relying on labeled training data, RL agents learn through trial-and-error exploration by observing environmental states, selecting actions, and receiving reward feedback. This learning mechanism enables the autonomous discovery of decision policies that maximize long-term cumulative rewards while adapting to dynamic and uncertain environments. Building upon RL, Deep RL integrates Deep Neural Networks (DNNs) to handle high-dimensional state and action spaces, resulting in improved learning ability and decision-making performance in complex environments [15,22]. The increasing deployment of multiple autonomous agents has further motivated the development of MARL, in which multiple interacting agents simultaneously learn cooperative, competitive, or mixed strategies. Compared with single-agent RL, MARL supports decentralized decision-making, collaborative adaptation, distributed task execution, and the emergence of collective behaviors through agent interactions, making it particularly suitable for large-scale autonomous systems involving multiple cooperating entities [31].
As AI systems become increasingly autonomous and operationally critical, transparency and interpretability have emerged as important requirements. Explainable AI seeks to improve the understanding of AI-generated decisions by providing human-interpretable explanations of model behavior and decision-making processes. By enhancing transparency, accountability, fairness, reliability, and trustworthiness, XAI facilitates the validation, monitoring, and deployment of AI systems in safety-critical environments. These characteristics are particularly important in defense systems, where autonomous decisions may have significant operational and ethical implications [2,3].
Although the above AI paradigms enable intelligent UAV-based defense systems, their successful deployment depends on the availability of suitable training data and learning environments. Supervised ML and DL approaches generally require large, representative, and high-quality datasets with reliable annotations to achieve robust generalization; in contrast, RL- and MARL-based methods rely on carefully designed reward functions, realistic simulation environments, and appropriate exploration strategies to learn effective decision policies. Since collecting large-scale defense data and conducting extensive real-world experiments are often challenging, many studies rely on synthetic datasets, digital twins, or high-fidelity simulators for training and evaluation. Thus, the quality and representativeness of training data, reward design, and simulation environments are critical to the performance and practical deployment of AI-enabled UAV-based defense systems.
Importantly, the above AI paradigms exhibit different trade-offs in terms of computational complexity, scalability, robustness, communication requirements, onboard implementation, and suitability for real-time operation. Conventional ML techniques generally require relatively low computational resources, making them well suited to resource-constrained UAV platforms. In contrast, architectures based on transformers and DL impose substantially higher computational, memory, and training requirements, making efficient onboard deployment more challenging. Approaches involving RL and DRL often require extensive training, and may exhibit slow convergence in complex and dynamic environments. MARL further introduces communication overhead and scalability challenges associated with multi-agent coordination, while FL depends on reliable communication links and periodic model synchronization among distributed UAVs. Finally, although XAI incurs relatively modest additional computational overhead, integrating explainability into complex AI models may increase implementation complexity and inference latency. Consequently, the selection of an appropriate AI paradigm should balance intelligence and autonomy against computational resources, communication constraints, and the real-time operational requirements of the target defense application.

3.3. Integrated System Architecture of AI-Enabled UAV-Based Defense Systems

The previous subsections have reviewed the principal UAV platforms, enabling communication and sensing technologies, and AI paradigms that underpin modern defense-oriented UAV-based systems. Building upon these observations, this subsection presents a conceptual architecture that synthesizes the recurring functional components and their interactions identified across the reviewed literature. Figure 2 illustrates this architecture and functional classification framework for AI-enabled UAV-based defense systems. The recurring functional components are organized into seven interconnected layers that provide a unified system-level perspective. The interactions among these layers illustrate how sensing, communication, intelligence, control, and security functions collectively support resilient, autonomous, and mission-aware UAV operation in contested and dynamic environments.
The UAV platform layer consists of heterogeneous platform infrastructures comprising fixed-wing, rotary-wing, and in some cases hybrid UAVs, each providing complementary capabilities in terms of endurance, maneuverability, payload capacity, and operational flexibility [27,28]. These aerial platforms constitute the physical foundation of the architecture by hosting the sensing payloads, communication subsystems, onboard computing resources, and mission execution functionalities. Through their direct interaction with the sensing and communication layers, the UAV platforms enable acquisition and dissemination of mission-critical information. Building upon the platform layer, the sensing and perception layer enables UAVs to acquire and interpret information about the surrounding operational environment. Modern UAV platforms increasingly integrate diverse mission-specific sensing technologies, including EO/IR cameras, radar systems, RF sensing modules, inertial navigation systems, and other onboard sensors to support environmental perception, target detection, navigation, and situational awareness [28,30]. A recurring trend identified across the reviewed literature is the adoption of multimodal sensing, radar-enabled perception, and multi-sensor data fusion techniques to improve robustness and operational effectiveness under uncertain and adversarial conditions. The heterogeneous observations collected by the sensing layer are subsequently forwarded to the AI-enabled intelligence and computing layer, where they are fused and transformed into actionable operational intelligence.
The communication layer provides reliable information exchange among UAVs, ground control stations, distributed operational nodes, and satellite networks [29]. As highlighted in previous works [27,32], resilient communication constitutes a key enabler of cooperative sensing, distributed intelligence, swarm coordination, and autonomous mission execution. However, military communication environments are frequently affected by interference, spectrum congestion, communication disruption, jamming, and spoofing [5]. Thus, recent research has increasingly adopted AI-driven communication mechanisms to support adaptive spectrum access, intelligent routing, transmission scheduling, anti-jamming communication, and secure network management [24]. Recent research has also investigated AI-driven event-triggered communication mechanisms [33], in which information is transmitted only when significant events or predefined conditions are detected rather than through periodic updates. Such strategies can substantially reduce communication overhead, conserve energy, and improve spectrum efficiency while maintaining timely situational awareness and coordination. These capabilities are particularly beneficial for cooperative UAV operations, swarm intelligence, and autonomous combat missions operating under bandwidth-constrained and contested environments. The communication layer continuously exchanges mission information, network status, and coordination data with the AI-enabled intelligence and computing layer to enable adaptive communication management, collaborative learning, and cooperative mission execution.
The AI-enabled intelligence and computing layer serves as the core of the architecture by integrating heterogeneous sensor observations with communication and mission information to generate high-level operational intelligence and support adaptive decision-making. Depending on the operational requirements, this functionality may be implemented through onboard edge intelligence or cloud-based computing infrastructures [34]. Edge intelligence enables real-time AI inference, low-latency processing, local decision-making, and energy-efficient operation directly onboard the UAV, thereby reducing dependence on remote computational resources and improving responsiveness in bandwidth-constrained and contested environments. In contrast, cloud infrastructures provide large-scale model training, data aggregation, intelligence fusion, and model synchronization capabilities. The adoption of advanced AI techniques enables intelligent perception, autonomous reasoning, adaptive learning, communication-aware operation, and collaborative mission execution [1,2,3,17]. Furthermore, distributed AI and collaborative learning mechanisms enable multiple UAVs to share information, coordinate actions, and collectively contribute to mission objectives while reducing reliance on centralized infrastructures [35]. Acting as the central integration layer, it exchanges information with the communication layer, provides high-level operational intelligence to the autonomous control and decision-making layer, cooperates with the cybersecurity layer to detect and mitigate cyber and electronic warfare threats, and continuously adapts its inference and decision-making processes according to changes in the surrounding operational environment. Building upon the intelligence generated by the AI layer, the autonomous control and decision-making layer is responsible for mission execution, adaptive operational management, and autonomous reasoning. It supports autonomous navigation, path planning, mission planning, task scheduling, decision support, target tracking and identification, swarm coordination, communication management, payload control, and fail-safe operation [15,22,31]. This layer plays a pivotal role in enabling autonomous and cooperative UAV operations within dynamic and contested battlefield environments.
Given the increasing exposure of UAV platforms to cyber–physical and electromagnetic threats, the layer constituting cybersecurity, electronic warfare protection, and operational resilience provides an indispensable architectural component. Contemporary architectures increasingly incorporate intrusion detection, anomaly detection, anti-jamming mechanisms, spoofing mitigation, authentication, encryption, secure communication protocols, spectrum-aware capabilities, electronic support and protection measures, data integrity protection, adversarial attack detection, and resilient fault-tolerant operation [13,14,15,16,17]. AI-driven cybersecurity and electronic warfare frameworks further enhance the ability to detect malicious communication behavior, abnormal UAV operation, cyber intrusions, adversarial attacks, jamming activities, spoofing attempts, and electromagnetic interference in real time. In addition, XAI techniques are increasingly employed to improve transparency, interpretability, and operator trust in AI-generated decisions as a way to facilitate trustworthy autonomous operation and system validation in mission-critical defense applications [2,3]. Through continuous interaction with the AI-enabled intelligence and computing layer, this component strengthens cybersecurity awareness, spectrum awareness, trust management, communication robustness, fault tolerance, and overall operational resilience.
Finally, the operational environment represents the external context within which all UAV operations are conducted. Changes in the operational environment continuously influence sensing observations, communication conditions, and mission objectives. Through bidirectional interaction with the AI-enabled intelligence and computing layer, the full architecture supports resilient and mission-aware UAV operation by continuously adapting perception, learning, resource allocation, and mission planning according to evolving operational conditions.

4. Autonomous Air Combat and Cooperative UAV Operations

This section reviews recent advances in AI-driven autonomous air combat and cooperative UAV operations that enable intelligent maneuver generation, adaptive tactical reasoning, autonomous engagement planning, and distributed combat coordination in dynamic and adversarial environments. Unlike conventional rule-based or human-centered approaches, AI-driven frameworks can enable UAVs to operate under incomplete situational awareness, adversarial actions, communication constraints, electronic warfare, and rapidly evolving battlefield conditions. These capabilities are particularly important in BVR engagements, close-range dogfighting, cooperative mission scenarios, and large-scale swarm operations, where rapid decision-making and adaptive coordination are essential for mission success.

4.1. Autonomous Air Combat Maneuver Decision-Making

Autonomous Air Combat Maneuver Decision-Making (ACMDM) represents one of the most demanding applications of AI in aerial warfare, enabling UAVs to autonomously assess combat situations, anticipate adversarial actions, and generate adaptive tactical maneuvers under highly dynamic and contested conditions. Recent studies have demonstrated that DRL-based approaches can enhance the ability of UAVs to learn complex combat strategies through continuous interaction with realistic air combat environments. This subsection reviews representative AI-driven approaches for ACMDM, which are summarized in Table 2.
The work in [36] proposed a DRL-based ACMDM framework based on a Motivational Curriculum Learning Distributed Proximal Policy Optimization (MCLDPPO) algorithm designed to mitigate the plasticity loss problem encountered in conventional curriculum learning approaches. A high-fidelity digital twin air combat environment incorporating realistic aerodynamics, radar sensing, infrared systems, missile guidance, and beyond-visual-range missile engagements was modeled as a Partially-Observable Markov Decision Process (POMDP). The framework integrated motivational curriculum learning rewards, distributed Proximal Policy Optimization (PPO) training, an LSTM-based actor–critic architecture, hierarchical maneuver actions, and an interruption mechanism that increased decision frequency during missile-threat situations. The observation space included aircraft kinematics, radar lock status, missile warning information, and several combat geometry metrics, namely, Antenna Train Angle (ATA), Aspect Angle (AA), and Horizontal Crossing Angle (HCA), whereas the action space comprised tracking, circling, attacking, and evasive maneuvers. Simulation results showed that MCLDPPO significantly outperformed Finite State Machine (FSM) expert systems, predictive game trees, and conventional RL algorithms, achieving an average win rate of approximately 66% against expert-level FSM opponents, compared with 60.7% for predictive game trees, 43% for Soft Actor–Critic (SAC), 40% for Deep Deterministic Policy Gradient (DDPG), and 35% for Deep Q-Network (DQN). Moreover, the trained agent exhibited sophisticated tactical behaviors such as proactive missile feints, adaptive diving attacks, guidance chain maintenance, and dynamic exploitation of posture advantages.
A similar direction was investigated in [37], where an improved PPO-based framework was developed for autonomous maneuver decision-making in one-to-one short-range UAV air combat. The considered combat environment had three degrees of freedom and incorporated maneuvering constraints, attack and advantage regions, probabilistic damage models, and blood value-based combat evaluation mechanisms. To improve learning performance, the framework employed Gated Recurrent Unit (GRU)-enhanced actor–critic networks, differentiated actor and critic observation spaces, phased curriculum training, and a reward structure combining dense, event, and terminal rewards. The action space comprised fifteen tactical maneuvers, while the observation space captured relative position, velocity, heading, attack angles, and combat geometry information. A prediction-and-decision-based adversarial policy further generated more challenging combat interactions during training. Experimental and ablation studies demonstrated substantially faster convergence, improved policy stability, enhanced temporal situational awareness, and superior combat effectiveness compared with conventional PPO implementations. As a result, the trained UAV learned adaptive offensive and evasive strategies capable of maintaining positional superiority, approaching attack regions, avoiding enemy threats, and autonomously executing complex combat maneuvers.
Beyond improving combat effectiveness, recent research has emphasized the explainability and transparency of AI-enabled air combat systems. The work in [38] proposed an XAI framework for interpreting the tactical decision-making behavior of RL-based air combat agents operating in a 2D close-range combat environment. The framework combined DQN, Double DQN with reward decomposition, semantically meaningful reward functions, and both local and global explainability mechanisms based on combat geometry metrics, including Line-of-Sight (LoS), ATA, and AA. By decomposing rewards into ATA-, AA-, and LoS-related components and employing tactical region-based analysis together with visual explanation maps, the framework enabled detailed interpretation of agent decisions and tactical preferences. Experimental results revealed tactical strengths, weaknesses, asymmetries, and inconsistent behavioral patterns that remained hidden in conventional black-box RL systems, leading to improved transparency, debugging capability, operator trust, and understanding of AI-generated combat maneuvers.
Complementing these studies, the work in [39] developed and implemented a complete DRL-based ACMDM framework for one-to-one UAV dogfight scenarios. The system employed a DQN operating in a 3D combat environment modeled through 6-DOF UAV kinematics, where the state space included velocity, spatial coordinates, attitude information, relative distance, and combat geometry parameters. Tactical behavior was represented through seven National Aeronautics and Space Administration (NASA)-inspired Basic Fighter Maneuvering (BFM) primitives. Prioritized experience replay, ϵ -greedy exploration, curriculum learning, and a reward mechanism combining angle-, distance-, speed-, height-, and outcome-based rewards were integrated to address sparse reward-related challenges and improve training efficiency. In addition to providing detailed implementation guidelines covering environment construction, state and action design, reward shaping, neural network architecture, training procedures, and performance evaluation, the study demonstrated that the trained UAV successfully learned target pursuit, attack positioning, maneuver adaptation, and autonomous engagement strategies through continuous interaction with the combat environment. The results further highlighted the effectiveness of DRL in addressing the large state space, nonlinear dynamics, and sequential decision-making requirements of autonomous air combat.
Overall, the reviewed studies demonstrate that DRL has become a widely explored AI paradigm for autonomous air combat maneuver decision-making, enabling UAVs to learn increasingly sophisticated tactical behaviors in complex combat environments. Recent research has enhanced learning efficiency, situational awareness, and policy interpretability through curriculum learning, recurrent neural networks, explainable AI, and high-fidelity digital twins. Nevertheless, common limitations remain across the reviewed approaches, including extensive reliance on simulation environments, partial observability, scalability and communication constraints in larger combat scenarios, and the need for more realistic operational validation before deployment.

4.2. Cooperative Multi-UAV Air Combat

This subsection investigates AI-driven approaches for cooperative multi-UAV air combat and MARL, which have emerged as major research directions in modern AI-enabled aerial combat systems and are summarized in Table 3. In such highly dynamic and adversarial environments, multiple UAVs must collaboratively perform coordinated target engagement, tactical maneuver synchronization, communication-aware cooperation, distributed battlefield perception, and adaptive mission execution under uncertainty, limited situational awareness, and rapidly evolving combat conditions.
The work in [40] proposed an evolutionary MARL framework for cooperative multi-UAV air combat in dynamic 3D environments. Specifically, the authors developed an Evolutionary Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (E-MATD3) algorithm to address strategic cycles arising from adversarial self-play and dynamically varying swarm sizes. A realistic within-visual-range combat environment involving two opposing teams of five UAVs operating in a continuous 3D battlefield was modeled using 3-DOF flight dynamics with realistic combat constraints. The framework integrated MATD3-based Centralized Training and Decentralized Execution (CTDE), evolutionary population-based training, attention-based policy networks, curriculum learning, death masking, crossover and mutation operators, and TrueSkill-based evolutionary selection. Experimental results demonstrated that E-MATD3 outperformed MADDPG [41], MATD3 [42], DYAN-SUM [43], and MATD3 with attention mechanism (ATT-MATD3) [40] baselines. Moreover, the learned policies exhibited sophisticated cooperative tactics, including nearest-target pursuit, coordinated cross-attacks, fire attraction maneuvers, and cooperative target neutralization, highlighting the benefits of combining evolutionary learning and attention mechanisms to improve robustness, scalability, and tactical cooperation.
Similarly, the work in [44] investigated hierarchical MARL for cooperative multi-UAV air combat in high-fidelity environments. The proposed framework combined virtual opponent-based hierarchical policies with value–attention decomposition policy gradients to address the challenges of high-order aircraft dynamics and multi-agent credit assignment. Evaluations were conducted using the Pyaircombat simulation platform, developed with the Unity game engine and the open-source aircraft dynamics model software JSBSim, which supports realistic flight dynamics, radar-based observations, weapons engagement modeling, and cooperative combat scenarios from 3 vs. 3 to 6 vs. 6. The framework integrated hierarchical RL (HRL) [45], CTDE, PPO [46], Generalized Advantage Estimation (GAE), attention-based value decomposition, and pretrained low-level maneuver policies. The upper layer generated 6-DOF virtual opponent poses, while the lower layer translated them into aircraft control commands. Experimental results demonstrated superior performance compared with Multi-Agent Proximal Policy Optimization (MAPPO) [46] and related MARL baselines across multiple combat scales. Furthermore, the learned policies exhibited coordinated target selection and collaborative engagement, demonstrating improved training efficiency, scalability, and cooperative decision-making.
In [47], the authors proposed a decomposed and Prioritized Experience Replay (PER)-enhanced Multi-Agent Deep Deterministic Policy Gradient (DP-MADDPG) framework for multi-UAV air combat operational scenario. The cooperative and adversarial UAV swarm combat was modeled as a Partially-Observable Markov Game (POMG) in which UAVs autonomously performed maneuvering, target tracking, cooperative attacks, and distributed battlefield coordination. In addition, a combat environment implemented using the Multi-agent Combat Arena (MaCA) involved two opposing swarms, each consisting of eight combat UAVs and two reconnaissance UAVs. The framework integrated CTDE, dual-critic MADDPG learning, PER, and local/global reward optimization to improve learning stability and cooperative coordination. Compared with conventional MADDPG [41] and Independent Learning Deep Deterministic Policy Gradient (ILDDPG), DP-MADDPG achieved a convergent reward of 695.5, compared with 592.9 for MADDPG and 344.8 for ILDDPG; it also attained a 96% win rate against rule-based opponents and an 80.5% win rate against ILDDPG-enabled adversaries over 1000 combat tests. The learned policies demonstrated coordinated encirclement attacks, tactical baiting, target tracking, evasive maneuvering, and cooperative protection of reconnaissance UAVs.
The work in [48] proposed a cooperative maneuver decision-making framework combining incomplete-information dynamic game theory, Bayesian inference, and MADDPG-based RL. The study addressed combat environments in which adversarial UAV intentions and tactical maneuvers were uncertain and only partially observable. A cooperative air combat situation assessment model was first developed using combat geometry metrics such as angle advantage, heading advantage, and distance advantage to evaluate battlefield superiority. The combat problem was then formulated as an incomplete-information dynamic game solved through Perfect Bayes-Nash Equilibrium (PBE) analysis, while a dynamic Bayesian network (DBN) inferred opponent tactical intentions from observable variables such as velocity, heading angle, and altitude evolution. The framework subsequently employed a MADDPG architecture with CTDE to learn adaptive cooperative maneuver strategies. Simulation results demonstrated faster convergence and improved exploitability characteristics compared with Twin Delayed DDPG (TD3)-based baselines [49]. The learned policies autonomously generated advanced tactical behaviors including outflanking maneuvers, coordinated pincer attacks, decoy tactics, altitude advantage exploitation, cooperative evasion, and adaptive pursuit strategies.
Similarly, the work in [50] proposed a Layer Delay Dual-Center MAPPO (MAPPO-LDC) framework to improve training stability and tactical performance in cooperative multi-UAV combat. The study addressed policy overestimation, cumulative training errors, and unstable convergence in complex air combat environments involving 2 vs. 2, 4 vs. 4, and 6 vs. 6 UAV operational scenarios. In particular, the proposed framework integrated CTDE, dual-center critic networks, delayed policy updates, curriculum-style training, and PPO-based multi-agent learning. The dual-center critic architecture alleviated overestimation problems through weighted averaging between critic networks, while delayed policy updates stabilized training by prioritizing critic updates during early learning stages. Simulation results demonstrated that MAPPO-LDC substantially outperformed MAPPO [46] and Independent PPO (IPPO) [51] baselines in convergence stability, combat performance, and tactical effectiveness. In particular, the framework achieved a maximum win rate of approximately 93% in the 2 vs. 2 scenario, corresponding to an 89% improvement over MAPPO, while also reducing draw rates by approximately 84%. The analysis further revealed advanced cooperative tactical behaviors including pincer attacks, coordinated flanking maneuvers, decoy tactics, and exploitation of advantageous positioning across multiple combat scenarios.
The work in [52] proposed a hierarchical game-theoretic and risk-aware predictive control framework for resilient cooperative multi-UAV air combat in A2/AD environments. The proposed architecture integrated Stackelberg game-based task allocation with hierarchical RL, where distributional RL based on Implicit Quantile Networks (IQN) and Conditional Value-at-Risk (CVaR) were employed to dynamically adapt tactical decisions according to battlefield threat levels. Specifically, the hierarchical framework separated strategic task allocation, risk-aware tactical decision-making, and resilient low-level flight control, enabling cooperative target engagement while accounting for communication uncertainty, disturbances, and cyber attacks. Experimental results demonstrated improved mission survivability, tactical coordination, and robustness compared with conventional MARL approaches, highlighting the potential of integrating risk-aware RL with hierarchical decision-making for cooperative UAV combat.
HRL has emerged as an effective approach for addressing the complexity of large-scale cooperative combat environments. In [53], the authors proposed an HRL-based framework using a hierarchical decision-making network and an empirical experience decomposition mechanism. The problem was modeled as a decentralized partially observable Markov decision process (DEC-POMDP) and evaluated using the high-fidelity JSBSim simulator in realistic 4 vs. 4 and 8 vs. 8 combat scenarios involving missile engagements, radar constraints, and cooperative maneuvering. The hierarchical structure separated UAV decisions into a flight decision layer for maneuver generation and an attack decision layer for target selection and firing decisions, which reduced the complexity in the action space and improved decision-making efficiency. Built upon the QMIX value decomposition architecture, the framework significantly outperformed Value Decomposition Network (VDN), Counterfactual Multi-Agent (COMA), and QMIX baselines [54] in win rate, convergence speed, combat loss rate, and training stability. Specifically, it achieved win rates of 64.7% and 82.9% in 4 vs. 4 and 8 vs. 8 combat scenarios, respectively, while demonstrating faster convergence, improved robustness under radar noise, and advanced cooperative behaviors such as encirclement tactics, missile timing optimization, and formation dispersion.
Communication-aware cooperation was investigated in [55], where the authors proposed an improved Multi-Agent DRL (MADRL) framework for cooperative maneuver decision-making in pursuit-evasion air combat scenarios. A 2vs1 operational scenario was considered in which two cooperative pursuit UAVs attempted to intercept an evasive adversarial UAV in a dynamic 3D battlefield. The proposed CommNet-Advantage Actor–Critic (CommNet-A2C) framework combined Advantage Actor–Critic RL with communication-aware multi-agent coordination to enhance cooperative decision-making. Specifically, the architecture integrated communication interaction layers, GRU-based memory mechanisms, shared rewards, and communication protocol learning to improve temporal situational awareness and inter-UAV information sharing. In addition, the framework employed cooperative tactical decision-making and NASA-inspired maneuver primitives, including acceleration, turning, pull-up, and dive-down actions. Simulation results demonstrated that CommNet-A2C significantly outperformed the baseline CommNet-REINFORCE algorithm, which corresponds to the original CommNet framework employing the REINFORCE policy-gradient algorithm, in both pursuit success rate and cumulative reward. The proposed framework converged after approximately 700 training episodes and achieved an average pursuit success rate of 92.7%, compared with 68.5% for CommNet-REINFORCE. Furthermore, the learned policies exhibited coordinated pincer attacks, adaptive encirclement maneuvers, and effective exploitation of altitude advantages.
Recent studies have also explored multimodal perception and intelligent task scheduling for cooperative UAV combat operations. In [56], the authors proposed an enhanced Multimodal DRL DQN (MDRL-DQN) for intelligent combat planning and cooperative UAV task scheduling in adversarial electronic warfare environments. The study considered dynamic multi-UAV combat operations involving target prioritization, electromagnetic interference, collision avoidance, and cooperative mission execution under red-team attack and blue-team defense scenarios. The proposed framework integrated multimodal data fusion, CNN-based image feature extraction, RNN-based sensor processing, adaptive reward mechanisms, self-attention-based task prioritization, actor–critic policy evaluation, and improved Q-learning. Experimental results demonstrated that MDRL-DQN significantly outperformed PPO, Asynchronous Advantage Actor–Critic (A3C) [57], and DDPG baselines in mission success rate, task completion time, UAV survivability, and convergence speed. In particular, the framework achieved task success rates of 89.6% and 94.8% in long-distance dispersed and concentrated defense scenarios, respectively, while also converging substantially faster than competing approaches.
Table 3. Overview of AI-driven approaches for cooperative multi-UAV air combat.
Table 3. Overview of AI-driven approaches for cooperative multi-UAV air combat.
ReferenceObjectiveAI/ML ModelsOperational ScenarioUAV ConfigurationMain ComponentsKey Results
Wang, B. et al., 2024 [40]Cooperative multi-UAV air combat maneuver decision-makingE-MATD3, evolutionary MARL, attention-based learning3D within-visual-range air combat5 vs. 5 cooperative combat UAV teamsCTDE, evolutionary population training, curriculum learning, attention-based policies, death masking, crossover and mutation operators, TrueSkill-based selectionOutperformed MADDPG, MATD3, DYAN-SUM, and ATT-MATD3 while enabling coordinated target pursuit and cooperative combat maneuvers
Xu et al., 2025 [44]Hierarchical cooperative multi-UAV air combat decision-makingHierarchical MARL, PPO, attention-based value decompositionHigh-fidelity multi-UAV air combat3 vs. 3–6 vs. 6 UAV operational scenariosHRL, CTDE, PPO, GAE, virtual-opponent modeling, attention-based value decomposition, pretrained low-level maneuver policiesOutperformed MAPPO and related MARL baselines while enabling coordinated target selection and collaborative target engagement
Yang et al., 2024 [47]Cooperative UAV swarm operational scenario and tactical coordinationDP-MADDPG, PER-enhanced MADDPGPOMG-based multi-UAV battlefieldEight combat UAVs and two reconnaissance UAVs per teamCTDE, dual-critic learning, PER, local/global reward optimization, cooperative coordinationAchieved a convergent reward of 695.5 vs. 592.9 (MADDPG) and 344.8 (ILDDPG), with 96% and 80.5% win rates against rule-based and ILDDPG opponents, respectively
Ren et al., 2023 [48]Cooperative maneuver generation under uncertaintyMADDPG, Bayesian inference, dynamic game theoryIncomplete-information multi-UAV air combat2 vs. 2, 2 vs. 3, and 3 vs. 2 combat scenariosPBE analysis, DBN-based intention inference, CTDE, cooperative situation assessmentImproved convergence and exploitability characteristics compared with TD3 while generating adaptive cooperative combat tactics
Ding et al., 2025 [50]Collaborative multi-UAV air combat decision-makingMAPPO-LDC, PPO-based MARLAdversarial 3D air combat2 vs. 2, 4 vs. 4, and 6 vs. 6 UAV teamsCTDE, dual-center critics, delayed policy updates, curriculum-style trainingAchieved a 93% win rate in 2 vs. 2 combat, improving over MAPPO by 89% while reducing draw rates by 84%
Tang et al., 2026 [52]Risk-aware cooperative multi-UAV air combat decision-makingHierarchical RL, IQN-based distributional RL, Stackelberg gameA2/AD cooperative air combatMulti-UAV cooperative combat swarmStackelberg-game task allocation, hierarchical RL, IQN, CVaR-based risk-aware decision-making, resilient predictive controlOutperformed MAPPO and CBBA while improving mission survivability, tactical coordination, and robustness, with over 60% lower computational load
Wang, H. et al., 2024 [53]Hierarchical cooperative multi-UAV combat decision-makingHRL, QMIX-based MARLJSBSim-based adversarial air combat4 vs. 4 and 8 vs. 8 UAV teamsDEC-POMDP modeling, hierarchical maneuver and attack decision layers, experience decomposition, QMIX value decompositionAchieved win rates of 64.7% and 82.9% in 4 vs. 4 and 8 vs. 8 combat, respectively, outperforming VDN, COMA, and QMIX
Luo et al., 2024 [55]Cooperative pursuit-evasion maneuver decision-makingCommNet-A2C, communication-aware MADRL3D pursuit-evasion air combatTwo pursuit UAVs versus one evasive UAVCommunication interaction layers, GRU memory modules, shared rewards, communication protocol learning, NASA-inspired maneuver primitivesConverged after approximately 700 episodes and achieved a pursuit success rate of 92.7%, compared with 68.5% for CommNet-REINFORCE
Jianhong et al., 2025 [56]Cooperative combat planning and UAV task schedulingMDRL-DQN, multimodal DRLElectronic warfare attack-defense environmentsMulti-UAV attack and defense teamsMultimodal fusion, CNN-based image processing, RNN-based sensor processing, adaptive rewards, self-attention task prioritization, actor–critic evaluationAchieved task success rates of 89.6% and 94.8% in dispersed and concentrated defense scenarios, respectively, outperforming PPO, A3C, and DDPG
Han et al., 2025 [58]Communication-efficient cooperative UAV air combatSIIS-MARL, ToM-based MARLCommunication-constrained multi-UAV combat2 vs. 2–16 vs. 16 UAV swarmsToM-based intention inference, sparse communication, LSTM temporal modeling, CTDEReduced communication overhead and achieved winning rates of 99.3% and 76.8% in 2 vs. 2 and 16 vs. 16 combat, respectively, surpassing MADDPG, COMA, DAACMP, and MASIA
In [58], the authors proposed a Sparse Inferred Intention Sharing MARL (SIIS-MARL) framework for communication-constrained multi-UAV combat environments. The study modeled cooperative air combat as a partially observable Markov game under a CTDE architecture in which UAVs autonomously performed reconnaissance, maneuvering, communication, target selection, and missile attacks in combat scenarios ranging from 2 vs. 2 to 16 vs. 16 UAV engagements. The proposed framework integrated a Theory-of-Mind (ToM)-based intention inference module [59], an attention-based sparse communication mechanism, LSTM-based temporal feature extraction, and actor–critic MARL. The ToM module inferred teammates’ future intentions and selection of enemy targets, whereas the sparse transmission mechanism selectively enabled communication only with tactically relevant UAVs, thereby reducing redundant communication overhead while preserving cooperative coordination. Experimental results demonstrated that SIIS-MARL significantly outperformed MARL baselines including MADDPG, COMA, Double Attentional Actor–Critic Message Processor (DAACMP) [60], and MASIA [61], particularly in large-scale combat environments. Specifically, the framework achieved winning rates of 99.3% in 2 vs. 2 combat and 76.8% in 16 vs. 16 combat. Ablation studies further confirmed the importance of both the intention inference and sparse communication mechanisms, as removing either degraded combat performance.
Overall, the reviewed studies demonstrate that MARL has emerged as a frequently adopted framework for cooperative multi-UAV air combat, enabling coordinated target engagement, distributed decision-making, and adaptive tactical cooperation in increasingly complex battlefield scenarios. Recent advances have enhanced scalability, training stability, communication efficiency, cooperative intelligence, and risk-aware tactical decision-making through CTDE, hierarchical learning, attention mechanisms, communication-aware coordination, multimodal perception, and distributional RL. Nevertheless, most existing approaches remain largely simulation-based and continue to face challenges related to scalability, communication overhead, partial observability, and realistic operational validation in large-scale adversarial environments.

4.3. Large-Scale Swarm Air Combat and Scalable Combat Intelligence

This subsection reviews AI-driven approaches for large-scale swarm air combat and scalable combat intelligence, which have emerged as key research directions in next-generation autonomous aerial warfare and are summarized in Table 4. Compared with small-scale cooperative air combat, large-scale UAV swarms introduce additional challenges associated with scalability, partial observability, communication constraints, dynamic formations, non-stationary interactions, and rapidly expanding state-action spaces.
The work in [62] investigated scalable swarm combat coordination using an RL-based framework for short-range UAV swarm air combat maneuver decision-making based on a Multi-Agent Transformer introducing Virtual Objects (MTVO) architecture. Specifically, the study addressed the challenge of processing complex and dynamically evolving battlefield information in large-scale cooperative swarm combat, where conventional fully connected neural networks often fail to capture tactically important interactions among multiple UAVs. The considered scenario involved two homogeneous UAV swarms engaging in short-range 3D air combat, where UAVs autonomously performed maneuvering, target tracking, cooperative attacks, and evasive operations without inter-UAV communication. The combat environment incorporated realistic attack cones, observation constraints, probabilistic destruction mechanisms, three-degree-of-freedom UAV flight dynamics, and a maneuver library comprising 27 actions. The proposed MTVO framework integrated transformer-based self-attention, virtual-object-assisted global situation representation, mask mechanisms for varying swarm sizes, PPO-based CTDE, attention-based information aggregation, reward shaping, situation assessment, and actor–critic RL. Specifically, the self-attention mechanism focused on tactically important UAVs, while the virtual object generated a global battlefield representation for weighted fusion of local tactical information. Simulation experiments ranging from 1 vs. 1 to 20 vs. 20 UAV operational scenarios demonstrated that MTVO significantly outperformed conventional fully connected neural-network approaches, particularly as swarm size increased. In 5 vs. 5 combat scenarios, the framework achieved win rates exceeding 70% while maintaining stable convergence and effective tactical learning. Unequal-number engagements further demonstrated strong robustness under numerical disadvantage, whereas ablation studies confirmed the importance of the self-attention, attention aggregation, and virtual object modules, as removing any component substantially degraded combat performance.
Similarly, the work in [63] investigated scalable multi-aircraft autonomous close-range air combat under a multi-sensor Unmanned Combat Aerial Vehicle (UCAV) platform. The framework considered variable-size cooperative UAV formations operating in partially observable and adversarial environments. It integrated QMIX-based MARL, CTDE, Multihead Attention (MHA)-based information embedding, Scalable Mixing Networks based on Attention (SMA), Population Curriculum Self-Play (PCSP), GRU-based recurrent Q-networks, and competitive self-play training. Combat scenarios involved homogeneous UCAV formations performing cooperative maneuvering, target engagement, threat assessment, and coordinated attacks in realistic close-range 3D environments incorporating Weapon Engagement Zones (WEZs), ATA, AA, probabilistic damage modeling, radar sensing, and three-degree-of-freedom flight dynamics. Multi-sensor information fusion enabled UAVs to exchange radar-based enemy information with nearby teammates, while MHA-based embedding and SMA allowed scalable learning with dynamically varying numbers of friendly and adversarial UAVs. Experimental results showed that the framework significantly outperformed QMIX, VDN, and MeanPool-based approaches in convergence stability, scalability, training efficiency, and combat effectiveness. In variable-size formations, it achieved winning rates exceeding 70%, while random-PCSP training attained the highest Elo scores among competitive self-play methods, indicating superior strategy quality and improved approximate Nash-equilibrium learning.
The work in [64] proposed a cooperative RL framework for autonomous military drones performing swarm-flight maintenance, adaptive target bombing, and energy-efficient mission execution under dynamic combat conditions. The proposed MADRL framework combined CommNet and CTDE, where a mobile command tower acted as a centralized critic while individual drones executed learned policies using local observations and inter-agent communication. Unlike independent RL approaches, the CommNet architecture enabled continuous information exchange through hidden-state communication mechanisms, thereby improving cooperative coordination and tactical synchronization. The considered environment modeled a large-scale battlefield of approximately 4 × 10 4 km 3 , while the UAV platform was based on the aerodynamic characteristics of the Shahed-136 military drone. The UAV model incorporated realistic flight dynamics, Euler angle-based maneuver control, aerodynamic energy consumption modeling, swarm distance maintenance, and dynamic target bombing objectives. The observation space included inter-drone distances, target zone distances, UAV coordinates, and residual energy information, whereas the reward function jointly optimized swarm cohesion, bombing effectiveness, and energy efficiency. Furthermore, the framework employed DEC-POMDP modeling to capture battlefield uncertainty and limited situational awareness. Experimental results demonstrated that the proposed CommNet-based MADRL framework significantly outperformed independent DQN, independent PPO, and MADDPG baselines in swarm coordination, bombing success rate, and residual energy preservation. In particular, the learned policies maintained stable swarm formations while achieving substantially higher bombing success rates and improved energy conservation.
In [65], the authors proposed an enhanced cooperative roundup framework based on an Improved Cross-Entropy Method with Intrinsic Curiosity-enhanced Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (I2C-MATD3) algorithm for autonomous UAV interception and collaborative encirclement. The considered scenario involved three cooperative UAVs attempting to intercept and surround a maneuvering adversarial UAV in a constrained 2D airspace containing obstacles and protected regions. In this context, the roundup problem as a Markov Decision Process (MDP) and integrated MATD3, an Improved Cross-Entropy Method (ICEM), an Intrinsic Curiosity Module (ICM), CTDE, replay buffer sharing, and global-elite evolutionary optimization. The state space incorporated local kinematics, formation-relative positioning, target azimuth estimation, and LiDAR-based obstacle perception, while the action space consisted of continuous acceleration–control commands. ICEM enhanced exploration through global-elite sample preservation, whereas the ICM generated intrinsic rewards from forward-state prediction errors to encourage exploration of previously unseen states. Performance was evaluated against Multi-Agent Twin Delayed Deep Deterministic Policy Gradient (MATD3), which follows the CTDE paradigm, and Intrinsic Motivation-enhanced Twin Delayed Deep Deterministic Policy Gradient (IMTD3), which employs a Decentralized Training and Decentralized Execution (DTDE) architecture. Simulation results demonstrated that I2C-MATD3 outperformed both baselines in convergence speed, roundup efficiency, robustness, and mission success rate. The framework achieved near-100% roundup success rates after approximately 1200 training episodes while maintaining superior reward stability and faster convergence. Robustness evaluations further demonstrated success rates exceeding 71% under wind disturbances and approximately 70% under severe communication degradation.
Overall, the reviewed studies demonstrate significant progress toward scalable AI-enabled swarm air combat through advanced DRL- and MARL-based frameworks capable of coordinating increasingly large UAV formations in dynamic and partially observable battlefield environments. Recent research has progressively shifted toward scalable architectures incorporating transformer-based representations, attention mechanisms, communication-aware coordination, curriculum learning, and variable-size multi-agent training to improve cooperative decision-making and combat effectiveness. Nevertheless, most existing approaches remain primarily validated in simulation and continue to face challenges related to communication uncertainty, incomplete situational awareness, electronic warfare, and the reliable deployment of large-scale UAV swarms in realistic operational environments.

5. Path Planning and Autonomous Navigation

Path planning and autonomous navigation constitute critical functionalities of AI-enabled UAV-based systems operating in contested and adversarial environments. Modern military UAVs are increasingly expected to perform missions autonomously in the presence of hostile threats, communication disruptions, GPS degradation, moving obstacles, limited situational awareness, and rapidly evolving battlefield conditions. These challenges often limit the effectiveness of conventional rule-based navigation strategies, particularly in unknown or partially observable environments. Consequently, significant research efforts have focused on AI-driven navigation frameworks capable of adaptive decision-making, resilient operation, and autonomous maneuvering. At the same time, the successful execution of AI-generated trajectories depends on robust low-level flight controllers that ensure accurate trajectory tracking despite model uncertainties, wind disturbances, and sensor imperfections. Therefore, robust nonlinear control techniques, including backstepping and sliding mode control can complement AI-based navigation by providing stability, disturbance rejection, and reliable flight execution under challenging operating conditions [66]. This section reviews representative AI-based approaches for path planning and autonomous navigation in defense-oriented UAV-based systems, which are summarized in Table 5.
In [67], the authors proposed a navigation framework integrated with digital twin technology for autonomous mobile target visitation in reconnaissance and surveillance environments. The scenario involved a UAV autonomously searching for and visiting mobile targets such as soldiers or vehicles while operating under uncertain mobility patterns and incomplete situational awareness. The target visitation problem was formulated as an MDP, while PPO-based DRL with Multi-Layer Perceptron (MLP) actor–critic networks enabled adaptive navigation and mobility-aware target tracking. The framework integrated Gazebo and Parrot Sphinx digital twin simulation, Quick Response (QR) code-based target identification, energy consumption estimates, collision-aware navigation, and real-world implementation using a Parrot ANAFI UAV and Raspberry Pi controller. The state space incorporated timestep information, grid cell positions, and target visitation indicators, whereas the action space consisted of nine navigation actions including directional movement and hovering. A reward structure encouraged successful target visitation while minimizing mission duration and energy consumption. Extensive simulations involving different area sizes, UAV altitudes, target densities, and mobility models demonstrated performance close to an ideal benchmark policy with full future target location knowledge while substantially outperforming baseline area-scanning strategies in accumulated reward, mission time, and total energy consumption. Experimental evaluations involving 1–50mobile targets distributed over up to 64 grid cells further demonstrated robust autonomous maneuver generation, adaptive target chasing, and efficient transfer of learned policies from the digital twin to the real UAV platform.
In addition to target-oriented navigation, the work in [68] proposed a DRL-based adaptive collision avoidance framework for UAVs operating in uncertain joint airspace environments involving both manned and unmanned aircraft. Unlike conventional geometry-based collision avoidance techniques, the proposed methodology formulated autonomous navigation as an MDP and employed a Hierarchical PER Double-Dueling DQN (HPER-D3QN). The framework integrated several important mechanisms, including sector-based partial observation, Dynamic Threat Prioritization Assessment (DTPA), uncertainty-aware wind field modeling, dual-layer safety protection zones, and hierarchical prioritized experience replay. Threat evaluation incorporated Time to Closest Point of Approach (TCPA), Distance at Closest Point of Approach (DCPA), and aircraft type information, while replay experiences were dynamically categorized into multiple priority layers to improve learning efficiency and convergence stability. The framework was evaluated in a simulated 30 × 30 km operational airspace supporting up to 25 heterogeneous aircraft. Extensive comparisons against DQN, Double DQN (DDQN), Dueling DQN, Dueling DDQN (D3QN), and PER-D3QN demonstrated that HPER-D3QN achieved the fastest convergence and the highest final reward value of approximately 2.36 after 25,000 training episodes. In dense operational scenarios involving 25 aircraft, the proposed method maintained a collision avoidance success rate of approximately 96.3 % while additionally reducing task completion time and Frequency of Hazardous Proximity (FHP). The results further demonstrated strong robustness against wind disturbances, heading uncertainties, and speed perturbations.
The work in [69] proposed a hybrid DL-enhanced navigation framework integrating neural network-based decision-making with Informed Rapidly-exploring Random Trees (Informed RRT) for UAV navigation in dynamic and adversarial environments. The proposed framework integrated CNNs, bidirectional LSTM networks, attention mechanisms, DL-enhanced Informed RRT, multi-sensor information fusion, adaptive informed sampling, and Bézier trajectory smoothing. CNN modules extracted spatial battlefield features from Geographic Information System (GIS), radar, and camera observations, whereas bidirectional LSTM layers captured temporal dependencies associated with UAV motion, moving adversarial threats, environmental evolution, and weather conditions. An attention mechanism was additionally employed to prioritize critical local and global battlefield information. Simulation experiments conducted in 100 × 100 grid-based battlefield environments demonstrated that the proposed framework significantly outperformed Informed RRT, Batch Informed Trees (BIT), DQN, DDPG, and Layered Recurrent Q-Network (RQN) baselines. In narrow-channel environments, the framework achieved approximately 91 % success rate compared with 59 % for Informed RRT and 32 % for BIT. In trap-obstacle scenarios, the proposed approach maintained approximately 88 % success rate while substantially reducing timeout occurrences and excessive node expansion. Furthermore, average path length was reduced by approximately 7.82 % relative to DDPG and 5.48 % relative to Layered-RQN, while simultaneously generating smoother trajectories with lower turning angles and improved obstacle-avoidance capability.
Recent studies have also investigated AI-driven navigation frameworks specifically designed for GPS-denied environments, where UAVs must maintain reliable autonomous flight despite GPS jamming, spoofing, or signal degradation. Garcia-Gascon et al. [70] proposed a hybrid Long LSTM and MLP-based navigation framework for fixed-wing UAVs operating in GPS-denied environments. The proposed methodology aimed to emulate both GPS functionality and conventional autopilot behavior using only onboard sensor measurements and historical telemetry data. The LSTM network estimated missing GPS coordinates during GPS-denied operation, whereas the MLP generated aerodynamic control commands including elevator, rudder, aileron, and throttle signals. Experimental evaluations involved six autonomous waypoint missions conducted in Valencia, Spain, using a fixed-wing UAV equipped with INertial NAVigation (INAV) autopilot firmware and multiple onboard sensors. Approximately 432,000 telemetry measurements and 49 input features were utilized during training and validation. The proposed framework achieved approximately 99 % prediction accuracy with coefficient-of-determination values approaching R 2 = 0.998 0.999 , while Mean Absolute Percentage Error (MAPE) remained approximately 0.14– 0.16 % . The results demonstrated highly accurate waypoint-following capability and strong resilience against GPS-denied operational conditions.
Similarly, the work in [71] proposed an RL-based decision-support framework for autonomous UAV navigation in GPS-denied and communication-constrained environments. The study addressed challenging operational conditions characterized by RF interference, intentional jamming, degraded situational awareness, and dynamic environmental uncertainty. To enable autonomous navigation under such conditions, the UAV control problem was formulated as a continuous-control MDP, and the performance of the classical REINFORCE algorithm was compared with that of the PPO actor–critic framework. Implementation was carried out in a Unity-based 3D simulation environment integrated with Unity ML-Agents and PyTorch. The observation space comprised a 146-dimensional feature vector containing obstacle-detection data, ground distance measurements, target location information, and UAV flight state variables, whereas the action space generated continuous thrust and attitude control commands. A reward function was designed to jointly promote target acquisition, obstacle avoidance, flight stability, operational area compliance, and energy-efficient maneuvering. To evaluate performance, three operational scenarios were considered, consisting of obstacle-free environments, partially obstructed environments, and dense obstacle configurations. During Stage 1 training, PPO achieved a cumulative reward of approximately 3.8658 , significantly exceeding the 1.5712 obtained by REINFORCE while reducing the average episode length from approximately 200 to 110 steps. Furthermore, PPO exhibited faster and more stable convergence, improved critic-value estimation accuracy, lower computational overhead, and higher mission success rates in complex obstacle-dense environments.
The work in [72] developed a PPO-based autonomous navigation system for Beyond-Visual Line-of-Sight (BVLOS) missions in unconstrained urban and rural environments. In contrast to conventional navigation approaches that rely on computationally demanding LiDAR or vision-based sensing, the proposed framework utilized only four low-cost distance sensors, an IMU, and GPS information combined with RL-based adaptive obstacle-aware navigation. The observation space included UAV position, heading angle, obstacle distance measurements, altitude information, and target coordinates, while the action space controlled altitude adjustments, forward velocity, and yaw angle rotation. Performance evaluation was conducted using Microsoft AirSim integrated with Unreal Engine under both easy area and hard area scenarios characterized by different obstacle densities. Results demonstrated that PPO consistently outperformed A2C and A* benchmarks in terms of target-reaching success rate, collision avoidance capability, and computational efficiency. In the easy area scenario, PPO achieved a success rate of approximately 99 % , while maintaining nearly zero collision probability and eliminating timeout failures. Robust navigation performance was also preserved in highly cluttered environments, highlighting the suitability of the proposed lightweight framework for resource-constrained autonomous missions.
In addition to shortest-path optimization, several studies have explored battlefield-aware navigation strategies emphasizing concealment, survivability, and tactical maneuver generation. The work in [73] proposed a Q-learning-based Strategic Artificial Potential Field (Q-SAPF) framework for autonomous battlefield navigation emphasizing concealment and tactical protection. Unlike conventional Artificial Potential Field (APF)-based navigation approaches that treat obstacles purely as repulsive regions, the proposed methodology strategically utilized terrain features and obstacles as concealment and cover sources. Battlefield environments were modeled using 50 × 50 grid maps containing forests, hills, rivers, roads, destructible buildings, and hazardous terrain regions. The framework integrated APF, SAPF, RL, concealment-aware reward design, and terrain-aware navigation scoring. Simulation experiments conducted across eight battlefield environments demonstrated that Q-SAPF significantly outperformed conventional APF, SAPF, and standard Q-learning methods. Compared with standard Q-learning, navigation path lengths were reduced by approximately 1.96 8.77 units, while learning and execution times decreased by approximately 14.57 69.29 s. Most importantly, mission success rates approached approximately 99.5 99.7 % , substantially outperforming conventional APF-based approaches.
The work in [74] presented an AI-driven tactical path-planning framework for battlefield environments that combines semantic segmentation and DRL to enable terrain-aware and threat-aware autonomous navigation. High-resolution aerial reconnaissance imagery was processed using a modified SwiftFormer transformer-based segmentation model to classify terrain into multiple operational categories, including roads, forests, wetlands, buildings, water bodies, and restricted areas. These terrain classes were subsequently mapped into mobility-aware categories and transformed into cost-weighted navigation maps reflecting terrain accessibility and mission risk. To improve survivability in contested environments, enemy regions were modeled as threat zones, and a DQN agent was employed to adaptively adjust threat avoidance parameters based on terrain characteristics, enemy-zone information, and path planning conditions, while an A* algorithm generated the final route. Experimental results demonstrated that the proposed framework achieved real-time segmentation performance with 93 frames per second (FPS), 93.0% pixel accuracy, and 72.8% mean Intersection over Union (mIoU). Furthermore, the DRL-assisted path planning strategy completely eliminated enemy-zone exposure during navigation, significantly outperforming conventional Dijkstra, Wavefront, Bug, and RRT algorithms in terms of tactical safety. In addition, the DQN-based adaptive decision mechanism achieved more than a 50-fold reduction in execution time compared with exhaustive coefficient-search approaches while maintaining equivalent threat avoidance performance.
The work in [75] investigated the maritime search-and-rescue path planning problem considering multiple rescue centers using UAV-assisted coordinated operations. The framework addressed large-scale maritime emergency scenarios requiring rapid search coverage, efficient workload balancing, and coordinated UAV allocation across geographically distributed rescue centers. The problem was formulated as a graph-based combinatorial optimization task involving UAV task assignments, flight duration constraints, energy limitations, and search region coverage requirements. To solve the resulting NP-hard optimization problem, the authors proposed an RL-enhanced genetic algorithm (GA-RL) integrating Q-learning-based adaptive population management, heuristic population initialization, adaptive crossover operators, mutation strategies, elite repositories, and population perturbation mechanisms. Extensive simulations involving between 100 and 1000 search tasks and between 2 and 10 rescue centers demonstrated that GA-RL consistently achieved superior convergence speed, lower search variability, improved optimization stability, and reduced search completion time compared with memetic algorithms, adaptive large neighborhood search, and heuristic baselines. Additional maritime case studies involving approximately 500 search tasks distributed over an 80-km2 search region further demonstrated improved rescue-center workload balancing and reduced mission duration.
Overall, the reviewed studies demonstrate that AI has substantially advanced autonomous UAV navigation by improving trajectory planning, obstacle avoidance, threat-aware navigation, and operation in dynamic and contested environments. Recent research has progressively integrated DRL, hybrid learning-based path planning, multi-sensor perception, digital twins, and terrain-aware decision-making to enhance navigation robustness and mission effectiveness across diverse operational scenarios. Nevertheless, many existing approaches continue to assume accurate localization and stable communication conditions. Robust navigation under GNSS-denied environments, dynamic adversarial interference, and cooperative multi-UAV missions therefore remains an open research challenge requiring tighter integration between perception, communication, and autonomous decision-making.

6. Target Tracking, Detection, and Classification

Target tracking, detection, and classification are fundamental capabilities of UAV-enabled defense systems. UAVs can be employed to detect, localize, identify, classify, and continuously monitor hostile or suspicious targets in dynamic and adversarial environments. These capabilities support situational awareness, threat assessment, border and maritime surveillance, counter-UAV operations, and precision targeting. To improve detection reliability, tracking accuracy, classification performance, and real-time situational awareness, recent research has increasingly integrated AI-driven techniques into UAV-based defense platforms. This section reviews representative AI-based approaches for target tracking, detection, and classification, as summarized in Table 6.
The work in [76] developed an integrated military reconnaissance framework combining object detection, RL, waypoint navigation, target tracking, and confidence-enhancement mechanisms within a Unity-based battlefield simulator. The considered ISR scenario involved reconnaissance UAVs tasked with detecting, tracking, and classifying hostile tanks concealed by trees and structures under limited onboard computational resources. The UAV platform incorporated cameras, Laser Range Finders (LRFs), Global Navigation Satellite System (GNSS) modules, and AI-based perception systems. To enable lightweight real-time operation, the framework employed YOLOv4-tiny trained on a hybrid dataset comprising 1180 real tank images, 700 Unity-generated synthetic images, and 400 augmented aerial images produced using conventional augmentation and StyleGAN3 techniques. Proportional-Integral-Differential (PID)-based flight control supported waypoint navigation, altitude stabilization, and target tracking, while PPO enabled autonomous UAV repositioning whenever target-detection confidence dropped below 90 % . The RL framework processed both image observations and vector-based state information through a ResNet visual encoder. After 5 million training steps, the image-and-vector configuration achieved an average cumulative reward of approximately 230 and episode length of 987, compared with 137 and 734, respectively, for the vector-only configuration. The results demonstrated that the UAV could autonomously patrol, detect partially occluded targets, track hostile tanks, and dynamically reposition itself to improve target visibility and detection confidence in complex battlefield environments.
In [77], an RL-based autonomous UAV surveillance framework was developed for dynamic target detection under uncertain target-mobility conditions. The considered system consisted of an autonomous UAV, a charging station, and a dynamic target moving among discretized surveillance subareas according to a stochastic Markov chain mobility model. The target-search problem was formulated as an MDP, where the system state incorporated the UAV battery level, UAV location, recharged energy, and target availability indicator. Both Q-learning and State-Action-Reward-State-Action (SARSA) algorithms were investigated to learn efficient search policies and maximize the expected cumulative search benefit. Simulation results obtained for surveillance scenarios consisting of five sub-areas and seven sub-areas demonstrated that the proposed RL-based approaches outperformed conventional random and circular search strategies. Under concentrated target mobility conditions, the RL-based framework achieved target detection rates of approximately 29–32%, compared with about 14–21% for the baseline methods, corresponding to gains of up to approximately 2.2 × . Moreover, the average target detection time was reduced from approximately 7 time units to about 3.3 time units in the scenario with seven sub-areas, while the required energy per successful detection decreased from approximately 36–41 units to about 18 units. The convergence-based RL algorithm consistently achieved the best performance and reduced energy consumption by up to 35 % compared with the adaptive ϵ -greedy approach, while the RL-based schemes provided up to 125 % and 118 % improvements in energy efficiency relative to the random and circular search strategies, respectively.
The work in [78] investigated secure UAV detection using adversarial ML integrated with AI-driven drone-detection frameworks. The study addressed the vulnerability of DL-based UAV detectors to adversarial attacks capable of inducing false positives, false negatives, and detection failures through imperceptible perturbations. The proposed framework combined CNNs, adaptive image segmentation, Q-learning-based trajectory optimization, and game-theoretic resource allocation strategies. Experimental evaluations were conducted using the VisDrone dataset under urban and rural operational conditions. The CNN-based detector employed convolutional, pooling, normalization, and fully connected layers for hierarchical feature extraction and UAV classification, while adaptive image segmentation dynamically adjusted thresholds according to local image characteristics. Experimental results revealed strong environmental sensitivity, achieving approximately 95 % detection accuracy in rural daytime scenarios but only 87 % under urban nighttime conditions. Precision–recall analysis showed that a recall of approximately 90 % corresponded to a precision of 55 % , whereas increasing precision to 95 % reduced recall to nearly 70 % . To improve robustness, the authors incorporated adversarial training, defensive distillation, and feature space regularization. These defenses increased precision and accuracy from approximately 80 % and 78 % in the baseline model to nearly 90 % in the adversarially enhanced framework. The robustness–complexity trade-off was also highlighted, as dense urban environments increased processing latency to approximately 0.5 0.6 s and memory consumption to nearly 300 MB.
The work in [79] studied AI-driven UAV detection and classification using millimeter-wave Frequency-Modulated Continuous Wave (FMCW) Multiple-Input Multiple-Output (MIMO) radar combined with full-wave electromagnetic simulations. Specifically, this work examined the impact of radar architecture and antenna field of view on classification accuracy under varying observation angles. Synthetic radar signatures and range-Doppler maps for seven UAV types, including quadcopters, fixed-wing UAVs, helicopters, and hexacopters, were generated using the Ansys high-frequency structure simulator (HFSS) [80]. The radar operated at 79 GHz with a bandwidth of 3.46 GHz, providing a range resolution of approximately 4.33 cm. Both Single-Input Single-Output (SISO) and 2Tx/4Rx MIMO configurations were evaluated, while UAV classification was performed using a DopplerNet-based CNN trained on synthetic and real radar measurements. Experimental results demonstrated classification accuracies of approximately 96.5 % using synthetic datasets and above 98 % using real measurements at 0° observation angle. However, performance degraded significantly at larger angles. At 80°, SISO accuracy dropped to approximately 13 % , whereas the MIMO radar maintained approximately 60 % accuracy in simulations and 38 % in real measurements. Receiver Operating Characteristic (ROC) analysis showed that the MIMO configuration achieved a detection probability near 0.9 at approximately 17.5 dB Signal-to-Noise Ratio (SNR) and false-alarm probability of 10 5 , while the SISO radar exhibited nearly zero detection probability under the same conditions.
An autonomous target-tracking and control framework for quadrotors based solely on onboard monocular cameras and lightweight embedded AI techniques was presented in [81]. The proposed framework integrated SSD-MobileNetV2 object detection, Kernelized Correlation Filter (KCF)-based tracking [82], Kalman Filter prediction, and Re-Identification (Re-ID) mechanisms to achieve robust real-time operation on platforms such as Raspberry Pi 5. SSD-MobileNetV2 enabled lightweight object detection and feature extraction, while MobileNetV2 generated Re-ID feature embeddings. KCF enabled efficient frame-by-frame tracking, whereas Kalman prediction improved robustness during target occlusion and temporary tracking loss. The control framework additionally employed Image-Based Visual Servoing (IBVS) and hierarchical Proportional-Integral-Derivative (PID) controllers to generate UAV motion commands directly from image-plane tracking errors. Experimental evaluation using the OTB50 dataset compared the proposed framework against MOSSE [83], KCF [82], CSRT (Discriminative Correlation Filter with Channel and Spatial Reliability) [84], DiMP50 [85], TransT [86], and OSTrack [87] trackers. Although MOSSE achieved the highest processing speed of 42-74 FPS, its IoU (Intersection over Union) performance remained below 0.1 . In contrast, the proposed framework achieved IoU values between 0.669 and 0.883 while maintaining processing speeds between 11 and 30 FPS, achieving a favorable balance between computational efficiency and tracking accuracy. The integrated Re-ID and Kalman prediction modules substantially improved robustness against target occlusion, rapid target motion, and appearance changes. Furthermore, Software-in-the-Loop (SIL) and Hardware-in-the-Loop (HIL) experiments demonstrated stable autonomous target following and distance control under GNSS-denied conditions.
Beyond real-time tracking, predictive localization has emerged as an effective approach for maintaining target tracking during communication failures. The work in [88] proposed a LSTM Autoencoder (LSTM-AE)-based UAV position prediction framework for scenarios involving communication loss between a Ground Data Terminal (GDT) and a UAV. Unlike conventional 6-DOF trajectory models requiring aerodynamic parameters that can be difficult to estimate, the proposed data-driven framework learned long-term motion patterns directly from flight telemetry data. The system integrated data preprocessing, LSTM-AE-based trajectory prediction, and radar look-angle estimation for autonomous GDT antenna steering. Approximately 36,000 trajectory samples collected from 13 real flight trials were used, with telemetry recorded every 10 s during flights lasting nearly two hours. The encoder employed LSTM layers with 128 and 64 neurons to extract latent trajectory representations, while the decoder reconstructed future UAV states, including latitude, longitude, altitude, pressure, and temperature. Predicted positions were subsequently transformed into Earth-Centered Earth-Fixed (ECEF) and North-East-Down (NED) coordinates to compute azimuth and elevation steering angles for radar tracking. Experimental evaluation using nested five-fold cross-validation demonstrated that LSTM-AE significantly outperformed Vanilla LSTM [89] and Multi-Output Regression (MOR) baselines [90]. In particular, the proposed framework achieved approximately 99 % prediction accuracy after 225 training epochs, compared with 96 % for Vanilla LSTM, while reducing the loss value to approximately 1.1 × 10 4 , compared with 2.0 × 10 2 . Quantitative evaluation further yielded Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Square Error (RMSE) values of approximately 0.079 , 0.012 , and 0.110 , respectively.
An AI-driven UAV-to-UAV Small Target Detection (UAV-STD) framework for detecting and localizing small aerial UAV targets was proposed in [91]. The UAV-STD architecture combined several convolutional modules, including the Convolution–Batch Normalization–SiLU (CBS) module, Cross-Stage Partial (CSP) module, and Spatial Pyramid Pooling Fast (SPPF) module. The UAV-STD model employed a Feature Pyramid Network (FPN) for multi-scale feature fusion and integrated an Attention Mechanism-Based Small Target Detection (AMSTD) module to enhance small-UAV feature extraction, retention, and shallow texture utilization. Specifically, the AMSTD module incorporated multilayer residual feature fusion and coordinate attention mechanisms to improve focus on small UAV targets, while a Spatial-aware and Scale-aware Prediction (SSP) Head utilized deformable convolution and scale-aware attention to enhance localization accuracy and scale perception. A normalized Wasserstein distance and Complete Intersection over Union (NWD-CIoU) loss function was also introduced to improve small-target localization performance. The framework was evaluated using a UAV-to-UAV dataset containing 13,270 high-resolution images collected across urban, mountain, sky, and field environments, where most UAV targets were smaller than 32 × 32 pixels. Experimental comparisons against SSD [92], Faster R-CNN [93], YOLOv5s, YOLOX, YOLOv6, YOLOv7, YOLOv8, RT-DETR [94], YOLOv9, and YOLOv10 demonstrated superior small-target detection performance. UAV-STD achieved AP50 and APs values of 83.7 % and 83.1 % , improving detection performance by 8.8 % and 9.3 % , respectively, compared with YOLOv5s. For extremely small targets below 10 × 10 pixels, the framework achieved AP50 and AP50:95 values of 70.4 % and 30.1 % , outperforming YOLOv5s by 11.9 % and 5.2 % , respectively. Moreover, UAV-STD maintained a lightweight architecture with only 8 million parameters and 18.9 Giga Floating-Point Operations Per Second (GFLOPs), substantially lower than YOLOv9c, which required 25.3 million parameters and 102.1 GFLOPs.
Table 6. Overview of AI-driven approaches for target tracking, detection, and classification in UAV-based defense systems.
Table 6. Overview of AI-driven approaches for target tracking, detection, and classification in UAV-based defense systems.
ReferenceObjectiveAI/ML ModelsOperational ScenarioUAV ConfigurationMain ComponentsKey Results
Lee et al., 2024 [76]Autonomous military reconnaissance and concealed target detectionYOLOv4-tiny, PPO, ResNet encoderBattlefield ISR and concealed tank reconnaissanceReconnaissance UAV with camera, LRF, and GNSSUnity-based simulator, StyleGAN3 data augmentation, PID flight control, RL-guided viewpoint optimization, confidence enhancementAfter 5M training steps, image + vector observations achieved reward 230 and episode length 987, compared with 137 and 734 for vector-only observations
Salameh et al., 2024 [77]Autonomous target detection under uncertain mobilityQ-learning, SARSADynamic surveillance environmentsSingle UAV with charging stationMDP formulation, Markov-chain target mobility, energy-aware navigation, adaptive surveillance optimizationDetection rates of 29– 32 % vs. 14– 21 % for baseline methods; detection time reduced from 7 to 3.3 units and energy consumption reduced by up to 35 %
Ud Din et al., 2025 [78]Robust UAV detection under adversarial attacksCNNs, Q-learning, adversarial MLUrban and rural UAV surveillanceAI-enabled UAV detection systemsVisDrone dataset, adaptive image segmentation, adversarial training, defensive distillation, feature-space regularizationDetection accuracy reached 95 % in rural daytime environments and 87 % in urban nighttime scenarios; adversarial defenses improved accuracy and precision to nearly  90 %
Sayed et al., 2024 [79]Radar-based UAV detection and classificationDopplerNet-based CNNCounter-UAV radar surveillanceSeven UAV types using SISO and MIMO radar sensing79 GHz FMCW radar, HFSS electromagnetic simulations, range-Doppler processing, 2Tx/4Rx MIMO architectureClassification accuracy of 96.5 % (synthetic) and >98% (real data); detection probability near 0.9 at −17.5 dB SNR and false-alarm probability 10 5
Han et al., 2025 [81]GNSS-denied autonomous visual target trackingSSD-MobileNetV2, KCF, Kalman Filter, Re-IDReconnaissance and autonomous tracking missionsQuadrotor UAV with Raspberry Pi 5IBVS, hierarchical PID control, Kalman prediction, Re-ID, SIL/HIL validationAchieved IoU values of 0.669 0.883 and processing speeds of 11–30 FPS while maintaining robust target tracking under GNSS-denied conditions
Kant et al., 2023 [88]Resilient UAV trajectory prediction during communication failuresLSTM-AECommunication-loss tracking and localizationFixed-wing UAV and GDT infrastructureTrajectory preprocessing, LSTM autoencoder, ECEF/NED transformations, radar look-angle estimationApproximately 99 % prediction accuracy versus 96 % for Vanilla LSTM; loss reduced to 0.0001121 with MAE 0.07862 and RMSE  0.110409
Zuo et al., 2025 [91]UAV-to-UAV small-target detection for counter-UAV operationsUAV-STD, attention-based DLDynamic air-to-air environmentsUAV surveillance and counter-UAV systemsAMSTD module, SSP Head, FPN feature fusion, NWD-CIoU loss, Coordinate AttentionAP50 = 83.7 % and APs = 83.1 % ; AP50 = 70.4 % for targets below 10 × 10 pixels while using only 8M parameters and 18.9 GFLOPs
Kim et al., 2025 [95]Detection, interception, and inspection of balloon-borne aerial threatsYOLOv8, CNNsHazardous-balloon interception and aerial defenseCooperative four-drone swarmYOLOv8 detection, swarm-net interception, CNN-based X-ray inspection, decentralized coordination, HITL supervision 91.3 % mAP@0.5, 88.9 % recall, 4.6 % false-positive rate, 92.7 % overall detection accuracy, 95.4 % capture success rate, and response time <7.8 s
To detect, track, intercept, and inspect balloon-borne aerial threats carrying hazardous or explosive payloads, the work in [95] proposed an AI-enabled autonomous drone defense architecture. This architecture integrated YOLOv8-based aerial object detection, cooperative swarm-drone interception, X-ray-based payload inspection, and real-time situational awareness visualization within a unified human-in-the-loop (HITL) defense system. The YOLOv8 detector was trained using approximately 10,000 images containing hazardous balloons and non-threat airborne objects captured under diverse weather and illumination conditions. To improve robustness, the dataset incorporated augmentation techniques including Gaussian noise, motion blur, brightness adjustment, and rotational transformations. Experimental evaluation demonstrated strong real-time detection performance with approximately 91.3 % mean Average Precision at IoU threshold 0.5 (mAP@0.5), 88.9 % recall, and false positive rate near 4.6 % . Under clear weather conditions, the framework achieved approximately 93.1 % precision and 92.4 % mAP@0.5, whereas rainy and nighttime environments slightly reduced recall toward 85.7–85.8%. Following target detection, the system employed a cooperative four-drone swarm carrying a butyl rubber-coated capture net for safe aerial interception and retrieval of hazardous balloons. The decentralized swarm coordination mechanism enabled adaptive trajectory planning, synchronized interception, and collision avoidance while maintaining low communication overhead through 10 Hz inter-drone synchronization. Additionally, the framework integrated a lightweight CNN-based X-ray inspection module for hazardous material classification together with a Flask-based situational awareness interface for real-time mission visualization and operator supervision. Extensive simulations demonstrated significant improvements over conventional manual response systems, achieving approximately 92.7 % overall detection accuracy, 95.4 % capture success rate, and average response time below 7.8 s.
Overall, the reviewed studies demonstrate that AI has substantially enhanced UAV-based target tracking, detection, and classification by improving detection accuracy, tracking robustness, situational awareness, and real-time battlefield perception across diverse defense applications. Recent research has progressively evolved from conventional vision-based approaches towards lightweight DL architectures, multimodal perception, radar-assisted sensing, adversarially robust learning, autonomous target tracking, and cooperative AI-enabled detection frameworks capable of operating in increasingly complex environments. Nevertheless, improving robustness under cluttered environments, adverse weather conditions, camouflage, small target detection, and electronic warfare remains an important challenge, motivating the continued development of more resilient and adaptive battlefield perception systems.

7. Cybersecurity, Electronic Warfare Protection, and Resilient UAV Operation

Cybersecurity, electronic warfare protection, and resiliency are all critical requirements for AI-enabled UAV-based defense systems operating in contested environments. UAVs employed in ISR missions, autonomous combat operations, and distributed multi-UAV networks are increasingly exposed to cyber–physical threats targeting communication links, navigation systems, onboard sensors, AI models, and control architectures as well as to electronic warfare threats such as GPS spoofing, GPS jamming, communication jamming, spectrum interference, and intelligent electromagnetic attacks. Such threats can degrade communication reliability, manipulate autonomous decision-making, disrupt mission execution, compromise situational awareness, and impair overall system performance. This section reviews recent AI-driven approaches for security surveillance, cybersecurity, electronic warfare protection, secure communications, and resilient UAV system operation, as summarized in Table 7.
An RL-based UAV surveillance framework for intrusion detection under stochastic target mobility conditions was proposed in [96]. The target-search mission was formulated as an MDP incorporating practical constraints such as finite battery capacity, energy causality, maximum flight distance, and battery recharging. SARSA and Q-learning were employed together with adaptive exploration strategies to optimize UAV movement, target-search efficiency, and energy-aware operation. Target mobility was modeled using a stochastic Markov chain process, enabling the UAV to adapt its surveillance actions according to learned mobility patterns and residual battery levels. Experimental results demonstrated that the RL-based framework significantly outperformed random and circular search strategies in terms of detection rate, detection time, energy efficiency, and cumulative return. Under concentrated target mobility conditions, the RL approaches achieved detection rates of approximately 29–32%, compared with 14–21% for the baseline methods, corresponding to gains of up to 2.2 × . In the scenario with seven sub-areas, the average detection time was reduced from approximately 7 to 3.3 time units, while the energy required per successful detection decreased from 36–41 units to about 18 units. The convergence-based RL algorithm consistently achieved the best performance, reducing energy consumption by up to 35 % compared with adaptive ϵ -greedy exploration [97]. Furthermore, the RL-based schemes improved energy efficiency by up to 125 % and 118 % relative to the random and circular search strategies, respectively, while achieving the highest cumulative discounted rewards across all target mobility patterns.
The work in [98] investigated resilient UAV navigation spoofing and covert cyber–physical attack strategies against UAVs employing integrated GPS and Inertial Navigation System (INS)-based navigation. The study considered a defense-oriented counter-UAV scenario in which a spoofing device attempted to manipulate the navigation and flight-control process of a noncooperative UAV without prior knowledge of its internal Kalman Filter (KF) parameters or predefined reference trajectory. The proposed framework combined Spatial Information Entropy (SIE) and Maximum Entropy RL (MERL) [99] to improve the concealment, stability, and effectiveness of deceptive navigation attacks. Specifically, the authors developed an SIE-enhanced Soft Actor–Critic (SIE-SAC) algorithm capable of learning covert spoofing trajectories by jointly optimizing deception/position concealment, trajectory smoothness, and successful UAV guidance toward a deceptive destination. The navigation spoofing process was modeled as an MDP, where the state space incorporated relative distances, azimuth and elevation angles, UAV velocities, and Normalized Innovation Square (NIS) [100] residuals derived from the GPS/INS-integrated KF navigation framework. Experimental results demonstrated that SIE-SAC substantially outperformed conventional SAC and existing navigation-spoofing methods, achieving return convergence values approaching 500, whereas conventional SAC converged near 300. Moreover, SIE-SAC generated smoother deceptive trajectories while maintaining NIS values below the chi-square detection thresholds corresponding to 99 % and 95 % confidence levels.
The work in [101] proposed a Decision Transformer (DT)-driven offline DRL framework for intelligent anti-jamming communications. The considered architecture comprised a centralized base station, multiple UAV communication and relay nodes, and AI-enabled adaptive jammers capable of spectrum sensing, waveform generation, power adaptation, and frequency-band targeting. Moreover, The anti-jamming problem was formulated as an MDP, where UAVs dynamically optimized communication modes, coding schemes, transmission power, and anti-jamming strategies based on environmental observations and base-station feedback. Unlike conventional online RL and DRL approaches that require extensive exploration and interaction with the environment, the proposed framework leveraged Transformer-based sequence modeling to learn robust anti-jamming policies directly from historical state–action–reward trajectories, enabling stable decision making under previously unseen jamming conditions while significantly reducing training complexity and convergence instability. A three-stage operational procedure was developed, including communication quality monitoring, jammer identification through waveform analysis, and dynamic transmission reconfiguration based on DT-assisted decision making. The proposed approach was evaluated under both Additive White Gaussian Noise (AWGN) and Rayleigh fading channels while considering ten distinct jamming types, including single-tone, multi-tone, comb-spectrum, noise-modulated, partial-band, wideband, narrow-band, periodic pulse, Binary Phase Shift Keying (BPSK), and Quadrature Phase Shift Keying (QPSK) jammers. Extensive simulations conducted over a 2 km × 2 km battlefield environment comprising a ground station, six UAV relay nodes, ten ground users, and AI-enabled adaptive jammers demonstrated the effectiveness of the proposed framework. The DT-based anti-jamming system achieved a BER of approximately (0.0025), an SNR of 25.5 dB, a Packet Delivery Ratio (PDR) of 97.1%, a jamming detection time of only 0.69 s, and an average communication latency of approximately 45 ms.
In [102], an explainable DRL (XDRL)-based framework for adversarial attack detection in autonomous UAV guidance systems was proposed. The framework combined DDPG, PER, APF, and SHapley Additive exPlanations (SHAP) [103] to improve both UAV navigation performance and adversarial robustness. The proposed UAV guidance architecture employed LiDAR-based depth sensing and generated continuous UAV control commands through an actor–critic DRL structure. To improve obstacle avoidance and training stability, APF-based attractive and repulsive force fields were integrated into the DDPG controller. The authors further introduced explainability-driven adversarial detection mechanisms using SHAP-value analysis to identify abnormal feature activations caused by adversarial perturbations. Fast Gradient Sign Method (FGSM) [104] and Basic Iterative Method (BIM) [105] attacks were evaluated within the AirSim simulation environment. The experimental results demonstrated that adversarial attacks significantly degraded UAV navigation performance, reducing obstacle course completion rates from approximately 97 % to nearly 35 % under sustained BIM attacks. To counter these threats, the authors developed CNN-based and LSTM-based adversarial detectors exploiting temporal and spatial variations in SHAP-value distributions. The CNN detector achieved approximately 80 % adversarial detection accuracy, while the LSTM-based detector achieved approximately 91 % accuracy with lower computational complexity and faster inference, demonstrating the effectiveness of explainability-driven UAV cyber defense.
Beyond adversarial perturbations, recent research has also focused on Trojan and backdoor attacks targeting autonomous UAV navigation systems. The work in [106] investigated Trojan attacks against the DroNet [107] autonomous UAV navigation architecture based on a ResNet-8 backbone. The study demonstrated that carefully designed trigger patterns inserted into monocular grayscale navigation images could manipulate UAV steering angle and collision prediction outputs while preserving high accuracy on benign data. Several trigger types were investigated, including square, X-shaped, left-turn, right-turn, and center-alignment triggers. The Trojan attacks forced UAVs to execute predefined steering maneuvers such as 45° left or right turns whenever trigger patterns appeared within the image. To defend against these attacks, the authors proposed a Trojan detection framework combining DeepFool-based Universal Adversarial Perturbations (UAPs) [108], MobileNetV3-based feature extraction [109], and a fully connected Trojan classifier. The experimental results demonstrated that the proposed detector substantially improved UAV security by reducing Trojan attack success rates from approximately 73.69 % to only 2.76 % , highlighting the importance of robust Trojan defense mechanisms for trustworthy autonomous UAV operation.
The work in [110] proposed UAV-CIDS, a real-time distributed Collaborative Intrusion Detection System (CIDS) for UAV communication networks that integrates deep learning, event-correlation mechanisms, and incident-response capabilities. The framework utilized the UAVIDS dataset [111], which contains encrypted WiFi traffic collected from three popular UAV platforms, namely, DJI Spark, DBPower UDI, and Parrot Bebop. To improve the detection of cyber attacks, including previously unseen threats, the authors developed a Feedforward CNN (FFCNN) architecture employing a hybrid Rectified Linear Unit (ReLU) and Tanh Linear Unit (TaLU) activation mechanism. The framework further incorporated Information Gain (IG)-based feature selection, distributed event validation, collaborative attack correlation, and trust-aware UAV node management. Unlike conventional standalone IDS solutions, UAV-CIDS enabled each UAV to operate simultaneously as a monitoring and analysis node, thereby eliminating single points of failure, improving network-wide situational awareness, and facilitating the identification of compromised UAVs. Furthermore, the proposed system integrated a real-time cyber incident response mechanism to mitigate attacks and prevent their propagation throughout the network. Experimental results demonstrated detection accuracy of up to 98.23 % together with high detection precision and sensitivity and low false alarm and detection error rates, highlighting the effectiveness of the proposed collaborative framework for securing UAV networks.
The work in [112] proposed a predictive queuing analysis framework to enhance the cybersecurity of Software-Defined UAV (SD-UAV) relay networks operating in dynamic urban battlefield environments. The considered architecture consisted of UAV relay nodes, a master software-defined networking (SDN) controller UAV, and a central station responsible for cybersecurity monitoring and ML-based analysis. The SDN controller maintained a global network view and dynamically managed routing through OpenFlow-based control mechanisms, while a non-preemptive dual-priority queuing scheme ensured efficient delivery of command-and-control and security-related traffic. To address the challenge of detecting previously unseen zero-day cyberattacks without requiring large labeled attack datasets, the framework combined SDN-based monitoring, predictive queuing analysis based on a modified open Jackson network model, and ML-based intrusion detection. The proposed framework was implemented and evaluated using Mininet-WiFi and the Open Network Operating System (ONOS) controller [113] in a simulated urban battlefield environment covering approximately one square mile. The simulation scenario consisted of 40 UAV relay nodes providing communication coverage for 80 mobile ground devices, with individual UAV communication ranges of approximately 120 m. Simulation results demonstrated that the SD-UAV architecture maintained dynamic communication connectivity despite frequent topology changes and mobility-induced disruptions. The generated theoretical estimates for inter-arrival time, transmission delay, packet count closely matched the measured simulation values across all evaluated UAV classes.
The work in [114] proposed Block-USB, a blockchain- and ML-based secure data dissemination framework for UAV-assisted battlefield operations over future 6G networks. The proposed architecture integrated blockchain security, InterPlanetary File System (IPFS)-based decentralized storage [115], ML-assisted intrusion detection, and ultra-low-latency sixth-generation (6G) communications to provide secure, scalable, and resilient UAV-enabled battlefield networking. The framework comprised five functional layers, namely data acquisition, analytics, communication, blockchain, and application layers, supporting secure reconnaissance-data collection, authentication, storage, and dissemination. Multiple ML classifiers were evaluated, including Random Forest (RF), Linear Discriminant Analysis (LDA), Gaussian, Ridge, Quadratic Discriminant Analysis (QDA), and Naive Bayes (NB). The RF classifier achieved the highest performance with classification accuracy approaching 100 % , substantially outperforming LDA (≈66%), QDA (≈33%), Ridge (≈28%), Gaussian (≈13%), and NB (≈12%). The proposed architecture further leveraged key 6G capabilities, including data rates up to 1 Tb / s , end-to-end latency below 1 ms , reliability of 99.99999 %, and connectivity densities approaching 10 7 devices per square kilometer. Performance evaluations demonstrated that the IPFS-enabled framework consistently achieved higher scalability than the non-IPFS blockchain implementation, reaching approximately 63 scalability units at 150 transactions compared with approximately 45 for the conventional blockchain architecture. The proposed framework reduced data processing complexity through AI-assisted filtering of malicious UAV traffic. For example, at the eighth message request, the processing delay was reduced from approximately 64.7 ms in the benchmark blockchain-only scheme to approximately 55.2 ms , corresponding to a reduction of nearly 15 % . At transaction loads approaching 10 6 , the proposed 6G-enabled architecture maintained a communication latency of approximately 37 ms , compared with approximately 55 ms for fifth generation (5G) and approximately 85 ms for fourth generation (4G) systems.
In [116], the authors proposed Flighter, a decentralized FL (DFL) framework for secure military aerial operations. Unlike conventional FL architectures relying on centralized aggregation servers, Flighter employed fully decentralized model aggregation, eliminating single points of failure while improving robustness against communication disruptions and adversarial attacks. The framework integrated asynchronous communications, mobility-aware networking, adaptive differential privacy, and a situational awareness defense mechanism that continuously assessed model similarity, flight formation consistency, geopositioning integrity, communication reliability, and onboard computational resources. The considered reconnaissance scenario involved four military aircraft equipped with Synthetic Aperture Radar (SAR) sensors collaboratively detecting and classifying ground and maritime targets. Target recognition was performed using a VGG16 CNN consisting of sixteen convolutional and fully connected layers trained and evaluated on the MSTAR, SAMPLE, and OpenSARShip SAR datasets [117]. Extensive battlefield simulations evaluated the framework under geopositioning deviation, collision-course manipulation, and adversarial data poisoning attacks. Under normal operating conditions, Flighter achieved F1-scores of approximately 95.8 % on MSTAR and 97.5 % on SAMPLE while maintaining low packet-loss rates and moderate resource consumption. During geopositioning deviation attacks, it preserved F1-scores above 90 % , whereas the unsecured decentralized learning baseline degraded to nearly 70 % . Similarly, under collision-course manipulation attacks, Flighter maintained F1-scores above 89 % , while the baseline dropped below 43 % .
To enable scalable and resilient autonomous UAV-swarm operation in mission-critical environments, the work in [118] investigated the integration of agentic AI, LLMs, and edge computing. Within the swarm, each UAV acted as an autonomous agent equipped with perception, reasoning, memory, coordination, and action-execution modules, enabling decentralized decision making and peer-to-peer collaboration. The framework employed YOLO-based object detection and multimodal sensing for situational awareness and TinyLLaMA-1.1B models running on Jetson Orin-class onboard processors for local path planning and reasoning. The proposed architecture was evaluated through a wildfire search-and-rescue scenario based on the Eaton wildfire, where satellite imagery was processed using a U-Net segmentation model to dynamically identify wildfire boundaries and generate survey regions. The simulated swarm consisted of Skydio X10-class UAVs equipped with Red–Green–Blue (RGB) and thermal sensors offering detection ranges up to 1500 m , operating at a cruise speed of 15 m / s . Approximately 300 dynamically generated survey points were distributed among UAVs through GPT-4.1-assisted mission planning, while each UAV independently optimized its flight path using TinyLLaMA-based reasoning. Experimental results demonstrated that the proposed edge-enabled architecture consistently maintained coverage rates close to 98–100% across dynamically changing wildfire conditions, significantly outperforming the centralized LLM baseline, whose coverage rate deteriorated to approximately 73–76% as the wildfire expanded. Furthermore, the proposed framework achieved substantially lower mission completion times than a greedy assignment strategy. With eight UAVs, missions were completed in less than 25 min compared with more than 30 min for the greedy baseline, while increasing the swarm size to twelve UAVs reduced mission completion time below 17 min compared with over 22 min for the benchmark approach.
Table 7. Overview of AI-driven approaches for cybersecurity, electronic-warfare protection, and resilient UAV operation.
Table 7. Overview of AI-driven approaches for cybersecurity, electronic-warfare protection, and resilient UAV operation.
ReferenceObjectiveAI/ML ModelsOperational ScenarioUAV ConfigurationMain ComponentsKey Results
Masadeh et al., 2024 [96]RL-based surveillance and intrusion detectionQ-learning, SARSAStochastic target-mobility environmentsSingle energy-constrained UAV with charging stationMDP formulation, Markov-chain mobility, adaptive exploration, energy-aware navigationDetection rates of 29–32% vs. 14–21%; detection time reduced from 7 to 3.3 units and energy per detection from 36–41 to 18 units
Ma et al., 2025 [98]Covert UAV navigation spoofing and deceptive trajectory manipulationSIE-SAC, SAC, MERLCounter-UAV GPS/INS spoofingSingle UAV with GPS/INS-integrated navigationSIE, NIS-based stealth constraints, GPS/INS-integrated KF, MDP-based spoofing optimizationReturn convergence approaching 500 vs. 300 for SAC; smoother deceptive trajectories while maintaining NIS below χ 2 detection thresholds
Zhao et al., 2026 [101]AI-driven anti-jamming communicationDT, DRLElectronic warfare and adaptive jammingMultiple UAV relays and base stationOffline trajectory learning, cloud-edge collaboration, adaptive reconfiguration, jammer detectionBER = 0.0025 , SNR = 25.5 dB, PDR = 97.1%, detection time = 0.69 s, latency = 45 ms
Hickling et al., 2023 [102]Adversarial attack detection for DRL-based navigationDDPG, PER, SHAP, CNN-AD, LSTM-ADAutonomous navigation under FGSM/BIM attacksLiDAR-equipped UAVAPF guidance, XAI, SHAP-value analysis, adversarial detectorsNavigation success reduced from 97 % to 35 % under BIM attacks; CNN and LSTM detectors achieved approximately 80 % and 91 % accuracy
Mynuddin et al., 2024 [106]Trojan attack detection and trustworthy UAV navigationMobileNetV3, DeepFool-UAPBackdoor attacks against autonomous navigationDroNet (ResNet-8) UAV navigation systemTrigger-pattern injection, UAP generation, feature extraction, Trojan classifierReduced Trojan attack success rate from 73.69 % to 2.76 %
Hadi et al., 2024 [110]Collaborative IDS for UAV communication networksFFCNNCyberattack-prone UAV networksMultiple UAVs and distributed monitoring nodesUAVIDS dataset, ReLU-TaLU activation, distributed event correlation, trust-aware management, incident responseDetection accuracy of 98.23 % with high precision and sensitivity and low false alarm rates
Agnew et al., 2024 [112]Cybersecurity enhancement of SD-UAV relay networksPredictive Queuing AnalysisUrban battlefield communication environments40 SD-UAV relay nodesSDN, Jackson Open Networks, predictive queuing analysis, OpenFlow routing, ONOS controllerTheoretical estimates of inter-arrival time, transmission delay, and packet count closely matched simulation results, validating the predictive model
Jadav et al., 2023 [114]Secure UAV communication over 6G networksRF, blockchainMilitary reconnaissance and battlefield communicationsUAV swarms with U2U and satellite linksBlock-USB, IPFS storage, 6G communications, ML-based intrusion detectionRF achieved near- 100 % classification accuracy; processing delay reduced from 64.7 ms to 55.2 ms and scalability improved from 45 to 63 units
Martínez Beltrán et al., 2025 [116]Resilient decentralized FL for military reconnaissanceDFL, VGG16Adversarial SAR reconnaissance and mosaic warfareFour military aircraft with SAR sensingFlighter, adaptive differential privacy, asynchronous communications, situational awareness defense, decentralized aggregationF1-scores of 95.8 % (MSTAR) and 97.5 % (SAMPLE); maintained >89–90% F1 under geopositioning and trajectory manipulation attacks
Nguyen et al., 2026 [118]Resilient autonomous UAV swarm coordinationAgentic AI, LLMs, TinyLLaMA, GPT-4.1Mission-critical edge-enabled swarm environmentsStandalone, edge-enabled, and hybrid UAV swarmsYOLO detection, U-Net segmentation, mesh networking, multimodal sensing, distributed reasoning, edge-assisted coordinationCoverage rates of 98–100%; 8-UAV missions completed in <25 min and 12-UAV missions in <17 min
Overall, the reviewed studies demonstrate that AI has become a key enabler of cybersecurity, electronic warfare protection, and resilient UAV operation by enhancing threat detection, secure communications, anti-jamming capability, intrusion detection, and autonomous mission resilience in contested environments. Recent research has progressively evolved from protecting individual UAV platforms toward securing distributed UAV networks through adversarially robust learning, XAI, collaborative intrusion detection, blockchain- and FL-based security, SDN, and edge-enabled autonomous coordination. Nevertheless, ensuring robust operation under sophisticated cyber–physical attacks, intelligent electronic warfare, communication disruptions, adversarial AI, and large-scale distributed deployments remains a major challenge for future autonomous UAV-based defense systems.

8. Lessons Learned and Research Insights

This section distills the major findings, recurring technological trends, operational observations, research insights, and key limitations identified across the four principal UAV-based defense application domains reviewed in this paper. In addition to these domain-specific observations, the reviewed literature reveals several cross-cutting lessons concerning evaluation and validation practices, deployment readiness, robustness, trustworthiness, and the future evolution of AI-enabled UAV-based defense systems. The key lessons learned are summarized as follows:
  • Autonomous Air Combat and Cooperative UAV Operations
    Prominent Role of DRL and MARL: The reviewed studies indicate that DRL and MARL have emerged as among the most widely investigated AI paradigms for autonomous air combat, maneuver decision-making, cooperative engagement, and swarm-level tactical coordination. PPO, DQN, DDPG, MATD3, MAPPO, QMIX, hierarchical RL, and transformer-assisted MARL frameworks were widely employed to enable UAVs to learn adaptive maneuvering, target engagement, pursuit, evasion, and cooperative attack strategies in adversarial environments.
    Emergent Cooperative Combat Behaviours: Several MARL-based studies demonstrated that cooperative UAV teams can autonomously develop advanced tactical behaviors such as pincer attacks, coordinated flanking, decoy maneuvers, cross-attacks, target allocation, cooperative encirclement, fire-attraction tactics, synchronized target engagement, and collaborative target elimination. These behaviors emerged through reward-driven learning rather than explicitly programmed combat rules.
    Increasing Importance of Opponent Modeling: Recent studies increasingly consider opponent capabilities, tactical behaviors, and threat levels as part of the decision-making process. Explicit modeling of adversarial agents enables more informed maneuver planning, adaptive target prioritization, and improved tactical coordination in dynamic air-combat scenarios.
    Scalability of Swarm Combat Intelligence: Large-scale swarm combat studies increasingly employed attention-assisted, transformer-based, communication-aware, and value-decomposition architectures to improve coordination across formations ranging from small teams to large-scale engagements such as 16 vs. 16 and 20 vs. 20 operational scenarios.
    Growing Adoption of Transformer and Attention Mechanisms: Recent studies increasingly incorporated transformer architectures and attention mechanisms for battlefield representation, scalable swarm coordination, target prioritization, temporal feature extraction, and adaptive tactical reasoning. Attention-based learning consistently improved situational awareness and enabled UAVs to focus on tactically relevant allies, opponents, and battlefield events.
    Importance of Hierarchical Learning: Hierarchical RL frameworks were repeatedly employed to decompose complex combat tasks into multiple decision layers. More recently, some studies have combined hierarchical RL with game-theoretic modeling to support cooperative task allocation, adversarial decision-making, and risk-aware tactical planning under battlefield uncertainty. High-level policies typically generated tactical objectives, virtual targets, or risk-aware tactical decisions, while low-level policies executed maneuver commands. This decomposition reduced exploration difficulty, improved training efficiency, enhanced scalability in large combat environments, and supported resilient decision-making under battlefield uncertainty.
    Communication-Efficient Cooperation: Communication-aware MARL frameworks increasingly adopted sparse communication, intention inference, CommNet-based interaction, and attention-assisted information exchange to improve cooperative coordination while reducing communication overhead.
    Role of Realistic Simulation Environments: The reviewed air combat studies increasingly relied on high-fidelity simulators and digital environments such as Unity3D, JSBSim, PyAirCombat, MaCA, and digital-twin battlefield platforms. These environments enabled the evaluation of realistic flight dynamics, radar sensing, missile engagements, probabilistic damage models, and cooperative combat tactics.
    Growing Interest in Explainability: Explainable RL approaches employed reward decomposition, tactical-region analysis, and visual explanation mechanisms to improve the interpretability of autonomous combat policies and facilitate operator understanding.
  • Autonomous Navigation and Path Planning
    Broad Adoption of DRL and Hybrid Navigation Frameworks: The reviewed navigation studies show that frameworks based on PPO, DQN, actor–critic, and Q-learning can substantially improve autonomous UAV navigation, obstacle and collision avoidance, target visitation, and tactical path planning. Several studies further combined DRL with optimization-based planners, semantic terrain analysis, and sampling-based path planning techniques to improve robustness and mission effectiveness.
    Navigation Beyond Shortest-Path Optimization: Several studies emphasized that defense-oriented UAV navigation must consider survivability, concealment, threat avoidance, semantic terrain information, situational awareness, and adversarial exposure rather than merely minimizing path length. AI-enabled navigation frameworks are increasingly able to exploit terrain cover, avoid regions that are under enemy threat, and adapt their trajectories according to evolving battlefield conditions.
    GNSS-Denied and Communication-Constrained Operation: Many reviewed studies investigated GNSS-denied, communication-constrained, and interference-prone environments using telemetry prediction, DRL-based autonomous control, and sensor-based navigation to improve operational resilience.
    Energy-Aware Mission Execution: Energy consumption emerged as an important design consideration in target visitation, surveillance, swarm operation, and path planning studies. Several frameworks incorporated residual energy, flight duration, mission time, and energy-aware reward mechanisms to improve mission efficiency and operational endurance.
    Hybrid AI and Traditional Planning Synergy: Multiple studies demonstrated that combining AI-driven decision-making with traditional path planning algorithms such as A*, RRT, and genetic optimization methods can improve convergence, navigation reliability, computational efficiency, and path quality in complex operational environments.
    Simulation-Based Validation and Transfer Challenges: Many navigation studies relied on Gazebo, Parrot Sphinx, AirSim, Unreal Engine, Unity ML-Agents, and battlefield simulation environments. Although several studies incorporated real UAV experiments and telemetry-based validation, simulation remained the main evaluation methodology.
  • Target Tracking, Detection, and Classification
    Advances in AI-Driven Battlefield Perception: The reviewed studies demonstrate that CNNs, YOLO-based detectors, SSD-MobileNetV2, DopplerNet, attention-assisted architectures, Kalman filtering, KCF tracking, and RL-based perception frameworks can improve UAV-based target detection, classification, tracking, reconnaissance, and situational awareness.
    Persistent Small-Target Detection Challenges: Detecting small UAVs and distant aerial targets remains particularly challenging due to sparse visual features, long observation distances, cluttered backgrounds, rapid motion, viewpoint changes, and adverse environmental conditions. Attention mechanisms, multiscale feature fusion, specialized detection heads, and enhanced loss functions were commonly employed to address these limitations.
    Trend Toward Multimodal Sensing: The reviewed studies increasingly combined multiple sensing modalities, including electro-optical imagery, radar sensing, MIMO FMCW radar, range-Doppler processing, telemetry information, LiDAR measurements, onboard sensors, and X-ray inspection. Multimodal sensing consistently improved robustness, detection reliability, classification accuracy, and operational awareness.
    Embedded and Real-Time AI Operation: Several studies demonstrated real-time or near-real-time operation using lightweight detectors, embedded AI hardware, Raspberry Pi platforms, onboard processors, and computationally efficient tracking pipelines. These studies highlight the growing emphasis on computationally efficient AI for onboard deployment.
    Adversarial Robustness of Perception Systems: The reviewed adversarial ML-based studies demonstrated that AI-enabled UAV perception systems remain vulnerable to adversarial perturbations, spoofing attacks, deceptive inputs, and cyber–physical manipulation. The reviewed literature explored adversarial training, defensive distillation, feature-space regularization, and explainability-based detection to improve perception robustness.
    HITL and Human-Supervised Decision-Making: Although AI has the potential to improve autonomous target detection, tracking, and classification, safety-critical defense applications continue to benefit from HITL and human-supervised operation. The reviewed studies indicate that combining AI-enabled perception with operator oversight enhances decision reliability, supports target-engagement verification, and increases trust in mission-critical environments.
  • Cybersecurity, Electronic Warfare Protection, and Resilient UAV Operation
    Cyber–Physical Threats to UAV-Based Systems: The reviewed cybersecurity studies addressed a broad spectrum of cyber–physical threats, including GPS spoofing, GPS jamming, communication attacks, actuator faults, adversarial attacks, intrusion attempts, intelligent electromagnetic jamming, and attacks targeting AI-driven decision-making mechanisms. These threats can significantly affect UAV navigation, sensing, communication, control, and mission execution.
    AI-Enabled Intrusion and Attack Detection: DL, CNNs, LSTMs, transformer-based architectures, FL, and RL were widely employed for intrusion detection, anomaly identification, spoofing detection, anti-jamming control, adversarial defense, and fault diagnosis. The reviewed studies consistently reported strong attack-detection and mitigation performance across diverse UAV security scenarios.
    Distributed and Decentralized Resilience: Collaborative intrusion detection, decentralized FL, trust-aware node management, distributed event validation, and SDN were investigated to reduce single points of failure and improve resilience in distributed UAV networks and swarms.
    Anti-Jamming and Secure Communication: The reviewed works showed that AI-driven anti-jamming methods, DRL-based communication adaptation, frequency hopping, spread-spectrum techniques, communication reconfiguration, and hybrid mitigation strategies can improve communication robustness under intelligent jamming and adverse wireless conditions.
    Explainability for Cyber Defense: Explainability-driven mechanisms, including SHAP-based analysis and feature-activation interpretation, demonstrated the ability to identify abnormal behaviors and adversarial manipulations. These approaches improve the interpretability of AI-enabled cyber defense mechanisms and support operator understanding of security decisions.
  • Simulation-Based Evaluation and Practical Validation
    Central Role of Simulation-Based Evaluation: Across the reviewed application domains, simulation environments, digital twins, synthetic datasets, physics-based simulators, battlefield simulators, and RL training platforms constituted the primary means of evaluating AI-enabled UAV autonomy, air combat, navigation, perception, and cybersecurity mechanisms.
    Increasing Emphasis on Practical Validation: Several reviewed studies complemented simulation-based evaluation with real UAV experiments, telemetry datasets, embedded AI hardware, onboard implementation, SIL, and HIL validation. These efforts improve practical relevance and increase confidence in the applicability of the proposed AI frameworks.
  • Key Lessons for Future AI-Enabled UAV-Based Defense Systems
    AI as an Enabler of Intelligent Defense Operations: The reviewed literature demonstrates that many defense problems, including autonomous air combat, cooperative multi-UAV coordination, navigation in dynamic and contested environments, multi-target tracking, and cyberattack mitigation, are difficult to address using traditional rule-based approaches due to the complexity, uncertainty, and rapidly changing nature of battlefield conditions. AI-based methods, particularly DRL, MARL, DL, and hybrid learning frameworks, enable adaptive decision-making, continuous learning, and improved generalization in scenarios where manually designed rules are difficult to develop, maintain, or scale.
    Robust Autonomous Operation Across Defense Applications: Across the reviewed domains, robust UAV operation increasingly relies on resilient perception, secure communications, reliable autonomous decision-making, and effective operation under uncertain battlefield conditions.
    Importance of Trustworthy and Interpretable AI: The reviewed literature highlights the increasing role of trustworthy and interpretable AI in supporting transparency, operator confidence, and responsible deployment of autonomous UAV-based defense systems.
    Toward Integrated UAV-Based Defense Ecosystems: An emerging trend is the convergence of perception, communications, autonomy, cybersecurity, and distributed intelligence into integrated UAV-based defense ecosystems capable of supporting increasingly complex missions.
Table 8 summarizes the distribution and relative prevalence of representative AI paradigms, architectural concepts, and enabling technologies across the UAV-based defense application domains reviewed in this paper. DRL and MARL are among the most widely investigated AI paradigms for autonomous air combat, cooperative engagement, swarm coordination, and tactical decision-making, whereas CNNs, attention-assisted DL, transformer-based architectures, radar-based learning, and RL-assisted tracking are widely adopted for battlefield perception, target detection, reconnaissance, counter-UAV monitoring, and target tracking. Across all domains, cross-cutting technologies such as attention mechanisms, transformer architectures, communication-aware AI, XAI, FL, digital twins, and multimodal sensing have become key enablers of next-generation UAV-based defense systems. Overall, the reviewed literature indicates a gradual convergence of perception, communications, edge intelligence, and autonomous decision-making towards integrated, mission-aware UAV defense ecosystems.

9. Challenges, Open Issues, and Future Research Directions

Despite the significant advances achieved in AI-enabled UAV-based defense systems, several important challenges remain unresolved. As illustrated in Figure 3, the reviewed literature reveals both domain-specific research challenges and cross-cutting issues that hinder the broader deployment and operational adoption of these systems. Based on the surveyed studies, the key challenges, research gaps, and promising future research directions are summarized as follows:
  • Autonomous Air Combat Intelligence and Cooperative UAV Operations: According to the reviewed literature, several important challenges remain, including partial observability, the design of effective reward functions, training instability, limited interpretability of learned policies, communication overhead in cooperative UAV networks, scalability to large-scale combat formations, and the limited operational validation of policies trained primarily in simulation. Moreover, increasing battlefield complexity and communication constraints further complicate cooperative decision-making in large-scale aerial engagements. Future research should focus on communication-aware and scalable MARL frameworks, hierarchical and risk-aware cooperative decision-making, distributional RL for uncertainty-aware planning, improved learning efficiency through curriculum learning and self-play strategies, explainable tactical reasoning, and more realistic battlefield validation under communication-constrained and adversarial operational environments.
  • Autonomous Navigation and Mission Planning: Despite the advances of recent AI-driven navigation frameworks, maintaining reliable navigation under localization uncertainty, communication degradation, adversarial interference, dynamic obstacles, uncertain environmental conditions, and long-duration missions remains challenging. In addition, robust autonomous decision-making under incomplete situational awareness continues to be an open problem. Future research should investigate resilient multi-sensor navigation, adaptive mission re-planning, learning-assisted navigation under degraded sensing and communication conditions, and validation in operationally representative battlefield environments.
  • Target Tracking, Detection, and Battlefield Perception: Based on the reviewed literature, robust perception remains challenging in cluttered environments, adverse weather, camouflage, low-visibility conditions, long-range observations, and low-SNR scenarios. In addition, limited defense-oriented datasets, computational constraints, and insufficient operational validation continue to hinder practical deployment. Future research should emphasize robust multi-modal perception, computationally efficient detection algorithms suitable for onboard deployment, improved radar-based recognition, more representative defense-oriented datasets, and evaluation under realistic battlefield conditions.
  • Cybersecurity, Electronic Warfare, and Resilient UAV Operations: The reviewed literature demonstrates increasing attention toward protecting UAV-based defense systems against spoofing, jamming, communication attacks, malicious data manipulation, adversarial AI attacks, and distributed cyber threats. Although AI-assisted intrusion detection, FL, and XAI have demonstrated promising capabilities, ensuring trustworthy and resilient autonomous operation under contested battlefield environments remains an important research challenge. Future work should focus on adversarially robust AI, secure collaborative learning, lightweight distributed intrusion detection, communication-aware cyber defense mechanisms, and XAI techniques that improve operator confidence and mission reliability.
  • Resource-Efficient Embedded AI: Many reviewed studies highlight the need to execute increasingly sophisticated AI algorithms directly onboard UAV platforms while operating under stringent computational, memory, energy, and communication constraints. Maintaining reliable perception, navigation, and autonomous decision-making within these resource limitations remains challenging, particularly for long-duration and communication-constrained missions. Future research should investigate lightweight neural architectures, computationally efficient onboard inference, communication-efficient distributed AI, energy-aware mission execution, and integrated hardware–software optimization for embedded UAV intelligence.
  • Simulation, Validation, and Operational Deployment: A recurring observation throughout the reviewed literature is the extensive reliance on simulation environments, synthetic datasets, digital twins, and relatively limited real-world experimentation. Although these approaches have accelerated algorithm development, they often fail to fully capture the complexity, uncertainty, and communication conditions encountered during real defense operations. Therefore, bridging the gap between laboratory-scale evaluation and operational deployment remains one of the most important challenges identified in this review. Future research should prioritize realistic defense-oriented datasets, digital twins, SIL and HIL validation, standardized benchmarking methodologies, and large-scale real-flight experimentation.
  • Ethical, Legal, Safety, and Governance Considerations: As AI-enabled UAV-based defense systems become increasingly autonomous, ethical, legal, safety, and governance considerations are becoming increasingly important alongside the technical challenges discussed in this review. While recent advances in AI have significantly enhanced battlefield perception, autonomous navigation, cooperative operations, and decision support, the growing autonomy of defense platforms raises broader concerns regarding meaningful human oversight, accountability, transparency, explainability, and operator trust, particularly when AI supports mission-critical functions. In addition, future operational deployment will require appropriate governance frameworks together with rigorous verification, validation, certification, and assurance processes to improve the reliability and trustworthiness of AI-enabled autonomous behaviors. Future research should investigate trustworthy AI, robust assurance and certification methodologies, governance frameworks, human–AI collaboration, and interdisciplinary approaches that support the responsible and safe deployment of AI-enabled UAV-based defense systems.

10. Conclusions

This paper has presented a review of AI-enabled UAV-based defense systems, examining the major operational domains. Overall, this paper highlights the significant research progress achieved by applying AI to battlefield autonomy, intelligent perception, adaptive decision-making, resilient communications, distributed coordination, and autonomous mission execution in dynamic and contested environments. AI paradigms such as ML, DL, RL, MARL, FL, XAI, and edge intelligence have emerged as key enabling technologies for next-generation UAV-based defense systems. Across the reviewed application domains, substantial progress has been achieved in autonomous air combat, cooperative multi-UAV coordination, swarm intelligence, resilient navigation, intelligent target tracking, secure and resilient UAV operation, and AI-enabled protection against cyber and electronic warfare threats. In addition, recent advances in multimodal sensing, radar-enabled perception, communication-aware learning, and distributed intelligence have shown considerable potential to enhance the performance of UAV-based systems under uncertainty, communication constraints, electromagnetic interference, electronic warfare, and adversarial conditions. Despite these advances, several important challenges remain, including scalability, communication and onboard resource constraints, adversarial robustness, cybersecurity, and the transfer of AI-enabled UAV-based systems from simulated to operational environments. Moreover, a large proportion of the reviewed studies rely primarily on simulations, synthetic datasets, digital twins, or controlled experimental environments, while comprehensive field validation and operational deployment remain comparatively limited. Addressing these challenges will require continued advances in trustworthy AI and XAI, resilient autonomous operation, secure distributed intelligence, rigorous operational validation, and appropriate ethical, legal, safety, and governance frameworks that support the responsible deployment of increasingly autonomous UAV-based defense systems.

Author Contributions

Conceptualization, E.T.M. and I.S.K.; methodology, E.T.M. and I.S.K.; investigation, E.T.M. and I.S.K.; writing—original draft preparation, E.T.M. and I.S.K.; writing—review and editing, E.T.M. and I.S.K.; visualization, E.T.M. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

DURC Statement

Current research is limited to the scholarly review of artificial intelligence (AI)-enabled unmanned aerial vehicle (UAV)-based defense systems, which is intended to synthesize recent scientific advances, identify research trends, discuss challenges, and highlight future research directions. This review does not present new offensive capabilities, operational procedures, software implementations, or technical details that could facilitate misuse and as such does not pose a threat to public health or national security. The authors acknowledge the dual-use potential of AI-enabled UAV technologies and confirm that all necessary precautions have been taken to ensure that the manuscript communicates the reviewed research responsibly. The authors strictly adhere to relevant national and international ethical principles concerning dual-use research and advocate the responsible development, regulatory compliance, ethical deployment, and transparent reporting of AI-enabled UAV technologies to mitigate potential misuse while maximizing their societal and scientific benefits.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AAAspect Angle
A2/ADAnti-Access/Area-Denial
ACMDMAutonomous Air Combat Maneuver Decision-Making
AIArtificial Intelligence
AMSTDAttention Mechanism-Based Small Target Detection
AP50Average Precision at IoU threshold 50%
APFArtificial Potential Field
APsAverage Precision for small objects
ARISAerial Reconfigurable Intelligent Surface
ATAAntenna Train Angle
BERBit Error Rate
BFMBasic Fighter Maneuvering
BIMBasic Iterative Method
BITBatch Informed Trees
BVLOSBeyond-Visual Line-of-Sight
BVRBeyond Visual Range
C2Command-and-Control
CBSConvolution-Batch Normalization-SiLU
CIDSCollaborative Intrusion Detection System
CNNConvolutional Neural Network
CSPCross-Stage Partial
CTDECentralized Training and Decentralized Execution
CVaRConditional Value-at-Risk
D3QNDueling Double DQN
DBNDynamic Bayesian Network
DCPADistance at Closest Point of Approach
DDPGDeep Deterministic Policy Gradient
DEC-POMDPDecentralized Partially Observable Markov Decision Process
DFLDecentralized Federated Learning
DLDeep Learning
DNNDeep Neural Network
DQNDeep Q-Network
DRLDeep Reinforcement Learning
DTDecision Transformer
DTDEDecentralized Training and Decentralized Execution
DTPADynamic Threat Prioritization Assessment
ECEFEarth-Centered Earth-Fixed
EO/IRElectro-Optical/Infrared
FFCNNFeedforward CNN
FGSMFast Gradient Sign Method
FHPFrequency of Hazardous Proximity
FLFederated Learning
FMCWFrequency-Modulated Continuous Wave
FPSFrames Per Second
FPNFeature Pyramid Network
FSMFinite State Machine
GAGenetic Algorithm
GAEGeneralized Advantage Estimation
GDTGround Data Terminal
GFLOPsGiga Floating-Point Operations per Second
GISGeographic Information System
GNSSGlobal Navigation Satellite System
GPSGlobal Positioning System
GRUGated Recurrent Unit
HALEHigh-Altitude Long-Endurance
HCAHorizontal Crossing Angle
HILHardware-in-the-Loop
HITLHuman-in-the-Loop
HRLHierarchical Reinforcement Learning
IBVSImage-Based Visual Servoing
ICEMImproved Cross-Entropy Method
ICMIntrinsic Curiosity Module
IDSIntrusion Detection Systems
IGInformation Gain
IMUInertial Measurement Unit
INAVINertial NAVigation
INSInertial Navigation System
IoDInternet of Drones
IoUIntersection over Union
IPFSInterPlanetary File System
IPPOIndependent PPO
IQNImplicit Quantile Networks
ISRIntelligence, Surveillance, and Reconnaissance
ISTARIntelligence, Surveillance, Target Acquisition, and Reconnaissance
I2C-MATD3Improved Cross-Entropy Method with Intrinsic Curiosity-enhanced Multi-Agent
Twin Delayed DDPG
KCFKernelized Correlation Filter
KFKalman Filter
LDALinear Discriminant Analysis
LiDARLight Detection and Ranging
LLMLarge Language Model
LoSLine-of-Sight
LRFLaser Range Finder
LSTMLong Short-Term Memory
LSTM-AELSTM Autoencoder
MALEMedium-Altitude Long-Endurance
MAEMean Absolute Error
MAPEMean Absolute Percentage Error
MAPPOMulti-Agent Proximal Policy Optimization
MARLMulti-Agent Reinforcement Learning
MaCAMulti-agent Combat Arena
MCLDPPOMotivational Curriculum Learning Distributed Proximal Policy Optimization
MECMobile Edge Computing
MDPMarkov Decision Process
MDRL-DQNMultimodal DRL DQN
MERLMaximum Entropy Reinforcement Learning
MHAMultihead Attention
MIMOMultiple-Input Multiple-Output
MLMachine Learning
MLPMulti-Layer Perceptron
mIoUmean Intersection over Union
mAP@0.5mean Average Precision at IoU threshold 0.5
MTVOMulti-Agent Transformer introducing Virtual Objects
NBNaive Bayes
NEDNorth-East-Down
NISNormalized Innovation Square
NWD-CIoUnormalized Wasserstein distance and Complete IoU
ONOSOpen Network Operating System
PBEPerfect Bayes–Nash Equilibrium
PCSPPopulation Curriculum Self-Play
PDRPacket Delivery Ratio
PERPrioritized Experience Replay
PIDProportional-Integral-Differential
POMGPartially Observable Markov Game
PPOProximal Policy Optimization
QRQuick Response
QDAQuadratic Discriminant Analysis
Q-SAPFQ-learning-based Strategic Artificial Potential Field
R2Coefficient of Determination (R-squared)
Re-IDRe-Identification
RFRadio-Frequency
RGBRed–Green–Blue
RLReinforcement Learning
RMSERoot Mean Square Error
RNNRecurrent Neural Network
ROCReceiver Operating Characteristic
RRTRapidly-exploring Random Trees
SAPFStrategic Artificial Potential Field
SARSynthetic Aperture Radar
SARSAState-Action-Reward-State-Action
SDNSoftware-Defined Networking
SD-UAVSoftware-Defined UAV
SHAPSHapley Additive exPlanations
SIESpatial Information Entropy
SIIS-MARLSparse Inferred Intention Sharing MARL
SILSoftware-in-the-Loop
SISOSingle-Input Single-Output
SLAMSimultaneous Localization and Mapping
SMAScalable Mixing Networks based on Attention
SNRSignal-to-Noise Ratio
SPPFSpatial Pyramid Pooling Fast
SSPSpatial-aware and Scale-aware Prediction
TaLUTanh Linear Unit
TCPATime to Closest Point of Approach
ToMTheory of Mind
U2GUAV-to-Ground
U2SUAV-to-Satellite
U2UUAV-to-UAV
UAVUnmanned Aerial Vehicle
UAV-STDUAV-to-UAV Small Target Detection
UCAVUnmanned Combat Aerial Vehicle
UAPUniversal Adversarial Perturbation
VTOLVertical Takeoff and Landing
WEZWeapon Engagement Zone
XAIExplainable AI

References

  1. Rashid, A.B.; Kausik, A.K.; Sunny, A.H.; Ahamed, B.; Hassan, M. Artificial Intelligence in the Military: An Overview of the Capabilities, Applications, and Challenges. Int. J. Intell. Syst. 2023, 8676366. [Google Scholar] [CrossRef]
  2. Bistron, M.; Piotrowski, Z. Artificial Intelligence Applications in Military Systems and Their Influence on Sense of Security of Citizens. Electronics 2021, 10, 871. [Google Scholar] [CrossRef]
  3. Alcántara Suárez, E.J.; Monzon Baeza, V. Evaluating the Role of Machine Learning in Defense Applications and Industry. Mach. Learn. Knowl. Extr. 2023, 5, 1557–1569. [Google Scholar] [CrossRef]
  4. Hadlington, L.; Binder, J.; Gardner, S.; Karanika-Murray, M.; Knight, S. The Use of Artificial Intelligence in a Military Context: Development of the Attitudes Toward AI in Defense (AAID) Scale. Front. Psychol. 2023, 14, 1164810. [Google Scholar] [CrossRef] [PubMed]
  5. Gargalakos, M. The Role of Unmanned Aerial Vehicles in Military Communications: Application Scenarios, Current Trends, and Beyond. J. Def. Model. Simul. Appl. Methodol. Technol. 2024, 21, 313–321. [Google Scholar] [CrossRef]
  6. Yu, A.; Kolotylo, I.; Hashim, H.A.; Eltoukhy, A.E.E.E. Electronic Warfare Cyberattacks, Countermeasures, and Modern Defensive Strategies of UAV Avionics: A Survey. IEEE Access 2025, 13, 68660–68681. [Google Scholar] [CrossRef]
  7. Michailidis, E.T.; Potirakis, S.M.; Kanatas, A.G. AI-Inspired Non-Terrestrial Networks for IIoT: Review on Enabling Technologies and Applications. IoT 2020, 1, 21–48. [Google Scholar] [CrossRef]
  8. Michailidis, E.T.; Maliatsos, K.; Skoutas, D.N.; Vouyioukas, D.; Skianis, C. Secure UAV-Aided Mobile Edge Computing for IoT: A Review. IEEE Access 2022, 10, 86353–86383. [Google Scholar] [CrossRef]
  9. Michailidis, E.T.; Vouyioukas, D. A Review on Software-Based and Hardware-Based Authentication Mechanisms for the Internet of Drones. Drones 2022, 6, 41. [Google Scholar] [CrossRef]
  10. Hadi, H.J.; Cao, Y.; Un Nisa, K.; Jamil, A.M.; Ni, Q. A Comprehensive Survey on Security, Privacy Issues and Emerging Defence Technologies for UAVs. J. Netw. Comput. Appl. 2023, 213, 103607. [Google Scholar] [CrossRef]
  11. Alsadie, D. Cybersecurity and Artificial Intelligence in Unmanned Aerial Vehicles: Emerging Challenges and Advanced Countermeasures. IET Inf. Secur. 2025, 2046868. [Google Scholar] [CrossRef]
  12. Oli, A.; Mahalal, E. UAV Security: Attacks, Defenses, and Open Challenges. IEEE Access 2025, 13, 215606–215635. [Google Scholar] [CrossRef]
  13. Ogab, M.; Zaidi, S.; Bourouis, A.; Calafate, C.T. Machine Learning-Based Intrusion Detection Systems for the Internet of Drones: A Systematic Literature Review. IEEE Access 2025, 13, 96681–96714. [Google Scholar] [CrossRef]
  14. AL-Syouf, R.; Bani-Hani, R.; AL-Jarrah, O.Y. Machine Learning Approaches to Intrusion Detection in Unmanned Aerial Vehicles (UAVs). Neural Comput. Appl. 2024, 36, 18009–18041. [Google Scholar] [CrossRef]
  15. Islam, M.S.; Mahmoud, A.S.; Sheltami, T.R. AI-Enhanced Intrusion Detection for UAV Systems: A Taxonomy and Comparative Review. Drones 2025, 9, 682. [Google Scholar] [CrossRef]
  16. Papathanasiou, D.; Zacharakis, E.; Liaperdos, J.; Kotsilieris, T.; Livieris, I.E.; Ioannou, K. Secure Communication Protocols and AI-Based Anomaly Detection in UAV-GCS. Appl. Sci. 2026, 16, 3339. [Google Scholar] [CrossRef]
  17. Kacem, T.; Benjamin, K. Artificial Intelligence Methods for Unmanned Aerial Vehicles Cybersecurity: A Comprehensive Survey. Drones 2026, 10, 400. [Google Scholar] [CrossRef]
  18. Skarka, W.; Ashfaq, R. Hybrid Machine Learning and Reinforcement Learning Framework for Adaptive UAV Obstacle Avoidance. Aerospace 2024, 11, 870. [Google Scholar] [CrossRef]
  19. Fagundes-Junior, L.A.; de Carvalho, K.B.; Ferreira, R.S.; Brandão, A.S. Machine Learning for Unmanned Aerial Vehicles Navigation: An Overview. SN Comput. Sci. 2024, 5, 256. [Google Scholar] [CrossRef]
  20. Meng, W.; Zhang, X.; Zhou, L.; Guo, H.; Hu, X. Advances in UAV Path Planning: A Comprehensive Review of Methods, Challenges, and Future Directions. Drones 2025, 9, 376. [Google Scholar] [CrossRef]
  21. Zhai, L.; Wu, H.; Lai, L.; Gao, Z. Intelligent Optimization Algorithms for Multi-UAV Path Planning: A Comprehensive Review. IEEE Access 2025, 13, 101106–101130. [Google Scholar] [CrossRef]
  22. Teixeira, K.; Miguel, G.; Silva, H.S.; Madeiro, F. A Survey on Applications of Unmanned Aerial Vehicles Using Machine Learning. IEEE Access 2023, 11, 117582–117621. [Google Scholar] [CrossRef]
  23. Caballero-Martin, D.; Lopez-Guede, J.M.; Estevez, J.; Graña, M. Artificial Intelligence Applied to Drone Control: A State of the Art. Drones 2024, 8, 296. [Google Scholar] [CrossRef]
  24. Yang, Z.; Zhang, Y.; Zeng, J.; Yang, Y.; Jia, Y.; Song, H.; Lv, T.; Sun, Q.; An, J. AI-Driven Safety and Security for UAVs: From Machine Learning to Large Language Models. Drones 2025, 9, 392. [Google Scholar] [CrossRef]
  25. Wu, P.; Li, Y.; Xue, D. UAV Target Tracking: A Survey. Artif. Intell. Rev. 2025, 58, 358. [Google Scholar] [CrossRef]
  26. Costa, A.N.; Dantas, J.P.A.; Scukins, E.; Medeiros, F.L.L.; Ögren, P. Simulation and Machine Learning in Beyond Visual Range Air Combat: A Survey. IEEE Access 2025, 13, 76755–76774. [Google Scholar] [CrossRef]
  27. Al-Kamali, F.; Chan, F.; Ammar, H.A.; Bayes, J.H.; D’Amours, C. UAV-Mounted Aerial Relays in Military Communications: A Comprehensive Survey. IEEE Open J. Commun. Soc. 2026, 7, 1096–1136. [Google Scholar] [CrossRef]
  28. Khawaja, W.; Ezuma, M.; Semkin, V.; Erden, F.; Ozdemir, O.; Guvenc, I. A Survey on Detection, Classification, and Tracking of AAVs Using Radar and Communications Systems. IEEE Commun. Surv. Tutor. 2026, 28, 3272–3310. [Google Scholar] [CrossRef]
  29. Michailidis, E.T.; Maliatsos, K.; Vouyioukas, D. Software-Defined Radio Deployments in UAV-Driven Applications: A Comprehensive Review. IEEE Open J. Veh. Technol. 2024, 5, 1545–1586. [Google Scholar] [CrossRef]
  30. Javed, S.; Hassan, A.; Ahmad, R.; Ahmed, W.; Ahmed, R.; Saadat, A.; Guizani, M. State-of-the-Art and Future Research Challenges in UAV Swarms. IEEE Internet Things J. 2024, 11, 19023–19045. [Google Scholar] [CrossRef]
  31. Ekechi, C.C.; Elfouly, T.; Alouani, A.; Khattab, T. A Survey on UAV Control with Multi-Agent Reinforcement Learning. Drones 2025, 9, 484. [Google Scholar] [CrossRef]
  32. Phadke, A.; Medrano, F.A. Towards Resilient UAV Swarms—A Breakdown of Resiliency Requirements in UAV Swarms. Drones 2022, 6, 340. [Google Scholar] [CrossRef]
  33. Chen, B.; Lin, B.; Li, M.; Li, Z.; Zhang, X.; Shi, M.; Qin, K. Event-Triggered-Based Neuroadaptive Bipartite Containment Tracking for Networked Unmanned Aerial Vehicles. Drones 2025, 9, 317. [Google Scholar] [CrossRef]
  34. McEnroe, P.; Wang, S.; Liyanage, M. A Survey on the Convergence of Edge Computing and AI for UAVs: Opportunities and Challenges. IEEE Internet Things J. 2022, 9, 15435–15459. [Google Scholar] [CrossRef]
  35. Ding, Y.; Yang, Z.; Pham, Q.-V.; Hu, Y.; Zhang, Z.; Shikh-Bahaei, M. Distributed Machine Learning for UAV Swarms: Computing, Sensing, and Semantics. IEEE Internet Things J. 2024, 11, 7447–7473. [Google Scholar] [CrossRef]
  36. Zhu, J.; Kuang, M.; Zhou, W.; Shi, H.; Zhu, J.; Han, X. Mastering Air Combat Game with Deep Reinforcement Learning. Def. Technol. 2024, 34, 295–312. [Google Scholar] [CrossRef]
  37. Zheng, Z.; Duan, H. UAV Maneuver Decision-Making via Deep Reinforcement Learning for Short-Range Air Combat. Intell. Robot. 2023, 3, 76–94. [Google Scholar] [CrossRef]
  38. Saldiran, E.; Hasanzade, M.; Inalhan, G.; Tsourdos, A. Towards Global Explainability of Artificial Intelligence Agent Tactics in Close Air Combat. Aerospace 2024, 11, 415. [Google Scholar] [CrossRef]
  39. Wang, X.; Wang, Y.; Su, X.; Wang, L.; Lu, C.; Peng, H.; Liu, J. Deep Reinforcement Learning-Based Air Combat Maneuver Decision-Making: Literature Review, Implementation Tutorial and Future Direction. Artif. Intell. Rev. 2024, 57, 1. [Google Scholar] [CrossRef]
  40. Wang, B.; Gao, X.; Xie, T. An Evolutionary Multi-Agent Reinforcement Learning Algorithm for Multi-UAV Air Combat. Knowl.-Based Syst. 2024, 299, 112000. [Google Scholar] [CrossRef]
  41. Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-Agent Actor–Critic for Mixed Cooperative-Competitive Environments. In Proceedings of the 31st International Conference on Neural Information Processing Systems (NIPS 2017); Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 6382–6393. [Google Scholar]
  42. Fujimoto, S.; van Hoof, H.; Meger, D. Addressing Function Approximation Error in Actor-Critic Methods. In Proceedings of the 35th International Conference on Machine Learning (ICML 2018); Proceedings of Machine Learning Research; PMLR: New York, NY, USA, 2018; Volume 80, pp. 1587–1596. [Google Scholar]
  43. Wang, W.; Yang, T.; Liu, Y.; Hao, J.; Hao, X.; Hu, Y.; Chen, Y.; Fan, C.; Gao, Y. From Few to More: Large-Scale Dynamic Multiagent Curriculum Learning. Proc. Aaai Conf. Artif. Intell. 2020, 34, 7293–7300. [Google Scholar] [CrossRef]
  44. Xu, X.; Wang, Y.; Guo, X.; Huang, K.; Zhang, X. Multi-UAV Air Combat Cooperative Game Based on Virtual Opponent and Value Attention Decomposition Policy Gradient. Expert Syst. Appl. 2025, 267, 126069. [Google Scholar] [CrossRef]
  45. Barto, A.G.; Mahadevan, S. Recent Advances in Hierarchical Reinforcement Learning. Discret. Event Dyn. Syst. 2003, 13, 341–379. [Google Scholar] [CrossRef]
  46. Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2022; pp. 24611–24624. [Google Scholar]
  47. Yang, J.; Yang, X.; Yu, T. Multi-Unmanned Aerial Vehicle Confrontation in Intelligent Air Combat: A Multi-Agent Deep Reinforcement Learning Approach. Drones 2024, 8, 382. [Google Scholar] [CrossRef]
  48. Ren, Z.; Zhang, D.; Tang, S.; Xiong, W.; Yang, S.-H. Cooperative Maneuver Decision Making for Multi-UAV Air Combat Based on Incomplete Information Dynamic Game. Def. Technol. 2023, 27, 308–317. [Google Scholar] [CrossRef]
  49. Dankwa, S.; Zheng, W. Twin-Delayed DDPG: A Deep Reinforcement Learning Technique to Model a Continuous Movement of an Intelligent Robot Agent. In Proceedings of the 3rd International Conference on Vision, Image and Signal Processing, Vancouver, BC, Canada, 26–28 August 2019; pp. 1–5. [Google Scholar]
  50. Ding, Z.; Wang, X.; Cai, C.; Jia, L.; Xu, Z. Multi-UAV Intelligent Decision-Making Method with Layer Delay Dual-Center MAPPO for Air Combat. Appl. Intell. 2025, 55, 811. [Google Scholar] [CrossRef]
  51. Song, Y.; Chen, H. Active Interception for Multi-Target Encirclement by Heterogeneous UAVs: An LSTM-Enhanced Independent PPO Algorithm. Designs 2026, 10, 26. [Google Scholar] [CrossRef]
  52. Tang, M.; Chen, R.; Zhu, J.; Jiang, Y.; Zhang, R.; Cui, Y.; Zhang, B. Hierarchical Game-Theoretic and Risk-Aware Predictive Control Framework for Resilient Multi-UAV Cooperative Combat. IEEE Access 2026, 14, 35678–35704. [Google Scholar] [CrossRef]
  53. Wang, H.; Wang, J. Enhancing Multi-UAV Air Combat Decision Making via Hierarchical Reinforcement Learning. Sci. Rep. 2024, 14, 4458. [Google Scholar] [CrossRef] [PubMed]
  54. Wu, Q.; Chen, L.; Liu, K.; Lü, J. A Simulation Platform for MARL Training and Evaluation in Swarm Confrontation. IEEE Robot. Autom. Lett. 2026, 11, 6050–6057. [Google Scholar] [CrossRef]
  55. Luo, D.; Fan, Z.; Yang, Z.; Xu, Y. Multi-UAV Cooperative Maneuver Decision-Making for Pursuit-Evasion Using Improved MADRL. Def. Technol. 2024, 35, 187–197. [Google Scholar] [CrossRef]
  56. Xie, J.; Li, G. Enhanced Q Learning and Deep Reinforcement Learning for Unmanned Combat Intelligence Planning in Adversarial Environments. Sci. Rep. 2025, 15, 28364. [Google Scholar] [CrossRef] [PubMed]
  57. Zhao, M.; Wang, G.; Fu, Q.; Quan, W.; Wen, Q.; Wang, X.; Li, T.; Chen, Y.; Xue, S.; Han, J. Intelligent Decision-Making System of Air Defense Resource Allocation via Hierarchical Reinforcement Learning. Int. J. Intell. Syst. 2024, 2024, 7777050. [Google Scholar] [CrossRef]
  58. Han, J.; Yan, Y.; Zhang, B. Towards Efficient Multi-UAV Air Combat: An Intention Inference and Sparse Transmission Based Multiagent Reinforcement Learning Algorithm. IEEE Trans. Artif. Intell. 2025, 6, 3441–3452. [Google Scholar] [CrossRef]
  59. Rabinowitz, N.; Perbet, F.; Song, F.; Zhang, C.; Eslami, S.M.A.; Botvinick, M. Machine Theory of Mind. In Proceedings of the International Conference on Machine Learning (ICML), Stockholm, Sweden, 10–15 July 2018; pp. 4218–4227. [Google Scholar]
  60. Mao, H.; Zhang, Z.; Xiao, Z.; Gong, Z.; Ni, Y. Learning Multi-Agent Communication with Double Attentional Deep Reinforcement Learning. Auton. Agents Multi-Agent Syst. 2020, 34, 32. [Google Scholar] [CrossRef]
  61. Guan, C.; Chen, F.; Yuan, L.; Wang, C.; Yin, H. Efficient Multiagent Communication via Self-Supervised Information Aggregation. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates, Inc.: Red Hook, NY, USA, 2022; Volume 35, pp. 1020–1033. [Google Scholar]
  62. Jiang, F.; Xu, M.; Li, Y.; Cui, H.; Wang, R. Short-Range Air Combat Maneuver Decision of UAV Swarm Based on Multi-Agent Transformer Introducing Virtual Objects. Eng. Appl. Artif. Intell. 2023, 123, 106358. [Google Scholar] [CrossRef]
  63. Kong, W.; Zhou, D.; Du, Y.; Zhou, Y.; Zhao, Y. Reinforcement Learning for Multiaircraft Autonomous Air Combat in Multisensor UCAV Platform. IEEE Sens. J. 2023, 23, 20596–20606. [Google Scholar] [CrossRef]
  64. Kim, G.S.; Lee, S.; Woo, T.; Park, S. Cooperative Reinforcement Learning for Military Drones over Large-Scale Battlefields. IEEE Trans. Intell. Veh. 2024, 1–11. [Google Scholar] [CrossRef]
  65. Li, B.; Huang, J.; Zhang, H.; Huai, L.; Neretin, E. Enhanced Exploration for Multi-UAV Cooperative Roundup: An I2C-MATD3 Reinforcement Learning Framework. Def. Technol. 2026, 58, 374–389. [Google Scholar] [CrossRef]
  66. Borja-Jaimes, V.; Valdez-Martínez, J.S.; Beltrán-Escobar, M.; Ramírez-Zúñiga, G.; Reyes-Mayer, A.; Calixto-Rodríguez, M. Robust Backstepping-Sliding Control of a Quadrotor UAV with Disturbance Compensation. Computation 2026, 14, 51. [Google Scholar] [CrossRef]
  67. Soliman, A.; Al-Ali, A.; Mohamed, A.; Gedawy, H.; Izham, D.; Bahri, M.; Erbad, A.; Guizani, M. AI-Based UAV Navigation Framework with Digital Twin Technology for Mobile Target Visitation. Eng. Appl. Artif. Intell. 2023, 123, 106318. [Google Scholar] [CrossRef]
  68. Shen, Y.; Zhang, X.; Li, Y.; Zhang, W. Deep Reinforcement Learning-Based Adaptive Collision Avoidance Method for UAV in Joint Operational Airspace. Def. Technol. 2026, 56, 142–159. [Google Scholar] [CrossRef]
  69. Zhang, J.; Xian, Y.; Zhu, X.; Deng, H. A Hybrid Deep Learning Model for UAV Path Planning in Dynamic Environments. IEEE Access 2025, 13, 67459–67475. [Google Scholar] [CrossRef]
  70. García-Gascón, C.; Castelló-Pedrero, P.; Chinesta, F.; García-Manrique, J.A. Artificial Intelligence-Driven Aircraft Systems to Emulate Autopilot and GPS Functionality in GPS-Denied Scenarios Through Deep Learning. Drones 2025, 9, 250. [Google Scholar] [CrossRef]
  71. Sujecki, P.; Frąszczak, D. Artificial Intelligence-Based Decision Support System for UAV Control in a Simulated Environment. Sensors 2026, 26, 2436. [Google Scholar] [CrossRef] [PubMed]
  72. Chronis, C.; Anagnostopoulos, G.; Politi, E.; Dimitrakopoulos, G.; Varlamis, I. Dynamic Navigation in Unconstrained Environments Using Reinforcement Learning Algorithms. IEEE Access 2023, 11, 117984–118001. [Google Scholar] [CrossRef]
  73. Lee, J.; Seo, Y. Q-Learning Based on Strategic Artificial Potential Field for Path Planning Enabling Concealment and Cover in Ground Battlefield Environments. Appl. Intell. 2024, 54, 7170–7200. [Google Scholar] [CrossRef]
  74. Kim, S.M.; Lee, J.; Kwon, H. Tactical Path Planning in Battlefield Environments Using Semantic Segmentation and Deep Reinforcement Learning. IEEE Access 2026, 14, 42838–42849. [Google Scholar] [CrossRef]
  75. Zhan, H.; Zhang, Y.; Huang, J.; Song, Y.; Xing, L.; Wu, J.; Gao, Z. A Reinforcement Learning-Based Evolutionary Algorithm for the Unmanned Aerial Vehicles Maritime Search and Rescue Path Planning Problem Considering Multiple Rescue Centers. Memetic Comput. 2024, 16, 373–386. [Google Scholar] [CrossRef]
  76. Lee, M.; Choi, M.; Yang, T.; Kim, J.; Kim, J.; Kwon, O.; Cho, N. A Study on the Advancement of Intelligent Military Drones: Focusing on Reconnaissance Operations. IEEE Access 2024, 12, 55964–55975. [Google Scholar] [CrossRef]
  77. Salameh, H.B.; Hussienat, A.; Alhafnawi, M.; Al-Ajlouni, A. Autonomous UAV-Based Surveillance System for Multi-Target Detection Using Reinforcement Learning. Clust. Comput. 2024, 27, 9381–9394. [Google Scholar] [CrossRef]
  78. Ud Din, I.; Almogren, A.; Rodrigues, J.J.P.C. Adversarial Machine Learning for Robust and Secure UAV Detection in Consumer Applications. IEEE Trans. Consum. Electron. 2025, 71, 3360–3367. [Google Scholar] [CrossRef]
  79. Sayed, A.N.; Abedi, H.; Ramahi, O.M.; Shaker, G. Enhanced UAV Detection and Classification Using Machine Learning and MIMO Radars. IEEE Trans. Microw. Theory Tech. 2024, 72, 6716–6727. [Google Scholar] [CrossRef]
  80. Ansys Inc. Ansys HFSS. 2022. Available online: https://www.ansys.com/products/electronics/ansys-hfss (accessed on 22 July 2026).
  81. Han, T.-T.; Duy, A.L.; Van, T.N.; Tan, H.D.; Trung, A.D. High Performance Autonomous Target Tracking and Control Method for Quadrotors Using Artificial Intelligence. IEEE Access 2025, 13, 214236–214252. [Google Scholar] [CrossRef]
  82. Yang, H.; Gao, S.; Wu, X.; Zhang, Y. Online Multi-Object Tracking Using KCF-Based Single-Object Tracker with Occlusion Analysis. Multimed. Syst. 2020, 26, 655–669. [Google Scholar] [CrossRef]
  83. Han, K. Image Object Tracking Based on Temporal Context and MOSSE. Clust. Comput. 2017, 20, 1259–1269. [Google Scholar] [CrossRef]
  84. Farkhodov, K.; Lee, S.; Kwon, K. Object Tracking Using CSRT Tracker and RCNN. In Proceedings of the 13th International Joint Conference on Biomedical Engineering Systems and Technologies (BIOSTEC), Valletta, Malt, 24–26 February 2020; pp. 209–212. [Google Scholar]
  85. Bhat, G.; Danelljan, M.; Van Gool, L.; Timofte, R. Learning Discriminative Model Prediction for Tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 6181–6190. [Google Scholar]
  86. Chen, X.; Yan, B.; Zhu, J.; Wang, D.; Yang, X.; Lu, H. Transformer Tracking. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 8122–8131. [Google Scholar]
  87. Ye, B.; Chang, H.; Ma, B.; Shan, S. Joint Feature Learning and Relation Modeling for Tracking: A One-Stream Framework. In Computer Vision—ECCV 2022; Springer: Cham, Switzerland, 2022; pp. 341–357. [Google Scholar]
  88. Kant, R.; Saini, P.; Kumari, J. Long Short-Term Memory Auto-Encoder-Based Position Prediction Model for Fixed-Wing UAV During Communication Failure. IEEE Trans. Artif. Intell. 2023, 4, 173–181. [Google Scholar] [CrossRef]
  89. Liu, Y.; Su, Z.; Li, H.; Zhang, Y. An LSTM Based Classification Method for Time Series Trend Forecasting. In Proceedings of the 14th IEEE Conference on Industrial Electronics and Applications (ICIEA), Xi’an, China, 19–21 June 2019; pp. 402–406. [Google Scholar]
  90. Li, C.; Wei, F.; Dong, W.; Wang, X.; Liu, Q.; Zhang, X. Dynamic Structure Embedded Online Multiple-Output Regression for Streaming Data. IEEE Trans. Pattern Anal. Mach. Intell. 2019, 41, 323–336. [Google Scholar] [CrossRef] [PubMed]
  91. Zuo, G.; Zhou, K.; Wang, Q. UAV-to-UAV Small Target Detection Method Based on Deep Learning in Complex Scenes. IEEE Sens. J. 2025, 25, 3806–3820. [Google Scholar] [CrossRef]
  92. Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.-Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Computer Vision—ECCV 2016; Springer: Cham, Switzerland, 2016; pp. 21–37. [Google Scholar]
  93. Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [PubMed]
  94. Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. DETRs Beat YOLOs on Real-Time Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar]
  95. Kim, J.; Joe, I. Deep Learning-Based Drone Defense System for Autonomous Detection and Mitigation of Balloon-Borne Threats. Electronics 2025, 14, 1553. [Google Scholar] [CrossRef]
  96. Masadeh, A.; Alhafnawi, M.; Salameh, H.A.B.; Musa, A.; Jararweh, Y. Reinforcement Learning-Based Security/Safety UAV System for Intrusion Detection Under Dynamic and Uncertain Target Movement. IEEE Trans. Eng. Manag. 2024, 71, 12498–12508. [Google Scholar] [CrossRef]
  97. Mignon, A.d.S.; Rocha, R.L.d.A.d. An Adaptive Implementation of ϵ-Greedy in Reinforcement Learning. Procedia Comput. Sci. 2017, 109, 1146–1151. [Google Scholar] [CrossRef]
  98. Ma, X.; Sun, T.; Gao, M. A Reinforcement-Learning-Enhanced Spoofing Algorithm for UAV With GPS/INS-Integrated Navigation. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 8659–8673. [Google Scholar] [CrossRef]
  99. Haarnoja, T.; Zhou, A.; Abbeel, P.; Levine, S. Soft Actor–Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning (ICML 2018), Proceedings of Machine Learning Research, Stockholm, Sweden, 10–15 July 2018; Volume 80, pp. 1861–1870. [Google Scholar]
  100. Ma, X.; Gao, M.; Zhao, Y.; Yu, M. A Novel Navigation Spoofing Algorithm for UAV Based on GPS/INS-Integrated Navigation. IEEE Trans. Veh. Technol. 2024, 73, 15424–15439. [Google Scholar] [CrossRef]
  101. Zhao, R.; Wen, H.; Hou, W.; Jiang, L.; Chen, Z.; Tang, T.; Feng, X. Anti-Jamming Communication Framework for UAV Based Defense Operations to Control AI-Enabled Attacks. IEEE Trans. Consum. Electron. 2026; early access. [CrossRef]
  102. Hickling, T.; Aouf, N.; Spencer, P. Robust Adversarial Attacks Detection Based on Explainable Deep Reinforcement Learning for UAV Guidance and Planning. IEEE Trans. Intell. Veh. 2023, 8, 4381–4394. [Google Scholar] [CrossRef]
  103. Lundberg, M.S.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems (NeurIPS); Curran Associates Inc.: Red Hook, NY, USA, 2017; pp. 4768–4777. [Google Scholar]
  104. Musa, A.; Vishi, K.; Rexha, B. Attack Analysis of Face Recognition Authentication Systems Using Fast Gradient Sign Method. Appl. Artif. Intell. 2021, 35, 1346–1360. [Google Scholar] [CrossRef]
  105. Kurakin, A.; Goodfellow, I.J.; Bengio, S. Adversarial Examples in the Physical World. In Artificial Intelligence Safety and Security; Chapman and Hall/CRC: Boca Raton, FL, USA, 2018; pp. 99–112. [Google Scholar]
  106. Mynuddin, M.; Khan, S.U.; Ahmari, R.; Landivar, L.; Mahmoud, M.N.; Homaifar, A. Trojan Attack and Defense for Deep Learning-Based Navigation Systems of Unmanned Aerial Vehicles. IEEE Access 2024, 12, 89887–89907. [Google Scholar] [CrossRef]
  107. Loquercio, A.; Maqueda, A.I.; del-Blanco, C.R.; Scaramuzza, D. DroNet: Learning to Fly by Driving. IEEE Robot. Autom. Lett. 2018, 3, 1088–1095. [Google Scholar] [CrossRef]
  108. Moosavi-Dezfooli, S.-M.; Fawzi, A.; Fawzi, O.; Frossard, P. Universal Adversarial Perturbations. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1765–1773. [Google Scholar]
  109. Howard, A.; Sandler, M.; Chen, B.; Wang, W.; Chen, L.-C.; Tan, M.; Chu, G.; Vasudevan, V.; Zhu, Y.; Pang, R.; et al. Searching for MobileNetV3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1314–1324. [Google Scholar]
  110. Hadi, H.J.; Cao, Y.; Li, S.; Hu, Y.; Wang, J.; Wang, S. Real-Time Collaborative Intrusion Detection System in UAV Networks Using Deep Learning. IEEE Internet Things J. 2024, 11, 33371–33391. [Google Scholar] [CrossRef]
  111. Unmanned Aerial Vehicle (UAV) Intrusion Detection [Dataset]. UCI Machine Learning Repository. 2020. Available online: https://archive.ics.uci.edu/dataset/564/unmanned+aerial+vehicle+uav+intrusion+detection (accessed on 22 July 2026).
  112. Agnew, D.; Aguila, A.D.; McNair, J. Enhanced Network Metric Prediction for Machine Learning-Based Cyber Security of a Software-Defined UAV Relay Network. IEEE Access 2024, 12, 54202–54219. [Google Scholar] [CrossRef]
  113. del Aguila, A.; Mendoza, J.V.; Mandavilli, S.B.; McNair, J. Remote and Rural Connectivity via Multi-Tier Systems Through SDN-Managed Drone Networks. In Proceedings of the IEEE Military Communications Conference (MILCOM), Rockville, MD, USA, 28 November–2 December 2022; pp. 956–961. [Google Scholar]
  114. Jadav, N.K.; Rathod, T.; Gupta, R.; Tanwar, S.; Kumar, N.; Iqbal, R.; Atalla, S.; Mohammad, H.; Al-Rubaye, S. Blockchain-Based Secure and Intelligent Data Dissemination Framework for UAVs in Battlefield Applications. IEEE Commun. Stand. Mag. 2023, 7, 16–23. [Google Scholar] [CrossRef]
  115. Gupta, R.; Tanwar, S.; Kumar, N. B-IoMV: Blockchain-Based Onion Routing Protocol for D2D Communication in an IoMV Environment Beyond 5G. Veh. Commun. 2021, 32, 100401. [Google Scholar] [CrossRef]
  116. Beltrán, E.T.M.; Sánchez, P.M.S.; Bovet, G.; Stiller, B.; Pérez, G.M.; Celdrán, A.H. Flighter: Decentralized Federated Learning and Situational Awareness for Secure Military Aerial Reconnaissance. IEEE Commun. Mag. 2025, 63, 136–142. [Google Scholar] [CrossRef]
  117. Beltrán, E.T.M. CyberDataLab/Flighter. GitHub Repository. 2024. Available online: https://github.com/CyberDataLab/flighter (accessed on 22 July 2026).
  118. Nguyen, T.M.; Truong, V.T.; Le, L.B. Agentic AI Meets Edge Computing in Autonomous UAV Swarms. IEEE Internet Things Mag. 2026, 9, 87–95. [Google Scholar] [CrossRef]
Figure 1. Conceptual ecosystem of AI-enabled UAV-based defense systems, illustrating key enabling capabilities and major operational application domains.
Figure 1. Conceptual ecosystem of AI-enabled UAV-based defense systems, illustrating key enabling capabilities and major operational application domains.
Drones 10 00602 g001
Figure 2. Integrated architecture and functional classification framework for AI-enabled UAV-based defense systems.
Figure 2. Integrated architecture and functional classification framework for AI-enabled UAV-based defense systems.
Drones 10 00602 g002
Figure 3. Open challenges, research gaps, and future research directions across major operational domains of AI-enabled UAV-based defense systems.
Figure 3. Open challenges, research gaps, and future research directions across major operational domains of AI-enabled UAV-based defense systems.
Drones 10 00602 g003
Table 1. Comparison of relevant review papers.
Table 1. Comparison of relevant review papers.
ReferenceUAVsAI ParadigmsDefense
Applications
Air CombatPath PlanningTarget TrackingCybersecurityElectronic Warfare
Rashid et al., 2023 [1]PartiallyPartiallyPartially
Yu et al., 2025 [6]Partially
Hadi et al., 2023 [10]PartiallyPartially
Alsadie et al., 2025 [11]Partially
Oli et al., 2025 [12]
Ogab et al., 2025 [13]PartiallyPartially
Al-Syouf et al., 2024 [14]Partially
Islam et al., 2025 [15]Partially
Papathanasiou et al., 2026 [16]Partially
Kacem et al., 2026 [17]Partially
Skarka et al., 2024 [18]Partially
Fagundes-Junior et al., 2024 [19]PartiallyPartially
Meng et al., 2025 [20]Partially
Zhai et al., 2025 [21]Partially
Teixeira et al., 2023 [22]PartiallyPartiallyPartiallyPartially
Caballero-Martin et al., 2024 [23]PartiallyPartiallyPartiallyPartially
Yang et al., 2025 [24]PartiallyPartially
Wu et al., 2025 [25]PartiallyPartially
Costa et al., 2025 [26]PartiallyPartially
Al-Kamali et al., 2026 [27]PartiallyPartiallyPartially
This Paper
Table 2. Overview of AI-driven approaches for autonomous air combat maneuver decision-making.
Table 2. Overview of AI-driven approaches for autonomous air combat maneuver decision-making.
ReferenceObjectiveAI/ML ModelsOperational ScenarioUAV
Configuration
Main ComponentsKey Results
Zhu et al., 2024 [36]Autonomous air combat maneuver and engagement decision-makingMCLDPPO, distributed PPO, LSTM-based actor–critic DRLDigital twin-enabled air combat with BVR missile engagementsSingle combat UAVPOMDP formulation, motivational curriculum learning, hierarchical maneuver actions, missile threat interruption mechanism, distributed battlefield simulationAchieved approximately 66% win rate against expert-level FSM opponents, outperforming predictive game trees, SAC, DDPG, and DQN while demonstrating advanced tactical combat behaviors
Zheng et al., 2023 [37]Autonomous maneuver decision-making for short-range UAV air combatGRU-enhanced PPO, actor–critic DRLOne-versus-one short-range aerial combatSingle combat UAVGRU-based temporal feature extraction, phased curriculum training, dense/event/terminal rewards, adversarial opponent policy, probabilistic damage modelingImproved convergence, policy stability, temporal situational awareness, and combat effectiveness enabling adaptive offensive and evasive maneuver generation
Saldiran et al., 2024 [38]Explainable decision-making for RL-based air combat agentsDQN, Double DQN, XAI-enabled RLTwo-dimensional (2D) close-range aerial combatSingle combat UAVReward decomposition, ATA/AA/LoS-based analysis, tactical-region explainability, visual explanation mapsIdentified tactical strengths, weaknesses, asymmetries, and hidden behavioral patterns while improving transparency, interpretability, and debugging capability
Wang et al., 2024 [39]Implementation and evaluation of DRL-based air combat maneuver decision-makingDQN-based DRL frameworkThree-dimensional (3D) 1 vs. 1 UAV dogfightSingle combat UAVSix-Degrees-of-Freedom (DOF) UAV kinematics, NASA-inspired BFM actions, prioritized experience replay, ϵ -greedy exploration, curriculum learning, reward shapingDemonstrated autonomous target pursuit, attack positioning, maneuver adaptation, and engagement decision-making while providing a practical DRL implementation framework for intelligent air combat systems
Table 4. Overview of AI-driven approaches for large-scale swarm air combat and scalable combat intelligence.
Table 4. Overview of AI-driven approaches for large-scale swarm air combat and scalable combat intelligence.
ReferenceObjectiveAI/ML ModelsOperational ScenarioUAV ConfigurationMain ComponentsKey Results
Jiang et al., 2023 [62]Large-scale UAV swarm combat maneuver decision-makingMTVO, transformer-based PPOShort-range 3D UAV swarm air combat1 vs. 1–20 vs. 20 UAV swarm operational scenariosTransformer self-attention, virtual object global representation, mask mechanisms, PPO-based CTDE, attention aggregation, reward shaping, situation assessment, actor–critic RLAchieved win rates exceeding 70% in 5 vs. 5 combat while maintaining stable convergence; demonstrated robustness under numerical disadvantage and verified the importance of self-attention, attention aggregation, and virtual object modules through ablation studies
Kong et al., 2023 [63]Scalable multi-aircraft autonomous close-range air combatQMIX-based MARL, MHA, SMA, PCSPPartially observable multi-sensor 3D air combatVariable-size cooperative UCAV formationsCTDE, MHA-based information embedding, SMA, PCSP, GRU-based recurrent Q-networks, multi-sensor information fusion, competitive self-playAchieved winning rates exceeding 70% and the highest Elo scores among self-play methods; outperformed QMIX, VDN, and MeanPool while improving convergence stability, scalability, and combat effectiveness
Kim et al., 2024 [64]Cooperative RL for large-scale military drone operationsCommNet, MADRL, CTDELarge-scale dynamic battlefield environmentsCooperative military drone swarms based on Shahed-136 UAVsCommunication-aware coordination, hidden state information sharing, DEC-POMDP modeling, swarm flight maintenance, energy-aware bombing, realistic flight dynamics and energy consumption modelingOutperformed independent DQN, PPO, and MADDPG baselines in swarm coordination, bombing success rate, and residual energy preservation; maintained stable formations and improved energy efficiency over a battlefield of approximately 4 × 10 4 km 3
Li et al., 2026 [65]Autonomous cooperative roundup and UAV interceptionI2C-MATD3, ICEM, ICMObstacle-constrained cooperative interception environmentThree cooperative UAVs versus one adversarial UAVMATD3, ICEM, ICM, CTDE, replay buffer sharing, global-elite evolutionary optimization, LiDAR sensing, curiosity-driven explorationAchieved near-100% roundup success after approximately 1200 training episodes; outperformed MATD3 and IMTD3 while maintaining success rates above 71% under wind disturbances and approximately 70% under severe communication degradation
Table 5. Overview of AI-driven approaches for path planning and autonomous UAV navigation.
Table 5. Overview of AI-driven approaches for path planning and autonomous UAV navigation.
ReferenceObjectiveAI/ML ModelsOperational ScenarioUAV ConfigurationMain ComponentsKey Results
Soliman et al., 2023 [67]Path planning for autonomous mobile-target search and reconnaissancePPO-based DRL, MLP Actor-CriticReconnaissance and surveillance with mobile targetsParrot ANAFI UAVMDP formulation, digital twin simulation, QR-code target identification, energy-aware navigation, mobility-aware target trackingPerformance close to an ideal benchmark policy and superior to area-scanning baselines for up to 50 mobile targets and 64 grid cells
Shen et al., 2026 [68]Adaptive collision avoidance under uncertaintyHPER-D3QNJoint manned-unmanned operational airspaceUp to 25 heterogeneous aircraftDTPA, TCPA/DCPA threat assessment, wind-aware navigation, hierarchical PER, dual-layer safety zones, partial observationAchieved final reward of approximately 2.36 and collision-avoidance success rate of 96.3% while reducing FHP and task-completion time
Zhang et al., 2025 [69]Hybrid intelligent UAV path planning in adversarial environmentsCNNs, Bi-LSTM, attention-assisted Informed RRTDynamic obstacle-rich battlefield environmentsSingle UAV navigation systemGIS/radar/camera fusion, attention mechanism, adaptive informed sampling, DL-enhanced Informed RRT, Bézier smoothingAchieved approximately 91% success rate in narrow channels and 88% in trap-obstacle scenarios while reducing path length by 7.82% versus DDPG
Garcia-Gascon et al., 2025 [70]Autonomous navigation in GPS-denied environmentsLSTM, MLP-based DLGPS-denied waypoint-following missionsFixed-wing UAVTelemetry-based GPS prediction, INAV autopilot emulation, onboard sensor fusion, 49 input featuresApproximately 99% prediction accuracy with R 2 = 0.998 –0.999 and MAPE of 0.14–0.16% using 432,000 telemetry measurements
Sujecki et al., 2026 [71]Autonomous UAV navigation in contested environmentsPPO, REINFORCEGPS-denied and communication-constrained environmentsQuadrotor UAVUnity ML-Agents, PyTorch, continuous-control learning, reward shaping, 146-dimensional observation spacePPO achieved cumulative reward of 3.8658 vs. 1.5712 for REINFORCE and reduced average episode length from 200 to 110 steps
Chronis et al., 2023 [72]Autonomous navigation in unknown dynamic environmentsPPO-based DRLUrban and rural BVLOS missionsSingle UAV with low-cost sensorsAirSim/Unreal Engine, actor–critic navigation, obstacle avoidance, lightweight sensingApproximately 99% success rate with near-zero collision probability, outperforming A2C and A*
Lee et al., 2024 [73]Concealment-aware battlefield navigationQ-learning, Q-SAPFBattlefield terrain environmentsAutonomous military UAV50 × 50 battlefield maps, concealment-aware rewards, terrain-aware scoring, strategic potential fieldsMission success rates of 99.5–99.7%; path-length reduction of 1.96–8.77 units and execution-time reduction of 14.57–69.29 s
Kim et al., 2026 [74]Threat-aware tactical battlefield navigationSwiftFormer, DQN-based DRLBattlefield mobility and adversarial terrain environmentsAutonomous military UAVSemantic segmentation, mobility-aware terrain mapping, threat zone modeling, adaptive DQN decision-making, A* routing93 FPS segmentation, 93.0% pixel accuracy, 72.8% mIoU, zero enemy-zone exposure, and over 50-fold execution-time reduction
Zhan et al., 2024 [75]Path planning for large-scale maritime search-and-rescue missionsQ-learning-enhanced genetic algorithmMaritime SAR and emergency-response operationsMultiple UAVs and rescue centersAdaptive population management, heuristic initialization, adaptive crossover, elite repositories, population perturbationImproved convergence, workload balancing, and reduced mission duration for 100–1000 tasks, 2–10 rescue centers, and 80 km2 search regions
Table 8. Distribution of representative AI paradigms, architectural concepts, and enabling technologies across UAV-based defense domains.
Table 8. Distribution of representative AI paradigms, architectural concepts, and enabling technologies across UAV-based defense domains.
UAV-Based Defense DomainDRLMARLCNNsAttention/
Transformers
Radar-Based LearningFLXAIDigital TwinsMulti-Modal SensingComm.-Aware AI
Autonomous Air Combat and Cooperative UAV Operations
Path Planning and Autonomous Navigation
Target Tracking, Detection, and Classification
Cybersecurity, Electronic Warfare Protection, and Resilient UAV Operation
Overall/
Cross-Cutting
Notation: = High prevalence/adoption. = Medium prevalence/adoption. • = Low prevalence/adoption.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Michailidis, E.T.; Karanasiou, I.S. A Review of AI-Enabled UAV-Based Systems for Defense Applications. Drones 2026, 10, 602. https://doi.org/10.3390/drones10080602

AMA Style

Michailidis ET, Karanasiou IS. A Review of AI-Enabled UAV-Based Systems for Defense Applications. Drones. 2026; 10(8):602. https://doi.org/10.3390/drones10080602

Chicago/Turabian Style

Michailidis, Emmanouel T., and Irene S. Karanasiou. 2026. "A Review of AI-Enabled UAV-Based Systems for Defense Applications" Drones 10, no. 8: 602. https://doi.org/10.3390/drones10080602

APA Style

Michailidis, E. T., & Karanasiou, I. S. (2026). A Review of AI-Enabled UAV-Based Systems for Defense Applications. Drones, 10(8), 602. https://doi.org/10.3390/drones10080602

Article Metrics

Back to TopTop