Next Article in Journal
A Dual-Modal Mixture-of-Experts Attention U-Net (DMoE-AttU-Net) for Change Detection Using Heterogeneous Optical and SAR Remote Sensing Images
Next Article in Special Issue
CAF-Net: A Unified Framework for Resolving Spatial–Frequency Representation Conflicts in Multimodal Remote Sensing Segmentation
Previous Article in Journal
A Continuous Cryosphere Index for Snow and Ice Reflectance
Previous Article in Special Issue
CLEAR: A Cognitive LLM-Empowered Adaptive Restoration Framework for Robust Ship Detection in Complex Maritime Scenarios
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Review

Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution

1
Key Laboratory of Target Cognition and Application Technology, Aerospace Information Research Institute, Chinese Academy of Sciences, Beijing 100190, China
2
School of Electronic, Electrical and Communication Engineering, University of Chinese Academy of Sciences, Beijing 100094, China
3
School of Information and Communication Engineering, Dalian University of Technology, Dalian 116024, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(10), 1509; https://doi.org/10.3390/rs18101509
Submission received: 26 March 2026 / Revised: 1 May 2026 / Accepted: 6 May 2026 / Published: 11 May 2026

Highlights

What are the main findings?
  • A nine-dimensional comparative framework establishes UAV Embodied AI as a distinct research regime, constrained by 6-DoF dynamics, large-scale semantic sparsity, and strict onboard computing limits.
  • Direct migration of terrestrial Visual-Language Models (VLMs) and Large Language Models (LLMs) to aerial platforms yields severe performance degradation, primarily caused by perspective mismatch and the detachment of cognitive reasoning from physical flight feasibility.
What are the implications of the main findings?
  • Autonomous remote sensing must transition from open-loop data acquisition to active, closed-loop architectures that couple real-time perception with spatial-semantic decision-making.
  • Future aerial autonomy requires hybrid neuro-symbolic pipelines, wherein high-level semantic planning is explicitly constrained by physics-based feasibility filters and domain-specific simulation infrastructure.

Abstract

Unmanned Aerial Vehicle (UAV), particularly rotary-wing platforms such as quadcopters and octocopters, has evolved from controlled remote sensing platforms into autonomous agents capable of active task execution. This evolution from collect-then-analyze workflows to closed-loop perception, reasoning, and action signifies a paradigm shift toward Embodied AI, unlocking opportunities for the low-altitude economy. However, current research on UAV Embodied AI (UAV-EAI) often implicitly frames the field as a direct extension of indoor robotics or autonomous driving, which overlooks the fundamental distinctions of aerial agents. To bridge this gap, we introduce a comparative framework contrasting UAV-EAI with Indoor-EAI and Autonomous Driving Embodied AI (AD-EAI). By systematically decomposing the domain into nine key dimensions, we (i) analyze core tasks such as perception, localization, and exploration; (ii) review enabling infrastructure, including simulators and datasets; and (iii) categorize modeling methods ranging from physics-centric control to cognition-centric models. Our analysis demonstrates that the convergence of 6-DoF motion space, kilometer-scale unstructured environments, and stringent on-device constraints establishes a research regime qualitatively different from ground-based agents. These factors significantly impede the migration of existing VLM/LLM-based embodied systems for UAVs. Finally, we summarize open challenges and outline promising directions for the next generation of UAV-EAI.

1. Introduction

Over the last decade, UAVs have revolutionized the field of Remote Sensing (RS). Offering unprecedented flexibility, cost-effectiveness, and high resolution, UAVs play a vital role in precision agriculture, terrain mapping, disaster assessment, and urban management [1]. However, in the traditional UAV remote sensing (UAV-RS) paradigm, the UAV functions primarily as a passive sensor platform. Its core mission restricts it to data collection tasks, such as acquiring photos, 3D point clouds, or multispectral data along predefined flight paths. This data is subsequently processed offline for map generation, feature classification, or change detection. While highly effective for cartography, this sequential acquisition and offline analysis model fails to meet the growing demand for real-time, dynamic, and interactive tasks. Emerging applications increasingly require capabilities beyond passive data logging, demanding autonomous planning and action driven by task-oriented reasoning.
These evolving demands are driving a paradigm shift in remote sensing from passive perception to active sensing and physical interaction. Future intelligent systems require capabilities beyond static identification; they must actively determine subsequent observation targets and formulate optimal execution strategies based on real-time environmental understanding [2]. This objective aligns closely with the central goal of Embodied AI, which aims to develop agents capable of perceiving, reasoning, and acting within the physical world [3]. The recent success of Large Language Models (LLMs) and Visual-Language Model (VLMs) in demonstrating zero-shot reasoning and complex instruction-following has provided agents with generalized cognitive capabilities for parsing high-level human intent [4]. Integrating this cognitive capacity with the physical attributes of a UAV transforms the platform from a mere sensor carrier into an aerial embodied agent. Accordingly, we define UAV-EAI as an intelligent system, physically embodied as a UAV, that fuses multimodal remote sensing information in real-time to autonomously interpret high-level commands, conduct long-horizon planning, and execute physical tasks such as navigation, search, and inspection within complex, unstructured 3D environments. This represents a promising frontier in remote sensing, advancing the field from environmental understanding to environmental interaction.
It is important to note that the term UAV encompasses a diverse range of platforms, including fixed-wing, flapping-wing, and rotary-wing aircraft. However, in the context of Embodied AI, the ability to hover, perform Vertical Take-Off and Landing (VTOL), and execute agile 6-DoF maneuvers is a prerequisite for autonomous perception and physical interaction in unstructured environments. Fixed-wing UAVs, constrained by non-holonomic kinematics and the inability to hover, are generally less suitable for the complex, close-range interaction tasks discussed in this review. Therefore, unless otherwise specified, the term “UAV” in this paper primarily refers to multi-rotor Micro Aerial Vehicles (MAVs), such as quadcopters, hexacopters, and octocopters, which currently dominate the research landscape of aerial embodied intelligence.
Despite the rapid growth of UAV-EAI, existing research exhibits significant path dependency and cognitive gaps. On one hand, many studies treat UAV-EAI as a direct extension of traditional UAV remote sensing computer vision problems, prioritizing the accuracy of offline object detection or semantic segmentation benchmarks while neglecting the closed-loop challenges of real-time decision-making and control [5]. On the other hand, numerous studies attempt to directly transplant task paradigms from Indoor-EAI, such as Object Navigation (ObjectNav) or Visual-Language Navigation (VLN) [6,7], or adopt perception and planning techniques from AD-EAI. We contend that such direct adaptation is suboptimal and often insufficient. UAV-EAI constitutes an independent research field with fundamental differences from its counterparts. While AD-EAI benefits significantly from high-definition maps and strong structural constraints like roadways, and indoor AI leverages rich semantic information within controlled 2.5D motion spaces, UAV-EAI requires long-horizon, high-risk autonomous decision-making in vast, semantically sparse, and true 6-DoF spaces. Furthermore, these operations often rely on lightweight onboard sensors amidst uncontrollable dynamic disturbances such as wind shear. Overlooking these distinctions has led to fragmented research and obscured the core challenges specific to the aerial domain.
To define the research scope of UAV-EAI and identify its core challenges, we introduce a horizontal comparative framework contrasting UAV-EAI with Indoor-EAI and AD-EAI. The rationale for selecting autonomous driving and indoor embodied agents as comparative benchmarks for UAVs is to construct a multidimensional evaluation framework encompassing spatial degrees of freedom, interaction logic, and resource constraints. First, from the perspective of spatial evolution, autonomous driving represents a 2.5D structured environment characterized by ground-plane limitations and rigid rule-based constraints, while indoor robots focus on fine-grained, near-field interactions within semi-structured settings. UAVs, conversely, operate in fully unstructured 3D spaces with six degrees of freedom (6-DoF); this comparison facilitates an in-depth exploration of how increased spatial dimensionality impacts the robustness of path planning and state estimation. Second, regarding core functional logic, autonomous driving emphasizes social game-theoretic interactions and normative compliance, indoor agents prioritize precise physical manipulation of the environment, and UAVs center on agile obstacle avoidance and survival navigation in highly dynamic contexts. Such a contrast allows for a granular analysis of the commonalities and disparities across ‘social intelligence,’ ‘physical contact,’ and ‘agile locomotion’ within embodied algorithmic architectures. Finally, from an engineering standpoint of hardware resource allocation, these three domains exhibit a distinct gradient in computational payload, power tolerance, and real-time requirements, ranging from the high-performance computing redundancy in autonomous vehicles to the extreme edge-computing constraints of UAVs. In summary, the inclusion of these two comparative agents provides a comprehensive mapping of the performance boundaries of embodied AI across diverse physical platforms and task scenarios, thereby rigorously validating the unique technical challenges and utilitarian value of UAVs in complex 3D environments. To the best of our knowledge, this is the first review to systematically contrast these three domains from an embodied perspective. Our contributions are summarized as follows:
  • We propose a nine-dimension analytical framework covering motion space, environmental structure, information density, and energy constraints. This framework qualitatively and quantitatively elucidates the essential differences and unique challenges of UAV-EAI compared to ground-based Embodied.
  • Guided by this framework, we review the core tasks, critical infrastructure, and modeling methods of UAV-EAI, highlighting the limitations of the existing toolchain.
  • Our analysis pinpoints key bottlenecks and future directions of existing UAV-EAI, specifically identifying the root causes of performance degradation in large-model reasoning and the unique Sim-to-Real gap of this emerging field.
As illustrated in Figure 1, the remainder of this review is organized as follows: Section 2 details our analytical framework. Section 3 reviews the core tasks of UAV-EAI under this framework. Section 4 evaluates key infrastructure, including simulators and datasets. Section 5 analyzes mainstream modeling paradigms and their applicability to the UAV domain. Section 6 discusses primary applications and core challenges. Finally, Section 7 concludes the paper and provides a future outlook.

2. Defining UAV Embodied AI: A Comparative Framework

In the introduction, we position UAV-EAI as the frontier of remote sensing, evolving from passive understanding to active interaction. However, academic research exploring this domain often falls into path dependency, attempting to directly migrate methodologies from two mature, related fields: Indoor Embodied AI (Indoor-EAI) and Autonomous Driving (AD-EAI). This direct transfer overlooks fundamental differences in embodiment, operational environment, and task objectives, leading to fragmented and siloed research.
Indoor Embodied AI platforms, including Habitat [2] and AI2-THOR [8], train agents to navigate semantically dense and constrained spaces. The core focus is translating natural language instructions into precise physical actions, such as picking up, placing, or manipulating objects [9]. The success of these algorithms relies heavily on a deep grounding in human common sense and environmental affordances.
Conversely, the field of AD-EAI [10] is dedicated to solving perception and decision-making problems in a quasi-2D space, governed by strong rules (traffic laws) and characterized by highly dynamic agents (traffic, pedestrians). The robustness and safety of AD-EAI systems are built upon heavy sensor redundancy (LiDAR, radar, high-fidelity cameras) and abundant prior information from maps [11,12].
Traditional UAV research [13,14] primarily addresses state estimation—such as Visual-Inertial Odometry (VIO) or Simultaneous Localization and Mapping (SLAM)—and obstacle avoidance within GPS-denied environments. These efforts prioritize localized precision and platform stability, often marginalizing the cognitive requirements of high-level task execution. Consequently, established frameworks do not adequately resolve the specific requirements of UAV-EAI, where agents must navigate expansive, unstructured 3D spaces while managing complex aerodynamic perturbations. Furthermore, these systems must operate under severe energy and sensing constraints while maintaining long-horizon reasoning within semantically sparse environments.
Therefore, this paper aims to delineate the research scope of UAV-EAI, identify its core bottlenecks, and support future algorithm design by introducing a systematic nine-dimension comparative framework that contrasts UAV-EAI with its ground-based counterparts. The framework, summarized in Table 1, captures the essential differences across the physical, environmental, and cognitive domains.

2.1. Framework Dimensionality Analysis

The framework presented in Table 1 characterizes the unique and composite challenge profile of UAV-EAI. Based on this analysis, the nine dimensions can be organized into three fundamental layers of differentiation that structure the comparison across domains. Despite these dimensional differences, UAV-EAI shares the fundamental Perception-Decision-Action cognitive architecture with Indoor-EAI and AD-EAI. This structural commonality allows UAVs to inherit mature methodological paradigms, such as hierarchical planning and end-to-end learning policies, as the foundational baseline for aerial autonomy.

2.1.1. Spatial and Dynamic Differences

The first core challenge of UAV-EAI stems from its true 3D (6-DoF) motion combined with critical energy constraints. While an indoor robot can perform complex manipulations in 2.5D space [15], its motion is ground-based and generally does not risk mission failure via energy depletion. AD-EAI, though high-speed, is confined to a quasi-2D plane defined by the road network. In contrast, a UAV must not only plan trajectories in full 6-DoF space but also actively counteract complex, unpredictable aerodynamic disturbances [16].

2.1.2. Environmental and Perceptual Differences

The second core challenge arises from UAV-EAI’s extreme reliance on lightweight perception in unknown, open, large-scale environments. The AD perception stack is “heavy and redundant” [17] and its decision-making relies heavily on strong priors from GPS and high-definition Maps [18]. Indoor-EAI benefits from rich geometric data via RGB-D sensors. The payload-constrained UAV, however, must rely primarily on vision + IMU (i.e., VIO/SLAM) for real-time localization and mapping in large-scale environments that are both unknown (mapless) and signal-denied (GPS-less) [19], all while mitigating natural disturbances [20].

2.1.3. Cognitive and Task Differences

The primary challenge, and a central theme of this review, involves the divergence in cognitive environmental structures. While Vision-Language Models (VLMs) and Large Language Models (LLMs) perform effectively in Indoor-EAI due to the “information-dense and semantically-logical” nature of those environments [4], they rely on human commonsense priors that are often absent in aerial contexts. Similarly, Autonomous Driving (AD) environments are “information-dense and rule-constrained,” providing structured cues such as traffic regulations [21]. In contrast, typical UAV-EAI operating theaters—such as expansive forests or repetitive mountainous terrain—are inherently semantically sparse and weakly structured. This sparsity frequently invalidates the commonsense reasoning capabilities of VLMs, which are predominantly trained on ground-based datasets [22]. Consequently, the lack of discernible inter-object relationships compromises task reasoning and self-localization, leading to significant cognitive degradation and rendering direct model migration ineffective.
The Target-to-FOV Ratio also makes a difference. In Indoor-EAI datasets such as AI2-THOR [8], the primary interaction targets often occupy 10% to 30% of the egocentric visual field due to close-range engagement. Similarly, AD-EAI datasets typically present target instances (vehicles, pedestrians) that occupy a moderate 1% to 10% of the FOV, supported by structured geometric priors. Instead, in UAV datasets, targets often undergo extreme scale degradation, yielding a Target-to-FOV ratio of less than 1% [23]. Furthermore, the high-risk profile and specialized nature of aerial missions necessitate a distinct task definition, where the requirement for deep spatial and semantic understanding results in elevated task complexity [24].

2.1.4. Challenges in Terrestrial Paradigm Migration

Despite shared algorithmic roots, several core methodologies from Indoor-EAI and AD-EAI encounter fundamental structural barriers when migrated to the UAV domain:
  • From Indoor-EAI to UAV-EAI: Terrestrial embodied models predominantly rely on object-centric affordance learning, which assumes canonical, human-eye-level viewpoints and stable object scales. In UAV-EAI, the nadir/oblique perspective and kilometer-scale operations lead to extreme scale collapse and perspective distortion, rendering indoor-trained semantic priors and interaction logic ineffective. Furthermore, the deterministic room-scale mapping typical of indoor agents fails in boundless outdoor environments, where metric state estimation must prioritize probabilistic robustness over rigid geometric reconstruction [25].
  • From AD-EAI to UAV-EAI: Autonomous driving systems are fundamentally predicated on HD-map-dependent predictive control and 2.5D road-based navigation rules. These priors are non-existent in roadless, true 3D aerial spaces. Moreover, the sensor-redundant decision-making paradigm—utilizing heavy LiDAR stacks and high-throughput compute units—is physically prohibited by the stringent SWaP constraints of UAVs. Consequently, UAV-EAI requires a transition from “compute-heavy redundancy” to “resource-aware efficiency,” where agents must achieve comparable safety using sparse, lightweight sensory streams [21,26].

2.2. Methodological Foundations of the UAV-EAI Paradigm

To transition from descriptive constraints to a rigorous research regime, UAV-EAI must be defined by distinct modeling and learning principles that address its fundamental divergence from ground-based agents. We identify two core pillars that constitute this independent paradigm:
  • Asynchronous Physical-Cognitive Decoupling: Unlike the monolithic, synchronous perception-action pipelines common in Indoor-EAI, the stringent Size, Weight, and Power (SWaP) limits of UAVs necessitate a structural decoupling of control and cognition. This principle mandates an asymmetric architecture where a high-frequency reactive loop guarantees flight stability, while a compute-intensive cognitive loop is triggered only upon detecting semantic uncertainty or task transitions. This shift from continuous to event-triggered inference is a methodological prerequisite for aerial embodiment.
  • Asymmetric Privileged Distillation: Stochastic aerodynamic disturbances create a pronounced dynamics gap in aerial environments, frequently invalidating the direct policy transfer methods established in terrestrial robotics. To address this, UAV-EAI architectures increasingly rely on asymmetric distillation frameworks. Policies are initially trained in simulation using exact, privileged environmental states, such as precise force vectors and global geometry. These high-fidelity representations are subsequently distilled into lightweight neural networks constrained to operate exclusively on noisy, partial onboard sensor data. By structurally integrating physical priors during the training phase, this method significantly reduces the sim-to-real discrepancy, enabling reliable autonomous execution in unmodeled physical environments [27,28].

3. Analysis of Core EAI Tasks

While Section 2 outlines the unique environmental and physical constraints of aerial agents, this section addresses their functional core. Traditional UAV-RS fundamentally relies on an open-loop, “collect-then-analyze” paradigm. The platform acts merely as a passive sensor carrier executing predefined flight paths, a model that is insufficient for dynamic missions requiring real-time reasoning and intervention.
In contrast, UAV-EAI operates on a closed-loop “perception-decision-action” cycle. Perception is no longer just logging a static data stream, but an active process dynamically driven by the agent’s spatial awareness and task objectives. To bridge the gap between algorithmic design and real-world deployment, we restructure the core tasks of UAV-EAI into three categories: (1) Autonomous Perception and Semantic Mapping, (2) Embodied Navigation and Instruction Following, and (3) Generalized Exploration. Throughout this section, we explicitly map these capabilities to their specific downstream applications, demonstrating how embodied tasks translate into practical utility.

3.1. Autonomous Perception and Semantic Mapping

Perception in UAV-EAI extends beyond the traditional remote sensing paradigm of passively recognizing and labeling objects in a static image. Instead, it requires the agent to actively construct a metrically consistent and semantically rich representation of its surroundings, using its own motion to resolve ambiguity.

3.1.1. Viewpoint Selection and Next-Best-View

In complex 3D environments, targets are frequently obscured by multi-scale occlusions, ranging from massive topographical features and urban canyons to porous forest canopies [29]. Traditional UAV workflows rely on greedy coverage or uniform random exploration. These methods are highly inefficient because they treat all spatial regions with equal importance, often leading to excessive energy expenditure in uninformative areas [30]. This creates a severe spatial imbalance, often characterized as a “needle-in-a-haystack” scenario where sparse targets are buried within weakly structured background data.
To overcome this, UAV-EAI fundamentally treats perception not as a passive data stream, but as an active, information-gathering query [31]. Through Next-Best-View planning, the agent dynamically determines “where to look next” to maximize the probability of target detection, minimize the variance of spatial state estimates, and reduce ambiguity in high-dimensional spaces [32]. When visual confidence is low due to environmental noise or poor viewing angles, uncertainty-aware perception modules distinguish between “absence of evidence” and “evidence of absence,” triggering active disoccluding maneuvers [33]. These autonomous behaviors include ascending to a nadir view to clear obstacles, performing lateral motions to induce parallax, or executing full circumnavigations to reconstruct a 3D volume.
Critically, aerial Next Best View (NBV) planning distinguishes itself from ground-based robotics by operating under strict dynamic and resource limitations. The planning algorithm cannot treat the UAV as a holonomic point mass; it must account for non-negligible inertia and the aerodynamic costs of maneuvers [34]. The optimization of the next viewpoint must mathematically balance the expected information gain—often quantified via Mutual Information (MI) or Fisher Information—against the battery expenditure required to reach that physical pose [35]. Consequently, modern UAV-EAI approaches formulate this active sensing problem as a Partially Observable Markov Decision Process (POMDP), where policies are optimized over belief spaces (representing joint probability distributions of target locations and occupancy) rather than deterministic geometric states [36]. This enables the UAV to proactively navigate physical trade-offs while continuously resolving its internal uncertainty.

3.1.2. Metric-Semantic Mapping

Traditional UAV perception relies heavily on Visual-Inertial Odometry and Simultaneous Localization and Mapping primarily to solve geometric state estimation and maintain flight stability in GPS-denied settings [14]. However, existing mapping benchmarks and traditional SLAM pipelines focus almost exclusively on geometric accuracy, lacking the high-level semantic annotations required for agents to comprehend abstract instructions such as “search for diseased areas in farmland” [37]. While pure metric maps are sufficient for basic obstacle avoidance and stabilization, they fall short in providing the semantic context necessary for embodied reasoning and long-horizon task planning.
Unlike indoor environments that feature fixed structures and provide abundant state constraints and well-defined semantic information [38], UAVs operate in large-scale, open environments characterized by substantially low scene density, sparse semantics, and weak inter-object relationships. In such settings, extracting critical information from vast, spatially discontinuous scenes significantly increases the difficulty of perception due to the absence of explicit semantic cues. Consequently, intelligent tasks for UAVs tend to exhibit higher semantic complexity, necessitating advanced semantic reasoning to assist in deep spatial understanding [39].
To achieve true embodiment, UAVs must move beyond geometric reconstructions and construct metric-semantic maps that continuously fuse 6-DoF pose estimates with high-level object and scene information. Recent advances in metric-semantic mapping systems enable UAVs to explicitly associate linguistic concepts with 3D spatial representations, which is critical for supporting downstream navigation and inspection tasks that require semantic grounding in large-scale environments [40]. This integration transforms a static point cloud into a dynamic, queryable spatial memory. By leveraging this semantic awareness, an embodied agent can, for example, proactively revisit landmark-rich regions to re-anchor its spatial map when localization stability degrades in feature-depleted environments, effectively prioritizing self-localization to ensure long-term systemic robustness [41].

3.2. Embodied Navigation and Instruction Following

While traditional UAV navigation relies on explicit geometric waypoints or GPS coordinates, embodied navigation requires the agent to interpret abstract, high-level commands and autonomously navigate towards semantic goals in unknown environments. This shifts the control paradigm from low-level direct control to semantic reasoning.

3.2.1. Vision-Language Navigation (VLN)

Vision-Language Navigation tasks agents with interpreting natural language directives to execute trajectories within continuous, physically realistic 3D environments. Distinct from traditional navigation frameworks that rely on explicit geometric waypoints, VLN necessitates the semantic grounding of linguistic cues directly into visual observations [6]. This shifts the navigational dependency from global positioning coordinates to local scene understanding and instruction following.
The adaptation of VLN from ground-based robotics to the aerial domain is complicated by a fundamental perspective mismatch. Natural language instructions are typically framed from a human-centric, ground-level viewpoint, where landmarks are characterized by their lateral profiles. In contrast, UAVs capture environmental data from high-altitude, nadir, or oblique angles [42]. This viewpoint divergence results in significant geometric distortions and distinct occlusion patterns, which frequently degrade the performance of conventional feature matching techniques utilized for landmark identification.
Aerial VLN occurs in unstructured, multi-scale outdoor environments, contrasting with the constrained, room-scale settings typical of indoor robotics. UAVs are required to isolate fine-grained semantic targets—such as specific vehicles or vegetation—within expansive, cluttered backgrounds, which introduces substantial visual ambiguity. To address these complexities, current research utilizes Vision-Language Models (VLMs) to develop robust cross-view representations. This approach facilitates the alignment of egocentric linguistic instructions with allocentric aerial perception, thereby enhancing model generalization across diverse viewpoints [43].

3.2.2. Object-Goal Navigation (ObjectNav)

ObjectNav entails navigating within unmapped environments to locate specific object instances without prior coordinate knowledge. Distinct from PointGoal navigation, which assumes known target coordinates, ObjectNav compels the agent to operate without a pre-built map, integrating autonomous exploration with real-time semantic recognition [7].
Within the domain of UAV-EAI, this task diverges sharply from indoor robotics. While indoor agents exploit structural semantic priors to guide search, aerial agents must navigate unstructured environments characterized by extreme scale and sparse features [44]. A critical challenge here is the scale imbalance: targets in aerial scenarios frequently occupy a negligible fraction of the sensor’s field of view relative to the kilometer-scale search space. Consequently, traditional geometric frontier-based exploration or greedy coverage strategies prove inefficient, as they assign uniform utility to the entire unmapped space.
Instead, robust aerial ObjectNav frames navigation as a semantic search problem. The agent must leverage contextual cues—such as following trails to locate hikers or tracking power lines to find pylons—to effectively prune the search space [45]. Success is thus measured not merely by trajectory length, but by the efficiency of semantic discovery, requiring policies to balance the exploitation of high-probability regions against the exploration of unknown sectors under strict energy constraints [46].
The transition from purely geometric navigation toward Vision-Language Navigation (VLN) and Vision-Language-Action (VLA) architectures has emerged as a primary focus in UAV autonomy. Given the 6-DoF dynamics and severe SWaP constraints inherent to aerial platforms, the direct application of terrestrial VLN models frequently yields suboptimal performance due to perspective mismatch and computational latency. To address this domain specific disparity, recent studies propose architectures explicitly optimized for aerial embodiment. For instance, the UAV-VLN framework [47] formulates an end-to-end VLN pipeline tailored for unconstrained oblique perspectives. In domain-specific applications such as urban air mobility, LogisticsVLN [48] utilizes multi-modal large language models (MLLMs) to guide autonomous terminal delivery, demonstrating the feasibility of language-conditioned control in low-altitude flight.
Beyond pure navigation, the extension into complex physical interaction is facilitated by emerging Vision-Language-Action (VLA) models. Recent empirical evaluations, such as DroneVLA [49], validate the integration of language priors into aerial manipulation tasks. Furthermore, the CognitiveDrone architecture [50] establishes a hierarchical VLA paradigm where a VLM reasoning module resolves textual ambiguity prior to generating real-time 4D action commands. This explicit structural decoupling enhances human-target recognition and spatial reasoning robustness during real-world deployments, marking a critical methodological shift from reactive obstacle avoidance to semantic-aware aerial task execution.

3.3. Generalized Exploration and Collaborative Interaction

3.3.1. Informative Path Planning

Informative Path Planning fundamentally shifts the exploration objective from exhaustive geometric coverage to the maximization of information gain subject to stringent resource constraints [51]. Unlike traditional Coverage Path Planning, which implicitly assumes a uniform distribution of interest across the domain, IPP formulates exploration as an optimization problem over a heterogeneous information field. The agent must continuously evaluate the trade-off between the expected reduction in map uncertainty (quantified via metrics such as Shannon entropy or Mutual Information) and the kinematic cost of trajectory execution [35].
In the context of UAV-EAI, IPP is distinct from ground-based exploration due to the tight coupling of 3D non-holonomic dynamics with critical energy limitations. The planning algorithm cannot abstract the UAV as a holonomic point mass; rather, it must account for aerodynamic drag and inertial constraints, ensuring that the potential informational value of a viewpoint justifies the specific energetic expenditure required to reach it [52]. To address the prohibitive computational cost of calculating volumetric mutual information online, recent methodologies increasingly leverage learned surrogate models or hierarchical planners. These approaches approximate information gain from raw sensor inputs, enabling real-time, horizon-limited optimization directly on resource-constrained embedded hardware [53].

3.3.2. Multi-Agent Coordination

The expansion of UAV-EAI to multi-agent systems is driven by the necessity for spatiotemporal efficiency and system-level robustness in large-scale missions. Unlike traditional remote sensing fleets that follow pre-separated flight corridors, embodied swarms require active, decentralized coordination to optimize collective perception and decision-making [54].
A critical advancement in this domain is the exploitation of heterogeneity. Effective exploration strategies often leverage a hierarchical fleet structure, where high-altitude UAVs generate coarse global occupancy maps to guide low-altitude UAVs toward regions of interest for fine-grained verification, as shown in Figure 2 [55]. This functional decomposition allows the swarm to optimize the trade-off between search breadth and inspection depth, a capability unavailable to homogeneous systems.
Furthermore, embodied coordination must rigorously account for communication constraints. In field environments, bandwidth is limited and network topology is time-varying. Consequently, agents cannot rely on sharing raw sensor streams; instead, coordination is often modeled through Decentralized Partially Observable Markov Decision Processes (Dec-POMDPs) [56]. In this framework, agents exchange compressed belief maps or latent intentions to maintain consensus, ensuring that the collective behavior remains coherent even under intermittent connectivity or individual node failure [57].

3.4. Summary

In summary, the core tasks analyzed in this section collectively define the functional scope of UAV-EAI. Unlike traditional remote sensing, which treats data acquisition and offline analysis as distinct, sequential stages, UAV-EAI necessitates a unified cognitive architecture where perception actively drives decision-making, and action continuously refines understanding. Table 2 summarizes the fundamental paradigm shifts across these core tasks.

4. Enabling Infrastructure: Simulators and Datasets

Research in UAV-EAI critically depends on simulation platforms that capture the coupled perception–action–dynamics loop characteristic of embodied systems. Unlike the development of Indoor-EAI, which benefited early from standardized interactive environments, the evolution of UAV simulation has followed a unique trajectory driven by the changing demands of aerial tasks—from basic control stability to visual perception, and finally to cognitive embodiment. In this section, we review the evolution of UAV simulation environments through three developmental phases, analyzing how infrastructure has progressed to meet the challenges of 6-DoF dynamics, large-scale unstructured environments, and data-driven policy learning.

4.1. Simulation Environments

We re-examine existing simulation platforms through an evolutionary lens, and categorize the progression of UAV simulation into three distinct phases, as shown in Figure 3: (1) the era of Physics-Centric Control, focused on dynamics and kinematics; (2) the era of Photorealistic Perception, driven by the rise of deep vision; and (3) the emerging era of Parallel Embodied Intelligence, characterized by high-throughput physical-cognitive interaction. This chronological analysis highlights how the limitations of each generation—specifically regarding visual fidelity, aerodynamic realism, and simulation throughput—have catalyzed the development of next-generation tools tailored for the challenges of UAV-EAI.

4.1.1. The Era of Physics-Centric Control and General-Purpose Robotics

Early work in autonomous aerial robotics primarily tackled fundamental challenges like low-level flight stability, trajectory tracking, and kinematics. To address these, researchers naturally turned to general-purpose simulators such as Gazebo [58] and Bullet Physics [59]. Built on reliable rigid-body physics engines like ODE and Bullet and offering easy plugin integration, these platforms became the standard environments for validating classical controllers (like Proportional-Integral-Derivative Control (PID) and Linear Quadratic Regulator (LQR)) and verifying kinematic designs [60].
While these tools provided the necessary “muscles” for early flight controllers, their architecture exhibits fundamental limitations when applied to the “cognitive” demands of modern UAV-EAI:
  • Deficiency in Visual Realism at Scale: Designed primarily for indoor laboratories or small-scale manipulation tasks, Phase I simulators lack the rendering pipelines required for long-range aerial perception. Constructing the vast, unstructured outdoor environments typical of UAV missions requires manually importing heavy meshes, which often leads to severe performance degradation and lacks the procedural generation capabilities needed for diverse terrain training [61].
  • Oversimplification of Aerodynamics: Although excellent at rigid-body dynamics, these platforms often abstract UAVs as simple floating masses, failing to model critical non-linear aerodynamic phenomena such as rotor-airflow interaction and stochastic wind turbulence. This omission creates a significant “Sim-to-Real” gap for embodied agents, as policies trained in these vacuums often fail to compensate for the complex disturbances characteristic of real-world flight [62].
  • Computational Bottlenecks for Learning: The architecture of Gazebo and similar tools is primarily CPU-bound, optimized for high-precision, single-instance simulation rather than massive parallelism. This presents a prohibitive barrier for Deep Reinforcement Learning (DRL) and active perception tasks, which require millions of interaction frames. Without native GPU parallelization, these simulators cannot provide the throughput necessary to train complex neural policies within a reasonable timeframe [63].
Thus, while indispensable for the era of control theory, Phase I simulators proved insufficient for the data-driven requirements of the subsequent deep learning revolution, necessitating the shift toward perceptually realistic and computationally scalable environments.

4.1.2. The Era of Photorealistic Perception

As deep learning began to dominate perception tasks, researchers quickly realized that early platforms like Gazebo lacked the visual detail needed to train robust computer vision models. Because their simplistic lighting and low-resolution textures struggled to support tasks like object detection and semantic segmentation, transferring these models to the real world proved difficult. This challenge drove a major transition in simulation design: rather than focusing solely on kinematics, the emphasis shifted to photorealism. To generate the rich visual data required, the community started bringing advanced game engines—most notably Unreal Engine and Unity—directly into their simulation pipelines:
  • Microsoft AirSim [61] emerged as a seminal platform in this era, leveraging Unreal Engine to provide varying weather conditions, realistic lighting, and detailed textures. Unlike its predecessors, AirSim supported multi-modal sensor simulation—including depth, segmentation masks, and LiDAR—enabling the end-to-end training of perception-driven policies for autonomous flight in unstructured outdoor environments.
  • FlightGoggles [64] offered a complementary approach by focusing on the tight coupling of visual realism with high-frequency inertial dynamics. By rendering photorealistic camera streams synchronized with rigorous vehicle dynamics, it facilitated the benchmarking of VIO and aggressive flight algorithms in virtual environments that closely mimicked real-world challenges, such as motion blur and dynamic lighting changes.
During the same period, platforms like Habitat [2] and AI2-THOR [8] emerged to support indoor Embodied AI research. While these environments provided fast rendering and rich semantic data for indoor navigation, their scope remained limited to relatively static, 2.5D layouts. Crucially for UAV research, they did not support 6-DoF flight dynamics or offer the expansive 3D workspaces and complex aerodynamic modeling—such as wind shear and ground effect—essential for aerial agents. Because of these structural limitations, the community had to build dedicated UAV simulators capable of handling both large-scale outdoor environments and sophisticated aerodynamics.
Despite the visual advancements, the simulators above faced significant computational hurdles. The reliance on heavy rendering engines meant that simulating a single agent often consumed substantial GPU resources, limiting the simulation speed to near real-time. This low throughput became a prohibitive barrier for Reinforcement Learning (RL) approaches, which require billions of interaction steps [65]. The inability to parallelize environments efficiently hindered the scalability of learning-based methods, paving the way for the next generation of GPU-accelerated, parallel simulation platforms.

4.1.3. The Era of Parallel Embodied Intelligence

The current frontier of UAV-EAI research is defined by a fundamental paradigm shift from high-fidelity visualization to high-throughput parallel interaction. While the photorealistic simulators discussed in the previous section successfully bridged the visual reality gap, their CPU-bound architectures imposed severe bottlenecks on data generation efficiency. Deep Reinforcement Learning (DRL) and evolutionary strategies typically require billions of interaction steps to converge on robust policies—a scale computationally prohibitive for serial, single-agent simulations [66]. To address the dual challenges of sample efficiency and physical fidelity, the field has transitioned toward GPU-accelerated, massively parallel simulation platforms. These environments leverage unified memory architectures to keep physics simulation, sensor rendering, and neural network training entirely within the GPU (VRAM), thereby eliminating the latency-heavy bandwidth bottlenecks associated with traditional CPU-GPU data transfer. This era is epitomized by the following parallel simulation ecosystems:
  • NVIDIA Isaac Sim and Isaac Gym: Running on the Omniverse platform, this environment pairs high-fidelity RTX rendering with GPU-accelerated PhysX 5. While older platforms like Gazebo or AirSim typically process environments serially, the tensor-based API in Isaac Gym enables the simultaneous execution of thousands of parallel environments on a single workstation. This capability allows researchers to rapidly randomize physical parameters, such as friction, mass, and lighting conditions, reducing the time required for domain randomization from weeks to mere minutes and generating the diverse data necessary to train reliable control policies.
  • Specialized Aerial Frameworks (Pegasus and OmniDrones): To bridge the gap between general-purpose physics and the specific aerodynamic requirements of UAVs, domain-specific extensions such as the Pegasus Simulator [67] and OmniDrones [68] have been developed within the Isaac ecosystem. These frameworks introduce modular vehicle backends and realistic aerodynamic plugins—modeling complex phenomena such as rotor drag, ground effect, and wind turbulence—thereby enabling the end-to-end training of aggressive flight maneuvers and multi-agent coordination that generic physics engines cannot support.
In summary, the transition toward massive parallelism represents the maturation of simulation from a verification tool into a primary data generator for embodied intelligence. By decoupling simulation speed from real-time constraints and introducing physically rigorous interaction models, these GPU-accelerated platforms have fundamentally altered the research workflow. They not only solve the problem of data scarcity in DRL but also provide the necessary infrastructure to narrow the Sim-to-Real gap, allowing policies trained in simulation to transfer zero-shot or few-shot to physical UAVs operating in complex, unstructured worlds [69].

4.1.4. The Multi-Faceted Roles of Simulation in UAV-EAI: From Offline Infrastructure to Online Cognition

The evolution of simulation through the three eras—from physics-centric control to photorealistic rendering, and finally to parallel embodied intelligence—highlights a persistent historical trend: the reliance on simulators primarily as offline infrastructure for policy training or perception validation. However, as UAV-EAI matures, this offline-only perspective proves fundamentally insufficient. Because UAV-EAI emphasizes real-time closed-loop decision-making in high-risk, 6-DoF environments, agents cannot rely on physical trial-and-error. Consequently, the role of simulation must evolve from a static training ground into a dynamic, structurally integrated component of the agent’s cognitive loop. We systematically categorize the necessity of simulation in UAV-EAI into three distinct functional roles:
  • Offline Policy Generation and Sim-to-Real Transfer: Corresponding to the current Era of Parallel Embodied Intelligence, modern simulators serve as the foundational data engines that provide the billions of interaction frames required for Deep Reinforcement Learning convergence. By leveraging massive GPU parallelization and extensive domain randomization, these platforms effectively bridge the reality gap, ensuring robust policy behavior prior to physical deployment [68].
  • Online Predictive Engine: During real-time execution, simulation must function as a forward-predictive world model. When an embodied agent formulates a long-horizon plan via Next-Best-View optimization or Informative Path Planning, it utilizes a lightweight, embedded simulator to predict the future physical states and aerodynamic costs of proposed trajectories. This predictive capability is critical because high-level semantic objectives must be continuously balanced against non-holonomic kinematic constraints and battery expenditure—complex trade-offs that static heuristics fail to capture [34].
  • Runtime Safety Verification: As will be further elaborated in the context of Large Language Models, high-level semantic planners frequently exhibit “feasibility hallucinations.” In this capacity, simulation serves as a strict runtime safety sandbox. Before any semantically generated way-point is transmitted to the low-level execution controller, it is rapidly evaluated within a localized simulation layer to verify aerodynamic stability, obstacle clearance, and Control Barrier Function compliance [44].
Integrating online simulation inherently introduces additional computational complexity, posing a significant challenge for SWaP-constrained aerial platforms. Nevertheless, this complexity is an unavoidable necessity; without an internal predictive model, a UAV cannot safely evaluate the physical consequences of its cognitive decisions in unstructured 3D space. To resolve the tension between real-time latency requirements and computational limits, future UAV-EAI architectures must transition from deploying monolithic physics engines onboard toward leveraging lightweight surrogate models—such as Physics-Informed Neural Networks or learned world models [70]. These surrogates are capable of executing high-frequency forward simulations directly on edge hardware, thereby enabling true physical-cognitive integration without overwhelming system resources.

4.2. Benchmark Datasets

UAV-EAI requires datasets that capture large-scale outdoor environments, viewpoint diversity, semantic sparsity, and embodied task structure. Existing datasets can be organized into three categories.

4.2.1. Remote Sensing and Mapping Datasets

Remote sensing and mapping datasets constitute an important data foundation for UAV-related perception research, particularly in tasks such as object detection, semantic segmentation, and large-scale scene understanding. In the context of UAV-EAI, however, their role is fundamentally limited by the passive nature of data collection and the absence of an explicit perception–action loop. Most existing datasets are designed to support offline visual recognition or mapping benchmarks, rather than closed-loop embodied tasks that require agents to reason, decide, and act based on their own motion and observations.
From an embodied intelligence perspective, these datasets are typically characterized by large-scale spatial coverage and high-resolution aerial imagery, but lack temporal continuity and action-conditioned structure. The data are usually collected along predefined flight trajectories, with annotations provided at the frame or image level, while the causal relationship between agent actions and sensory feedback is not modeled. As a result, such datasets are well suited for evaluating passive perception capabilities, yet insufficient for studying navigation, exploration, or task-driven decision-making in UAV-EAI.
Representative datasets in this category focus primarily on urban or land-cover understanding from aerial viewpoints. For instance, UAVid [71] provides high-resolution UAV imagery with pixel-level semantic annotations for urban scene segmentation, enabling fine-grained classification of buildings, roads, vegetation, and vehicles. VisDrone [72] targets object detection and multi-object tracking in complex urban environments, offering diverse viewpoints and dynamic scenes captured from UAV platforms. While these datasets have significantly advanced visual perception research in UAV remote sensing, they generally omit ego-motion states, long-horizon trajectories, and goal-conditioned task definitions that are essential for embodied intelligence.
Beyond UAVid and VisDrone, numerous aerial semantic segmentation datasets provide high-resolution top-down imagery for land-cover classification and scene understanding. These datasets are widely used in remote sensing applications such as urban planning, environmental monitoring, and disaster management. However, compared with conventional natural-image semantic segmentation datasets, remote sensing imagery presents unique challenges, including large variations in object scale, high intra-class variability, low inter-class separability, frequent interference from shadows and occlusions, and stringent requirements for precise boundary delineation. To address these challenges, researchers have proposed various innovative approaches, such as multi-scale strategies and feature fusion, the adoption of Transformer-based models, self-supervised and few-shot learning techniques, and multi-source data fusion.
Despite providing rich data and benchmarks for semantic segmentation, detection, and tracking, these datasets exhibit several common limitations when considered from the perspective of embodied applications.
  • Lack of ego-motion information: Rather than capturing the dynamic states of the sensor platform, these datasets typically supply only static images or raw video frames. They omit critical kinematic details, including the UAV’s position, velocity, and orientation. Because true embodied agents depend heavily on self-motion feedback for spatial perception and navigation, the absence of this telemetry data significantly limits their utility for training interactive systems.
  • Lack of dynamic trajectories: Although datasets such as VisDrone include target tracking annotations, they mainly focus on 2D trajectories and lack complete descriptions of dynamic objects in 3D space. Moreover, these trajectories are generally observed rather than generated through agent–environment interaction.
  • Lack of goal-conditioned tasks: Existing datasets primarily provide labeled categories or boundaries but rarely include semantic information or reward signals associated with specific task goals. Embodied agents must understand task objectives and make decisions accordingly, which requires datasets to incorporate task-relevant annotations or scenario designs.
Therefore, to better support the development of Embodied, future remote sensing and mapping datasets should integrate richer spatiotemporal dynamics, embodied agent kinematic data, and task-goal–related semantic and behavioral feedback. Such integration would enable the construction of multimodal, multi-dimensional datasets that are better aligned with the requirements of Embodied research.

4.2.2. UAV Navigation and Inspection Datasets

While remote sensing datasets primarily address environmental perception, navigation, and inspection datasets provide the essential “proprioception” for embodied agents. These benchmarks typically aggregate flight trajectories, synchronized onboard imagery, and high-frequency inertial measurement unit (IMU) data, serving as the foundational ground truth for developing robust Visual–Inertial Odometry and Simultaneous Localization and Mapping algorithms [37]. To cover the spectrum from precise industrial inspection to agile flight, existing research relies heavily on two complementary paradigms:
  • Precision in GPS-Denied Environments (EuroC MAV [73]): As a standard benchmark for MAV-scale operations, EuroC focuses on the challenges of precise localization within confined, GPS-denied industrial and indoor settings [73]. By capturing visual-inertial data along complex trajectories, it simulates critical operational stressors such as rapid rotations and close-range object interactions. Although originally designed for state estimation, it has evolved into a rigorous testbed for evaluating algorithmic robustness, enabling improvements in feature tracking and corner detection for systems like OV2SLAM [74].
  • Agility in Dynamic Clutter (UZH-FPV [28]): In contrast to the stability-focused EuroC, the UZH-FPV dataset targets the regime of high-speed, aggressive flight [28]. It addresses the acute challenges of motion blur and extreme optical flow experienced during rapid acceleration and maneuvering through cluttered environments [19]. This dataset is instrumental for validating perception-action loops where the agent must maintain state estimation integrity under significant dynamic stress, a prerequisite for autonomous racing and emergency response.
Despite the success of these datasets in advancing metric state estimation, the transition to UAV-EAI exposes fundamental limitations in the current data landscape. The primary deficiency lies in the decoupling of metric states from semantic reasoning. Existing benchmarks focus almost exclusively on geometric accuracy, lacking the high-level semantic annotations required for agents to comprehend abstract instructions such as “search for diseased areas in farmland” [2]. Furthermore, current datasets fail to support long-horizon task planning. They typically feature short, continuous flight segments rather than the multi-stage, temporally extended missions necessary for power line inspection or agricultural monitoring, where environmental conditions evolve over time [6]. Finally, there is a notable absence of data supporting belief reasoning under uncertainty. To achieve true embodiment, datasets must go beyond raw sensor streams to include metadata on uncertainty factors and belief updates, enabling agents to learn robust decision-making strategies in partially observable or adverse weather environments [75].
In all, these directions highlight the need for next-generation UAV-EAI datasets that move beyond pure perception and toward richer representations of semantics, temporality, and decision-making under uncertainty [76].

4.2.3. UAV-EAI Task Datasets: Sparse and Emerging

Embodied task datasets for UAVs are critical to advancing UAV capabilities in embodied applications. At present, datasets that directly target UAV-EAI tasks remain relatively scarce, and related research is still at an early and emerging stage [76]. Existing datasets can be broadly grouped into several representative categories.
  • AerialVLN/CityNav-like datasets: These datasets primarily focus on language-conditioned navigation tasks in aerial scenarios. Recent efforts have explored Vision-Language Navigation (VLN) for UAV-EAI by adapting instruction-following paradigms originally developed for ground-based embodied agents [6]. VLMs have been increasingly adopted to address multimodal fusion, generalization, and interpretability challenges in navigation tasks, which are particularly relevant for UAV applications such as disaster response, logistics delivery, and urban inspection [77]. In parallel, advances in metric–semantic mapping systems enable UAVs to associate linguistic concepts with 3D spatial representations, supporting downstream navigation and inspection tasks that require semantic grounding in large-scale environments [38].
  • UAV-Search/OpenFly-like datasets: These datasets emphasize long-horizon target search and exploration tasks, often involving extended trajectories and large-scale environments. Cooperative multi-UAV search has been studied under conditions of partial observability and limited communication, where decentralized coordination strategies are required to maintain system performance [78]. Prior work has formulated multi-UAV search as probabilistic coverage and pursuit–evasion problems, emphasizing objectives such as minimizing time-to-detection or capture rather than simple area coverage. Other studies investigate formation-based and topology-aware search strategies that adapt sensing resolution and spatial abstraction to cope with sparse observations and large environments [79].
  • VLA and Cognitive Benchmarks: As aerial autonomy progresses from passive navigation to active environmental interaction, there is a critical requirement for datasets that tightly couple natural language instructions with continuous 6-DoF control outputs. Recent platforms such as CognitiveDroneBench [50] address this gap by providing dedicated evaluation frameworks for aerial VLA models. Encompassing extensive annotated flight trajectories, these emerging benchmarks explicitly target complex spatial reasoning and cognitive task execution, expanding the evaluation paradigm beyond traditional geometric metrics.
Despite these advancements, existing navigation and inspection benchmarks primarily focus on geometric stability, leaving a significant gap in evaluating semantic-aware task execution and long-horizon reasoning.
  • Viewpoint–information coupling. There is currently no dataset that explicitly captures the coupling between viewpoint selection and information gain, which is fundamental for tasks like NBV selection and IPP. Existing NBV and active perception benchmarks primarily focus on static scenes or ground-based robots, leaving aerial viewpoint optimization underexplored [80]. In cooperative multi-UAV perception, formation geometry directly influences target observability, sensor complementarity, and occlusion patterns, highlighting the need for datasets that systematically vary viewpoint configurations and sensor baselines [81].
  • Large-scale belief state annotations. For probabilistic search and exploration tasks, large-scale datasets with explicit belief state annotations remain largely absent. A belief state represents an agent’s probabilistic estimate of environmental states under uncertainty, and is central to POMDP-based planning frameworks [75]. Many UAV exploration systems still rely on single-modality sensing and deterministic maps, limiting robustness in cluttered or dynamic environments [82]. Multimodal sensing and belief-space planning have been shown to significantly improve exploration efficiency, yet corresponding annotated datasets are rare.
  • Limitations in multimodal sensing. Current datasets remain limited in terms of multimodal sensing modalities, such as thermal imaging, acoustics, and hyperspectral data. However, multimodal sensing is increasingly critical for UAV-based perception in challenging environments [83]. For example:
    Thermal imaging: Thermal-visible sensor fusion has been shown to improve detection robustness under low-illumination or adverse weather conditions, yet large-scale, task-oriented datasets remain limited [84].
    LiDAR and hyperspectral sensing: The integration of LiDAR with hyperspectral or multispectral imagery enables richer geometric–semantic reasoning, particularly for inspection and environmental monitoring tasks [85].
    Multispectral imagery: Multimodal fusion across spectral bands provides complementary cues for robust feature extraction and classification, but introduces challenges in sensor calibration, spatial alignment, and temporal synchronization [86].
    Despite their promise, most existing UAV datasets treat these modalities independently rather than as components of a unified embodied sensing pipeline. To mitigate the severe data scarcity inherent to these specialized modalities, current research is increasingly leveraging advanced data augmentation strategies. For instance, the SFBDA framework [87] introduces a semantic-decoupled data augmentation approach that significantly enhances infrared few-shot object detection capabilities on UAVs. Such methodologies offer a computationally viable alternative to the prohibitive costs associated with large-scale, real-world multi-modal data collection.

4.2.4. Summary

Current UAV datasets are highly fragmented, rarely satisfying the integrated requirements of Embodied AI for perception, decision-making, and control. While remote sensing datasets like VisDrone and UAVid offer excellent Visual Realism, Scale, and Diversity, they are fundamentally static. They omit the 3D Dynamics and ego-motion data necessary for true embodiment. In contrast, navigation benchmarks such as EuRoC and UZH-FPV supply the rigorous 6-DoF trajectories and IMU data needed for aggressive flight, but they lack rich Semantic Annotation and high-level Task/Goal Interaction. To clarify these tradeoffs and guide future benchmark design, Figure 4 provides a qualitative comparison across these five core dimensions.

5. Modeling Methods: From Control to Cognition

The modeling approaches used in UAV-EAI reflect the field’s transition from “stability-centric” UAV autonomy to “cognition-centric” embodied agents. Traditional UAV systems prioritize reliable control and state estimation; contemporary systems must additionally perform semantic reasoning, long-horizon planning, and embodied interaction in complex outdoor environments. This section organizes modeling methods for UAV-EAI into three layers: physics-centric control and estimation, learning-centric perception and decision making, and cognition-centric reasoning using large models. For each layer we highlight the key assumptions, identify the failure modes in UAV-EAI settings, and discuss opportunities for hybrid integration.

5.1. Conventional Model-Based Methods

5.1.1. Classical Control and Model-Based Estimation

The foundation of modern aerial robotics rests upon a rigorous tradition of classical control theory and probabilistic state estimation. Control architectures such as Proportional-Integral-Derivative, Linear Quadratic Regulator, backstepping, and geometric control on the special Euclidean group [88,89] provide mathematically rigorous stability guarantees. These methods treat the UAV as a dynamic system with well-defined equations of motion, ensuring precise tracking of aggressive trajectories even under high-maneuverability conditions. Complementing these controllers, state estimation pipelines typically leverage Extended or Unscented Kalman Filters [90] to fuse asynchronous data from Inertial Measurement Units, GPS, and visual odometry.
The maturity of these model-based approaches makes them highly attractive for field deployment; they are computationally lightweight, deterministic, and crucially certifiable for safety-critical operations on resource-constrained onboard hardware. However, the transition toward Embodied AI exposes a fundamental bottleneck: while these methods excel at low-level stabilization, they lack the semantic awareness and adaptive reasoning required for complex, open-world missions.
The limitations of relying solely on classical model-based frameworks become evident in the presence of uncertainties inherent to unstructured environments.
  • Brittleness of Dynamic Assumptions: Most model-based controllers depend on nominal dynamics, which often fail to hold in real-world Embodied AI applications. Environmental disturbances such as wind gusts, ground-effect turbulence, and sudden payload shifts during tasks like package delivery or physical sampling create complex non-linearities that rigid models struggle to accommodate.
  • Vulnerability to Estimation Drift: Reliable state estimation assumes a consistent stream of clean sensor data. In weakly structured environments—such as vast forests or fog-covered canyons—visual features may disappear and GPS signals may be obstructed. In these scenarios, the EKF/UKF-based pose tracking can suffer from catastrophic drift, yet the controller remains “blind” to the loss of situational context.
  • Fixed Trajectory Paradigms: Classical methods typically follow a predefined reference trajectory or a geometric path. This “blind following” is fundamentally incompatible with embodied tasks like Next-Best-View planning or Informative Path Planning, where the flight path must be dynamically re-synthesized in real-time based on the semantic content of the visual scene.
Ultimately, while classical control provides the necessary “muscles” for flight, it does not provide the “brain” for intelligent interaction. The challenge for UAV-EAI lies in bridging the gap between stable low-level stabilization and high-level semantic search, complex industrial inspection, and language-directed missions, where the agent must reason about what it sees, not just how it moves.

5.1.2. Trajectory Optimization and Model Predictive Control

Beyond static control laws, Model Predictive Control (MPC) and numerical trajectory optimization have emerged as the standard for achieving predictive, constraint-aware flight in complex scenes [91]. By solving a constrained optimization problem over a finite time horizon, MPC allows UAV-EAI to simultaneously account for non-linear dynamics, actuator saturation limits, and obstacle avoidance maneuvers. This predictive capability provides a natural bridge to embodied navigation, as the agent can anticipate future states and adjust its control inputs to maintain safety while pursuing mission-critical objectives.
However, the efficacy of traditional MPC is often predicated on the assumption of a “perfect world” model—specifically, access to smooth, differentiable cost maps, high-fidelity environment representations, and noise-free state estimates. In the deployment of UAV-EAI for real-world exploration, these assumptions are frequently challenged:
  • Uncertain and Partially Observed Cost Manifolds: Unlike industrial settings, UAV-EAI agents operate in environments where cost maps (representing risk or occupancy) are incrementally built and inherently uncertain. The agent must optimize its trajectory over a “belief map” where the cost of a certain path is a stochastic variable rather than a deterministic scalar.
  • Scale-Induced Localization Drift: While MPC relies on a stable feedback loop, localization accuracy tends to degrade over kilometer-scale missions in unstructured terrain. This drift introduces a mismatch between the optimized plan and the physical execution, potentially leading to catastrophic collisions if the optimization does not account for spatial uncertainty.
  • Dynamic and Epistemic Disturbances: External perturbations, such as sudden wind gusts or aerodynamic interactions in urban canyons, cause unmodeled deviations that exceed the rejection capabilities of standard MPC. These disturbances represent a mix of aleatory noise and epistemic model gaps that can destabilize the predictive loop.
To bridge these gaps, modern UAV-EAI architectures must embed MPC within more robust safety frameworks. This includes the integration of Control Barrier Functions (CBFs) [92] and Safety Shields, which act as a supervisory layer to filter the optimizer’s commands. By defining forward-invariant safe sets, CBFs can guarantee that the UAV remains within a safe operating envelope even when the high-level IPP or NBV planner encounters epistemic uncertainty. Such “Safety-Critical MPC” approaches ensure that the embodiment remains intact—preventing physical damage—while the agent’s “mind” navigates the complexities of semantic search and information gathering.

5.2. Learning-Centric Methods: Deep Learning and Reinforcement Learning

Learning-based methods have transformed perception and navigation, but their applicability to UAV-EAI depends heavily on environmental structure and data availability.

5.2.1. Deep Perception: Recognition, Depth, and Scene Understanding

Modern deep visual networks such as ResNet, EfficientNet, and Vision Transformers (ViTs), have substantially advanced the capabilities of UAV imagery in detection, segmentation, and depth estimation tasks [93,94,95]. However, UAV visual data exhibit pronounced out-of-distribution (OOD) differences compared with conventional ground-based or indoor datasets, posing significant challenges to the generalization ability of existing vision models [96,97]. Understanding these differences and developing corresponding mitigation strategies is essential for building robust UAV vision systems. The distinctive characteristics of UAV imagery can be attributed to several underlying factors.
  • Perspective mismatch: Most existing vision models—typically trained on datasets such as COCO or large-scale web datasets—are learned from a “human eye-level” perspective. This object-centric viewpoint contains rich semantic cues, including occlusion relationships and well-defined object boundaries [98]. In contrast, UAV images are often captured from top-down, oblique, or steep overhead viewpoints, which are inherently scene-centric and lack the fine-grained semantic cues common in ground-level imagery [99]. Background textures are frequently dominated by large-scale terrain or homogeneous vegetation, making it difficult for models to exploit local features learned from traditional datasets. Such viewpoint-induced geometric distortions and weak-texture regions significantly complicate depth reasoning and scene understanding in aerial imagery [100].
  • Scale and altitude variation: Targets in UAV images—such as pedestrians or vehicles in search-and-rescue or surveillance missions—may occupy only a few pixels, rendering small-object detection particularly challenging [101]. Moreover, changes in flight altitude lead to drastic variations in object scale and image resolution, further increasing the difficulty of detection and segmentation [102]. These large scale variations require models to maintain robustness across orders of magnitude in spatial resolution, which is rarely encountered in conventional ground-level vision benchmarks.
  • Changes in depth perception cues: Geometric edge cues that are effective in close-range indoor environments become unreliable from high-altitude UAV perspectives [103]. UAV depth estimation relies more heavily on texture gradients, shading, or illumination variations across terrain [104]. Depth discontinuities induced by perspective distortion and the prevalence of weakly textured regions pose persistent challenges for both supervised and self-supervised depth estimation in aerial scenarios [105].
These observations indicate that directly applying depth vision models pretrained on generic datasets may lead to severe scale misestimation, missed detections, and erroneous depth predictions. To visually elucidate this fundamental challenge, Figure 5 illustrates the mechanism of perspective mismatch and the necessity of domain adaptation. As depicted, standard feature extractors pre-trained on canonical ground-view datasets often fail to extract robust representations from nadir-view aerial imagery (Target Domain) due to severe feature misalignment. To bridge this gap, the architecture incorporates a Domain Adaptation Module—such as cross-view pretraining or adversarial learning—to project disparate visual features into a shared, view-invariant space, thereby recovering recognition accuracy. Accordingly, the design of UAV-EAI perception systems must account for a set of strategies tailored to aerial sensing conditions.
  • cross-view pretraining (BEV-aware data): To align feature spaces, pretraining on datasets that include bird’s-eye-view (BEV) perspectives is necessary [106]. Recent studies in BEV-based representation learning demonstrate that multi-view and multi-geometry alignment can significantly reduce domain gaps across viewpoints, offering valuable insights for UAV vision systems [107]. Domain Generalization aims to train models that can extract transferable knowledge from one or multiple source domains and generalize to unseen target domains [108]. ViTs have demonstrated strong potential in domain generalization, as cross-domain self-supervised pretraining and fine-tuning enable models to learn more generalizable geometric representations [109].
  • multi-altitude augmentation and domain randomization: Robustness can be enhanced by simulating different flight altitudes, illumination conditions, and environmental variations during training through data augmentation and domain randomization techniques [110,111]. In aerial and outdoor depth estimation, augmentation strategies that explicitly perturb scale and viewpoint have been shown to improve robustness and generalization [112]. Furthermore, incorporating synthetic or simulated aerial data during training is an effective way to improve performance when real-world UAV data are scarce or collected under extreme conditions [113].
  • hybrid geometric–learning depth estimators: Combining geometric constraints with learning-based approaches enables the design of depth estimators better suited for long-range outdoor scenarios [114,115]. For example, self-supervised monocular depth estimation methods often rely on view synthesis and image warping as supervision signals, eliminating the need for dense ground-truth depth annotations [116]. Such hybrid formulations improve geometric consistency and robustness under large viewpoint changes, which are common in UAV imagery [105]. In addition, lightweight attention-based architectures are increasingly explored to enable real-time depth inference on resource-constrained UAV platforms [117].
In recent years, deep learning models have achieved remarkable progress in UAV image processing. For example, modern object detection and tracking pipelines originally developed for real-time vision tasks have been successfully adapted to aerial scenarios, demonstrating robustness to illumination variation and partial occlusion [118,119]. Vision Transformers, owing to their global feature modeling capability, have shown advantages over traditional convolutional neural networks (CNNs) in depth estimation, domain adaptation, and domain generalization tasks [95,120]. Their ability to capture long-range dependencies is particularly beneficial for large-scale scenes observed from high-altitude UAV viewpoints [121]. Meanwhile, to address limited onboard memory and computational constraints, recent work explores multi-task networks that jointly predict depth and semantic information to improve efficiency and inference speed [122]. Such designs are especially relevant for wide field-of-view and fisheye camera setups commonly used in autonomous UAV platforms [123].
In summary, by tailoring model architectures, data augmentation strategies, and training paradigms to the characteristics of UAV viewpoints with cross-domain pretraining and hybrid depth estimators, it is possible to significantly enhance the performance of UAV perception systems in complex environments. These advances will further promote the deployment of UAVs in applications such as search and rescue, environmental monitoring, precision agriculture, and autonomous navigation [124,125]. By explicitly addressing viewpoint diversity, scale variation, and geometric sparsity, UAV-EAI perception systems can better meet the demands of real-world deployment and long-term autonomy [126].

5.2.2. RL/IL for Navigation and Active Perception

Reinforcement learning (RL) and imitation learning (IL) have demonstrated substantial potential in UAV racing and agile flight, as expert demonstrations can provide efficient policy samples that enable autonomous decision-making and navigation [19,127]. However, in long-horizon embodied tasks such as search-and-rescue and inspection, these methods face three fundamental structural challenges [44].
A primary challenge in UAV-EAI reinforcement learning arises from the extreme reward sparsity inherent in vast aerial search spaces. In contrast to indoor navigation, UAVs tasked with locating small targets in expansive environments encounter highly infrequent successful outcomes, resulting in a scarcity of positive reward signals. Under such conditions, conventional reinforcement learning algorithms often exhibit poor convergence behavior or fail to learn effective policies [128]. Although reward shaping can alleviate this issue, the design of informative dense reward functions for complex aerial tasks typically demands substantial domain expertise and extensive manual tuning. Conversely, while sparse reward formulations simplify task specification, they significantly reduce learning efficiency. Consequently, existing studies have explored a range of methods aimed at improving learning performance under sparse-reward conditions in UAV-EAI.
  • Experience Replay and Curiosity Mechanisms: Curiosity-driven experience replay methods introduce intrinsic motivation to encourage exploration of unseen states, enabling agents to acquire relatively effective policies under sparse rewards [129,130]. Empirical results in sparse-reward benchmarks indicate that curiosity-based methods can significantly improve exploration efficiency compared with naive baselines.
  • Hierarchical Reinforcement Learning (HRL): By decomposing complex tasks into a sequence of sub-tasks, HRL effectively mitigates sparse-reward issues in long-horizon navigation problems [131]. This approach is particularly suitable for mixed action spaces, where agents first select abstract goals and then execute low-level control policies.
  • Combining Imitation Learning with Reinforcement Learning: Integrating imitation learning with reinforcement learning leverages expert demonstrations to guide policy learning and accelerate convergence under sparse rewards [132]. Self-imitation learning further improves performance by reinforcing previously successful trajectories, even without additional expert supervision [133].
  • Model Predictive Control: MPC can be incorporated into reinforcement learning frameworks to alleviate sparse rewards by predicting future system states and optimizing control sequences, thereby providing denser learning signals and improved stability [134].
  • Parameterized and Hybrid Action Spaces: To cope with the complexity of state and action spaces in UAV navigation, parameterized and hybrid action-space formulations have been introduced, enabling more efficient learning and structured exploration under sparse-reward conditions [135].
  • Predictive Coding: By learning predictive representations offline and using them for reward shaping, predictive coding supplies reward signals that reflect higher-level understanding of environmental structure and dynamics [136].
Building upon the training complexity, the Sim-to-Real gap constitutes a critical challenge that is particularly pronounced in aerial robotics. Small modeling inaccuracies in simulation such as imperfect representations of aerodynamic drag, wind disturbances, or actuator dynamics, can be significantly amplified in real-world deployment [137]. As a result, policies that perform well in simulation may exhibit instability or catastrophic failures when transferred to real UAV platforms. To address the Sim-to-Real gap, existing studies have investigated several complementary strategies.
  • Domain Randomization: Randomizing physical parameters and visual properties during simulation training improves policy robustness to real-world variability and enhances generalization from simulation to reality.
  • Advanced Low-Level Control with Reinforcement Learning: Combining learning-based high-level decision-making with robust classical control techniques reduces sensitivity to modeling errors and improves flight stability.
  • Curriculum Learning: Gradually increasing task difficulty during training enables agents to acquire more robust policies in simulation, facilitating transfer to real-world scenarios.
  • Physics-Aware Simulators: Developing simulators that more accurately capture real-world physics reduces discrepancies between simulated and real environments.
  • Zero-Shot Safety: Zero-shot safety approaches aim to ensure constraint satisfaction during deployment, guaranteeing safe behavior even when real-world conditions differ from those encountered during training.
Bridging the aerial Sim-to-Real gap requires more than generic domain randomization; it necessitates a transition toward asymmetric learning architectures. By employing privileged environmental distillation, agents can learn to internalize complex aerodynamic disturbances that are otherwise unobservable through onboard sensors [68]. Furthermore, the implementation of asynchronous physical-cognitive decoupling ensures that the high-level reasoning of foundation models does not compromise the deterministic latency required for low-level flight stability [138]. These methodologies collectively define a specialized execution framework for UAV-EAI, enabling high-speed autonomous operation in unstructured 3D spaces where traditional terrestrial learning paradigms remain computationally and physically inapplicable [139].
Furthermore, the practical deployment of RL/IL policies is heavily constrained by the limited onboard computational resources characteristic of UAV platforms. Embodied AI systems typically rely on deep neural networks with substantial computational demands. However, UAV platforms are subject to strict onboard resource constraints, and learned policies must operate alongside high-frequency control loops. Embedded hardware often struggles to execute large models at the update rates required for stable flight, limiting the applicability of computationally intensive architectures [140].
Addressing onboard computational constraints requires the development of lightweight policy architectures. Model compression techniques—including pruning, quantization, and knowledge distillation—can significantly reduce computational load while preserving policy performance [141]. In addition, improving algorithmic efficiency through sample-efficient reinforcement learning methods is essential for practical UAV deployment [142]. In the context of UAV-EAI, successful reinforcement learning pipelines reported in the literature typically emphasize several key principles.
  • Belief-Space Formulations: Planning in belief space enables robust decision-making under partial observability by explicitly reasoning over uncertainty, rather than relying solely on geometric state representations.
  • Lightweight Policy Architectures: Through compression and architectural optimization, policies can be adapted to limited onboard computational resources while maintaining real-time inference capability.
  • Physically Aware Domain Randomization: Extensive randomization of physical parameters during training improves robustness and increases the likelihood of successful simulation-to-reality transfer.
  • Safety Filters: Incorporating low-level safety mechanisms—such as control barrier functions—ensures constraint satisfaction and prevents catastrophic failures even when high-level policies are unreliable.
In summary, although reinforcement learning and imitation learning provide powerful tools for UAV-EAI, their deployment in complex embodied tasks requires systematic solutions to sparse rewards, Sim-to-Real gaps, and onboard computational constraints. By integrating belief-space planning, lightweight models, domain randomization, and safety filters, UAV systems can achieve improved robustness, safety, and real-world reliability.

5.3. Cognition-Centric Methods: Large Language and Vision Language Models

The emergence of VLMs and LLMs—such as Flamingo [143], GPT-4V, LLaVA [144], and BLIP-2 [145]—has introduced a transformative paradigm for embodied agents. These models empower agents to interpret complex natural language instructions, execute zero-shot reasoning, and generate semantically grounded decisions. However, the direct deployment of these terrestrial-pretrained models onto UAV platforms frequently results in significant performance degradation. This subsection investigates the underlying factors contributing to this decline, specifically focusing on the perspective mismatch, the domain gap in semantic representation, and the onboard computational constraints that necessitate a specialized cognitive architecture for aerial embodied intelligence.

5.3.1. Perspective Mismatch and Domain Gap

The deployment of VLMs in UAV-EAI exposes fundamental challenges that extend beyond a general lack of situational “common sense.” The dominant failure modes are primarily attributed to a structured perspective mismatch and a pronounced dataset domain gap between the pretraining corpora and the operational conditions of aerial robotics [146,147]. Although VLMs demonstrate strong zero-shot capabilities in terrestrial environments, their performance often degrades when applied to aerial settings involving a transition from horizontal to vertical viewpoints [148]. These limitations can be traced to several forms of systematic misalignment, which are discussed below.
  • Canonical View vs. Nadir Perspective: Dominant VLM training sets are composed almost exclusively of human-eye-level imagery, reflecting canonical horizontal viewpoints [98]. In contrast, UAV imagery is predominantly top-down (nadir) or oblique, leading to substantial geometric and appearance shifts that undermine feature reuse in pretrained visual encoders [99].
  • Object-Centric vs. Scene-Centric Representation: Large-scale vision-language datasets are typically object-centric, featuring high-resolution, centered subjects. UAV data, by contrast, is inherently scene-centric, where targets often occupy only a few pixels within large, cluttered landscapes. This mismatch exacerbates the difficulty of grounding language queries in low-saliency aerial scenes [149].
  • Semantic Feature Disparity: The semantic cues leveraged by VLMs for contextual reasoning—such as object co-occurrence patterns or human-centric interactions—are largely absent in wilderness or remote environments. In such scenarios, interpretation relies on geomorphological structure and texture statistics that are underrepresented in web-scale pretraining corpora [97].
The consequence of these gaps is a marked increase in semantic hallucinations, where models misinterpret natural terrain features or fail to detect critical targets in low-information-density scenes [150]. To mitigate these failures and ground VLMs in the aerial domain, several emerging research directions are currently being explored.
  • Cross-View Contrastive Pretraining: Designing contrastive objectives that explicitly align ground-level and aerial viewpoints, encouraging the learned representation to remain invariant under large changes in camera pitch, altitude, and scale [151].
  • Parameter-Efficient Adapter Tuning: Leveraging lightweight adaptation mechanisms, such as low-rank adapters, to specialize pretrained VLM backbones using limited aerial datasets while preserving general multimodal knowledge [152].
  • Geometric Prior Injection: Incorporating explicit geometric metadata—such as camera pose, altitude, or GPS-derived priors—into the attention mechanisms of VLMs, enabling pose-aware reasoning and improved spatial grounding [153].
  • Multi-Altitude and Multi-Scale Pipelines: Employing hierarchical, multi-resolution processing pipelines that jointly capture global scene context and fine-grained local details, an approach shown to be effective in aerial perception and active vision systems [154].
By addressing the domain gap not as a generic data scarcity issue but as a problem of geometric and semantic misalignment, UAV-EAI can more effectively leverage foundation VLMs to enable robust, language-guided aerial autonomy [155].

5.3.2. LLMs for Instruction Following and High-Level Planning

Recent advancements in LLMs have facilitated the translation of natural language into machine-executable commands [156,157]. Within UAV-EAI, LLMs function as cognitive interfaces that interpret complex, underspecified mission objectives—such as “Search for a missing hiker near the north ridgeline, prioritizing areas with dense forest cover”—and decompose them into actionable subtasks [9]. Drawing on broad pretraining corpora, these models infer semantic priors to support mission specification and goal formulation [158].
However, translating linguistic reasoning into physical action remains limited by a lack of embodied grounding. While LLMs excel at symbolic decomposition and commonsense inference, their detachment from physical reality frequently causes feasibility hallucinations during robotic deployment [159]. Consequently, planners relying exclusively on LLMs generally fail to account for the fundamental constraints of physical embodiment and environmental interaction.
  • Resource and Physical Constraints: LLMs do not inherently model finite onboard resources such as battery capacity or flight time, nor do they explicitly reason about geometric constraints including camera field-of-view, sensing range, and occlusion [160].
  • Stochastic Map Uncertainty: Unlike static textual environments, real-world UAV operations are characterized by partial observability and evolving uncertainty. LLMs struggle to reason over probabilistic occupancy maps or belief-space representations common in robotic exploration.
  • Aerodynamic and Kinematic Feasibility: Linguistically valid action descriptions may violate vehicle dynamics or safety envelopes. High-level plans produced by LLMs rarely account for aerodynamic disturbances, non-holonomic motion constraints, or stability margins inherent to aerial platforms [24].
  • Long-horizon Informative Coupling: LLMs are poorly suited for fine-grained, long-horizon viewpoint planning problems such as Next-Best-View or Informative Path Planning, where action selection depends on continuous optimization of expected information gain rather than symbolic sequencing [161].
To mitigate these grounding failures and bridge the gap between semantic reasoning and physical reality, recent architectures increasingly adopt neuro-symbolic hierarchical frameworks. As illustrated in Figure 6, we conceptualize this process as a closed-loop “Generate-Verify-Revise” pipeline.
In this architecture, the LLM serves as a high-level planner, transforming abstract user instructions and environmental context into a Draft Plan. Crucially, unlike purely text-based agents, this draft is not executed directly. Instead, it is passed through a Physics-Based Feasibility Filter (or Dynamics Checker). This critical module evaluates the proposed action sequence against rigid UAV constraints—such as remaining battery life, kinematic limits, and environmental disturbances like wind shear.
If the plan is deemed infeasible, the system triggers a feedback loop. The feasibility verifier translates the physical violation into a linguistic prompt, enabling the LLM to perform iterative re-planning with updated constraints. Only plans that survive this filtration process are transmitted to the Execution Controller, ensuring that the agent’s high-level intent is always bounded by its physical capabilities.
  • Physics-Based Feasibility Filters: A verification layer that evaluates LLM-generated waypoints or subgoals against vehicle dynamics, safety constraints, and control limits, often implemented with control-theoretic tools [162].
  • Hierarchical Task Abstraction: A structured decomposition in which the LLM specifies semantic goals and task priorities, while mid-level planners compute dynamically feasible trajectories that maximize information gain under uncertainty [35].
  • Uncertainty-Aware Feedback Loops: Closed-loop architectures that feed perception uncertainty, belief updates, or task execution failures back into the LLM prompt context—using mechanisms such as Chain-of-Thought or ReAct—enabling adaptive re-planning when physical reality deviates from the original mission description [163].
By explicitly grounding LLM-based reasoning within the physical, stochastic, and resource-constrained nature of aerial robotics, UAV-EAI systems can transform abstract human instructions into safe, efficient, and semantically rich autonomous behaviors [155].
Beyond single-agent physical constraints, scaling LLM-driven autonomy to swarm deployments exposes the system to multi-agent cognitive divergence. When multiple independent LLMs interpret shared, ambiguous mission objectives, the absence of rigid constraint mechanisms frequently generates misaligned subgoals or conflicting spatial trajectories, undermining swarm coordination.
To enforce strategic consistency and mitigate cognitive divergence, multi-agent LLM orchestration necessitates robust hierarchical organizational structures and collaborative workflows. Rather than relying on flat communication, future architectures must implement explicit role-based frameworks. By deterministically assigning specific cognitive functions, authority levels, and task scopes to designated agents within a hierarchy, the reasoning space of individual LLMs is strictly bounded. This structured orchestration ensures that decentralized semantic executions remain globally coherent, preventing logic conflicts and maintaining absolute strategic alignment across the swarm.

5.3.3. Onboard Inference and Embedded Deployment

While foundation models provide necessary spatial and semantic reasoning capabilities, transitioning these architectures from server clusters to the edge introduces significant SWaP constraints. Current UAV-EAI research often overlooks the specific mechanisms of onboard computational trade-offs. Achieving real-time autonomy requires the strict partitioning of a heavily constrained thermal design power envelope, which typically ranges from 10–40 W for embedded systems like the NVIDIA Jetson series. This limited power budget must be carefully distributed across three competing subsystems: control, perception, and cognition. The challenge of sharing these finite resources directly dictates two critical engineering considerations for system deployment.
  • Hierarchical Resource Partitioning: Computational resources must be allocated asymmetrically to prevent subsystem starvation. Low-level flight control algorithms like Model Predictive Control require kilohertz-level update rates but consume minimal power, usually under 1 W. Mid-level perception tasks, encompassing Visual-Inertial Odometry and dense geometric mapping, necessitate deterministic latency operating at 20 to 30 Hz. These tasks continuously occupy a significant baseline fraction of the GPU and memory bandwidth, demanding around 10–15 W. Consequently, high-level cognitive models are restricted entirely to the residual compute budget. Attempting continuous and synchronous inference of large language or vision models can deprive the perception module of required memory bandwidth, leading to state estimation drift and flight instability [164].
  • The Complexity-Latency Trade-off: The relationship between model complexity and real-time performance on edge devices is primarily governed by memory access speeds rather than pure floating-point operations. For instance, executing a 7B-parameter foundation model in half-precision format requires approximately 14 GB of video memory. Moving these weights across the limited memory bus of a mobile chip frequently pushes inference latency beyond 500 milliseconds per query. However, safe aerial obstacle avoidance dictates that the entire perception-to-action cycle must remain below a critical threshold of 100 to 200 milliseconds. Latency exceeding this margin induces prolonged dead-reckoning during the model’s forward pass, compromising the predictive control loop [165].
To resolve these hardware limitations, current engineering practices must pivot from continuous inference toward asynchronous and event-triggered architectures. By decoupling the high-frequency perception and control loops from the low-frequency cognitive processes, systems can maintain physical stability while querying large foundation models only upon detecting significant semantic uncertainty or explicit task transitions.
  • Aggressive Quantization and Pruning: Moving beyond FP16, recent work explores INT8, FP8, and sub-8-bit quantization combined with structured pruning to dramatically reduce memory footprint and increase throughput-per-watt [166,167].
  • Sparsity and Mixture-of-Experts (MoE): Sparse activation mechanisms enable only a subset of parameters to be evaluated per input, significantly reducing FLOPs while retaining large-model capacity [168,169].
  • Edge–Cloud Hybrid Pipelines: When communication permits, split-inference architectures assign safety-critical control and perception to onboard systems while offloading high-level semantic reasoning to edge-cloud infrastructure, forming a hierarchical intelligence loop [170].
  • Event-Triggered and Asynchronous Inference: Rather than continuous VLM evaluation, event-driven inference pipelines query models only when perceptual novelty or confidence degradation is detected, reducing unnecessary energy expenditure.
In summary, the successful deployment of UAV-EAI hinges on co-designing algorithms with their target hardware, ensuring that the cognitive core of the agent remains compatible with the physical and electrical realities of flight [171].

5.3.4. Hierarchical System Architecture: Edge-Cloud Collaborative Orchestration

A fundamental architectural consideration in the deployment of UAV-EAI is the spatial distribution of computational workloads. While theoretical embodied frameworks often premise on fully self-contained, isolated agents, the stringent SWaP constraints inherent to aerial platforms render the localized execution of large-scale foundation models and complex multi-agent coordination algorithms computationally prohibitive. Consequently, advancing UAV-EAI from simulation to real-world application necessitates a transition toward a hierarchical edge-cloud collaborative architecture. This paradigm strategically allocates computational responsibilities based on latency tolerances, algorithmic complexity, and system-level orchestration requirements [31].
Within this hybrid architecture, tasks are functionally decomposed into two distinct layers:
  • Edge Node Processing: Latency-Bounded Reactive Autonomy. The UAV’s onboard computational payload is strictly reserved for high-frequency, safety-critical execution. This encompasses low-level attitude control, real-time VIO, dynamic obstacle avoidance, and the physics-based feasibility verification discussed in Section 5.3.2. By isolating these processes at the edge, the system ensures that the agent maintains basic physical stability and reactive survival capabilities independent of external network states [172].
  • Cloud and Ground Control Station (GCS): Global Semantic Orchestration. Computationally intensive, long-horizon cognitive functions are offloaded to centralized cloud infrastructure or local GCS units. This layer executes LLMs for abstract instruction decomposition and maintains global metric-semantic maps. Critically, in cooperative multi-agent operations, this layer serves as the centralized orchestration hub. It facilitates hierarchical role assignments, manages task allocation, and aligns collaborative workflows across the swarm, ensuring that localized edge behaviors serve a coherent global objective [78].
A defining vulnerability of edge-cloud hybrid architectures is their dependency on communication integrity. In real-world field applications, network topologies are frequently dynamic, and bandwidth is severely constrained. To mitigate this, robust UAV-EAI systems must employ asymmetric communication protocols—prioritizing the transmission of compressed semantic tokens and probabilistic belief states over raw, high-bandwidth sensory streams [57]. Furthermore, the architecture must incorporate mechanisms for “degraded autonomy.” In the event of connection loss with the global orchestrator, the agent must seamlessly transition from cloud-guided semantic planning to localized, rule-based survival policies (e.g., autonomous loitering or returning to a predefined communication relay). This guarantees that the hierarchical system remains fail-safe even when the cognitive communication link is compromised.

5.4. Toward Hybrid Physical–Cognitive UAV Models

The maturation of UAV-EAI increasingly calls for a departure from decoupled modular designs toward hybrid architectures that integrate physical dynamics, deep perception, and high-level cognitive reasoning within a unified framework [173]. In contrast to prevailing linear pipeline paradigms, emerging aerial autonomy systems are expected to feature bidirectional information flow, enabling physical constraints and semantic objectives to be jointly considered during decision-making and control. From this perspective, the convergence of physical modeling, perception, and cognition delineates a set of critical research directions that collectively shape the development of hybrid physical–cognitive models for UAV-EAI.
  • Physics-aware Cognition: Future models must explicitly project high-level semantic intent into reachable sets and flight envelopes, ensuring that cognitive decisions respect aerodynamic feasibility and safety constraints [162].
  • Cognitive-driven Adaptive Control: Hybrid architectures should allow semantic uncertainty and task priority to modulate low-level control strategies, enabling context-aware transitions between aggressive maneuvering and precision flight modes [174].
  • Task-aware Active Perception: Integrating NBV and IPP directly with reasoning backbones allows perception to be guided by semantic objectives rather than uniform coverage.
  • Belief-space Hybrid Planning: A unified belief-space formulation that simultaneously captures pose uncertainty, map entropy, and detection confidence enables principled trade-offs between exploration and localization [75,175].
This hybrid perspective moves UAV autonomy away from the traditional rigid perception-to-control pipeline toward a more fluid and unified embodied intelligence. In this paradigm, motion and perception are inseparable: every action refines belief, and every observation informs intent, enabling aerial agents that are not only autonomous but genuinely intelligent [176].

6. Applications and Challenges

As UAV-EAI transitions from theoretical modeling to real-world deployment, it encounters highly diverse operational requirements across different domains. The severe Size, Weight, and Power (SWaP) constraints discussed in Section 5 dictate that equipping every UAV with a monolithic, “all-in-one” foundation model is neither computationally feasible nor practically efficient for end-users.
Instead, the successful application of UAV-EAI fundamentally relies on task-driven modularity and heterogeneous orchestration. In this paradigm, complex missions are not solved by single-agent gigantism, but rather by deploying tailored, lightweight cognitive modules that address specific operational pain points. Advanced capabilities are achieved by distributing specific sensing, reasoning, and physical intervention roles across a collaborative swarm. Guided by this application-centric philosophy, this section examines key domains by strictly following a structured analysis: identifying domain-specific pain points, proposing tailored EAI developments, and outlining the remaining engineering challenges [70].

6.1. Key Application Domains

Embodied intelligent UAVs are applicable across a wide range of scenarios that demand high levels of autonomy, precision, and adaptability, as shown in Figure 7 [177].

6.1.1. Critical Infrastructure Inspection

Autonomous inspection of high-value infrastructure—including bridges, transmission lines, and urban logistics corridors—demands high-precision navigation within geometrically complex environments. In these scenarios, the agent is required to maintain close-range proximity to large-scale structures to perform fine-grained anomaly detection or safe payload delivery. Unlike open-field remote sensing, these missions occur in “urban canyons” or industrial sites where the environment imposes strict constraints on the agent’s maneuverability and sensing fields.
The primary operational pain points in this domain arise from the inherent brittleness of traditional waypoint-based navigation. For end-users, static flight paths are inadequate for addressing signal shadowing (GPS-denied) and proximity-induced aerodynamic disturbances, such as the ground effect or wall effect, which compromise flight stability. Furthermore, identifying localized structural defects requires adaptive viewpoints that cannot be pre-defined, leading to significant data incompleteness when relying on open-loop, pre-programmed trajectories.
To address these limitations, tailored UAV-EAI development must pivot toward active, perception-driven workflows rather than monolithic foundation models. By implementing asynchronous physical-cognitive decoupling, the platform ensures low-latency reactive control while leveraging NBV planning to autonomously resolve occlusions. This enables the agent to prioritize informative regions—such as potential fracture points or safe landing zones—through active 6-DoF pose optimization. Such modular intelligence ensures that high-level reasoning is strictly bounded by the physical realities of the inspection site.
Despite these advancements, the reliable fusion of heterogeneous sensor data under strict energy constraints remains a persistent challenge. Maintaining high-frequency stability while executing compute-heavy semantic reasoning for real-time anomaly detection continues to strain embedded hardware payloads. Future research must prioritize the development of lightweight, task-specific cognitive filters that can operate within the deterministic latency requirements of safety-critical infrastructure environments.

6.1.2. Emergency Search and Rescue

Emergency search and rescue (SAR) operations in post-disaster environments demand rapid, large-scale spatial exploration to locate highly sparse targets. In these time-critical missions, agents must navigate through vast, unstructured terrains—such as dense forests or collapsed urban zones—where prior maps are obsolete and GPS signals may be degraded. The primary objective shifts from high-resolution mapping to minimizing the time-to-detection of specific semantic targets (e.g., survivors or vehicles) across expansive areas.
The critical pain point for SAR end-users lies in the severe inefficiency of traditional, pre-programmed sweep patterns. Standard “lawnmower” coverage strategies assign uniform importance to all regions, wasting finite battery life on vast empty areas. When target signatures are minimal (e.g., occupying <1% of the field of view), these rigid trajectories fail to adapt to intermediate discoveries, such as footprints or debris. Furthermore, the spatiotemporal limitations and strict SWaP constraints of a single UAV make it fundamentally incapable of scanning kilometer-scale regions within the critical rescue window.
To overcome these single-agent limitations, EAI development in SAR focuses on deploying decentralized, homogeneous swarms governed by shared probabilistic belief-spaces. Instead of relying on a monolithic cognitive model, each agent utilizes lightweight policies to actively construct local semantic maps. By exchanging latent intentions and confidence scores rather than raw trajectories, the swarm dynamically optimizes its area coverage based on real-time discoveries. This decentralized, intent-driven orchestration allows the fleet to autonomously converge on high-probability regions, significantly accelerating target localization without requiring human intervention.
Scaling these cognitive capabilities to multi-agent deployments, however, imposes strict quantitative constraints on both SWaP and communication throughput. For instance, continuous processing of 30 Hz RGB-D streams for metric-semantic mapping consumes 15–20 W per edge node, heavily saturating typical embedded power budgets. Transmitting uncompressed sensory point clouds rapidly exceeds the bandwidth capacities of ad-hoc wireless networks in disaster zones [52]. Consequently, robust swarm coordination necessitates a transition to highly compressed semantic representations. Experimental deployments demonstrate that exchanging localized semantic belief maps reduces per-agent communication bandwidth to <100 KB/s, which remains an absolute engineering prerequisite for maintaining collaborative search workflows under severely degraded network topologies [57].

6.1.3. Precision Agriculture and Aerial Intervention

Precision agriculture necessitates continuous aerial monitoring of crop vitality and the subsequent execution of targeted physical interventions, such as localized pesticide application or precision seeding. Unlike broad-scale remote sensing, this domain requires a seamless transition from macro-level vegetation analysis to micro-level physical interaction within complex, unstructured canopies.
The fundamental limitation of current agricultural UAVs stems from their reliance on rigid, pre-programmed waypoint trajectories that cannot adapt to dynamic environmental anomalies. Static grid-based coverage fails to facilitate immediate, fine-grained inspection upon detecting localized disease outbreaks, relegating critical analysis to offline processing. Furthermore, deploying heavy payload drones for close-canopy intervention introduces severe aerodynamic turbulence and ground effects, significantly elevating collision risks in environments with unpredictable geometric boundaries.
Addressing these domain-specific inefficiencies requires replacing monolithic cognitive models with tailored, heterogeneous multi-agent orchestration. In this modular framework, a high-altitude UAV operates as a global planner, utilizing lightweight semantic segmentation to identify macro-scale vegetation anomalies. Upon detection, it transmits compressed semantic belief states to a low-altitude, rotary-wing executor. This intervention agent eschews heavy foundation models in favor of specialized, physics-informed reinforcement learning policies, enabling it to navigate canopy turbulence and execute targeted spraying with strict computational and energy efficiency.
The primary engineering bottleneck in this heterogeneous deployment lies in cross-platform spatiotemporal alignment and multi-modal data fusion. Accurately projecting high-altitude semantic anomalies onto the local 3D coordinate frame of an intervention drone requires deterministic latency and robust state estimation. Achieving this seamless collaborative orchestration under the severely degraded ad-hoc network topologies typical of rural environments remains a critical focus for future agricultural EAI research.

6.2. Core Challenges and Open Problems

Despite their significant potential, embodied intelligent UAVs face several fundamental challenges and open research problems [177].

6.2.1. Sim-to-Real Gap

UAVs operating in real environments are affected by complex aerodynamic effects and real-world disturbances, such as wind, illumination changes, humidity, and terrain variability, which are difficult to model accurately in simulation [178]. Control policies trained in simulated environments often experience performance degradation when deployed in real-world settings [179]. In complex scenarios, reliance on remote human control limits UAV capabilities and reduces system efficiency, highlighting the urgency of achieving autonomous navigation through AI-driven approaches. To mitigate this gap, learning frameworks that combine reinforcement learning with advanced low-level control have been proposed to enhance real-world navigation and obstacle avoidance performance [180].
In autonomous swarm target search, the efficacy of UAV-EAI depends fundamentally on robust Sim-to-Real perception transfer. Current empirical studies reveal that the continuous ’dynamics gap’—arising from unmodeled aerodynamic disturbances and visual sparsity—causes severe performance degradation upon physical deployment. While baseline end-to-end policies often achieve >90% target acquisition success in parallel simulations, real-world deployments typically suffer a 40–50% drop in success rates [28,68]. To address this, state-of-the-art zero-shot Sim-to-Real transferring frameworks increasingly utilize asymmetric actor-critic architectures and privileged environmental distillation. By structurally integrating physical priors into the training pipeline, these methods mitigate the dynamics gap, sustaining real-world search success rates above 80% while bounding inference latency to under 50 ms [27].

6.2.2. Model Generalization

VLMs face significant generalization challenges in sparse and unstructured outdoor environments, where visual data are often incomplete and highly variable [181]. UAVs must recognize and interpret objects and environmental features across diverse scenarios, requiring strong cross-domain generalization capabilities that current models still lack [182]. Moreover, many existing autonomous exploration approaches rely on a single sensor modality, limiting information acquisition and reducing success rates in search tasks. Multimodal autonomous exploration systems aim to overcome these limitations by fusing heterogeneous sensor data [183].

6.2.3. Safety, Robustness, and Ethical Considerations in the LLM/VLM Era

The integration of foundation models shifts the safety bottleneck in UAV-EAI from low-level physical control to high-level cognitive robustness. Traditional safety paradigms must be expanded to address the systemic vulnerabilities introduced by VLMs and LLMs.
  • Semantic Hallucinations and Feasibility Mismatch: In unstructured aerial environments, foundation models are highly susceptible to out-of-distribution errors and semantic hallucinations. A critical failure mode is the disconnect between abstract planning and physical reality, where linguistically valid plans inadvertently violate kinematic envelopes or energy constraints. Mitigating these risks requires coupling cognitive models with physics-based feasibility filters to strictly bound symbolic reasoning within safe flight parameters [184].
  • Multi-Agent Cognitive Consistency: Swarm deployments relying on independent cognitive backbones face the risk of “cognitive divergence,” where ambiguous local observations lead to conflicting task prioritizations. Maintaining strategic coherence requires moving beyond basic spatial deconfliction. Implementing hierarchical organizational structures and explicit, role-based collaborative workflows within the swarm can enforce semantic consensus and align decentralized behaviors, even amidst local perception failures [185].
  • Algorithmic Bias and Information Security: VLMs pretrained predominantly on terrestrial data may propagate systematic biases during aerial target recognition, compromising fairness in critical missions like search-and-rescue. Furthermore, semantic-driven autonomy amplifies privacy risks. Robust UAV-EAI frameworks must therefore incorporate domain-specific debiasing, privacy-preserving perception, and dynamic cybersecurity protocols to secure decentralized cognitive networks [186].

6.2.4. Data Scarcity

Training high-performance embodied agents requires large-scale, diverse, and high-quality datasets [160]. However, UAV datasets for complex embodied tasks—such as manipulation, interaction, and fine-grained inspection—remain scarce, particularly in real-world scenarios where data collection and annotation are costly [172]. This data scarcity significantly constrains progress in Embodied. Although reinforcement learning can acquire complex skills through extensive trial and error, it typically demands carefully designed environments and human supervision [187]. Addressing data collection scalability, for example through reset-free reinforcement learning and related approaches, remains an open problem [188].
A fundamental impediment to aerial embodiment is the scarcity of task-specific datasets and benchmarks. While terrestrial EAI frameworks frequently leverage platforms such as AI2-THOR or Habitat, these environments generally lack the aerodynamic fidelity and perspective-driven scale variance essential for UAV operations. Recent developments have begun to address this deficit. For instance, CognitiveDroneBench [50] provides a dedicated evaluation framework for language-conditioned aerial agents, encompassing a broad range of simulated trajectories and cognitive tasks. To mitigate data scarcity in specialized sensing modalities, the SFBDA framework [87] introduces semantic-decoupled data augmentation, significantly improving infrared few-shot detection performance. Furthermore, bridging the dynamics gap necessitates advanced learning paradigms such as privileged environmental distillation. Recent studies demonstrate that utilizing privileged state information during training allows for the extraction of robust policies that account for unmodeled aerodynamic disturbances, thereby facilitating a more stable Sim-to-Real transition in unstructured 3D environments [27].
To address these challenges, future research directions include developing more advanced simulation environments and transfer learning techniques to narrow the Sim-to-Real gap; exploring multimodal and self-supervised learning methods to enhance generalization in unstructured environments; designing more robust control algorithms and safety assurance mechanisms to improve reliability in complex missions; and investigating more efficient data collection, annotation, and synthesis strategies to alleviate data scarcity. Furthermore, the integration of UAVs with next-generation communication networks, such as 5G and beyond, is expected to enhance communication and computational capabilities, providing a critical foundation for the continued advancement of UAV-EAI [189].

7. Conclusions

This survey argues that UAV-EAI constitutes a fundamentally distinct research paradigm rather than a simple extension of indoor Embodied or AD-EAI. Through a comparative, multi-dimensional analysis of perception, localization, navigation, and decision-making, we show that the combination of true 3D unstructured environments, full 6-DoF dynamics, high-speed motion, and severe payload and power constraints imposes challenges that are qualitatively different from those faced by ground-based embodied agents. Unlike indoor robots operating in structured 2.5D spaces or autonomous vehicles relying on quasi-2D road networks and rich infrastructure support, UAVs must perceive, reason, and act under sparse semantics, long-range uncertainty, aerodynamic disturbances, and limited sensing and computation.
These characteristics fundamentally reshape the requirements for perception and mapping, task semantics, simulation platforms, and learning paradigms. In particular, UAV-EAI demands lightweight yet robust multi-modal sensing, autonomy without reliance on dense prior maps, simulators that accurately capture aerial dynamics and long-horizon 3D interaction, and learning frameworks capable of generalization under data scarcity and Sim-to-Real gaps. Emerging vision-language and foundation models offer promising opportunities, but their integration must explicitly account for physical constraints, real-time decision-making, and safety-critical operation.
Overall, UAV-EAI should be recognized as an independent discipline centered on closed-loop cognition and action in open 3D worlds. Progress in this field requires moving beyond ground-centric assumptions and developing dedicated theories, datasets, benchmarks, and evaluation metrics tailored to aerial embodiment. Such advances are essential for enabling reliable and intelligent UAV systems in real-world applications, including search and rescue, infrastructure inspection, environmental monitoring, logistics, and precision agriculture.

Author Contributions

Conceptualization, Y.Z. and E.Z.; methodology, Y.Z. and Z.C.; validation, Z.C. and X.Z.; investigation, Y.Z., E.Z. and Z.C.; resources, X.Z. and Y.C.; writing—original draft preparation, Y.Z. and W.H.; writing—review and editing, E.Z., W.H. and Y.C.; visualization, Y.Z. and E.Z.; supervision, X.Z. and Y.C.; project administration, Z.C.; funding acquisition, X.Z. and B.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the Key Laboratory of Target Cognition and Application Technology under Grant 2023-CXPT-LC-005.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Zhang, Z.; Zhu, L. A review on unmanned aerial vehicle remote sensing: Platforms, sensors, data processing methods, and applications. Drones 2023, 7, 398. [Google Scholar] [CrossRef]
  2. Savva, M.; Kadian, A.; Maksymets, O.; Zhao, Y.; Wijmans, E.; Jain, B.; Straub, J.; Liu, J.; Koltun, V.; Malik, J.; et al. Habitat: A platform for embodied AI research. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2019; pp. 9339–9347. [Google Scholar]
  3. Duan, J.; Yu, S.; Tan, H.L.; Zhu, H.; Tan, C. A survey of embodied ai: From simulators to research tasks. IEEE Trans. Emerg. Top. Comput. Intell. 2022, 6, 230–244. [Google Scholar] [CrossRef]
  4. Yang, R.; Chen, H.; Zhang, J.; Zhao, M.; Qian, C.; Wang, K.; Wang, Q.; Koripella, T.V.; Movahedi, M.; Li, M.; et al. Embodiedbench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents. arXiv 2025, arXiv:2502.09560. [Google Scholar]
  5. Zhao, C.; Liu, R.W.; Qu, J.; Gao, R. Deep learning-based object detection in maritime unmanned aerial vehicle imagery: Review and experimental comparisons. Eng. Appl. Artif. Intell. 2024, 128, 107513. [Google Scholar] [CrossRef]
  6. Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; Van Den Hengel, A. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 3674–3683. [Google Scholar]
  7. Batra, D.; Gokaslan, A.; Kembhavi, A.; Maksymets, O.; Mottaghi, R.; Savva, M.; Toshev, A.; Wijmans, E. Objectnav revisited: On evaluation of embodied agents navigating to objects. arXiv 2020, arXiv:2006.13171. [Google Scholar] [CrossRef]
  8. Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Deitke, M.; Ehsani, K.; Gordon, D.; Zhu, Y.; et al. AI2-THOR: An interactive 3D environment for visual AI. arXiv 2017, arXiv:1712.05474. [Google Scholar]
  9. Ahn, M.; Brohan, A.; Brown, N.; Chebotar, Y.; Cortes, O.; David, B.; Finn, C.; Fu, C.; Gopalakrishnan, K.; Hausman, K.; et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv 2022, arXiv:2204.01691. [Google Scholar] [CrossRef]
  10. Grigorescu, S.; Trasnea, B.; Cocias, T.; Macesanu, G. A survey of deep learning techniques for autonomous driving. J. Field Robot. 2020, 37, 362–386. [Google Scholar] [CrossRef]
  11. Levinson, J.; Askeland, J.; Becker, J.; Dolson, J.; Held, D.; Kammel, S.; Kolter, J.Z.; Langer, D.; Pink, O.; Pratt, V.; et al. Towards fully autonomous driving: Systems and algorithms. In 2011 IEEE Intelligent Vehicles Symposium (IV); IEEE: Piscataway, NJ, USA, 2011; pp. 163–168. [Google Scholar]
  12. Pendleton, S.D.; Andersen, H.; Du, X.; Shen, X.; Meghjani, M.; Eng, Y.H.; Rus, D.; Ang, M.H., Jr. Perception, planning, control, and coordination for autonomous vehicles. Machines 2017, 5, 6. [Google Scholar] [CrossRef]
  13. Faessler, M.; Fontana, F.; Forster, C.; Mueggler, E.; Pizzoli, M.; Scaramuzza, D. Autonomous, vision-based flight and live dense 3D mapping with a quadrotor micro aerial vehicle. J. Field Robot. 2016, 33, 431–450. [Google Scholar] [CrossRef]
  14. Qin, T.; Li, P.; Shen, S. Vins-mono: A robust and versatile monocular visual-inertial state estimator. IEEE Trans. Robot. 2018, 34, 1004–1020. [Google Scholar] [CrossRef]
  15. Gupta, S.; Davidson, J.; Levine, S.; Sukthankar, R.; Malik, J. Cognitive mapping and planning for visual navigation. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2017; pp. 2616–2625. [Google Scholar]
  16. Shi, G.; Shi, X.; O’Connell, M.; Yu, R.; Azizzadenesheli, K.; Anandkumar, A.; Yue, Y.; Chung, S.J. Neural lander: Stable drone landing control using learned dynamics. In 2019 International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2019; pp. 9784–9790. [Google Scholar]
  17. Arnold, E.; Al-Jarrah, O.Y.; Dianati, M.; Fallah, S.; Oxtoby, D.; Mouzakitis, A. A survey on 3d object detection methods for autonomous driving applications. IEEE Trans. Intell. Transp. Syst. 2019, 20, 3782–3795. [Google Scholar] [CrossRef]
  18. Wang, M. High definition map for autonomous driving: Overview and analysis. Geomat. World 2020, 27, 109–114. [Google Scholar]
  19. Loquercio, A.; Kaufmann, E.; Ranftl, R.; Müller, M.; Koltun, V.; Scaramuzza, D. Learning high-speed flight in the wild. Sci. Robot. 2021, 6, eabg5810. [Google Scholar] [CrossRef] [PubMed]
  20. Zhou, B.; Zhang, Y.; Chen, X.; Shen, S. Fuel: Fast uav exploration using incremental frontier structure and hierarchical planning. IEEE Robot. Autom. Lett. 2021, 6, 779–786. [Google Scholar] [CrossRef]
  21. Cui, C.; Ma, Y.; Cao, X.; Ye, W.; Zhou, Y.; Liang, K.; Chen, J.; Lu, J.; Yang, Z.; Liao, K.D.; et al. A survey on multimodal large language models for autonomous driving. In 2024 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW); IEEE: Piscataway, NJ, USA, 2024; pp. 958–979. [Google Scholar]
  22. Chen, Z.; Wu, J.; Wang, W.; Su, W.; Chen, G.; Xing, S.; Zhong, M.; Zhang, Q.; Zhu, X.; Lu, L.; et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024); IEEE: Piscataway, NJ, USA, 2024; pp. 24185–24198. [Google Scholar]
  23. Xu, C.; Zhang, R.; Yang, W.; Zhu, H.; Xu, F.; Ding, J.; Xia, G.S. Oriented tiny object detection: A dataset, benchmark, and dynamic unbiased learning. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 48, 3167–3184. [Google Scholar] [CrossRef] [PubMed]
  24. Gao, Y.; Wang, Z.; Jing, L.; Wang, D.; Li, X.; Zhao, B. Aerial vision-and-language navigation via semantic-topo-metric representation guided LLM reasoning. arXiv 2024, arXiv:2410.08500. [Google Scholar]
  25. Liu, S.; Zhang, H.; Qi, Y.; Wang, P.; Zhang, Y.; Wu, Q. Aerialvln: Vision-and-language navigation for uavs. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV 2023); IEEE: Piscataway, NJ, USA, 2023; pp. 15384–15394. [Google Scholar]
  26. Kaufmann, E.; Bauersfeld, L.; Loquercio, A.; Müller, M.; Koltun, V.; Scaramuzza, D. Champion-level drone racing using deep reinforcement learning. Nature 2023, 620, 982–987. [Google Scholar] [CrossRef]
  27. Wang, J.; Yu, Z.; Zhou, D.; Shi, J.; Deng, R. Vision-based deep reinforcement learning of Unmanned Aerial Vehicle (UAV) autonomous navigation using privileged information. Drones 2024, 8, 782. [Google Scholar] [CrossRef]
  28. Hanover, D.; Loquercio, A.; Bauersfeld, L.; Romero, A.; Penicka, R.; Song, Y.; Cioffi, G.; Kaufmann, E.; Scaramuzza, D. Autonomous drone racing: A survey. IEEE Trans. Robot. 2024, 40, 3044–3067. [Google Scholar] [CrossRef]
  29. Strand, S.H.; Wiedemann, T.; Burczek, B.; Shutin, D. Enhancing UAV Search Under Occlusion Using Next Best View Planning. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 19, 1085–1096. [Google Scholar] [CrossRef]
  30. Datsko, D.; Nekovar, F.; Penicka, R.; Saska, M. Energy-aware multi-uav coverage mission planning with optimal speed of flight. IEEE Robot. Autom. Lett. 2024, 9, 2893–2900. [Google Scholar] [CrossRef]
  31. Li, Z.; Wu, W.; Guo, Y.; Sun, J.; Han, Q.L. Embodied multi-agent systems: A review. IEEE/CAA J. Autom. Sin. 2025, 12, 1095–1116. [Google Scholar] [CrossRef]
  32. Batinovic, A.; Ivanovic, A.; Petrovic, T.; Bogdan, S. A shadowcasting-based next-best-view planner for autonomous 3D exploration. IEEE Robot. Autom. Lett. 2022, 7, 2969–2976. [Google Scholar] [CrossRef]
  33. Li, Y.; Guo, X.; Zhang, H.; Li, S.; Dai, X. Active Visual Perception: Opportunities and Challenges. arXiv 2025, arXiv:2512.03687. [Google Scholar] [CrossRef]
  34. Zhai, Y.; Reiter, R.; Scaramuzza, D. Pa-mppi: Perception-aware model predictive path integral control for quadrotor navigation in unknown environments. IEEE Robot. Autom. Lett. 2026, 11, 3804–3811. [Google Scholar] [CrossRef]
  35. Charrow, B.; Kahn, G.; Patil, S.; Liu, S.; Goldberg, K.; Abbeel, P.; Michael, N.; Kumar, V. Information-Theoretic Planning with Trajectory Optimization for Dense 3D Mapping. In Proceedings of the Robotics: Science and Systems, Rome, Italy, 13–17 July 2015; Volume 11, pp. 3–12. [Google Scholar]
  36. Vanegas, F.; Campbell, D.; Eich, M.; Gonzalez, F. UAV based target finding and tracking in GPS-denied and cluttered environments. In 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2016; pp. 2307–2313. [Google Scholar]
  37. Cadena, C.; Carlone, L.; Carrillo, H.; Latif, Y.; Scaramuzza, D.; Neira, J.; Reid, I.; Leonard, J.J. Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age. IEEE Trans. Robot. 2017, 32, 1309–1332. [Google Scholar] [CrossRef]
  38. Rosinol, A.; Abate, M.; Chang, Y.; Carlone, L. Kimera: An open-source library for real-time metric-semantic localization and mapping. In 2020 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2020; pp. 1689–1696. [Google Scholar]
  39. Garg, S.; Sünderhauf, N.; Dayoub, F.; Morrison, D.; Cosgun, A.; Carneiro, G.; Wu, Q.; Chin, T.J.; Reid, I.; Gould, S.; et al. Semantics for robotic mapping, perception and interaction: A survey. Found. Trends® Robot. 2020, 8, 1–224. [Google Scholar] [CrossRef]
  40. Jatavallabhula, K.M.; Kuwajerwala, A.; Gu, Q.; Omama, M.; Chen, T.; Maalouf, A.; Li, S.; Iyer, G.; Saryazdi, S.; Keetha, N.; et al. Conceptfusion: Open-set multimodal 3d mapping. arXiv 2023, arXiv:2302.07241. [Google Scholar]
  41. Blum, H.; Rohrbach, S.; Popovic, M.; Bartolomei, L.; Siegwart, R. Active learning for UAV-based semantic mapping. arXiv 2019, arXiv:1908.11157. [Google Scholar] [CrossRef]
  42. Liu, X.; Liu, Y.; Qiu, H.; Yang, Q.; Lian, Z. Indooruav: Benchmarking vision-language uav navigation in continuous indoor environments. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2026; Volume 40, pp. 23864–23872. [Google Scholar]
  43. Shah, D.; Osiński, B.; Ichter, B.; Levine, S. Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2023; pp. 492–504. [Google Scholar]
  44. Yue, P.; Xin, J.; Zhang, Y.; Lu, Y.; Shan, M. Semantic-driven autonomous visual navigation for unmanned aerial vehicles. IEEE Trans. Ind. Electron. 2024, 71, 14853–14863. [Google Scholar] [CrossRef]
  45. Habibi, I.; Msadaa, I.C.; Grayaa, K. Adaptive UAV Inspection of PV Panels Using Goal-Conditioned Reinforcement Learning and Zigzag Coverage Planning. Proc. Mach. Learn. Res. 2025, 302, 1–17. [Google Scholar]
  46. Asgharivaskasi, A.; Atanasov, N. Semantic octree mapping and shannon mutual information computation for robot exploration. IEEE Trans. Robot. 2023, 39, 1910–1928. [Google Scholar] [CrossRef]
  47. Saxena, P.; Raghuvanshi, N.; Goveas, N. Uav-vln: End-to-end vision language guided navigation for uavs. In 2025 European Conference on Mobile Robots (ECMR); IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar]
  48. Zhang, X.; Tian, Y.; Lin, F.; Liu, Y.; Ma, J.; Wang, X.; Szatmáry, K.S.; Wang, F.Y. LogisticsVLN: Vision-language navigation for low-altitude terminal delivery based on agentic UAVs. In 2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC); IEEE: Piscataway, NJ, USA, 2025; pp. 4437–4442. [Google Scholar]
  49. Mehboob, F.; James, M.W.; Habel, A.A.; Sam, J.; Altamirano Cabrera, M.; Tsetserukou, D. DroneVLA: VLA-Based Aerial Manipulation. In Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction; Association for Computing Machinery: New York, NY, USA, 2026; pp. 1135–1139. [Google Scholar]
  50. Lykov, A.; Serpiva, V.; Khan, M.H.; Sautenkov, O.; Myshlyaev, A.; Tadevosyan, G.; Yaqoot, Y.; Tsetserukou, D. Cognitivedrone: A vla model and evaluation benchmark for real-time cognitive task solving and reasoning in uavs. arXiv 2025, arXiv:2503.01378. [Google Scholar]
  51. Rückin, J.; Magistri, F.; Stachniss, C.; Popović, M. An Informative Path Planning Framework for Active Learning in UAV-based Semantic Mapping. IEEE Trans. Robot. 2023, 39, 4279–4296. [Google Scholar] [CrossRef]
  52. Merei, A.; Mcheick, H.; Ghaddar, A.; Beltrame, G. EA-IPP2n: Energy-Aware Informative Path Planning Algorithm for UAVs. IEEE Access 2026, 14, 42381–42392. [Google Scholar] [CrossRef]
  53. Zheng, T.; Jin, Y.; Zhao, H.; Ma, Z.; Chen, Y.; Xu, K. Deep reinforcement learning based coverage path planning in unknown environments. In 2024 6th International Conference on Frontier Technologies of Information and Computer (ICFTIC); IEEE: Piscataway, NJ, USA, 2024; pp. 1608–1611. [Google Scholar]
  54. Song, F.; Zeng, Q.; Zhang, R.; Zhu, X.; Ye, X.; Zhang, Z. Multi-UAV Cooperative Navigation Based on Multi-Source Information Fusion: A Review. IEEE Sens. J. 2025, 26, 3460–3490. [Google Scholar] [CrossRef]
  55. Rizk, Y.; Awad, M.; Tunstel, E.W. Cooperative heterogeneous multi-robot systems: A survey. ACM Comput. Surv. (CSUR) 2019, 52, 29. [Google Scholar] [CrossRef]
  56. Amato, C.; Chowdhary, G.; Geramifard, A.; Üre, N.K.; Kochenderfer, M.J. Decentralized control of partially observable Markov decision processes. In 52nd IEEE Conference on Decision and Control; IEEE: Piscataway, NJ, USA, 2013; pp. 2398–2405. [Google Scholar]
  57. Zhai, Z.; Ni, W.; Wang, X.; Niyato, D.; Hossain, E. Integrated Sensing and Communication with UAV Swarms via Decentralized Consensus ADMM. arXiv 2025, arXiv:2511.03283. [Google Scholar] [CrossRef]
  58. Tang, H.K.; Lake, M.J.; Foster, R.J.; Bezombes, F.A. Design, development and pilot of a realistic virtual reality application to analyse quick directional change in sport: Avatar cutting scenario with alterable parameters. PLoS ONE 2025, 20, e0324941. [Google Scholar] [CrossRef]
  59. Coumans, E. Bullet physics simulation. In ACM SIGGRAPH 2015 Courses; Association for Computing Machinery: New York, NY, USA, 2015. [Google Scholar]
  60. Mellinger, D.; Kumar, V. Minimum snap trajectory generation and control for quadrotors. In 2011 IEEE International Conference on Robotics and Automation; IEEE: Piscataway, NJ, USA, 2011; pp. 2520–2525. [Google Scholar]
  61. Shah, S.; Dey, D.; Lovett, C.; Kapoor, A. Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and Service Robotics: Results of the 11th International Conference; Springer: Cham, Switzerland, 2017; pp. 621–635. [Google Scholar]
  62. Faessler, M.; Franchi, A.; Scaramuzza, D. Differential flatness of quadrotor dynamics subject to rotor drag for accurate tracking of high-speed trajectories. IEEE Robot. Autom. Lett. 2017, 3, 620–626. [Google Scholar] [CrossRef]
  63. Makoviychuk, V.; Wawrzyniak, L.; Guo, Y.; Lu, M.; Storey, K.; Macklin, M.; Hoeller, D.; Rudin, N.; Allshire, A.; Handa, A.; et al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv 2021, arXiv:2108.10470. [Google Scholar] [CrossRef]
  64. Guerra, W.; Tal, E.; Murali, V.; Ryou, G.; Karaman, S. Flightgoggles: Photorealistic sensor simulation for perception-driven robotics using photogrammetry and virtual reality. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2019; pp. 6941–6948. [Google Scholar]
  65. Zhu, Y.; Wong, J.; Mandlekar, A.; Martín-Martín, R.; Joshi, A.; Lin, K.; Maddukuri, A.; Nasiriany, S.; Zhu, Y. robosuite: A modular simulation framework and benchmark for robot learning. arXiv 2020, arXiv:2009.12293. [Google Scholar]
  66. Rudin, N.; Hoeller, D.; Reist, P.; Hutter, M. Learning to walk in minutes using massively parallel deep reinforcement learning. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2022; pp. 91–100. [Google Scholar]
  67. Jacinto, M.; Pinto, J.; Patrikar, J.; Keller, J.; Cunha, R.; Scherer, S.; Pascoal, A. Pegasus simulator: An isaac sim framework for multiple aerial vehicles simulation. In 2024 International Conference on Unmanned Aircraft Systems (ICUAS); IEEE: Piscataway, NJ, USA, 2024; pp. 917–922. [Google Scholar]
  68. Xu, B.; Gao, F.; Yu, C.; Zhang, R.; Wu, Y.; Wang, Y. Omnidrones: An efficient and flexible platform for reinforcement learning in drone control. IEEE Robot. Autom. Lett. 2024, 9, 2838–2844. [Google Scholar] [CrossRef]
  69. Handa, A.; Allshire, A.; Makoviychuk, V.; Petrenko, A.; Singh, R.; Liu, J.; Makoviichuk, D.; Van Wyk, K.; Zhurkevich, A.; Sundaralingam, B.; et al. Dextreme: Transfer of agile in-hand manipulation from simulation to reality. In 2023 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2023; pp. 5977–5984. [Google Scholar]
  70. Wang, M.; Niu, Y.; Wang, B.; Zhang, W.; Wang, C. A Survey on Learning Motion Planning and Control for Mobile Robots: Toward Embodied Intelligence. IEEE Trans. Neural Netw. Learn. Syst. 2026. Early Access. [Google Scholar] [CrossRef]
  71. Lyu, Y.; Vosselman, G.; Xia, G.S.; Yilmaz, A.; Yang, M.Y. UAVid: A semantic segmentation dataset for UAV imagery. ISPRS J. Photogramm. Remote Sens. 2020, 165, 108–119. [Google Scholar] [CrossRef]
  72. Zhu, P.; Wen, L.; Du, D.; Bian, X.; Ling, H.; Hu, Q.; Nie, Q.; Cheng, H.; Liu, C.; Liu, X.; et al. Visdrone-det2018: The vision meets drone object detection in image challenge results. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany, 8–14 September 2018. [Google Scholar]
  73. Burri, M.; Nikolic, J.; Gohl, P.; Schneider, T.; Rehder, J.; Omari, S.; Achtelik, M.W.; Siegwart, R. The EuRoC micro aerial vehicle datasets. Int. J. Robot. Res. 2016, 35, 1157–1163. [Google Scholar] [CrossRef]
  74. Ferrera, M.; Eudes, A.; Moras, J.; Sanfourche, M.; Le Besnerais, G. OV2SLAM: A fully online and versatile visual SLAM for real-time applications. IEEE Robot. Autom. Lett. 2021, 6, 1399–1406. [Google Scholar] [CrossRef]
  75. Thrun, S. Probabilistic robotics. Commun. ACM 2002, 45, 52–57. [Google Scholar] [CrossRef]
  76. Lu, Y.; Xue, Z.; Xia, G.S.; Zhang, L. A survey on vision-based UAV navigation. Geo-Spat. Inf. Sci. 2018, 21, 21–32. [Google Scholar] [CrossRef]
  77. Popović, M.; Vidal-Calleja, T.; Chung, J.J.; Nieto, J.; Siegwart, R. Informative path planning for active field mapping under localization uncertainty. In 2020 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2020; pp. 10751–10757. [Google Scholar]
  78. Wang, C.; Yu, C.; Xu, X.; Gao, Y.; Yang, X.; Tang, W.; Yu, S.; Chen, Y.; Gao, F.; Jian, Z.; et al. Multi-robot system for cooperative exploration in unknown environments: A survey. arXiv 2025, arXiv:2503.07278. [Google Scholar] [CrossRef]
  79. Schwager, M.; Rus, D.; Slotine, J.J. Decentralized, adaptive coverage control for networked robots. Int. J. Robot. Res. 2009, 28, 357–375. [Google Scholar] [CrossRef]
  80. Low, K.H.; Dolan, J.; Khosla, P. Information-theoretic approach to efficient adaptive path planning for mobile robotic environmental sensing. In Proceedings of the International Conference on Automated Planning and Scheduling; AAAI Press: Washington, DC, USA, 2009; Volume 19, pp. 233–240. [Google Scholar]
  81. Schwager, M.; Julian, B.J.; Angermann, M.; Rus, D. Eyes in the sky: Decentralized control for the deployment of robotic camera networks. Proc. IEEE 2011, 99, 1541–1561. [Google Scholar] [CrossRef]
  82. Bourgault, F.; Makarenko, A.A.; Williams, S.B.; Grocholsky, B.; Durrant-Whyte, H.F. Information based adaptive robotic exploration. In IEEE/RSJ International Conference on Intelligent Robots and Systems; IEEE: Piscataway, NJ, USA, 2002; Volume 1, pp. 540–545. [Google Scholar]
  83. Ye, X.; Song, F.; Zhang, Z.; Zeng, Q. A review of small UAV navigation system based on multisource sensor fusion. IEEE Sens. J. 2023, 23, 18926–18948. [Google Scholar] [CrossRef]
  84. Lee, J.H.; Choi, J.S.; Jeon, E.S.; Kim, Y.G.; Thanh Le, T.; Shin, K.Y.; Lee, H.C.; Park, K.R. Robust pedestrian detection by combining visible and thermal infrared cameras. Sensors 2015, 15, 10580–10615. [Google Scholar] [CrossRef] [PubMed]
  85. Manfreda, S.; McCabe, M.F.; Miller, P.E.; Lucas, R.; Pajuelo Madrigal, V.; Mallinis, G.; Ben Dor, E.; Helman, D.; Estes, L.; Ciraolo, G.; et al. On the use of unmanned aerial systems for environmental monitoring. Remote Sens. 2018, 10, 641. [Google Scholar] [CrossRef]
  86. Samadzadegan, F.; Toosi, A.; Dadrass Javan, F. A critical review on multi-sensor and multi-platform remote sensing data fusion approaches: Current status and prospects. Int. J. Remote Sens. 2025, 46, 1327–1402. [Google Scholar] [CrossRef]
  87. Weng, Z.; He, W.; Lv, J.; Zhou, D.; Yu, Z. SFBDA: A Semantic-Decoupled Data Augmentation Framework for Infrared Few-Shot Object Detection on UAVs. IEEE Geosci. Remote Sens. Lett. 2025, 22, 7002205. [Google Scholar] [CrossRef]
  88. Lee, T.; Leok, M.; McClamroch, N.H. Geometric tracking control of a quadrotor UAV on SE (3). In 49th IEEE Conference on Decision and Control (CDC); IEEE: Piscataway, NJ, USA, 2010; pp. 5420–5425. [Google Scholar]
  89. Lee, T. Geometric control of quadrotor UAVs transporting a cable-suspended rigid body. IEEE Trans. Control Syst. Technol. 2017, 26, 255–264. [Google Scholar] [CrossRef]
  90. Mellinger, D.; Michael, N.; Kumar, V. Trajectory generation and control for precise aggressive maneuvers with quadrotors. Int. J. Robot. Res. 2012, 31, 664–674. [Google Scholar] [CrossRef]
  91. Hewing, L.; Wabersich, K.P.; Menner, M.; Zeilinger, M.N. Learning-based model predictive control: Toward safe learning in control. Annu. Rev. Control Robot. Auton. Syst. 2020, 3, 269–296. [Google Scholar] [CrossRef]
  92. Ames, A.D.; Coogan, S.; Egerstedt, M.; Notomista, G.; Sreenath, K.; Tabuada, P. Control barrier functions: Theory and applications. In 2019 18th European Control Conference (ECC); IEEE: Piscataway, NJ, USA, 2019; pp. 3420–3431. [Google Scholar]
  93. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016); IEEE: Piscataway, NJ, USA, 2016; pp. 770–778. [Google Scholar]
  94. Tan, M.; Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2019; pp. 6105–6114. [Google Scholar]
  95. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
  96. Gevaert, C.M.; Belgiu, M. Assessing the generalization capability of deep learning networks for aerial image classification using landscape metrics. Int. J. Appl. Earth Obs. Geoinf. 2022, 114, 103054. [Google Scholar] [CrossRef]
  97. Vivone, G.; Deng, L.J.; Deng, S.; Hong, D.; Jiang, M.; Li, C.; Li, W.; Shen, H.; Wu, X.; Xiao, J.L.; et al. Deep learning in remote sensing image fusion: Methods, protocols, data, and future perspectives. IEEE Geosci. Remote Sens. Mag. 2024, 13, 269–310. [Google Scholar] [CrossRef]
  98. Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Computer Vision—ECCV 2014; Springer: Cham, Switzerland, 2014; pp. 740–755. [Google Scholar]
  99. Zhu, X.X.; Tuia, D.; Mou, L.; Xia, G.S.; Zhang, L.; Xu, F.; Fraundorfer, F. Deep learning in remote sensing: A comprehensive review and list of resources. IEEE Geosci. Remote Sens. Mag. 2017, 5, 8–36. [Google Scholar] [CrossRef]
  100. Li, Z.; Wang, Y.; Zhang, N.; Zhang, Y.; Zhao, Z.; Xu, D.; Ben, G.; Gao, Y. Deep learning-based object detection techniques for remote sensing images: A survey. Remote Sens. 2022, 14, 2385. [Google Scholar] [CrossRef]
  101. Gao, G.; Liu, Q.; Wang, Y. Counting from sky: A large-scale data set for remote sensing object counting and a benchmark method. IEEE Trans. Geosci. Remote Sens. 2020, 59, 3642–3655. [Google Scholar] [CrossRef]
  102. Jiao, L.; Zhang, F.; Liu, F.; Yang, S.; Li, L.; Feng, Z.; Qu, R. A survey of deep learning-based object detection. IEEE Access 2019, 7, 128837–128868. [Google Scholar] [CrossRef]
  103. Xu, D.; Ricci, E.; Ouyang, W.; Wang, X.; Sebe, N. Monocular depth estimation using multi-scale continuous crfs as sequential deep networks. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 41, 1426–1440. [Google Scholar] [CrossRef]
  104. Kendall, A.; Martirosyan, H.; Dasgupta, S.; Henry, P.; Kennedy, R.; Bachrach, A.; Bry, A. End-to-end learning of geometry and context for deep stereo regression. In 2017 IEEE International Conference on Computer Vision (ICCV 2017); IEEE: Piscataway, NJ, USA, 2017; pp. 66–75. [Google Scholar]
  105. Arafat, M.Y.; Alam, M.M.; Moh, S. Vision-based navigation techniques for unmanned aerial vehicles: Review and challenges. Drones 2023, 7, 89. [Google Scholar] [CrossRef]
  106. Gosala, N.; Petek, K.; Drews, P.L., Jr.; Burgard, W.; Valada, A. Skyeye: Self-supervised bird’s-eye-view semantic mapping using monocular frontal view images. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2023); IEEE: Piscataway, NJ, USA, 2023; pp. 14901–14910. [Google Scholar]
  107. Liu, Z.; Tang, H.; Amini, A.; Yang, X.; Mao, H.; Rus, D.L.; Han, S. Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation. In 2023 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2023; pp. 2774–2781. [Google Scholar]
  108. Gulrajani, I.; Lopez-Paz, D. In search of lost domain generalization. arXiv 2020, arXiv:2007.01434. [Google Scholar] [CrossRef]
  109. Bucci, S.; D’Innocente, A.; Liao, Y.; Carlucci, F.M.; Caputo, B.; Tommasi, T. Self-supervised learning across domains. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 5516–5528. [Google Scholar] [CrossRef]
  110. Tobin, J.; Fong, R.; Ray, A.; Schneider, J.; Zaremba, W.; Abbeel, P. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2017; pp. 23–30. [Google Scholar]
  111. Shorten, C.; Khoshgoftaar, T.M. A survey on image data augmentation for deep learning. J. Big Data 2019, 6, 60. [Google Scholar] [CrossRef]
  112. Rajapaksha, U.; Sohel, F.; Laga, H.; Diepeveen, D.; Bennamoun, M. Deep learning-based depth estimation methods from monocular image and videos: A comprehensive survey. ACM Comput. Surv. 2024, 56, 315. [Google Scholar] [CrossRef]
  113. Qi, Q.; Wang, G.; Pan, Y.; Fan, H.; Li, B. MCS-Sim: A Photo-Realistic Simulator for Multi-Camera UAV Visual Perception Research. Drones 2025, 9, 656. [Google Scholar] [CrossRef]
  114. Zhou, T.; Brown, M.; Snavely, N.; Lowe, D.G. Unsupervised learning of depth and ego-motion from video. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2017; pp. 1851–1858. [Google Scholar]
  115. Garg, R.; Bg, V.K.; Carneiro, G.; Reid, I. Unsupervised cnn for single view depth estimation: Geometry to the rescue. In Computer Vision—ECCV 2016; Springer: Cham, Switzerland, 2016; pp. 740–756. [Google Scholar]
  116. Godard, C.; Mac Aodha, O.; Firman, M.; Brostow, G.J. Digging into self-supervised monocular depth estimation. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2019; pp. 3828–3838. [Google Scholar]
  117. Maslov, D.; Makarov, I. Online supervised attention-based recurrent depth estimation from monocular video. PeerJ Comput. Sci. 2020, 6, e317. [Google Scholar] [CrossRef]
  118. Bochkovskiy, A.; Wang, C.Y.; Liao, H.Y.M. Yolov4: Optimal speed and accuracy of object detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef]
  119. Wojke, N.; Bewley, A.; Paulus, D. Simple online and realtime tracking with a deep association metric. In 2017 IEEE International Conference on Image Processing (ICIP); IEEE: Piscataway, NJ, USA, 2017; pp. 3645–3649. [Google Scholar]
  120. Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. Transunet: Transformers make strong encoders for medical image segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef]
  121. Zhao, C.; Zhang, Y.; Poggi, M.; Tosi, F.; Guo, X.; Zhu, Z.; Huang, G.; Tang, Y.; Mattoccia, S. Monovit: Self-supervised monocular depth estimation with a vision transformer. In 2022 International Conference on 3D Vision (3DV); IEEE: Piscataway, NJ, USA, 2022; pp. 668–678. [Google Scholar]
  122. Kendall, A.; Gal, Y.; Cipolla, R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018); IEEE: Piscataway, NJ, USA, 2018; pp. 7482–7491. [Google Scholar]
  123. Teed, Z.; Deng, J. Deepv2d: Video to depth with differentiable structure from motion. arXiv 2018, arXiv:1812.04605. [Google Scholar]
  124. Tsouros, D.C.; Bibi, S.; Sarigiannidis, P.G. A review on UAV-based applications for precision agriculture. Information 2019, 10, 349. [Google Scholar] [CrossRef]
  125. Eskandari, R.; Mahdianpari, M.; Mohammadimanesh, F.; Salehi, B.; Brisco, B.; Homayouni, S. Meta-analysis of unmanned aerial vehicle (UAV) imagery for agro-environmental monitoring using machine learning and statistical models. Remote Sens. 2020, 12, 3511. [Google Scholar] [CrossRef]
  126. Kunze, L.; Hawes, N.; Duckett, T.; Hanheide, M.; Krajník, T. Artificial intelligence for long-term robot autonomy: A survey. IEEE Robot. Autom. Lett. 2018, 3, 4023–4030. [Google Scholar] [CrossRef]
  127. Loquercio, A.; Kaufmann, E.; Ranftl, R.; Dosovitskiy, A.; Koltun, V.; Scaramuzza, D. Deep drone racing: From simulation to reality with domain randomization. IEEE Trans. Robot. 2019, 36, 1–14. [Google Scholar] [CrossRef]
  128. Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Pieter Abbeel, O.; Zaremba, W. Hindsight experience replay. Adv. Neural Inf. Process. Syst. 2017, 30, 5055–5065. [Google Scholar]
  129. Pathak, D.; Agrawal, P.; Efros, A.A.; Darrell, T. Curiosity-driven exploration by self-supervised prediction. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2017; pp. 2778–2787. [Google Scholar]
  130. Burda, Y.; Edwards, H.; Storkey, A.; Klimov, O. Exploration by random network distillation. arXiv 2018, arXiv:1810.12894. [Google Scholar] [CrossRef]
  131. Nachum, O.; Gu, S.S.; Lee, H.; Levine, S. Data-efficient hierarchical reinforcement learning. Adv. Neural Inf. Process. Syst. 2018, 31, 3307–3317. [Google Scholar]
  132. Chaysri, P.; Spatharis, C.; Blekas, K.; Vlachos, K. Unmanned surface vehicle navigation through generative adversarial imitation learning. Ocean Eng. 2023, 282, 114989. [Google Scholar] [CrossRef]
  133. Oh, J.; Guo, Y.; Singh, S.; Lee, H. Self-imitation learning. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; pp. 3878–3887. [Google Scholar]
  134. Williams, G.; Wagener, N.; Goldfain, B.; Drews, P.; Rehg, J.M.; Boots, B.; Theodorou, E.A. Information theoretic MPC for model-based reinforcement learning. In 2017 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2017; pp. 1714–1721. [Google Scholar]
  135. Hausknecht, M.; Stone, P. Deep reinforcement learning in parameterized action space. arXiv 2015, arXiv:1511.04143. [Google Scholar]
  136. Hafner, D.; Lillicrap, T.; Fischer, I.; Villegas, R.; Ha, D.; Lee, H.; Davidson, J. Learning latent dynamics for planning from pixels. In Proceedings of the 36th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2019; pp. 2555–2565. [Google Scholar]
  137. Molchanov, A.; Chen, T.; Hönig, W.; Preiss, J.A.; Ayanian, N.; Sukhatme, G.S. Sim-to-(multi)-real: Transfer of low-level robust control policies to multiple quadrotors. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2019; pp. 59–66. [Google Scholar]
  138. Kim, M.J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; et al. Openvla: An open-source vision-language-action model. arXiv 2024, arXiv:2406.09246. [Google Scholar]
  139. Zitkovich, B.; Yu, T.; Xu, S.; Xu, P.; Xiao, T.; Xia, F.; Wu, J.; Wohlhart, P.; Welker, S.; Wahid, A.; et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2023; pp. 2165–2183. [Google Scholar]
  140. Han, S.; Mao, H.; Dally, W.J. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv 2015, arXiv:1510.00149. [Google Scholar]
  141. Hinton, G.; Vinyals, O.; Dean, J. Distilling the knowledge in a neural network. arXiv 2015, arXiv:1503.02531. [Google Scholar] [CrossRef]
  142. Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P.; et al. Soft actor-critic algorithms and applications. arXiv 2018, arXiv:1812.05905. [Google Scholar]
  143. Alayrac, J.B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. Flamingo: A visual language model for few-shot learning. Adv. Neural Inf. Process. Syst. 2022, 35, 23716–23736. [Google Scholar]
  144. Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual instruction tuning. Adv. Neural Inf. Process. Syst. 2023, 36, 34892–34916. [Google Scholar]
  145. Li, J.; Li, D.; Savarese, S.; Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2023; pp. 19730–19742. [Google Scholar]
  146. Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
  147. Bommasani, R.; Hudson, D.A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M.S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. On the opportunities and risks of foundation models. arXiv 2021, arXiv:2108.07258. [Google Scholar] [CrossRef]
  148. Li, J.; Li, D.; Xiong, C.; Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2022; pp. 12888–12900. [Google Scholar]
  149. Cheng, G.; Han, J.; Lu, X. Remote sensing image scene classification: Benchmark and state of the art. Proc. IEEE 2017, 105, 1865–1883. [Google Scholar] [CrossRef]
  150. Cossio, M. A comprehensive taxonomy of hallucinations in large language models. arXiv 2025, arXiv:2508.01781. [Google Scholar] [CrossRef]
  151. Huynh, A.V.; Gillespie, L.E.; Lopez-Saucedo, J.; Tang, C.; Sikand, R.; Expósito-Alonso, M. Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery. In Computer Vision—ECCV 2024; Springer: Cham, Switzerland, 2024; pp. 173–190. [Google Scholar]
  152. Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. In Proceedings of the Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, 25–29 April 2022; Volume 1, p. 3. [Google Scholar]
  153. Wang, Y.; Guizilini, V.C.; Zhang, T.; Wang, Y.; Zhao, H.; Solomon, J. Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2022; pp. 180–191. [Google Scholar]
  154. Zhou, L.; Zhao, S.; Wan, Z.; Liu, Y.; Wang, Y.; Zuo, X. MFEFNet: A multi-scale feature information extraction and fusion network for multi-scale object detection in UAV aerial images. Drones 2024, 8, 186. [Google Scholar] [CrossRef]
  155. Driess, D.; Xia, F.; Sajjadi, M.S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. Palm-e: An embodied multimodal language model. arXiv 2023, arXiv:2303.03378. [Google Scholar]
  156. Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901. [Google Scholar]
  157. Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H.P.D.O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. Evaluating large language models trained on code. arXiv 2021, arXiv:2107.03374. [Google Scholar] [CrossRef]
  158. Yang, S.; Nachum, O.; Du, Y.; Wei, J.; Abbeel, P.; Schuurmans, D. Foundation models for decision making: Problems, methods, and opportunities. arXiv 2023, arXiv:2303.04129. [Google Scholar] [CrossRef]
  159. Valmeekam, K.; Marquez, M.; Sreedharan, S.; Kambhampati, S. On the planning abilities of large language models-a critical investigation. Adv. Neural Inf. Process. Syst. 2023, 36, 75993–76005. [Google Scholar]
  160. Xu, Z.; Wu, K.; Wen, J.; Li, J.; Liu, N.; Che, Z.; Tang, J. A survey on robotics with foundation models: Toward embodied ai. arXiv 2024, arXiv:2402.02385. [Google Scholar] [CrossRef]
  161. Rana, K.; Haviland, J.; Garg, S.; Abou-Chakra, J.; Reid, I.; Suenderhauf, N. Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning. arXiv 2023, arXiv:2307.06135. [Google Scholar] [CrossRef]
  162. Xiao, W.; Cassandras, C.G.; Belta, C. Safe Autonomy with Control Barrier Functions: Theory and Applications; Springer: Cham, Switzerland, 2023. [Google Scholar]
  163. Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.R.; Cao, Y. React: Synergizing reasoning and acting in language models. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2022. [Google Scholar]
  164. Xia, X.; Fattah, S.M.M.; Babar, M.A. A survey on UAV-enabled edge computing: Resource management perspective. ACM Comput. Surv. 2023, 56, 78. [Google Scholar] [CrossRef]
  165. Lin, J.; Tang, J.; Tang, H.; Yang, S.; Xiao, G.; Han, S. Awq: Activation-aware weight quantization for on-device llm compression and acceleration. GetMob. Mob. Comput. Commun. 2025, 28, 12–17. [Google Scholar] [CrossRef]
  166. Dettmers, T.; Lewis, M.; Belkada, Y.; Zettlemoyer, L. Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale. Adv. Neural Inf. Process. Syst. 2022, 35, 30318–30332. [Google Scholar]
  167. Micikevicius, P.; Stosic, D.; Burgess, N.; Cornea, M.; Dubey, P.; Grisenthwaite, R.; Ha, S.; Heinecke, A.; Judd, P.; Kamalu, J.; et al. Fp8 formats for deep learning. arXiv 2022, arXiv:2209.05433. [Google Scholar] [CrossRef]
  168. Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv 2017, arXiv:1701.06538. [Google Scholar]
  169. Fedus, W.; Zoph, B.; Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res. 2022, 23, 1–39. [Google Scholar]
  170. Shi, W.; Cao, J.; Zhang, Q.; Li, Y.; Xu, L. Edge computing: Vision and challenges. IEEE Internet Things J. 2016, 3, 637–646. [Google Scholar] [CrossRef]
  171. Ponzina, F.; Machetti, S.; Rios, M.; Denkinger, B.W.; Levisse, A.; Ansaloni, G.; Peón-Quirós, M.; Atienza, D. A hardware/software co-design vision for deep learning at the edge. IEEE Micro 2022, 42, 48–54. [Google Scholar] [CrossRef]
  172. Xiao, J.; Zhang, R.; Zhang, Y.; Feroskhan, M. Vision-based learning for drones: A survey. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 15601–15621. [Google Scholar] [CrossRef] [PubMed]
  173. Kormushev, P.; Calinon, S.; Saegusa, R.; Metta, G. Learning the skill of archery by a humanoid robot iCub. In 2010 10th IEEE-RAS International Conference on Humanoid Robots; IEEE: Piscataway, NJ, USA, 2010; pp. 417–423. [Google Scholar]
  174. Qi, P.; Zhao, X. Flight control for very flexible aircraft using model-free adaptive control. J. Guid. Control Dyn. 2020, 43, 608–619. [Google Scholar] [CrossRef]
  175. Van Den Berg, J.; Abbeel, P.; Goldberg, K. LQG-MP: Optimized path planning for robots with motion uncertainty and imperfect state information. Int. J. Robot. Res. 2011, 30, 895–913. [Google Scholar] [CrossRef]
  176. Ha, D.; Schmidhuber, J. World models. arXiv 2018, arXiv:1803.10122. [Google Scholar]
  177. Ollero, A.; Tognon, M.; Suarez, A.; Lee, D.; Franchi, A. Past, present, and future of aerial robotic manipulators. IEEE Trans. Robot. 2021, 38, 626–645. [Google Scholar] [CrossRef]
  178. Ware, J.; Roy, N. An analysis of wind field estimation and exploitation for quadrotor flight in the urban canopy layer. In 2016 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2016; pp. 1507–1514. [Google Scholar]
  179. Jakobi, N.; Husbands, P.; Harvey, I. Noise and the reality gap: The use of simulation in evolutionary robotics. In Advances in Artificial Life: Third European Conference on Artificial Life, Granada, Spain, June 4–6, 1995 Proceedings; Springer: Berlin/Heidelberg, Germany, 1995; pp. 704–720. [Google Scholar]
  180. Hwangbo, J.; Sa, I.; Siegwart, R.; Hutter, M. Control of a quadrotor with reinforcement learning. IEEE Robot. Autom. Lett. 2017, 2, 2096–2103. [Google Scholar] [CrossRef]
  181. Miyai, A.; Yang, J.; Zhang, J.; Ming, Y.; Lin, Y.; Yu, Q.; Irie, G.; Joty, S.; Li, Y.; Li, H.; et al. Generalized out-of-distribution detection and beyond in vision language model era: A survey. arXiv 2024, arXiv:2407.21794. [Google Scholar]
  182. Li, S.; Chaplot, D.S.; Tsai, Y.H.H.; Wu, Y.; Morency, L.P.; Salakhutdinov, R. Unsupervised domain adaptation for visual navigation. arXiv 2020, arXiv:2010.14543. [Google Scholar] [CrossRef]
  183. Zhang, Y.; Liu, Y.; Liu, S.; Liang, W.; Wang, C.; Wang, K. Multimodal perception for indoor mobile robotics navigation and safe manipulation. IEEE Trans. Cogn. Dev. Syst. 2024, 17, 1074–1086. [Google Scholar] [CrossRef]
  184. Ahn, M.; Dwibedi, D.; Finn, C.; Arenas, M.G.; Gopalakrishnan, K.; Hausman, K.; Ichter, B.; Irpan, A.; Joshi, N.; Julian, R.; et al. Autort: Embodied foundation models for large scale orchestration of robotic agents. arXiv 2024, arXiv:2401.12963. [Google Scholar] [CrossRef]
  185. Mandi, Z.; Jain, S.; Song, S. Roco: Dialectic multi-robot collaboration with large language models. In 2024 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2024; pp. 286–299. [Google Scholar]
  186. Huang, Y.; Sun, L.; Wang, H.; Wu, S.; Zhang, Q.; Li, Y.; Gao, C.; Huang, Y.; Lyu, W.; Zhang, Y.; et al. Trustllm: Trustworthiness in large language models. arXiv 2024, arXiv:2401.05561. [Google Scholar] [CrossRef]
  187. Dulac-Arnold, G.; Levine, N.; Mankowitz, D.J.; Li, J.; Paduraru, C.; Gowal, S.; Hester, T. Challenges of real-world reinforcement learning: Definitions, benchmarks and analysis. Mach. Learn. 2021, 110, 2419–2468. [Google Scholar] [CrossRef]
  188. Ecoffet, A.; Huizinga, J.; Lehman, J.; Stanley, K.O.; Clune, J. First return, then explore. Nature 2021, 590, 580–586. [Google Scholar] [CrossRef]
  189. Mozaffari, M.; Saad, W.; Bennis, M.; Nam, Y.H.; Debbah, M. A tutorial on UAVs for wireless networks: Applications, challenges, and open problems. IEEE Commun. Surv. Tutor. 2019, 21, 2334–2360. [Google Scholar] [CrossRef]
Figure 1. Main sections and the structure of this review.
Figure 1. Main sections and the structure of this review.
Remotesensing 18 01509 g001
Figure 2. An example of Schematic Diagram of Drone Collaboration at Different Altitudes, where a high-altitude UAV acquires macroscopic aerial imagery for target detection, guiding the low-altitude UAV in precise target localization.
Figure 2. An example of Schematic Diagram of Drone Collaboration at Different Altitudes, where a high-altitude UAV acquires macroscopic aerial imagery for target detection, guiding the low-altitude UAV in precise target localization.
Remotesensing 18 01509 g002
Figure 3. The timeline of the development of Simulation Environments.
Figure 3. The timeline of the development of Simulation Environments.
Remotesensing 18 01509 g003
Figure 4. Comparison of representative UAV datasets across five core dimensions. Scores range from 0 (lowest) to 5 (highest) to highlight structural disparities in dataset capabilities. An ideal dataset should perform well in all the aspects.
Figure 4. Comparison of representative UAV datasets across five core dimensions. Scores range from 0 (lowest) to 5 (highest) to highlight structural disparities in dataset capabilities. An ideal dataset should perform well in all the aspects.
Remotesensing 18 01509 g004
Figure 5. Illustration of Perspective Mismatch and Domain Adaptation in UAV Perception. Domain adaptation techniques bridge this gap by learning a shared, view-invariant feature space for robust task prediction.
Figure 5. Illustration of Perspective Mismatch and Domain Adaptation in UAV Perception. Domain adaptation techniques bridge this gap by learning a shared, view-invariant feature space for robust task prediction.
Remotesensing 18 01509 g005
Figure 6. The Neuro-Symbolic Planning Pipeline with a Physics-Based Feasibility Filter.
Figure 6. The Neuro-Symbolic Planning Pipeline with a Physics-Based Feasibility Filter.
Remotesensing 18 01509 g006
Figure 7. Overview of UAV-EAI Application Domains and Core Capabilities.
Figure 7. Overview of UAV-EAI Application Domains and Core Capabilities.
Remotesensing 18 01509 g007
Table 1. Comparison of dimensions across UAV-EAI, AD-EAI and Indoor-EAI. This framework highlights the unique “physical-cognitive” duality of aerial agents.
Table 1. Comparison of dimensions across UAV-EAI, AD-EAI and Indoor-EAI. This framework highlights the unique “physical-cognitive” duality of aerial agents.
AspectUAV-EAIAD-EAIIndoor-EAI
Motion SpaceTrue 3D (6-DoF)Quasi 2D2.5D
Scale and OpennessKilometer-scale, UnstructuredCity-scale, Semi-structuredRoom-scale, Structured
Prior InfoInformation-ScarcePrior-RichPartially Unknown
DisturbancesNatural ForcesTraffic DynamicsMinimal
SensorsSWaP ConstrainedRedundant StackRGB-D Rich
Info-DensitySparseRule-BasedDense
1D Target-to-FOV Ratio<1%5–10%10–30%
Safety RiskHighHighMedium
Task ComplexityHighLow–MediumMedium–High
Table 2. Comparison of Core Tasks between Traditional UAV-RS and UAV-EAI Paradigms.
Table 2. Comparison of Core Tasks between Traditional UAV-RS and UAV-EAI Paradigms.
Core Task DomainTraditional UAV-RS ParadigmUAV-EAI ParadigmKey Paradigm Shift
Perception & MappingPassive sensor logging along predefined flight paths; purely geometric VIO/SLAM.Active sensing via Next-Best-View planning; Metric-Semantic mapping.From open-loop geometric reconstruction to closed-loop, real-time spatial-semantic awareness.
Navigation & ExecutionBlind tracking of explicit GPS coordinates and predefined geometric waypoints.Vision-Language Navigation and Object-Goal Navigation.From low-level direct flight control to abstract semantic reasoning and goal-driven search.
Exploration & CollaborationExhaustive Coverage Path Planning; independent or pre-separated flight corridors.Informative Path Planning; Decentralized multi-agent coordination (Dec-POMDPs).From uniform area coverage to uncertainty-aware information maximization and active teaming.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhao, Y.; Zhu, E.; Chen, Z.; Zhang, B.; Huo, W.; Zhao, X.; Chang, Y. Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution. Remote Sens. 2026, 18, 1509. https://doi.org/10.3390/rs18101509

AMA Style

Zhao Y, Zhu E, Chen Z, Zhang B, Huo W, Zhao X, Chang Y. Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution. Remote Sensing. 2026; 18(10):1509. https://doi.org/10.3390/rs18101509

Chicago/Turabian Style

Zhao, Yihao, Enze Zhu, Zhan Chen, Benkui Zhang, Wenxiang Huo, Xinyu Zhao, and Ying Chang. 2026. "Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution" Remote Sensing 18, no. 10: 1509. https://doi.org/10.3390/rs18101509

APA Style

Zhao, Y., Zhu, E., Chen, Z., Zhang, B., Huo, W., Zhao, X., & Chang, Y. (2026). Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution. Remote Sensing, 18(10), 1509. https://doi.org/10.3390/rs18101509

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop