Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution
Highlights
- A nine-dimensional comparative framework establishes UAV Embodied AI as a distinct research regime, constrained by 6-DoF dynamics, large-scale semantic sparsity, and strict onboard computing limits.
- Direct migration of terrestrial Visual-Language Models (VLMs) and Large Language Models (LLMs) to aerial platforms yields severe performance degradation, primarily caused by perspective mismatch and the detachment of cognitive reasoning from physical flight feasibility.
- Autonomous remote sensing must transition from open-loop data acquisition to active, closed-loop architectures that couple real-time perception with spatial-semantic decision-making.
- Future aerial autonomy requires hybrid neuro-symbolic pipelines, wherein high-level semantic planning is explicitly constrained by physics-based feasibility filters and domain-specific simulation infrastructure.
Abstract
1. Introduction
- We propose a nine-dimension analytical framework covering motion space, environmental structure, information density, and energy constraints. This framework qualitatively and quantitatively elucidates the essential differences and unique challenges of UAV-EAI compared to ground-based Embodied.
- Guided by this framework, we review the core tasks, critical infrastructure, and modeling methods of UAV-EAI, highlighting the limitations of the existing toolchain.
- Our analysis pinpoints key bottlenecks and future directions of existing UAV-EAI, specifically identifying the root causes of performance degradation in large-model reasoning and the unique Sim-to-Real gap of this emerging field.
2. Defining UAV Embodied AI: A Comparative Framework
2.1. Framework Dimensionality Analysis
2.1.1. Spatial and Dynamic Differences
2.1.2. Environmental and Perceptual Differences
2.1.3. Cognitive and Task Differences
2.1.4. Challenges in Terrestrial Paradigm Migration
- From Indoor-EAI to UAV-EAI: Terrestrial embodied models predominantly rely on object-centric affordance learning, which assumes canonical, human-eye-level viewpoints and stable object scales. In UAV-EAI, the nadir/oblique perspective and kilometer-scale operations lead to extreme scale collapse and perspective distortion, rendering indoor-trained semantic priors and interaction logic ineffective. Furthermore, the deterministic room-scale mapping typical of indoor agents fails in boundless outdoor environments, where metric state estimation must prioritize probabilistic robustness over rigid geometric reconstruction [25].
- From AD-EAI to UAV-EAI: Autonomous driving systems are fundamentally predicated on HD-map-dependent predictive control and 2.5D road-based navigation rules. These priors are non-existent in roadless, true 3D aerial spaces. Moreover, the sensor-redundant decision-making paradigm—utilizing heavy LiDAR stacks and high-throughput compute units—is physically prohibited by the stringent SWaP constraints of UAVs. Consequently, UAV-EAI requires a transition from “compute-heavy redundancy” to “resource-aware efficiency,” where agents must achieve comparable safety using sparse, lightweight sensory streams [21,26].
2.2. Methodological Foundations of the UAV-EAI Paradigm
- Asynchronous Physical-Cognitive Decoupling: Unlike the monolithic, synchronous perception-action pipelines common in Indoor-EAI, the stringent Size, Weight, and Power (SWaP) limits of UAVs necessitate a structural decoupling of control and cognition. This principle mandates an asymmetric architecture where a high-frequency reactive loop guarantees flight stability, while a compute-intensive cognitive loop is triggered only upon detecting semantic uncertainty or task transitions. This shift from continuous to event-triggered inference is a methodological prerequisite for aerial embodiment.
- Asymmetric Privileged Distillation: Stochastic aerodynamic disturbances create a pronounced dynamics gap in aerial environments, frequently invalidating the direct policy transfer methods established in terrestrial robotics. To address this, UAV-EAI architectures increasingly rely on asymmetric distillation frameworks. Policies are initially trained in simulation using exact, privileged environmental states, such as precise force vectors and global geometry. These high-fidelity representations are subsequently distilled into lightweight neural networks constrained to operate exclusively on noisy, partial onboard sensor data. By structurally integrating physical priors during the training phase, this method significantly reduces the sim-to-real discrepancy, enabling reliable autonomous execution in unmodeled physical environments [27,28].
3. Analysis of Core EAI Tasks
3.1. Autonomous Perception and Semantic Mapping
3.1.1. Viewpoint Selection and Next-Best-View
3.1.2. Metric-Semantic Mapping
3.2. Embodied Navigation and Instruction Following
3.2.1. Vision-Language Navigation (VLN)
3.2.2. Object-Goal Navigation (ObjectNav)
3.3. Generalized Exploration and Collaborative Interaction
3.3.1. Informative Path Planning
3.3.2. Multi-Agent Coordination
3.4. Summary
4. Enabling Infrastructure: Simulators and Datasets
4.1. Simulation Environments
4.1.1. The Era of Physics-Centric Control and General-Purpose Robotics
- Deficiency in Visual Realism at Scale: Designed primarily for indoor laboratories or small-scale manipulation tasks, Phase I simulators lack the rendering pipelines required for long-range aerial perception. Constructing the vast, unstructured outdoor environments typical of UAV missions requires manually importing heavy meshes, which often leads to severe performance degradation and lacks the procedural generation capabilities needed for diverse terrain training [61].
- Oversimplification of Aerodynamics: Although excellent at rigid-body dynamics, these platforms often abstract UAVs as simple floating masses, failing to model critical non-linear aerodynamic phenomena such as rotor-airflow interaction and stochastic wind turbulence. This omission creates a significant “Sim-to-Real” gap for embodied agents, as policies trained in these vacuums often fail to compensate for the complex disturbances characteristic of real-world flight [62].
- Computational Bottlenecks for Learning: The architecture of Gazebo and similar tools is primarily CPU-bound, optimized for high-precision, single-instance simulation rather than massive parallelism. This presents a prohibitive barrier for Deep Reinforcement Learning (DRL) and active perception tasks, which require millions of interaction frames. Without native GPU parallelization, these simulators cannot provide the throughput necessary to train complex neural policies within a reasonable timeframe [63].
4.1.2. The Era of Photorealistic Perception
- Microsoft AirSim [61] emerged as a seminal platform in this era, leveraging Unreal Engine to provide varying weather conditions, realistic lighting, and detailed textures. Unlike its predecessors, AirSim supported multi-modal sensor simulation—including depth, segmentation masks, and LiDAR—enabling the end-to-end training of perception-driven policies for autonomous flight in unstructured outdoor environments.
- FlightGoggles [64] offered a complementary approach by focusing on the tight coupling of visual realism with high-frequency inertial dynamics. By rendering photorealistic camera streams synchronized with rigorous vehicle dynamics, it facilitated the benchmarking of VIO and aggressive flight algorithms in virtual environments that closely mimicked real-world challenges, such as motion blur and dynamic lighting changes.
4.1.3. The Era of Parallel Embodied Intelligence
- NVIDIA Isaac Sim and Isaac Gym: Running on the Omniverse platform, this environment pairs high-fidelity RTX rendering with GPU-accelerated PhysX 5. While older platforms like Gazebo or AirSim typically process environments serially, the tensor-based API in Isaac Gym enables the simultaneous execution of thousands of parallel environments on a single workstation. This capability allows researchers to rapidly randomize physical parameters, such as friction, mass, and lighting conditions, reducing the time required for domain randomization from weeks to mere minutes and generating the diverse data necessary to train reliable control policies.
- Specialized Aerial Frameworks (Pegasus and OmniDrones): To bridge the gap between general-purpose physics and the specific aerodynamic requirements of UAVs, domain-specific extensions such as the Pegasus Simulator [67] and OmniDrones [68] have been developed within the Isaac ecosystem. These frameworks introduce modular vehicle backends and realistic aerodynamic plugins—modeling complex phenomena such as rotor drag, ground effect, and wind turbulence—thereby enabling the end-to-end training of aggressive flight maneuvers and multi-agent coordination that generic physics engines cannot support.
4.1.4. The Multi-Faceted Roles of Simulation in UAV-EAI: From Offline Infrastructure to Online Cognition
- Offline Policy Generation and Sim-to-Real Transfer: Corresponding to the current Era of Parallel Embodied Intelligence, modern simulators serve as the foundational data engines that provide the billions of interaction frames required for Deep Reinforcement Learning convergence. By leveraging massive GPU parallelization and extensive domain randomization, these platforms effectively bridge the reality gap, ensuring robust policy behavior prior to physical deployment [68].
- Online Predictive Engine: During real-time execution, simulation must function as a forward-predictive world model. When an embodied agent formulates a long-horizon plan via Next-Best-View optimization or Informative Path Planning, it utilizes a lightweight, embedded simulator to predict the future physical states and aerodynamic costs of proposed trajectories. This predictive capability is critical because high-level semantic objectives must be continuously balanced against non-holonomic kinematic constraints and battery expenditure—complex trade-offs that static heuristics fail to capture [34].
- Runtime Safety Verification: As will be further elaborated in the context of Large Language Models, high-level semantic planners frequently exhibit “feasibility hallucinations.” In this capacity, simulation serves as a strict runtime safety sandbox. Before any semantically generated way-point is transmitted to the low-level execution controller, it is rapidly evaluated within a localized simulation layer to verify aerodynamic stability, obstacle clearance, and Control Barrier Function compliance [44].
4.2. Benchmark Datasets
4.2.1. Remote Sensing and Mapping Datasets
- Lack of ego-motion information: Rather than capturing the dynamic states of the sensor platform, these datasets typically supply only static images or raw video frames. They omit critical kinematic details, including the UAV’s position, velocity, and orientation. Because true embodied agents depend heavily on self-motion feedback for spatial perception and navigation, the absence of this telemetry data significantly limits their utility for training interactive systems.
- Lack of dynamic trajectories: Although datasets such as VisDrone include target tracking annotations, they mainly focus on 2D trajectories and lack complete descriptions of dynamic objects in 3D space. Moreover, these trajectories are generally observed rather than generated through agent–environment interaction.
- Lack of goal-conditioned tasks: Existing datasets primarily provide labeled categories or boundaries but rarely include semantic information or reward signals associated with specific task goals. Embodied agents must understand task objectives and make decisions accordingly, which requires datasets to incorporate task-relevant annotations or scenario designs.
4.2.2. UAV Navigation and Inspection Datasets
- Precision in GPS-Denied Environments (EuroC MAV [73]): As a standard benchmark for MAV-scale operations, EuroC focuses on the challenges of precise localization within confined, GPS-denied industrial and indoor settings [73]. By capturing visual-inertial data along complex trajectories, it simulates critical operational stressors such as rapid rotations and close-range object interactions. Although originally designed for state estimation, it has evolved into a rigorous testbed for evaluating algorithmic robustness, enabling improvements in feature tracking and corner detection for systems like OV2SLAM [74].
- Agility in Dynamic Clutter (UZH-FPV [28]): In contrast to the stability-focused EuroC, the UZH-FPV dataset targets the regime of high-speed, aggressive flight [28]. It addresses the acute challenges of motion blur and extreme optical flow experienced during rapid acceleration and maneuvering through cluttered environments [19]. This dataset is instrumental for validating perception-action loops where the agent must maintain state estimation integrity under significant dynamic stress, a prerequisite for autonomous racing and emergency response.
4.2.3. UAV-EAI Task Datasets: Sparse and Emerging
- AerialVLN/CityNav-like datasets: These datasets primarily focus on language-conditioned navigation tasks in aerial scenarios. Recent efforts have explored Vision-Language Navigation (VLN) for UAV-EAI by adapting instruction-following paradigms originally developed for ground-based embodied agents [6]. VLMs have been increasingly adopted to address multimodal fusion, generalization, and interpretability challenges in navigation tasks, which are particularly relevant for UAV applications such as disaster response, logistics delivery, and urban inspection [77]. In parallel, advances in metric–semantic mapping systems enable UAVs to associate linguistic concepts with 3D spatial representations, supporting downstream navigation and inspection tasks that require semantic grounding in large-scale environments [38].
- UAV-Search/OpenFly-like datasets: These datasets emphasize long-horizon target search and exploration tasks, often involving extended trajectories and large-scale environments. Cooperative multi-UAV search has been studied under conditions of partial observability and limited communication, where decentralized coordination strategies are required to maintain system performance [78]. Prior work has formulated multi-UAV search as probabilistic coverage and pursuit–evasion problems, emphasizing objectives such as minimizing time-to-detection or capture rather than simple area coverage. Other studies investigate formation-based and topology-aware search strategies that adapt sensing resolution and spatial abstraction to cope with sparse observations and large environments [79].
- VLA and Cognitive Benchmarks: As aerial autonomy progresses from passive navigation to active environmental interaction, there is a critical requirement for datasets that tightly couple natural language instructions with continuous 6-DoF control outputs. Recent platforms such as CognitiveDroneBench [50] address this gap by providing dedicated evaluation frameworks for aerial VLA models. Encompassing extensive annotated flight trajectories, these emerging benchmarks explicitly target complex spatial reasoning and cognitive task execution, expanding the evaluation paradigm beyond traditional geometric metrics.
- Viewpoint–information coupling. There is currently no dataset that explicitly captures the coupling between viewpoint selection and information gain, which is fundamental for tasks like NBV selection and IPP. Existing NBV and active perception benchmarks primarily focus on static scenes or ground-based robots, leaving aerial viewpoint optimization underexplored [80]. In cooperative multi-UAV perception, formation geometry directly influences target observability, sensor complementarity, and occlusion patterns, highlighting the need for datasets that systematically vary viewpoint configurations and sensor baselines [81].
- Large-scale belief state annotations. For probabilistic search and exploration tasks, large-scale datasets with explicit belief state annotations remain largely absent. A belief state represents an agent’s probabilistic estimate of environmental states under uncertainty, and is central to POMDP-based planning frameworks [75]. Many UAV exploration systems still rely on single-modality sensing and deterministic maps, limiting robustness in cluttered or dynamic environments [82]. Multimodal sensing and belief-space planning have been shown to significantly improve exploration efficiency, yet corresponding annotated datasets are rare.
- Limitations in multimodal sensing. Current datasets remain limited in terms of multimodal sensing modalities, such as thermal imaging, acoustics, and hyperspectral data. However, multimodal sensing is increasingly critical for UAV-based perception in challenging environments [83]. For example:
- –
- Thermal imaging: Thermal-visible sensor fusion has been shown to improve detection robustness under low-illumination or adverse weather conditions, yet large-scale, task-oriented datasets remain limited [84].
- –
- LiDAR and hyperspectral sensing: The integration of LiDAR with hyperspectral or multispectral imagery enables richer geometric–semantic reasoning, particularly for inspection and environmental monitoring tasks [85].
- –
- Multispectral imagery: Multimodal fusion across spectral bands provides complementary cues for robust feature extraction and classification, but introduces challenges in sensor calibration, spatial alignment, and temporal synchronization [86].
Despite their promise, most existing UAV datasets treat these modalities independently rather than as components of a unified embodied sensing pipeline. To mitigate the severe data scarcity inherent to these specialized modalities, current research is increasingly leveraging advanced data augmentation strategies. For instance, the SFBDA framework [87] introduces a semantic-decoupled data augmentation approach that significantly enhances infrared few-shot object detection capabilities on UAVs. Such methodologies offer a computationally viable alternative to the prohibitive costs associated with large-scale, real-world multi-modal data collection.
4.2.4. Summary
5. Modeling Methods: From Control to Cognition
5.1. Conventional Model-Based Methods
5.1.1. Classical Control and Model-Based Estimation
- Brittleness of Dynamic Assumptions: Most model-based controllers depend on nominal dynamics, which often fail to hold in real-world Embodied AI applications. Environmental disturbances such as wind gusts, ground-effect turbulence, and sudden payload shifts during tasks like package delivery or physical sampling create complex non-linearities that rigid models struggle to accommodate.
- Vulnerability to Estimation Drift: Reliable state estimation assumes a consistent stream of clean sensor data. In weakly structured environments—such as vast forests or fog-covered canyons—visual features may disappear and GPS signals may be obstructed. In these scenarios, the EKF/UKF-based pose tracking can suffer from catastrophic drift, yet the controller remains “blind” to the loss of situational context.
- Fixed Trajectory Paradigms: Classical methods typically follow a predefined reference trajectory or a geometric path. This “blind following” is fundamentally incompatible with embodied tasks like Next-Best-View planning or Informative Path Planning, where the flight path must be dynamically re-synthesized in real-time based on the semantic content of the visual scene.
5.1.2. Trajectory Optimization and Model Predictive Control
- Uncertain and Partially Observed Cost Manifolds: Unlike industrial settings, UAV-EAI agents operate in environments where cost maps (representing risk or occupancy) are incrementally built and inherently uncertain. The agent must optimize its trajectory over a “belief map” where the cost of a certain path is a stochastic variable rather than a deterministic scalar.
- Scale-Induced Localization Drift: While MPC relies on a stable feedback loop, localization accuracy tends to degrade over kilometer-scale missions in unstructured terrain. This drift introduces a mismatch between the optimized plan and the physical execution, potentially leading to catastrophic collisions if the optimization does not account for spatial uncertainty.
- Dynamic and Epistemic Disturbances: External perturbations, such as sudden wind gusts or aerodynamic interactions in urban canyons, cause unmodeled deviations that exceed the rejection capabilities of standard MPC. These disturbances represent a mix of aleatory noise and epistemic model gaps that can destabilize the predictive loop.
5.2. Learning-Centric Methods: Deep Learning and Reinforcement Learning
5.2.1. Deep Perception: Recognition, Depth, and Scene Understanding
- Perspective mismatch: Most existing vision models—typically trained on datasets such as COCO or large-scale web datasets—are learned from a “human eye-level” perspective. This object-centric viewpoint contains rich semantic cues, including occlusion relationships and well-defined object boundaries [98]. In contrast, UAV images are often captured from top-down, oblique, or steep overhead viewpoints, which are inherently scene-centric and lack the fine-grained semantic cues common in ground-level imagery [99]. Background textures are frequently dominated by large-scale terrain or homogeneous vegetation, making it difficult for models to exploit local features learned from traditional datasets. Such viewpoint-induced geometric distortions and weak-texture regions significantly complicate depth reasoning and scene understanding in aerial imagery [100].
- Scale and altitude variation: Targets in UAV images—such as pedestrians or vehicles in search-and-rescue or surveillance missions—may occupy only a few pixels, rendering small-object detection particularly challenging [101]. Moreover, changes in flight altitude lead to drastic variations in object scale and image resolution, further increasing the difficulty of detection and segmentation [102]. These large scale variations require models to maintain robustness across orders of magnitude in spatial resolution, which is rarely encountered in conventional ground-level vision benchmarks.
- Changes in depth perception cues: Geometric edge cues that are effective in close-range indoor environments become unreliable from high-altitude UAV perspectives [103]. UAV depth estimation relies more heavily on texture gradients, shading, or illumination variations across terrain [104]. Depth discontinuities induced by perspective distortion and the prevalence of weakly textured regions pose persistent challenges for both supervised and self-supervised depth estimation in aerial scenarios [105].
- cross-view pretraining (BEV-aware data): To align feature spaces, pretraining on datasets that include bird’s-eye-view (BEV) perspectives is necessary [106]. Recent studies in BEV-based representation learning demonstrate that multi-view and multi-geometry alignment can significantly reduce domain gaps across viewpoints, offering valuable insights for UAV vision systems [107]. Domain Generalization aims to train models that can extract transferable knowledge from one or multiple source domains and generalize to unseen target domains [108]. ViTs have demonstrated strong potential in domain generalization, as cross-domain self-supervised pretraining and fine-tuning enable models to learn more generalizable geometric representations [109].
- multi-altitude augmentation and domain randomization: Robustness can be enhanced by simulating different flight altitudes, illumination conditions, and environmental variations during training through data augmentation and domain randomization techniques [110,111]. In aerial and outdoor depth estimation, augmentation strategies that explicitly perturb scale and viewpoint have been shown to improve robustness and generalization [112]. Furthermore, incorporating synthetic or simulated aerial data during training is an effective way to improve performance when real-world UAV data are scarce or collected under extreme conditions [113].
- hybrid geometric–learning depth estimators: Combining geometric constraints with learning-based approaches enables the design of depth estimators better suited for long-range outdoor scenarios [114,115]. For example, self-supervised monocular depth estimation methods often rely on view synthesis and image warping as supervision signals, eliminating the need for dense ground-truth depth annotations [116]. Such hybrid formulations improve geometric consistency and robustness under large viewpoint changes, which are common in UAV imagery [105]. In addition, lightweight attention-based architectures are increasingly explored to enable real-time depth inference on resource-constrained UAV platforms [117].
5.2.2. RL/IL for Navigation and Active Perception
- Experience Replay and Curiosity Mechanisms: Curiosity-driven experience replay methods introduce intrinsic motivation to encourage exploration of unseen states, enabling agents to acquire relatively effective policies under sparse rewards [129,130]. Empirical results in sparse-reward benchmarks indicate that curiosity-based methods can significantly improve exploration efficiency compared with naive baselines.
- Hierarchical Reinforcement Learning (HRL): By decomposing complex tasks into a sequence of sub-tasks, HRL effectively mitigates sparse-reward issues in long-horizon navigation problems [131]. This approach is particularly suitable for mixed action spaces, where agents first select abstract goals and then execute low-level control policies.
- Combining Imitation Learning with Reinforcement Learning: Integrating imitation learning with reinforcement learning leverages expert demonstrations to guide policy learning and accelerate convergence under sparse rewards [132]. Self-imitation learning further improves performance by reinforcing previously successful trajectories, even without additional expert supervision [133].
- Model Predictive Control: MPC can be incorporated into reinforcement learning frameworks to alleviate sparse rewards by predicting future system states and optimizing control sequences, thereby providing denser learning signals and improved stability [134].
- Parameterized and Hybrid Action Spaces: To cope with the complexity of state and action spaces in UAV navigation, parameterized and hybrid action-space formulations have been introduced, enabling more efficient learning and structured exploration under sparse-reward conditions [135].
- Predictive Coding: By learning predictive representations offline and using them for reward shaping, predictive coding supplies reward signals that reflect higher-level understanding of environmental structure and dynamics [136].
- Domain Randomization: Randomizing physical parameters and visual properties during simulation training improves policy robustness to real-world variability and enhances generalization from simulation to reality.
- Advanced Low-Level Control with Reinforcement Learning: Combining learning-based high-level decision-making with robust classical control techniques reduces sensitivity to modeling errors and improves flight stability.
- Curriculum Learning: Gradually increasing task difficulty during training enables agents to acquire more robust policies in simulation, facilitating transfer to real-world scenarios.
- Physics-Aware Simulators: Developing simulators that more accurately capture real-world physics reduces discrepancies between simulated and real environments.
- Zero-Shot Safety: Zero-shot safety approaches aim to ensure constraint satisfaction during deployment, guaranteeing safe behavior even when real-world conditions differ from those encountered during training.
- Belief-Space Formulations: Planning in belief space enables robust decision-making under partial observability by explicitly reasoning over uncertainty, rather than relying solely on geometric state representations.
- Lightweight Policy Architectures: Through compression and architectural optimization, policies can be adapted to limited onboard computational resources while maintaining real-time inference capability.
- Physically Aware Domain Randomization: Extensive randomization of physical parameters during training improves robustness and increases the likelihood of successful simulation-to-reality transfer.
- Safety Filters: Incorporating low-level safety mechanisms—such as control barrier functions—ensures constraint satisfaction and prevents catastrophic failures even when high-level policies are unreliable.
5.3. Cognition-Centric Methods: Large Language and Vision Language Models
5.3.1. Perspective Mismatch and Domain Gap
- Canonical View vs. Nadir Perspective: Dominant VLM training sets are composed almost exclusively of human-eye-level imagery, reflecting canonical horizontal viewpoints [98]. In contrast, UAV imagery is predominantly top-down (nadir) or oblique, leading to substantial geometric and appearance shifts that undermine feature reuse in pretrained visual encoders [99].
- Object-Centric vs. Scene-Centric Representation: Large-scale vision-language datasets are typically object-centric, featuring high-resolution, centered subjects. UAV data, by contrast, is inherently scene-centric, where targets often occupy only a few pixels within large, cluttered landscapes. This mismatch exacerbates the difficulty of grounding language queries in low-saliency aerial scenes [149].
- Semantic Feature Disparity: The semantic cues leveraged by VLMs for contextual reasoning—such as object co-occurrence patterns or human-centric interactions—are largely absent in wilderness or remote environments. In such scenarios, interpretation relies on geomorphological structure and texture statistics that are underrepresented in web-scale pretraining corpora [97].
- Cross-View Contrastive Pretraining: Designing contrastive objectives that explicitly align ground-level and aerial viewpoints, encouraging the learned representation to remain invariant under large changes in camera pitch, altitude, and scale [151].
- Parameter-Efficient Adapter Tuning: Leveraging lightweight adaptation mechanisms, such as low-rank adapters, to specialize pretrained VLM backbones using limited aerial datasets while preserving general multimodal knowledge [152].
- Geometric Prior Injection: Incorporating explicit geometric metadata—such as camera pose, altitude, or GPS-derived priors—into the attention mechanisms of VLMs, enabling pose-aware reasoning and improved spatial grounding [153].
- Multi-Altitude and Multi-Scale Pipelines: Employing hierarchical, multi-resolution processing pipelines that jointly capture global scene context and fine-grained local details, an approach shown to be effective in aerial perception and active vision systems [154].
5.3.2. LLMs for Instruction Following and High-Level Planning
- Resource and Physical Constraints: LLMs do not inherently model finite onboard resources such as battery capacity or flight time, nor do they explicitly reason about geometric constraints including camera field-of-view, sensing range, and occlusion [160].
- Stochastic Map Uncertainty: Unlike static textual environments, real-world UAV operations are characterized by partial observability and evolving uncertainty. LLMs struggle to reason over probabilistic occupancy maps or belief-space representations common in robotic exploration.
- Aerodynamic and Kinematic Feasibility: Linguistically valid action descriptions may violate vehicle dynamics or safety envelopes. High-level plans produced by LLMs rarely account for aerodynamic disturbances, non-holonomic motion constraints, or stability margins inherent to aerial platforms [24].
- Long-horizon Informative Coupling: LLMs are poorly suited for fine-grained, long-horizon viewpoint planning problems such as Next-Best-View or Informative Path Planning, where action selection depends on continuous optimization of expected information gain rather than symbolic sequencing [161].
- Physics-Based Feasibility Filters: A verification layer that evaluates LLM-generated waypoints or subgoals against vehicle dynamics, safety constraints, and control limits, often implemented with control-theoretic tools [162].
- Hierarchical Task Abstraction: A structured decomposition in which the LLM specifies semantic goals and task priorities, while mid-level planners compute dynamically feasible trajectories that maximize information gain under uncertainty [35].
- Uncertainty-Aware Feedback Loops: Closed-loop architectures that feed perception uncertainty, belief updates, or task execution failures back into the LLM prompt context—using mechanisms such as Chain-of-Thought or ReAct—enabling adaptive re-planning when physical reality deviates from the original mission description [163].
5.3.3. Onboard Inference and Embedded Deployment
- Hierarchical Resource Partitioning: Computational resources must be allocated asymmetrically to prevent subsystem starvation. Low-level flight control algorithms like Model Predictive Control require kilohertz-level update rates but consume minimal power, usually under 1 W. Mid-level perception tasks, encompassing Visual-Inertial Odometry and dense geometric mapping, necessitate deterministic latency operating at 20 to 30 Hz. These tasks continuously occupy a significant baseline fraction of the GPU and memory bandwidth, demanding around 10–15 W. Consequently, high-level cognitive models are restricted entirely to the residual compute budget. Attempting continuous and synchronous inference of large language or vision models can deprive the perception module of required memory bandwidth, leading to state estimation drift and flight instability [164].
- The Complexity-Latency Trade-off: The relationship between model complexity and real-time performance on edge devices is primarily governed by memory access speeds rather than pure floating-point operations. For instance, executing a 7B-parameter foundation model in half-precision format requires approximately 14 GB of video memory. Moving these weights across the limited memory bus of a mobile chip frequently pushes inference latency beyond 500 milliseconds per query. However, safe aerial obstacle avoidance dictates that the entire perception-to-action cycle must remain below a critical threshold of 100 to 200 milliseconds. Latency exceeding this margin induces prolonged dead-reckoning during the model’s forward pass, compromising the predictive control loop [165].
- Edge–Cloud Hybrid Pipelines: When communication permits, split-inference architectures assign safety-critical control and perception to onboard systems while offloading high-level semantic reasoning to edge-cloud infrastructure, forming a hierarchical intelligence loop [170].
- Event-Triggered and Asynchronous Inference: Rather than continuous VLM evaluation, event-driven inference pipelines query models only when perceptual novelty or confidence degradation is detected, reducing unnecessary energy expenditure.
5.3.4. Hierarchical System Architecture: Edge-Cloud Collaborative Orchestration
- Edge Node Processing: Latency-Bounded Reactive Autonomy. The UAV’s onboard computational payload is strictly reserved for high-frequency, safety-critical execution. This encompasses low-level attitude control, real-time VIO, dynamic obstacle avoidance, and the physics-based feasibility verification discussed in Section 5.3.2. By isolating these processes at the edge, the system ensures that the agent maintains basic physical stability and reactive survival capabilities independent of external network states [172].
- Cloud and Ground Control Station (GCS): Global Semantic Orchestration. Computationally intensive, long-horizon cognitive functions are offloaded to centralized cloud infrastructure or local GCS units. This layer executes LLMs for abstract instruction decomposition and maintains global metric-semantic maps. Critically, in cooperative multi-agent operations, this layer serves as the centralized orchestration hub. It facilitates hierarchical role assignments, manages task allocation, and aligns collaborative workflows across the swarm, ensuring that localized edge behaviors serve a coherent global objective [78].
5.4. Toward Hybrid Physical–Cognitive UAV Models
- Physics-aware Cognition: Future models must explicitly project high-level semantic intent into reachable sets and flight envelopes, ensuring that cognitive decisions respect aerodynamic feasibility and safety constraints [162].
- Cognitive-driven Adaptive Control: Hybrid architectures should allow semantic uncertainty and task priority to modulate low-level control strategies, enabling context-aware transitions between aggressive maneuvering and precision flight modes [174].
- Task-aware Active Perception: Integrating NBV and IPP directly with reasoning backbones allows perception to be guided by semantic objectives rather than uniform coverage.
6. Applications and Challenges
6.1. Key Application Domains
6.1.1. Critical Infrastructure Inspection
6.1.2. Emergency Search and Rescue
6.1.3. Precision Agriculture and Aerial Intervention
6.2. Core Challenges and Open Problems
6.2.1. Sim-to-Real Gap
6.2.2. Model Generalization
6.2.3. Safety, Robustness, and Ethical Considerations in the LLM/VLM Era
- Semantic Hallucinations and Feasibility Mismatch: In unstructured aerial environments, foundation models are highly susceptible to out-of-distribution errors and semantic hallucinations. A critical failure mode is the disconnect between abstract planning and physical reality, where linguistically valid plans inadvertently violate kinematic envelopes or energy constraints. Mitigating these risks requires coupling cognitive models with physics-based feasibility filters to strictly bound symbolic reasoning within safe flight parameters [184].
- Multi-Agent Cognitive Consistency: Swarm deployments relying on independent cognitive backbones face the risk of “cognitive divergence,” where ambiguous local observations lead to conflicting task prioritizations. Maintaining strategic coherence requires moving beyond basic spatial deconfliction. Implementing hierarchical organizational structures and explicit, role-based collaborative workflows within the swarm can enforce semantic consensus and align decentralized behaviors, even amidst local perception failures [185].
- Algorithmic Bias and Information Security: VLMs pretrained predominantly on terrestrial data may propagate systematic biases during aerial target recognition, compromising fairness in critical missions like search-and-rescue. Furthermore, semantic-driven autonomy amplifies privacy risks. Robust UAV-EAI frameworks must therefore incorporate domain-specific debiasing, privacy-preserving perception, and dynamic cybersecurity protocols to secure decentralized cognitive networks [186].
6.2.4. Data Scarcity
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Zhang, Z.; Zhu, L. A review on unmanned aerial vehicle remote sensing: Platforms, sensors, data processing methods, and applications. Drones 2023, 7, 398. [Google Scholar] [CrossRef]
- Savva, M.; Kadian, A.; Maksymets, O.; Zhao, Y.; Wijmans, E.; Jain, B.; Straub, J.; Liu, J.; Koltun, V.; Malik, J.; et al. Habitat: A platform for embodied AI research. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2019; pp. 9339–9347. [Google Scholar]
- Duan, J.; Yu, S.; Tan, H.L.; Zhu, H.; Tan, C. A survey of embodied ai: From simulators to research tasks. IEEE Trans. Emerg. Top. Comput. Intell. 2022, 6, 230–244. [Google Scholar] [CrossRef]
- Yang, R.; Chen, H.; Zhang, J.; Zhao, M.; Qian, C.; Wang, K.; Wang, Q.; Koripella, T.V.; Movahedi, M.; Li, M.; et al. Embodiedbench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents. arXiv 2025, arXiv:2502.09560. [Google Scholar]
- Zhao, C.; Liu, R.W.; Qu, J.; Gao, R. Deep learning-based object detection in maritime unmanned aerial vehicle imagery: Review and experimental comparisons. Eng. Appl. Artif. Intell. 2024, 128, 107513. [Google Scholar] [CrossRef]
- Anderson, P.; Wu, Q.; Teney, D.; Bruce, J.; Johnson, M.; Sünderhauf, N.; Reid, I.; Gould, S.; Van Den Hengel, A. Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: Piscataway, NJ, USA, 2018; pp. 3674–3683. [Google Scholar]
- Batra, D.; Gokaslan, A.; Kembhavi, A.; Maksymets, O.; Mottaghi, R.; Savva, M.; Toshev, A.; Wijmans, E. Objectnav revisited: On evaluation of embodied agents navigating to objects. arXiv 2020, arXiv:2006.13171. [Google Scholar] [CrossRef]
- Kolve, E.; Mottaghi, R.; Han, W.; VanderBilt, E.; Weihs, L.; Herrasti, A.; Deitke, M.; Ehsani, K.; Gordon, D.; Zhu, Y.; et al. AI2-THOR: An interactive 3D environment for visual AI. arXiv 2017, arXiv:1712.05474. [Google Scholar]
- Ahn, M.; Brohan, A.; Brown, N.; Chebotar, Y.; Cortes, O.; David, B.; Finn, C.; Fu, C.; Gopalakrishnan, K.; Hausman, K.; et al. Do as i can, not as i say: Grounding language in robotic affordances. arXiv 2022, arXiv:2204.01691. [Google Scholar] [CrossRef]
- Grigorescu, S.; Trasnea, B.; Cocias, T.; Macesanu, G. A survey of deep learning techniques for autonomous driving. J. Field Robot. 2020, 37, 362–386. [Google Scholar] [CrossRef]
- Levinson, J.; Askeland, J.; Becker, J.; Dolson, J.; Held, D.; Kammel, S.; Kolter, J.Z.; Langer, D.; Pink, O.; Pratt, V.; et al. Towards fully autonomous driving: Systems and algorithms. In 2011 IEEE Intelligent Vehicles Symposium (IV); IEEE: Piscataway, NJ, USA, 2011; pp. 163–168. [Google Scholar]
- Pendleton, S.D.; Andersen, H.; Du, X.; Shen, X.; Meghjani, M.; Eng, Y.H.; Rus, D.; Ang, M.H., Jr. Perception, planning, control, and coordination for autonomous vehicles. Machines 2017, 5, 6. [Google Scholar] [CrossRef]
- Faessler, M.; Fontana, F.; Forster, C.; Mueggler, E.; Pizzoli, M.; Scaramuzza, D. Autonomous, vision-based flight and live dense 3D mapping with a quadrotor micro aerial vehicle. J. Field Robot. 2016, 33, 431–450. [Google Scholar] [CrossRef]
- Qin, T.; Li, P.; Shen, S. Vins-mono: A robust and versatile monocular visual-inertial state estimator. IEEE Trans. Robot. 2018, 34, 1004–1020. [Google Scholar] [CrossRef]
- Gupta, S.; Davidson, J.; Levine, S.; Sukthankar, R.; Malik, J. Cognitive mapping and planning for visual navigation. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2017; pp. 2616–2625. [Google Scholar]
- Shi, G.; Shi, X.; O’Connell, M.; Yu, R.; Azizzadenesheli, K.; Anandkumar, A.; Yue, Y.; Chung, S.J. Neural lander: Stable drone landing control using learned dynamics. In 2019 International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2019; pp. 9784–9790. [Google Scholar]
- Arnold, E.; Al-Jarrah, O.Y.; Dianati, M.; Fallah, S.; Oxtoby, D.; Mouzakitis, A. A survey on 3d object detection methods for autonomous driving applications. IEEE Trans. Intell. Transp. Syst. 2019, 20, 3782–3795. [Google Scholar] [CrossRef]
- Wang, M. High definition map for autonomous driving: Overview and analysis. Geomat. World 2020, 27, 109–114. [Google Scholar]
- Loquercio, A.; Kaufmann, E.; Ranftl, R.; Müller, M.; Koltun, V.; Scaramuzza, D. Learning high-speed flight in the wild. Sci. Robot. 2021, 6, eabg5810. [Google Scholar] [CrossRef] [PubMed]
- Zhou, B.; Zhang, Y.; Chen, X.; Shen, S. Fuel: Fast uav exploration using incremental frontier structure and hierarchical planning. IEEE Robot. Autom. Lett. 2021, 6, 779–786. [Google Scholar] [CrossRef]
- Cui, C.; Ma, Y.; Cao, X.; Ye, W.; Zhou, Y.; Liang, K.; Chen, J.; Lu, J.; Yang, Z.; Liao, K.D.; et al. A survey on multimodal large language models for autonomous driving. In 2024 IEEE/CVF Winter Conference on Applications of Computer Vision Workshops (WACVW); IEEE: Piscataway, NJ, USA, 2024; pp. 958–979. [Google Scholar]
- Chen, Z.; Wu, J.; Wang, W.; Su, W.; Chen, G.; Xing, S.; Zhong, M.; Zhang, Q.; Zhu, X.; Lu, L.; et al. Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2024); IEEE: Piscataway, NJ, USA, 2024; pp. 24185–24198. [Google Scholar]
- Xu, C.; Zhang, R.; Yang, W.; Zhu, H.; Xu, F.; Ding, J.; Xia, G.S. Oriented tiny object detection: A dataset, benchmark, and dynamic unbiased learning. IEEE Trans. Pattern Anal. Mach. Intell. 2025, 48, 3167–3184. [Google Scholar] [CrossRef] [PubMed]
- Gao, Y.; Wang, Z.; Jing, L.; Wang, D.; Li, X.; Zhao, B. Aerial vision-and-language navigation via semantic-topo-metric representation guided LLM reasoning. arXiv 2024, arXiv:2410.08500. [Google Scholar]
- Liu, S.; Zhang, H.; Qi, Y.; Wang, P.; Zhang, Y.; Wu, Q. Aerialvln: Vision-and-language navigation for uavs. In 2023 IEEE/CVF International Conference on Computer Vision (ICCV 2023); IEEE: Piscataway, NJ, USA, 2023; pp. 15384–15394. [Google Scholar]
- Kaufmann, E.; Bauersfeld, L.; Loquercio, A.; Müller, M.; Koltun, V.; Scaramuzza, D. Champion-level drone racing using deep reinforcement learning. Nature 2023, 620, 982–987. [Google Scholar] [CrossRef]
- Wang, J.; Yu, Z.; Zhou, D.; Shi, J.; Deng, R. Vision-based deep reinforcement learning of Unmanned Aerial Vehicle (UAV) autonomous navigation using privileged information. Drones 2024, 8, 782. [Google Scholar] [CrossRef]
- Hanover, D.; Loquercio, A.; Bauersfeld, L.; Romero, A.; Penicka, R.; Song, Y.; Cioffi, G.; Kaufmann, E.; Scaramuzza, D. Autonomous drone racing: A survey. IEEE Trans. Robot. 2024, 40, 3044–3067. [Google Scholar] [CrossRef]
- Strand, S.H.; Wiedemann, T.; Burczek, B.; Shutin, D. Enhancing UAV Search Under Occlusion Using Next Best View Planning. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 19, 1085–1096. [Google Scholar] [CrossRef]
- Datsko, D.; Nekovar, F.; Penicka, R.; Saska, M. Energy-aware multi-uav coverage mission planning with optimal speed of flight. IEEE Robot. Autom. Lett. 2024, 9, 2893–2900. [Google Scholar] [CrossRef]
- Li, Z.; Wu, W.; Guo, Y.; Sun, J.; Han, Q.L. Embodied multi-agent systems: A review. IEEE/CAA J. Autom. Sin. 2025, 12, 1095–1116. [Google Scholar] [CrossRef]
- Batinovic, A.; Ivanovic, A.; Petrovic, T.; Bogdan, S. A shadowcasting-based next-best-view planner for autonomous 3D exploration. IEEE Robot. Autom. Lett. 2022, 7, 2969–2976. [Google Scholar] [CrossRef]
- Li, Y.; Guo, X.; Zhang, H.; Li, S.; Dai, X. Active Visual Perception: Opportunities and Challenges. arXiv 2025, arXiv:2512.03687. [Google Scholar] [CrossRef]
- Zhai, Y.; Reiter, R.; Scaramuzza, D. Pa-mppi: Perception-aware model predictive path integral control for quadrotor navigation in unknown environments. IEEE Robot. Autom. Lett. 2026, 11, 3804–3811. [Google Scholar] [CrossRef]
- Charrow, B.; Kahn, G.; Patil, S.; Liu, S.; Goldberg, K.; Abbeel, P.; Michael, N.; Kumar, V. Information-Theoretic Planning with Trajectory Optimization for Dense 3D Mapping. In Proceedings of the Robotics: Science and Systems, Rome, Italy, 13–17 July 2015; Volume 11, pp. 3–12. [Google Scholar]
- Vanegas, F.; Campbell, D.; Eich, M.; Gonzalez, F. UAV based target finding and tracking in GPS-denied and cluttered environments. In 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2016; pp. 2307–2313. [Google Scholar]
- Cadena, C.; Carlone, L.; Carrillo, H.; Latif, Y.; Scaramuzza, D.; Neira, J.; Reid, I.; Leonard, J.J. Past, present, and future of simultaneous localization and mapping: Toward the robust-perception age. IEEE Trans. Robot. 2017, 32, 1309–1332. [Google Scholar] [CrossRef]
- Rosinol, A.; Abate, M.; Chang, Y.; Carlone, L. Kimera: An open-source library for real-time metric-semantic localization and mapping. In 2020 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2020; pp. 1689–1696. [Google Scholar]
- Garg, S.; Sünderhauf, N.; Dayoub, F.; Morrison, D.; Cosgun, A.; Carneiro, G.; Wu, Q.; Chin, T.J.; Reid, I.; Gould, S.; et al. Semantics for robotic mapping, perception and interaction: A survey. Found. Trends® Robot. 2020, 8, 1–224. [Google Scholar] [CrossRef]
- Jatavallabhula, K.M.; Kuwajerwala, A.; Gu, Q.; Omama, M.; Chen, T.; Maalouf, A.; Li, S.; Iyer, G.; Saryazdi, S.; Keetha, N.; et al. Conceptfusion: Open-set multimodal 3d mapping. arXiv 2023, arXiv:2302.07241. [Google Scholar]
- Blum, H.; Rohrbach, S.; Popovic, M.; Bartolomei, L.; Siegwart, R. Active learning for UAV-based semantic mapping. arXiv 2019, arXiv:1908.11157. [Google Scholar] [CrossRef]
- Liu, X.; Liu, Y.; Qiu, H.; Yang, Q.; Lian, Z. Indooruav: Benchmarking vision-language uav navigation in continuous indoor environments. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2026; Volume 40, pp. 23864–23872. [Google Scholar]
- Shah, D.; Osiński, B.; Ichter, B.; Levine, S. Lm-nav: Robotic navigation with large pre-trained models of language, vision, and action. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2023; pp. 492–504. [Google Scholar]
- Yue, P.; Xin, J.; Zhang, Y.; Lu, Y.; Shan, M. Semantic-driven autonomous visual navigation for unmanned aerial vehicles. IEEE Trans. Ind. Electron. 2024, 71, 14853–14863. [Google Scholar] [CrossRef]
- Habibi, I.; Msadaa, I.C.; Grayaa, K. Adaptive UAV Inspection of PV Panels Using Goal-Conditioned Reinforcement Learning and Zigzag Coverage Planning. Proc. Mach. Learn. Res. 2025, 302, 1–17. [Google Scholar]
- Asgharivaskasi, A.; Atanasov, N. Semantic octree mapping and shannon mutual information computation for robot exploration. IEEE Trans. Robot. 2023, 39, 1910–1928. [Google Scholar] [CrossRef]
- Saxena, P.; Raghuvanshi, N.; Goveas, N. Uav-vln: End-to-end vision language guided navigation for uavs. In 2025 European Conference on Mobile Robots (ECMR); IEEE: Piscataway, NJ, USA, 2025; pp. 1–6. [Google Scholar]
- Zhang, X.; Tian, Y.; Lin, F.; Liu, Y.; Ma, J.; Wang, X.; Szatmáry, K.S.; Wang, F.Y. LogisticsVLN: Vision-language navigation for low-altitude terminal delivery based on agentic UAVs. In 2025 IEEE 28th International Conference on Intelligent Transportation Systems (ITSC); IEEE: Piscataway, NJ, USA, 2025; pp. 4437–4442. [Google Scholar]
- Mehboob, F.; James, M.W.; Habel, A.A.; Sam, J.; Altamirano Cabrera, M.; Tsetserukou, D. DroneVLA: VLA-Based Aerial Manipulation. In Companion Proceedings of the 21st ACM/IEEE International Conference on Human-Robot Interaction; Association for Computing Machinery: New York, NY, USA, 2026; pp. 1135–1139. [Google Scholar]
- Lykov, A.; Serpiva, V.; Khan, M.H.; Sautenkov, O.; Myshlyaev, A.; Tadevosyan, G.; Yaqoot, Y.; Tsetserukou, D. Cognitivedrone: A vla model and evaluation benchmark for real-time cognitive task solving and reasoning in uavs. arXiv 2025, arXiv:2503.01378. [Google Scholar]
- Rückin, J.; Magistri, F.; Stachniss, C.; Popović, M. An Informative Path Planning Framework for Active Learning in UAV-based Semantic Mapping. IEEE Trans. Robot. 2023, 39, 4279–4296. [Google Scholar] [CrossRef]
- Merei, A.; Mcheick, H.; Ghaddar, A.; Beltrame, G. EA-IPP2n: Energy-Aware Informative Path Planning Algorithm for UAVs. IEEE Access 2026, 14, 42381–42392. [Google Scholar] [CrossRef]
- Zheng, T.; Jin, Y.; Zhao, H.; Ma, Z.; Chen, Y.; Xu, K. Deep reinforcement learning based coverage path planning in unknown environments. In 2024 6th International Conference on Frontier Technologies of Information and Computer (ICFTIC); IEEE: Piscataway, NJ, USA, 2024; pp. 1608–1611. [Google Scholar]
- Song, F.; Zeng, Q.; Zhang, R.; Zhu, X.; Ye, X.; Zhang, Z. Multi-UAV Cooperative Navigation Based on Multi-Source Information Fusion: A Review. IEEE Sens. J. 2025, 26, 3460–3490. [Google Scholar] [CrossRef]
- Rizk, Y.; Awad, M.; Tunstel, E.W. Cooperative heterogeneous multi-robot systems: A survey. ACM Comput. Surv. (CSUR) 2019, 52, 29. [Google Scholar] [CrossRef]
- Amato, C.; Chowdhary, G.; Geramifard, A.; Üre, N.K.; Kochenderfer, M.J. Decentralized control of partially observable Markov decision processes. In 52nd IEEE Conference on Decision and Control; IEEE: Piscataway, NJ, USA, 2013; pp. 2398–2405. [Google Scholar]
- Zhai, Z.; Ni, W.; Wang, X.; Niyato, D.; Hossain, E. Integrated Sensing and Communication with UAV Swarms via Decentralized Consensus ADMM. arXiv 2025, arXiv:2511.03283. [Google Scholar] [CrossRef]
- Tang, H.K.; Lake, M.J.; Foster, R.J.; Bezombes, F.A. Design, development and pilot of a realistic virtual reality application to analyse quick directional change in sport: Avatar cutting scenario with alterable parameters. PLoS ONE 2025, 20, e0324941. [Google Scholar] [CrossRef]
- Coumans, E. Bullet physics simulation. In ACM SIGGRAPH 2015 Courses; Association for Computing Machinery: New York, NY, USA, 2015. [Google Scholar]
- Mellinger, D.; Kumar, V. Minimum snap trajectory generation and control for quadrotors. In 2011 IEEE International Conference on Robotics and Automation; IEEE: Piscataway, NJ, USA, 2011; pp. 2520–2525. [Google Scholar]
- Shah, S.; Dey, D.; Lovett, C.; Kapoor, A. Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and Service Robotics: Results of the 11th International Conference; Springer: Cham, Switzerland, 2017; pp. 621–635. [Google Scholar]
- Faessler, M.; Franchi, A.; Scaramuzza, D. Differential flatness of quadrotor dynamics subject to rotor drag for accurate tracking of high-speed trajectories. IEEE Robot. Autom. Lett. 2017, 3, 620–626. [Google Scholar] [CrossRef]
- Makoviychuk, V.; Wawrzyniak, L.; Guo, Y.; Lu, M.; Storey, K.; Macklin, M.; Hoeller, D.; Rudin, N.; Allshire, A.; Handa, A.; et al. Isaac gym: High performance gpu-based physics simulation for robot learning. arXiv 2021, arXiv:2108.10470. [Google Scholar] [CrossRef]
- Guerra, W.; Tal, E.; Murali, V.; Ryou, G.; Karaman, S. Flightgoggles: Photorealistic sensor simulation for perception-driven robotics using photogrammetry and virtual reality. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2019; pp. 6941–6948. [Google Scholar]
- Zhu, Y.; Wong, J.; Mandlekar, A.; Martín-Martín, R.; Joshi, A.; Lin, K.; Maddukuri, A.; Nasiriany, S.; Zhu, Y. robosuite: A modular simulation framework and benchmark for robot learning. arXiv 2020, arXiv:2009.12293. [Google Scholar]
- Rudin, N.; Hoeller, D.; Reist, P.; Hutter, M. Learning to walk in minutes using massively parallel deep reinforcement learning. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2022; pp. 91–100. [Google Scholar]
- Jacinto, M.; Pinto, J.; Patrikar, J.; Keller, J.; Cunha, R.; Scherer, S.; Pascoal, A. Pegasus simulator: An isaac sim framework for multiple aerial vehicles simulation. In 2024 International Conference on Unmanned Aircraft Systems (ICUAS); IEEE: Piscataway, NJ, USA, 2024; pp. 917–922. [Google Scholar]
- Xu, B.; Gao, F.; Yu, C.; Zhang, R.; Wu, Y.; Wang, Y. Omnidrones: An efficient and flexible platform for reinforcement learning in drone control. IEEE Robot. Autom. Lett. 2024, 9, 2838–2844. [Google Scholar] [CrossRef]
- Handa, A.; Allshire, A.; Makoviychuk, V.; Petrenko, A.; Singh, R.; Liu, J.; Makoviichuk, D.; Van Wyk, K.; Zhurkevich, A.; Sundaralingam, B.; et al. Dextreme: Transfer of agile in-hand manipulation from simulation to reality. In 2023 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2023; pp. 5977–5984. [Google Scholar]
- Wang, M.; Niu, Y.; Wang, B.; Zhang, W.; Wang, C. A Survey on Learning Motion Planning and Control for Mobile Robots: Toward Embodied Intelligence. IEEE Trans. Neural Netw. Learn. Syst. 2026. Early Access. [Google Scholar] [CrossRef]
- Lyu, Y.; Vosselman, G.; Xia, G.S.; Yilmaz, A.; Yang, M.Y. UAVid: A semantic segmentation dataset for UAV imagery. ISPRS J. Photogramm. Remote Sens. 2020, 165, 108–119. [Google Scholar] [CrossRef]
- Zhu, P.; Wen, L.; Du, D.; Bian, X.; Ling, H.; Hu, Q.; Nie, Q.; Cheng, H.; Liu, C.; Liu, X.; et al. Visdrone-det2018: The vision meets drone object detection in image challenge results. In Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Munich, Germany, 8–14 September 2018. [Google Scholar]
- Burri, M.; Nikolic, J.; Gohl, P.; Schneider, T.; Rehder, J.; Omari, S.; Achtelik, M.W.; Siegwart, R. The EuRoC micro aerial vehicle datasets. Int. J. Robot. Res. 2016, 35, 1157–1163. [Google Scholar] [CrossRef]
- Ferrera, M.; Eudes, A.; Moras, J.; Sanfourche, M.; Le Besnerais, G. OV2SLAM: A fully online and versatile visual SLAM for real-time applications. IEEE Robot. Autom. Lett. 2021, 6, 1399–1406. [Google Scholar] [CrossRef]
- Thrun, S. Probabilistic robotics. Commun. ACM 2002, 45, 52–57. [Google Scholar] [CrossRef]
- Lu, Y.; Xue, Z.; Xia, G.S.; Zhang, L. A survey on vision-based UAV navigation. Geo-Spat. Inf. Sci. 2018, 21, 21–32. [Google Scholar] [CrossRef]
- Popović, M.; Vidal-Calleja, T.; Chung, J.J.; Nieto, J.; Siegwart, R. Informative path planning for active field mapping under localization uncertainty. In 2020 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2020; pp. 10751–10757. [Google Scholar]
- Wang, C.; Yu, C.; Xu, X.; Gao, Y.; Yang, X.; Tang, W.; Yu, S.; Chen, Y.; Gao, F.; Jian, Z.; et al. Multi-robot system for cooperative exploration in unknown environments: A survey. arXiv 2025, arXiv:2503.07278. [Google Scholar] [CrossRef]
- Schwager, M.; Rus, D.; Slotine, J.J. Decentralized, adaptive coverage control for networked robots. Int. J. Robot. Res. 2009, 28, 357–375. [Google Scholar] [CrossRef]
- Low, K.H.; Dolan, J.; Khosla, P. Information-theoretic approach to efficient adaptive path planning for mobile robotic environmental sensing. In Proceedings of the International Conference on Automated Planning and Scheduling; AAAI Press: Washington, DC, USA, 2009; Volume 19, pp. 233–240. [Google Scholar]
- Schwager, M.; Julian, B.J.; Angermann, M.; Rus, D. Eyes in the sky: Decentralized control for the deployment of robotic camera networks. Proc. IEEE 2011, 99, 1541–1561. [Google Scholar] [CrossRef]
- Bourgault, F.; Makarenko, A.A.; Williams, S.B.; Grocholsky, B.; Durrant-Whyte, H.F. Information based adaptive robotic exploration. In IEEE/RSJ International Conference on Intelligent Robots and Systems; IEEE: Piscataway, NJ, USA, 2002; Volume 1, pp. 540–545. [Google Scholar]
- Ye, X.; Song, F.; Zhang, Z.; Zeng, Q. A review of small UAV navigation system based on multisource sensor fusion. IEEE Sens. J. 2023, 23, 18926–18948. [Google Scholar] [CrossRef]
- Lee, J.H.; Choi, J.S.; Jeon, E.S.; Kim, Y.G.; Thanh Le, T.; Shin, K.Y.; Lee, H.C.; Park, K.R. Robust pedestrian detection by combining visible and thermal infrared cameras. Sensors 2015, 15, 10580–10615. [Google Scholar] [CrossRef] [PubMed]
- Manfreda, S.; McCabe, M.F.; Miller, P.E.; Lucas, R.; Pajuelo Madrigal, V.; Mallinis, G.; Ben Dor, E.; Helman, D.; Estes, L.; Ciraolo, G.; et al. On the use of unmanned aerial systems for environmental monitoring. Remote Sens. 2018, 10, 641. [Google Scholar] [CrossRef]
- Samadzadegan, F.; Toosi, A.; Dadrass Javan, F. A critical review on multi-sensor and multi-platform remote sensing data fusion approaches: Current status and prospects. Int. J. Remote Sens. 2025, 46, 1327–1402. [Google Scholar] [CrossRef]
- Weng, Z.; He, W.; Lv, J.; Zhou, D.; Yu, Z. SFBDA: A Semantic-Decoupled Data Augmentation Framework for Infrared Few-Shot Object Detection on UAVs. IEEE Geosci. Remote Sens. Lett. 2025, 22, 7002205. [Google Scholar] [CrossRef]
- Lee, T.; Leok, M.; McClamroch, N.H. Geometric tracking control of a quadrotor UAV on SE (3). In 49th IEEE Conference on Decision and Control (CDC); IEEE: Piscataway, NJ, USA, 2010; pp. 5420–5425. [Google Scholar]
- Lee, T. Geometric control of quadrotor UAVs transporting a cable-suspended rigid body. IEEE Trans. Control Syst. Technol. 2017, 26, 255–264. [Google Scholar] [CrossRef]
- Mellinger, D.; Michael, N.; Kumar, V. Trajectory generation and control for precise aggressive maneuvers with quadrotors. Int. J. Robot. Res. 2012, 31, 664–674. [Google Scholar] [CrossRef]
- Hewing, L.; Wabersich, K.P.; Menner, M.; Zeilinger, M.N. Learning-based model predictive control: Toward safe learning in control. Annu. Rev. Control Robot. Auton. Syst. 2020, 3, 269–296. [Google Scholar] [CrossRef]
- Ames, A.D.; Coogan, S.; Egerstedt, M.; Notomista, G.; Sreenath, K.; Tabuada, P. Control barrier functions: Theory and applications. In 2019 18th European Control Conference (ECC); IEEE: Piscataway, NJ, USA, 2019; pp. 3420–3431. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016); IEEE: Piscataway, NJ, USA, 2016; pp. 770–778. [Google Scholar]
- Tan, M.; Le, Q. Efficientnet: Rethinking model scaling for convolutional neural networks. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2019; pp. 6105–6114. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Gevaert, C.M.; Belgiu, M. Assessing the generalization capability of deep learning networks for aerial image classification using landscape metrics. Int. J. Appl. Earth Obs. Geoinf. 2022, 114, 103054. [Google Scholar] [CrossRef]
- Vivone, G.; Deng, L.J.; Deng, S.; Hong, D.; Jiang, M.; Li, C.; Li, W.; Shen, H.; Wu, X.; Xiao, J.L.; et al. Deep learning in remote sensing image fusion: Methods, protocols, data, and future perspectives. IEEE Geosci. Remote Sens. Mag. 2024, 13, 269–310. [Google Scholar] [CrossRef]
- Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Computer Vision—ECCV 2014; Springer: Cham, Switzerland, 2014; pp. 740–755. [Google Scholar]
- Zhu, X.X.; Tuia, D.; Mou, L.; Xia, G.S.; Zhang, L.; Xu, F.; Fraundorfer, F. Deep learning in remote sensing: A comprehensive review and list of resources. IEEE Geosci. Remote Sens. Mag. 2017, 5, 8–36. [Google Scholar] [CrossRef]
- Li, Z.; Wang, Y.; Zhang, N.; Zhang, Y.; Zhao, Z.; Xu, D.; Ben, G.; Gao, Y. Deep learning-based object detection techniques for remote sensing images: A survey. Remote Sens. 2022, 14, 2385. [Google Scholar] [CrossRef]
- Gao, G.; Liu, Q.; Wang, Y. Counting from sky: A large-scale data set for remote sensing object counting and a benchmark method. IEEE Trans. Geosci. Remote Sens. 2020, 59, 3642–3655. [Google Scholar] [CrossRef]
- Jiao, L.; Zhang, F.; Liu, F.; Yang, S.; Li, L.; Feng, Z.; Qu, R. A survey of deep learning-based object detection. IEEE Access 2019, 7, 128837–128868. [Google Scholar] [CrossRef]
- Xu, D.; Ricci, E.; Ouyang, W.; Wang, X.; Sebe, N. Monocular depth estimation using multi-scale continuous crfs as sequential deep networks. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 41, 1426–1440. [Google Scholar] [CrossRef]
- Kendall, A.; Martirosyan, H.; Dasgupta, S.; Henry, P.; Kennedy, R.; Bachrach, A.; Bry, A. End-to-end learning of geometry and context for deep stereo regression. In 2017 IEEE International Conference on Computer Vision (ICCV 2017); IEEE: Piscataway, NJ, USA, 2017; pp. 66–75. [Google Scholar]
- Arafat, M.Y.; Alam, M.M.; Moh, S. Vision-based navigation techniques for unmanned aerial vehicles: Review and challenges. Drones 2023, 7, 89. [Google Scholar] [CrossRef]
- Gosala, N.; Petek, K.; Drews, P.L., Jr.; Burgard, W.; Valada, A. Skyeye: Self-supervised bird’s-eye-view semantic mapping using monocular frontal view images. In 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2023); IEEE: Piscataway, NJ, USA, 2023; pp. 14901–14910. [Google Scholar]
- Liu, Z.; Tang, H.; Amini, A.; Yang, X.; Mao, H.; Rus, D.L.; Han, S. Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation. In 2023 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2023; pp. 2774–2781. [Google Scholar]
- Gulrajani, I.; Lopez-Paz, D. In search of lost domain generalization. arXiv 2020, arXiv:2007.01434. [Google Scholar] [CrossRef]
- Bucci, S.; D’Innocente, A.; Liao, Y.; Carlucci, F.M.; Caputo, B.; Tommasi, T. Self-supervised learning across domains. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 44, 5516–5528. [Google Scholar] [CrossRef]
- Tobin, J.; Fong, R.; Ray, A.; Schneider, J.; Zaremba, W.; Abbeel, P. Domain randomization for transferring deep neural networks from simulation to the real world. In 2017 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2017; pp. 23–30. [Google Scholar]
- Shorten, C.; Khoshgoftaar, T.M. A survey on image data augmentation for deep learning. J. Big Data 2019, 6, 60. [Google Scholar] [CrossRef]
- Rajapaksha, U.; Sohel, F.; Laga, H.; Diepeveen, D.; Bennamoun, M. Deep learning-based depth estimation methods from monocular image and videos: A comprehensive survey. ACM Comput. Surv. 2024, 56, 315. [Google Scholar] [CrossRef]
- Qi, Q.; Wang, G.; Pan, Y.; Fan, H.; Li, B. MCS-Sim: A Photo-Realistic Simulator for Multi-Camera UAV Visual Perception Research. Drones 2025, 9, 656. [Google Scholar] [CrossRef]
- Zhou, T.; Brown, M.; Snavely, N.; Lowe, D.G. Unsupervised learning of depth and ego-motion from video. In 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: Piscataway, NJ, USA, 2017; pp. 1851–1858. [Google Scholar]
- Garg, R.; Bg, V.K.; Carneiro, G.; Reid, I. Unsupervised cnn for single view depth estimation: Geometry to the rescue. In Computer Vision—ECCV 2016; Springer: Cham, Switzerland, 2016; pp. 740–756. [Google Scholar]
- Godard, C.; Mac Aodha, O.; Firman, M.; Brostow, G.J. Digging into self-supervised monocular depth estimation. In 2019 IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: Piscataway, NJ, USA, 2019; pp. 3828–3838. [Google Scholar]
- Maslov, D.; Makarov, I. Online supervised attention-based recurrent depth estimation from monocular video. PeerJ Comput. Sci. 2020, 6, e317. [Google Scholar] [CrossRef]
- Bochkovskiy, A.; Wang, C.Y.; Liao, H.Y.M. Yolov4: Optimal speed and accuracy of object detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef]
- Wojke, N.; Bewley, A.; Paulus, D. Simple online and realtime tracking with a deep association metric. In 2017 IEEE International Conference on Image Processing (ICIP); IEEE: Piscataway, NJ, USA, 2017; pp. 3645–3649. [Google Scholar]
- Chen, J.; Lu, Y.; Yu, Q.; Luo, X.; Adeli, E.; Wang, Y.; Lu, L.; Yuille, A.L.; Zhou, Y. Transunet: Transformers make strong encoders for medical image segmentation. arXiv 2021, arXiv:2102.04306. [Google Scholar] [CrossRef]
- Zhao, C.; Zhang, Y.; Poggi, M.; Tosi, F.; Guo, X.; Zhu, Z.; Huang, G.; Tang, Y.; Mattoccia, S. Monovit: Self-supervised monocular depth estimation with a vision transformer. In 2022 International Conference on 3D Vision (3DV); IEEE: Piscataway, NJ, USA, 2022; pp. 668–678. [Google Scholar]
- Kendall, A.; Gal, Y.; Cipolla, R. Multi-task learning using uncertainty to weigh losses for scene geometry and semantics. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018); IEEE: Piscataway, NJ, USA, 2018; pp. 7482–7491. [Google Scholar]
- Teed, Z.; Deng, J. Deepv2d: Video to depth with differentiable structure from motion. arXiv 2018, arXiv:1812.04605. [Google Scholar]
- Tsouros, D.C.; Bibi, S.; Sarigiannidis, P.G. A review on UAV-based applications for precision agriculture. Information 2019, 10, 349. [Google Scholar] [CrossRef]
- Eskandari, R.; Mahdianpari, M.; Mohammadimanesh, F.; Salehi, B.; Brisco, B.; Homayouni, S. Meta-analysis of unmanned aerial vehicle (UAV) imagery for agro-environmental monitoring using machine learning and statistical models. Remote Sens. 2020, 12, 3511. [Google Scholar] [CrossRef]
- Kunze, L.; Hawes, N.; Duckett, T.; Hanheide, M.; Krajník, T. Artificial intelligence for long-term robot autonomy: A survey. IEEE Robot. Autom. Lett. 2018, 3, 4023–4030. [Google Scholar] [CrossRef]
- Loquercio, A.; Kaufmann, E.; Ranftl, R.; Dosovitskiy, A.; Koltun, V.; Scaramuzza, D. Deep drone racing: From simulation to reality with domain randomization. IEEE Trans. Robot. 2019, 36, 1–14. [Google Scholar] [CrossRef]
- Andrychowicz, M.; Wolski, F.; Ray, A.; Schneider, J.; Fong, R.; Welinder, P.; McGrew, B.; Tobin, J.; Pieter Abbeel, O.; Zaremba, W. Hindsight experience replay. Adv. Neural Inf. Process. Syst. 2017, 30, 5055–5065. [Google Scholar]
- Pathak, D.; Agrawal, P.; Efros, A.A.; Darrell, T. Curiosity-driven exploration by self-supervised prediction. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2017; pp. 2778–2787. [Google Scholar]
- Burda, Y.; Edwards, H.; Storkey, A.; Klimov, O. Exploration by random network distillation. arXiv 2018, arXiv:1810.12894. [Google Scholar] [CrossRef]
- Nachum, O.; Gu, S.S.; Lee, H.; Levine, S. Data-efficient hierarchical reinforcement learning. Adv. Neural Inf. Process. Syst. 2018, 31, 3307–3317. [Google Scholar]
- Chaysri, P.; Spatharis, C.; Blekas, K.; Vlachos, K. Unmanned surface vehicle navigation through generative adversarial imitation learning. Ocean Eng. 2023, 282, 114989. [Google Scholar] [CrossRef]
- Oh, J.; Guo, Y.; Singh, S.; Lee, H. Self-imitation learning. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2018; pp. 3878–3887. [Google Scholar]
- Williams, G.; Wagener, N.; Goldfain, B.; Drews, P.; Rehg, J.M.; Boots, B.; Theodorou, E.A. Information theoretic MPC for model-based reinforcement learning. In 2017 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2017; pp. 1714–1721. [Google Scholar]
- Hausknecht, M.; Stone, P. Deep reinforcement learning in parameterized action space. arXiv 2015, arXiv:1511.04143. [Google Scholar]
- Hafner, D.; Lillicrap, T.; Fischer, I.; Villegas, R.; Ha, D.; Lee, H.; Davidson, J. Learning latent dynamics for planning from pixels. In Proceedings of the 36th International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2019; pp. 2555–2565. [Google Scholar]
- Molchanov, A.; Chen, T.; Hönig, W.; Preiss, J.A.; Ayanian, N.; Sukhatme, G.S. Sim-to-(multi)-real: Transfer of low-level robust control policies to multiple quadrotors. In 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2019; pp. 59–66. [Google Scholar]
- Kim, M.J.; Pertsch, K.; Karamcheti, S.; Xiao, T.; Balakrishna, A.; Nair, S.; Rafailov, R.; Foster, E.; Lam, G.; Sanketi, P.; et al. Openvla: An open-source vision-language-action model. arXiv 2024, arXiv:2406.09246. [Google Scholar]
- Zitkovich, B.; Yu, T.; Xu, S.; Xu, P.; Xiao, T.; Xia, F.; Wu, J.; Wohlhart, P.; Welker, S.; Wahid, A.; et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2023; pp. 2165–2183. [Google Scholar]
- Han, S.; Mao, H.; Dally, W.J. Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding. arXiv 2015, arXiv:1510.00149. [Google Scholar]
- Hinton, G.; Vinyals, O.; Dean, J. Distilling the knowledge in a neural network. arXiv 2015, arXiv:1503.02531. [Google Scholar] [CrossRef]
- Haarnoja, T.; Zhou, A.; Hartikainen, K.; Tucker, G.; Ha, S.; Tan, J.; Kumar, V.; Zhu, H.; Gupta, A.; Abbeel, P.; et al. Soft actor-critic algorithms and applications. arXiv 2018, arXiv:1812.05905. [Google Scholar]
- Alayrac, J.B.; Donahue, J.; Luc, P.; Miech, A.; Barr, I.; Hasson, Y.; Lenc, K.; Mensch, A.; Millican, K.; Reynolds, M.; et al. Flamingo: A visual language model for few-shot learning. Adv. Neural Inf. Process. Syst. 2022, 35, 23716–23736. [Google Scholar]
- Liu, H.; Li, C.; Wu, Q.; Lee, Y.J. Visual instruction tuning. Adv. Neural Inf. Process. Syst. 2023, 36, 34892–34916. [Google Scholar]
- Li, J.; Li, D.; Savarese, S.; Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2023; pp. 19730–19742. [Google Scholar]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning transferable visual models from natural language supervision. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2021; pp. 8748–8763. [Google Scholar]
- Bommasani, R.; Hudson, D.A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M.S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. On the opportunities and risks of foundation models. arXiv 2021, arXiv:2108.07258. [Google Scholar] [CrossRef]
- Li, J.; Li, D.; Xiong, C.; Hoi, S. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In Proceedings of the International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2022; pp. 12888–12900. [Google Scholar]
- Cheng, G.; Han, J.; Lu, X. Remote sensing image scene classification: Benchmark and state of the art. Proc. IEEE 2017, 105, 1865–1883. [Google Scholar] [CrossRef]
- Cossio, M. A comprehensive taxonomy of hallucinations in large language models. arXiv 2025, arXiv:2508.01781. [Google Scholar] [CrossRef]
- Huynh, A.V.; Gillespie, L.E.; Lopez-Saucedo, J.; Tang, C.; Sikand, R.; Expósito-Alonso, M. Contrastive ground-level image and remote sensing pre-training improves representation learning for natural world imagery. In Computer Vision—ECCV 2024; Springer: Cham, Switzerland, 2024; pp. 173–190. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. Lora: Low-rank adaptation of large language models. In Proceedings of the Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, 25–29 April 2022; Volume 1, p. 3. [Google Scholar]
- Wang, Y.; Guizilini, V.C.; Zhang, T.; Wang, Y.; Zhao, H.; Solomon, J. Detr3d: 3d object detection from multi-view images via 3d-to-2d queries. In Proceedings of the Conference on Robot Learning; PMLR: Cambridge, MA, USA, 2022; pp. 180–191. [Google Scholar]
- Zhou, L.; Zhao, S.; Wan, Z.; Liu, Y.; Wang, Y.; Zuo, X. MFEFNet: A multi-scale feature information extraction and fusion network for multi-scale object detection in UAV aerial images. Drones 2024, 8, 186. [Google Scholar] [CrossRef]
- Driess, D.; Xia, F.; Sajjadi, M.S.; Lynch, C.; Chowdhery, A.; Ichter, B.; Wahid, A.; Tompson, J.; Vuong, Q.; Yu, T.; et al. Palm-e: An embodied multimodal language model. arXiv 2023, arXiv:2303.03378. [Google Scholar]
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901. [Google Scholar]
- Chen, M.; Tworek, J.; Jun, H.; Yuan, Q.; Pinto, H.P.D.O.; Kaplan, J.; Edwards, H.; Burda, Y.; Joseph, N.; Brockman, G.; et al. Evaluating large language models trained on code. arXiv 2021, arXiv:2107.03374. [Google Scholar] [CrossRef]
- Yang, S.; Nachum, O.; Du, Y.; Wei, J.; Abbeel, P.; Schuurmans, D. Foundation models for decision making: Problems, methods, and opportunities. arXiv 2023, arXiv:2303.04129. [Google Scholar] [CrossRef]
- Valmeekam, K.; Marquez, M.; Sreedharan, S.; Kambhampati, S. On the planning abilities of large language models-a critical investigation. Adv. Neural Inf. Process. Syst. 2023, 36, 75993–76005. [Google Scholar]
- Xu, Z.; Wu, K.; Wen, J.; Li, J.; Liu, N.; Che, Z.; Tang, J. A survey on robotics with foundation models: Toward embodied ai. arXiv 2024, arXiv:2402.02385. [Google Scholar] [CrossRef]
- Rana, K.; Haviland, J.; Garg, S.; Abou-Chakra, J.; Reid, I.; Suenderhauf, N. Sayplan: Grounding large language models using 3d scene graphs for scalable robot task planning. arXiv 2023, arXiv:2307.06135. [Google Scholar] [CrossRef]
- Xiao, W.; Cassandras, C.G.; Belta, C. Safe Autonomy with Control Barrier Functions: Theory and Applications; Springer: Cham, Switzerland, 2023. [Google Scholar]
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.R.; Cao, Y. React: Synergizing reasoning and acting in language models. In Proceedings of the Eleventh International Conference on Learning Representations, Kigali, Rwanda, 1–5 May 2022. [Google Scholar]
- Xia, X.; Fattah, S.M.M.; Babar, M.A. A survey on UAV-enabled edge computing: Resource management perspective. ACM Comput. Surv. 2023, 56, 78. [Google Scholar] [CrossRef]
- Lin, J.; Tang, J.; Tang, H.; Yang, S.; Xiao, G.; Han, S. Awq: Activation-aware weight quantization for on-device llm compression and acceleration. GetMob. Mob. Comput. Commun. 2025, 28, 12–17. [Google Scholar] [CrossRef]
- Dettmers, T.; Lewis, M.; Belkada, Y.; Zettlemoyer, L. Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale. Adv. Neural Inf. Process. Syst. 2022, 35, 30318–30332. [Google Scholar]
- Micikevicius, P.; Stosic, D.; Burgess, N.; Cornea, M.; Dubey, P.; Grisenthwaite, R.; Ha, S.; Heinecke, A.; Judd, P.; Kamalu, J.; et al. Fp8 formats for deep learning. arXiv 2022, arXiv:2209.05433. [Google Scholar] [CrossRef]
- Shazeer, N.; Mirhoseini, A.; Maziarz, K.; Davis, A.; Le, Q.; Hinton, G.; Dean, J. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv 2017, arXiv:1701.06538. [Google Scholar]
- Fedus, W.; Zoph, B.; Shazeer, N. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. J. Mach. Learn. Res. 2022, 23, 1–39. [Google Scholar]
- Shi, W.; Cao, J.; Zhang, Q.; Li, Y.; Xu, L. Edge computing: Vision and challenges. IEEE Internet Things J. 2016, 3, 637–646. [Google Scholar] [CrossRef]
- Ponzina, F.; Machetti, S.; Rios, M.; Denkinger, B.W.; Levisse, A.; Ansaloni, G.; Peón-Quirós, M.; Atienza, D. A hardware/software co-design vision for deep learning at the edge. IEEE Micro 2022, 42, 48–54. [Google Scholar] [CrossRef]
- Xiao, J.; Zhang, R.; Zhang, Y.; Feroskhan, M. Vision-based learning for drones: A survey. IEEE Trans. Neural Netw. Learn. Syst. 2025, 36, 15601–15621. [Google Scholar] [CrossRef] [PubMed]
- Kormushev, P.; Calinon, S.; Saegusa, R.; Metta, G. Learning the skill of archery by a humanoid robot iCub. In 2010 10th IEEE-RAS International Conference on Humanoid Robots; IEEE: Piscataway, NJ, USA, 2010; pp. 417–423. [Google Scholar]
- Qi, P.; Zhao, X. Flight control for very flexible aircraft using model-free adaptive control. J. Guid. Control Dyn. 2020, 43, 608–619. [Google Scholar] [CrossRef]
- Van Den Berg, J.; Abbeel, P.; Goldberg, K. LQG-MP: Optimized path planning for robots with motion uncertainty and imperfect state information. Int. J. Robot. Res. 2011, 30, 895–913. [Google Scholar] [CrossRef]
- Ha, D.; Schmidhuber, J. World models. arXiv 2018, arXiv:1803.10122. [Google Scholar]
- Ollero, A.; Tognon, M.; Suarez, A.; Lee, D.; Franchi, A. Past, present, and future of aerial robotic manipulators. IEEE Trans. Robot. 2021, 38, 626–645. [Google Scholar] [CrossRef]
- Ware, J.; Roy, N. An analysis of wind field estimation and exploitation for quadrotor flight in the urban canopy layer. In 2016 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2016; pp. 1507–1514. [Google Scholar]
- Jakobi, N.; Husbands, P.; Harvey, I. Noise and the reality gap: The use of simulation in evolutionary robotics. In Advances in Artificial Life: Third European Conference on Artificial Life, Granada, Spain, June 4–6, 1995 Proceedings; Springer: Berlin/Heidelberg, Germany, 1995; pp. 704–720. [Google Scholar]
- Hwangbo, J.; Sa, I.; Siegwart, R.; Hutter, M. Control of a quadrotor with reinforcement learning. IEEE Robot. Autom. Lett. 2017, 2, 2096–2103. [Google Scholar] [CrossRef]
- Miyai, A.; Yang, J.; Zhang, J.; Ming, Y.; Lin, Y.; Yu, Q.; Irie, G.; Joty, S.; Li, Y.; Li, H.; et al. Generalized out-of-distribution detection and beyond in vision language model era: A survey. arXiv 2024, arXiv:2407.21794. [Google Scholar]
- Li, S.; Chaplot, D.S.; Tsai, Y.H.H.; Wu, Y.; Morency, L.P.; Salakhutdinov, R. Unsupervised domain adaptation for visual navigation. arXiv 2020, arXiv:2010.14543. [Google Scholar] [CrossRef]
- Zhang, Y.; Liu, Y.; Liu, S.; Liang, W.; Wang, C.; Wang, K. Multimodal perception for indoor mobile robotics navigation and safe manipulation. IEEE Trans. Cogn. Dev. Syst. 2024, 17, 1074–1086. [Google Scholar] [CrossRef]
- Ahn, M.; Dwibedi, D.; Finn, C.; Arenas, M.G.; Gopalakrishnan, K.; Hausman, K.; Ichter, B.; Irpan, A.; Joshi, N.; Julian, R.; et al. Autort: Embodied foundation models for large scale orchestration of robotic agents. arXiv 2024, arXiv:2401.12963. [Google Scholar] [CrossRef]
- Mandi, Z.; Jain, S.; Song, S. Roco: Dialectic multi-robot collaboration with large language models. In 2024 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2024; pp. 286–299. [Google Scholar]
- Huang, Y.; Sun, L.; Wang, H.; Wu, S.; Zhang, Q.; Li, Y.; Gao, C.; Huang, Y.; Lyu, W.; Zhang, Y.; et al. Trustllm: Trustworthiness in large language models. arXiv 2024, arXiv:2401.05561. [Google Scholar] [CrossRef]
- Dulac-Arnold, G.; Levine, N.; Mankowitz, D.J.; Li, J.; Paduraru, C.; Gowal, S.; Hester, T. Challenges of real-world reinforcement learning: Definitions, benchmarks and analysis. Mach. Learn. 2021, 110, 2419–2468. [Google Scholar] [CrossRef]
- Ecoffet, A.; Huizinga, J.; Lehman, J.; Stanley, K.O.; Clune, J. First return, then explore. Nature 2021, 590, 580–586. [Google Scholar] [CrossRef]
- Mozaffari, M.; Saad, W.; Bennis, M.; Nam, Y.H.; Debbah, M. A tutorial on UAVs for wireless networks: Applications, challenges, and open problems. IEEE Commun. Surv. Tutor. 2019, 21, 2334–2360. [Google Scholar] [CrossRef]







| Aspect | UAV-EAI | AD-EAI | Indoor-EAI |
|---|---|---|---|
| Motion Space | True 3D (6-DoF) | Quasi 2D | 2.5D |
| Scale and Openness | Kilometer-scale, Unstructured | City-scale, Semi-structured | Room-scale, Structured |
| Prior Info | Information-Scarce | Prior-Rich | Partially Unknown |
| Disturbances | Natural Forces | Traffic Dynamics | Minimal |
| Sensors | SWaP Constrained | Redundant Stack | RGB-D Rich |
| Info-Density | Sparse | Rule-Based | Dense |
| 1D Target-to-FOV Ratio | <1% | 5–10% | 10–30% |
| Safety Risk | High | High | Medium |
| Task Complexity | High | Low–Medium | Medium–High |
| Core Task Domain | Traditional UAV-RS Paradigm | UAV-EAI Paradigm | Key Paradigm Shift |
|---|---|---|---|
| Perception & Mapping | Passive sensor logging along predefined flight paths; purely geometric VIO/SLAM. | Active sensing via Next-Best-View planning; Metric-Semantic mapping. | From open-loop geometric reconstruction to closed-loop, real-time spatial-semantic awareness. |
| Navigation & Execution | Blind tracking of explicit GPS coordinates and predefined geometric waypoints. | Vision-Language Navigation and Object-Goal Navigation. | From low-level direct flight control to abstract semantic reasoning and goal-driven search. |
| Exploration & Collaboration | Exhaustive Coverage Path Planning; independent or pre-separated flight corridors. | Informative Path Planning; Decentralized multi-agent coordination (Dec-POMDPs). | From uniform area coverage to uncertainty-aware information maximization and active teaming. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhao, Y.; Zhu, E.; Chen, Z.; Zhang, B.; Huo, W.; Zhao, X.; Chang, Y. Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution. Remote Sens. 2026, 18, 1509. https://doi.org/10.3390/rs18101509
Zhao Y, Zhu E, Chen Z, Zhang B, Huo W, Zhao X, Chang Y. Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution. Remote Sensing. 2026; 18(10):1509. https://doi.org/10.3390/rs18101509
Chicago/Turabian StyleZhao, Yihao, Enze Zhu, Zhan Chen, Benkui Zhang, Wenxiang Huo, Xinyu Zhao, and Ying Chang. 2026. "Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution" Remote Sensing 18, no. 10: 1509. https://doi.org/10.3390/rs18101509
APA StyleZhao, Y., Zhu, E., Chen, Z., Zhang, B., Huo, W., Zhao, X., & Chang, Y. (2026). Embodied AI in the Sky: A Comparative Review of UAV Embodied AI, from Autonomous Remote Sensing to Task Execution. Remote Sensing, 18(10), 1509. https://doi.org/10.3390/rs18101509

