Next Article in Journal
High-Precision Modeling of UAV Electric Propulsion for Improving Endurance Estimation
Next Article in Special Issue
Joint Optimization for Energy Efficiency in UAV-Enabled Networks
Previous Article in Journal
Robust Optimal Consensus Control for Multi-Agent Systems with Disturbances
Previous Article in Special Issue
Real-Time Long-Range Control of an Autonomous UAV Using 4G LTE Network
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

From Human Teams to Autonomous Swarms: A Reinforcement Learning-Based Benchmarking Framework for Unmanned Aerial Vehicle Search and Rescue Missions

by
Julian Bialas
1,2,†,
Mohammad Reza Mohebbi
1,2,†,
Michiel J. van Veelen
3,4,*,†,
Abraham Mejia-Aguilar
5,
Robert Kathrein
1,2 and
Mario Döller
1
1
Department of Data Science, FH Kufstein Tirol, 6330 Kufstein, Austria
2
Department of Mathematics and Informatics, University of Passau, 94030 Passau, Germany
3
Institute of Mountain Emergency Medicine, Eurac Research, 39100 Bolzano, Italy
4
Department of Sport Science, Medical Section, University of Innsbruck, 6020 Innsbruck, Austria
5
terraXcube, Eurac Research, 39100 Bolzano, Italy
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Drones 2026, 10(2), 79; https://doi.org/10.3390/drones10020079
Submission received: 2 December 2025 / Revised: 15 January 2026 / Accepted: 20 January 2026 / Published: 23 January 2026

Highlights

What are the main findings?
  • A unified benchmarking framework was developed to systematically compare human-only, UAV-assisted, and fully autonomous UAV swarm Search and Rescue (SAR) missions.
  • The Reinforcement Learning-based autonomous swarm, trained using Proximal Policy Optimization, achieved significantly faster target localization times than both ground and piloted UAV teams.
What are the implications of the main findings?
  • Autonomous UAV swarms can substantially reduce search times in SAR missions, potentially improving victim survival by shortening the treatment-free interval.
  • The results support future integration of Reinforcement Learning-driven swarm autonomy into operational SAR workflows, offering scalable, efficient, and safer rescue operations.

Abstract

The adoption of novel technologies such as Unmanned Aerial Vehicles (UAVs) in Search and Rescue (SAR) operations remains limited. As a result, their full potential is not yet realized. Although UAVs have been deployed on an ad hoc basis, typically under manual control by dedicated operators, assisted and fully autonomous configurations remain largely unexplored. In this study, three SAR frameworks are systematically evaluated within a unified benchmarking framework: conventional ground missions, UAV-assisted missions, and fully autonomous UAV operations. As the key performance indicator, the target localization time was quantified and used as the means of comparison amongst frameworks. The conventional and assisted frameworks were experimentally tested through physical hardware in a controlled outdoor setting, wherein simulated callouts occurred via rescue teams. The autonomous swarm framework was simulated in the form of a multi-agent Reinforcement Learning (RL) method via the use of the Proximal Policy Optimization (PPO) algorithm. This enabled the optimization of the decentralized cooperative actions that could occur for efficient exploration of a partially observed three-dimensional environment. Our results demonstrated that the autonomous swarm significantly outperformed the conventional and assisted approaches in terms of speed and coverage. Finally, a detailed depiction of the framework’s integration into an operational system is provided.

1. Introduction

Search and Rescue (SAR) operations are highly time-critical emergency missions conducted in uncertain and often hazardous environments following remote medical incidents, natural disasters, or man-made catastrophes. Victims in remote or mountainous regions frequently experience prolonged out-of-hospital times. These delays correlate strongly with increased morbidity and mortality, making rapid response paramount [1]. Traditional SAR missions rely primarily on ground or helicopter teams, whose performance is constrained by visibility, environmental hazards, and human endurance [2]. The growing frequency of climate-driven disasters and the rise in adventure tourism in remote areas further underscore the urgent need for faster, safer, and more adaptive SAR frameworks [3].
Unmanned Aerial Vehicles (UAVs) have emerged as force multipliers in SAR missions, owing to their real-time aerial imaging, thermal sensing, and precise target-localization capabilities [4,5]. By providing extensive coverage of otherwise inaccessible or unsafe terrains—such as collapsed structures, flood zones, landslides, or glacial areas—UAVs substantially enhance situational awareness and reduce risk to human rescuers [6]. Nevertheless, most existing UAV deployments remain manually piloted or semi-autonomous, requiring constant operator supervision and predefined flight plans [7]. These constraints limit scalability, responsiveness, and adaptability in rapidly evolving rescue missions. The transition toward fully autonomous and cooperative UAV systems—capable of decentralized decision-making and real-time mission adaptation—remains a major operational and research challenge [8].
Within this evolving landscape, three distinct UAV integration frameworks can be identified in SAR contexts: (i) human-only missions, (ii) human-led missions augmented by UAV support, and (iii) fully autonomous UAV swarms. Each paradigm introduces trade-offs between efficiency, scalability, complexity, and safety [9]. However, direct empirical comparisons across these approaches are rare. Prior studies have largely emphasized isolated technical aspects such as path planning or target detection rather than evaluating complete mission effectiveness [10,11]. Furthermore, most research on swarm autonomy remains confined to simulation environments, with limited real-world validation due to safety, cost, and logistical constraints [12]. This scarcity of benchmarking studies impedes the quantitative understanding of how autonomy and inter-vehicle cooperation can transform operational SAR performance.
To address this research gap, this study presents a unified benchmarking framework for systematically evaluating SAR performance across human-only, UAV-assisted, and fully autonomous swarm configurations. The first two frameworks were empirically tested through field experiments involving trained rescue teams and manually controlled UAVs. The autonomous swarm scenario was evaluated in a physics-based simulation to ensure repeatability and safety. The proposed autonomous framework employs a decentralized multi-agent Reinforcement Learning (RL) approach—specifically the Proximal Policy Optimization (PPO) algorithm—to enable cooperative exploration and target detection within a three-dimensional (3D) partially observable environment.
The main contributions of this study are as follows:
  • Development of a unified SAR benchmarking framework that enables direct comparison of human-only, UAV-assisted, and autonomous swarm missions.
  • Design of a fully decentralized RL-based swarm architecture leveraging PPO for efficient cooperative exploration and victim localization.
  • Quantitative performance comparison demonstrating the superior scalability, speed, and coverage efficiency of autonomous swarms relative to traditional and human-aided missions.
The remainder of this paper is organized as follows: Section 2 provides a comprehensive Literature Review covering UAV-assisted SAR, autonomous swarm systems, and multi-agent RL algorithms; Section 3 details the Methodology, describing the benchmarking framework, environmental modeling, and PPO-based learning scheme; Section 4 presents the Experimental Setup and performance analysis; Section 5 offers a detailed Discussion of results; Section 6 outlines the Outlook and Limitations of the study; and, finally, Section 7 concludes the paper and discusses directions for future research and real-world implementation.

2. Literature Review

Recent advances in UAVs have significantly expanded their use in SAR operations. To position this study within the state of the art, the literature is reviewed across three interrelated themes: (i) traditional and UAV-assisted SAR missions, where drones mainly serve as remote sensing and situational awareness tools; (ii) autonomous UAV systems for SAR, including single-agent and swarm-based approaches; and (iii) the integration of RL into UAV autonomy, emphasizing single- and multi-agent frameworks. Together, these domains trace the evolution from human-centered to fully autonomous, learning-driven SAR systems.

2.1. Traditional and UAV-Assisted SAR Missions

Traditional SAR operations rely heavily on human teams that navigate hazardous and unpredictable environments such as mountainous regions, collapsed infrastructure, or flood zones [13]. These missions are physically demanding and constrained by limited visibility, coverage capacity, and high exposure to danger, all of which impede timely rescues and increase risk to responders [14,15]. In large-scale disasters, environmental factors such as debris, unstable terrain, and weather extremes further complicate operations [16,17]. Consequently, modern SAR strategies increasingly integrate robotic and aerial technologies to enhance mission speed, safety, and situational awareness [18].
UAVs have become indispensable tools for these purposes, providing real-time imagery, thermal vision, and geospatial data via the Global Positioning System (GPS) [19,20]. They have proven effective for tasks such as post-earthquake reconnaissance, wildfire monitoring, and flood mapping [4,5]. Moreover, UAVs have supported the delivery of essential supplies such as defibrillators and flotation devices in remote areas [6,21,22]. Despite these benefits, several operational constraints persist, including limited endurance, payload capacity, and dependence on trained pilots [23]. Current UAV-based SAR missions remain predominantly manual or semi-autonomous, employing predefined routes that reduce adaptability in fast-changing disaster environments [24]. These limitations underline the need for greater autonomy, adaptability, and coordination capabilities within aerial rescue systems.

2.2. Autonomous UAV Systems in SAR

Autonomous UAV systems extend beyond passive sensing toward independent decision-making for exploration, victim localization, and navigation under uncertain conditions [25]. Leveraging onboard perception and computation, these systems can adapt to unstructured terrains and dynamic obstacles [26]. Early research focused on single-UAV autonomy through techniques such as Simultaneous Localization and Mapping (SLAM), deep visual recognition, and navigation in GPS-denied environments [27,28]. These studies achieved promising results in simulations and small-scale trials, demonstrating that learned autonomy improves mapping accuracy and obstacle avoidance but still struggles with robustness and computational efficiency in complex terrains.
To scale autonomy, multi-UAV or swarm configurations have gained prominence as they distribute exploration tasks and enhance fault tolerance through decentralized control [29,30]. Swarm systems apply collective intelligence concepts, using distributed communication and formation control to optimize coverage and resource utilization [31]. Leader–Follower and task-allocation architectures further improve scalability and cooperative behavior [32,33]. Nonetheless, real-world validation remains scarce, as most swarm frameworks operate in simulated settings and face challenges such as communication latency, sensor noise, and partial observability [34]. Recent frameworks advocate for hybrid models that merge algorithmic autonomy with limited human oversight to ensure safety and accountability [12,35]. The literature thus highlights an urgent need for integrated frameworks that jointly evaluate algorithmic performance, hardware constraints, and operational feasibility.

2.3. Reinforcement Learning in UAV-Based SAR

Reinforcement Learning (RL) has emerged as a transformative approach for enabling UAVs to learn adaptive control strategies in complex and uncertain environments [36]. By interacting with the environment and receiving feedback through a reward function, RL agents optimize policies for long-term mission goals such as exploration efficiency, obstacle avoidance, and victim detection [37]. Foundational algorithms such as Deep Q-Networks (DQNs), Deep Deterministic Policy Gradient (DDPG), and Proximal Policy Optimization (PPO) have demonstrated robust control performance in continuous action spaces [38,39]. PPO, in particular, has become widely adopted for UAV path planning due to its sample efficiency and stability in continuous domains [38]. However, most implementations remain limited to single-agent settings or simulation-based validation without considering communication constraints or heterogeneous agent coordination.
To overcome these limitations, Multi-Agent Reinforcement Learning (MARL) frameworks extend RL to cooperative UAV swarms, enabling decentralized policy learning and collective decision-making in partially observable environments [40,41,42]. Advanced algorithms such as Multi-Agent Deep Deterministic Policy Gradient (MADDPG) [43] and Q-learning with Monotonic Value Function Mixing (QMIX) [44] have been proposed to address inter-agent coordination and credit assignment problems in cooperative tasks. However, these methods often depend on centralized critics or communication channels, which limit scalability and increase computational overhead in swarm deployments. In contrast, decentralized PPO has recently gained attention for its ability to balance stability and independence among agents while maintaining global cooperation through implicit reward sharing [45,46]. Despite promising progress, hardware implementation of MARL systems remains limited, and the gap between simulation performance and field applicability persists [47,48]. Hence, there is a clear research opportunity to evaluate decentralized RL algorithms under realistic SAR mission scenarios, bridging the divide between theoretical learning efficiency and operational deployment.

2.4. Multi-Agent Reinforcement Learning for UAV-Based SAR Missions

Multi-Agent Reinforcement Learning (MARL) enables UAV teams to cooperatively navigate, explore, and allocate tasks in dynamic Search and Rescue environments through decentralized decision-making and shared learning. Su and Qian used Multi-Agent Proximal Policy Optimization (MAPPO) to enhance the cooperative search capabilities and tracking efficiency under limited sensing and communication capabilities and achieved superior exploration and observation performance [49]. Liao et al. introduced their Escape Target Search-MAPPO (ETS-MAPPO) approach for dynamic target search as well as achieving faster detection and better coverage in uncertain terrains [50]. Collectively, these studies demonstrate MARL’s potential to enhance UAV autonomy, coordination, and resilience in real-world SAR missions. Building on this foundation, the proposed framework in this work further contributes to the MARL-UAV SAR trend.

3. Methodology

This section defines the methodology used to evaluate the performance of the UAV swarm on SAR missions. The approach includes two main components: (i) applying the proposed autonomous framework to a standardized SAR scenario for benchmarking and (ii) developing a decentralized multi-agent system using Reinforcement Learning to support collaborative exploration and target detection. The experimental hardware and software are defined in a simulated environment to represent realistic mission conditions with terrain data and realistic operational constraints. Furthermore, the framework architecture, which outlines the formulation of the problem, the execution of policy learning, and the system integration pipeline, is described in detail.

3.1. Application of an Autonomous UAV Model to Established SAR Missions

In a recent randomized controlled crossover trial, van Veelen et al. [51] explored how UAVs could improve SAR efforts in a remote and mountainous area. The experiment brought together six trained mountain rescue teams who carried out 24 simulated missions in the Bletterbach canyon in South Tyrol, Italy. Each team performed two types of operations: traditional ground-based searches and UAV-assisted missions, where manually piloted drones carried RGB cameras and small medical kits. The findings were striking—UAV support cut down both search time and the delay before treatment could begin, demonstrating clear benefits for rescue operations in complex terrain.
This study expands on the work [51] by incorporating an autonomous UAV system trained through RL. To ensure fidelity to the real-world environment, the mission parameters and terrain of the Bletterbach canyon were recreated within a physics-based simulation setting. Figure 1 illustrates the reconstructed mission setting, including the designated headquarters and multiple target locations within the canyon. Terrain occlusions and starting zones were modeled to replicate real-world conditions. UAV velocities were set to a commonly referenced operational speed of 5   m / s [52]. The environment was discretized into a three-dimensional grid of size 64 × 64 × 15 , enabling spatial reasoning across the terrain, and three autonomous UAV agents were deployed for the simulated missions.
Target detection was evaluated using two criteria: (i) maintaining an unobstructed line of sight between the UAV and the target, and (ii) achieving a minimum pixel count across pedestrian height in imagery from an optical sensor (Sony FS700 (Sony Corporation, Tokyo, Japan) 18 mm focal length). Following Dollar et al., a threshold of 50 pixels is required [54]. With this camera configuration, the requirement is satisfied when the UAV–target distance is <125 m. These measures ensured alignment between simulated and real-world detection capabilities. The main performance indicator was the time taken to achieve successful localization. The next subsection provides an introduction to the proposed architecture of the autonomous UAV framework and how to integrate an approach for environmental modeling, the policy learning methodology, and the integration pipeline going forward.

3.2. Overview of the Framework

The proposed autonomous UAV system facilitates multi-agent collaborative exploration and target localization in SAR missions. The framework builds on a Decentralized Partially Observable Markov Decision Process (Dec-POMDP) formulation that allows the UAVs to make decisions under uncertainty while having limited observability. The main goal is to provide fast coverage of target areas while also satisfying operational constraints such as energy budgets and environmental complexity.

3.2.1. Environment and Problem Formulation

Let M = B w × d × h be the set of three-dimensional binary gridmaps, with w N , d N and h N representing the width depth and hight. The environment is represented as a discrete three-dimensional grid m M , where each voxel corresponds to a spatial unit with six visibility directions. Terrain information is derived from point cloud data and processed into height and priority layers to guide exploration according to [55]. The coverage state evolves over time, tracked by binary occupancy maps updated at each timestep t T = { t 0 , , t terminal } :
y ( t + 1 ) = y ( t ) ¬ v ( t ) .
Here, y : T M 6 represents the uncovered target cells at timestep t T and v : T M 6 represents the covered cells by the onboard sensor. The optimization goal is to minimize the time required to find a target y ˜ N 3 :
Minimize : t terminal subject to : y ˜ y ( t terminal 1 ) , y ˜ y ( t terminal ) .

3.2.2. Dec-POMDP and Policy Learning

The decision-making model is formulated as a Decentralized Partially Observable Markov Decision Process (Dec-POMDP). It consists of the following:
  • A global state S, formalized as
    S = M map   m × N w × d × h × 6 Ψ × B w × d SLZ × M Φ × N n power   levels .
    Here, the tensor Ψ is created in a three-stage process. First, a binary map is made by marking each grid cell as containing a trail or not. The trail information is retrieved from an overpass query. This map is then blurred with a Gaussian filter to spread the influence of nearby trails. The result is rescaled to priority values between one and five. These values are then mapped to the 3D gridmap m and restricted to the user-defined target area shown in Section 3.2.6. The Start And Landing zone (SLZ) is a 2D gridmap, which is mission-dependent and defined by the user as also described in Section 3.2.6. Φ represents the positions of the agents. The power levels or movement budget area is represented by a natural number, and displays how many movement steps in the environment can be taken before the battery is empty. It is initiated by an estimation of the total travel distance divided by cell width.
  • Local action sets A i = { 0 , 1 , 2 , 3 , 4 , 5 } , corresponding to six possible movement directions.
  • Observation sets Ω i for each agent explained in Section 3.2.3.
  • A reward function displayed in Equation (6).
Each UAV follows a decentralized stochastic policy:
π ( a i s ; θ ) ,
which is trained offline using PPO [56] to maximize the expected sum of discounted rewards:
R t = k = 0 t terminal t γ k r t + k .
Here, a i is the action performed by agent i, being in state s with policy parameters θ . r t is the reward gained in timestep t T .
The PPO algorithm optimizes the following clipped surrogate objective:
L CLIP ( θ ) = E ^ t [ min ( r t ( θ ) A ( a t , s t ) , clip ( r t ( θ ) , 1 ϵ , 1 + ϵ ) A ( a t , s t ) ) ] .
where r t ( θ ) is the probability ratio, A ( a t , s t ) denotes the advantage function and ϵ is a clipping hyperparameter.

3.2.3. Inputs and Observations

Each UAV works off a predefined observational set that will provide both local and global situational awareness. In particular, the following observations are processed at every timestep for each agent:
  • A local 9 × 9 × 9 × 6 volumetric map ( L 3 D ) containing the map layer and the priority surroundings of the agent;
  • A max-pooled global two-dimensional overviewn ( G 2 D ) centered on the agent containing the priority layer, the start landing zone, and the position layer;
  • An unpooled 17 × 17 local 2D map ( L 2 D ) for fine-grained navigation containing the same layer as the global 2D map;
  • The remaining movement and energy budget.
These observations leverage a shared convolutional neural network used within both the actor and critic.
The actor network is displayed in Figure 2. The actor returns action probabilities using a softmax activation function (all other layers use ReLu), while the critic determines the value of that state. In addition, to increase the resilience of the UAV systems, a fallback mechanism is established that makes use of Dijkstra’s algorithm to provide safe emergency landings.

3.2.4. Environment Generation

The mission environment is synthetically created by both geospatial and semantic inputs. Target regions are obtained by means of Overpass queries, while SLZs are defined with polygon annotations. The starting UAV positions are embedded in a position tensor ϕ ; acquired and interpolated terrain elevations are supplemented from processed point cloud data.

3.2.5. Reward Function and Training Procedure

The reward function is designed to encourage rapid, efficient coverage of high-priority areas while penalizing collisions and prolonged mission duration:
R ( s t , s t + 1 ) = i = 1 d j = 1 w 1 2 ( Ψ ^ ( t ) i , j Ψ ^ ( t + 1 ) i , j ) 2 c c c m , .
where c c and c m denote penalties associated with collisions and time consumption, respectively. d and w represent the depth and width of the map and Ψ ^ ( t ) contains the sum of priority cells over the height and the cube faces. Priority cells contribute quadratically to the reward, emphasizing focused exploration in critical regions.
Training involves a series of randomized episodes with different terrain configurations, agent starting points, and priority distributions. In each run, three UAVs are initialized with energy budgets randomly drawn from [ 60 , 90 ] , allowing the policy to adapt to a wide range of mission scenarios. The movement budget is a natural number, which is simply decremented with each timestep.
We train the policy based on a single agent with access to the global state information. All agents then operate using this learned policy. Although execution is fully decentralized, the agents exchange information following the procedures described in Section 3.2.6. Hence, our framework is based on centralized learning and decentralized execution.
The model is trained by taking 500,000 movement steps that are sampled from the replay memory. Other than PPO-based update clipping, no more gradient clipping is performed.

3.2.6. System Integration

The framework is deployed using a modular client–server architecture to facilitate real-time mission planning, visualization, and decentralized agent control. The system comprises the following: (i) a graphical user interface for mission definition, (ii) a centralized server for spatial preprocessing and task allocation, and (iii) onboard processes for executing learned policies. Figure 3 illustrates the core components: (a) the user interface, (b) the user workflow, and (c) the server-side logic for spatial processing and mission management.
Hence, the agents execute the framework decentralized on their companion computers and only share the coverage status with the centralized server. The server distributes the shared coverage with a predefined frequency with all agents.
A central monitoring module visualizes UAV trajectories, mission progress, and system logs in real time, ensuring operational transparency during both research simulations and deployment scenarios.

3.2.7. Mission Initialization and Execution

Mission configuration begins with the user defining the target area (TA) and SLZ through the graphical interface (Figure 3a). Each region is represented as a closed polygon:
P = { ( φ 1 , λ 1 ) , , ( φ n , λ n ) , ( φ 1 , λ 1 ) } ,
where ( φ , λ ) denote geographic coordinates in latitude and longitude. These polygonal inputs are essential for specifying the operational constraints of the mission.
Once submitted by the user client (Figure 3b), the server (Figure 3c) processes these boundaries to calculate the mission’s bounding rectangle:
φ min = min ( P T P S ) , λ min = min ( P T P S ) ,
φ max = max ( P T P S ) , λ max = max ( P T P S ) .
A four-dimensional coverage matrix C M 6 is initialized to track voxel-level visibility from six sensor perspectives.
If newly provided mission boundaries exceed the original bounding box, the system dynamically recalculates spatial limits and resamples the coverage matrix using linear transformations.
Finally, during the execution stage, each UAV activates onboard processes that apply the trained decentralized policy derived from PPO, allowing agents to make local decisions while navigating the environment. To safeguard mission continuity, the system incorporates additional safety measures, including a Dijkstra-based fallback path planner, which ensures emergency recovery if communication links fail or the policy becomes unreliable. For implementation details of the agent-side processes, readers are referred to [55].

4. Results

4.1. Benchmarking

The used hyperparameters of the final model are displayed in Table 1.
To assess the model’s performance, we compared it against three well-known optimization algorithms: Particle Swarm Optimization, Ant Colony Optimization, and Genetic Algorithms. For each method, three agents with 100 movement steps were evaluated and compared to the proposed model, which had not been fine-tuned for this mission. Each algorithm was tested on three different maps.
Figure 4 presents the results. The dashed line denotes the mean coverage rate achieved by the proposed model. For the optimization algorithms, the coverage rates obtained on the three maps were averaged over the full optimization period of three hours. All computations were performed on a machine equipped with an i7-1165G7 CPU and 16 GB of RAM. The results indicate that even after three hours of optimization, the proposed model substantially outperforms the benchmark algorithms.
The PPO backend was also evaluated against other RL frameworks: DQN and multi-agent SAC (MASAC). The first 500 episodes of the training process regarding the coverage rate are shown in Figure 5. Although the training endpoint has not yet been reached, PPO demonstrates a clear advantage in the steepness of the learning curve.
For the real-life missions, the performance was benchmarked by measuring the time from mission start to the first successful visual line of sight to a victim. This metric reflects the earliest operationally relevant point of victim localization. For ground and UAV-assisted missions, times were recorded during field trials; for autonomous swarms, they were logged in simulation. Each scenario was repeated (two trials for human missions; five runs for swarms), and mean values were used to ensure fair comparison across approaches, thereby accounting for potential stochastic variability in both the training process and the resulting agent behavior.

4.2. Analysis

The mean search times derived from these multiple simulations, which serve as a representative indicator of efficiency across scenarios, are summarized in Table 2. We fitted a linear mixed model with method (Ground, Assisted, and Auto) as a fixed effect and a random intercept for Scenario to account for repeated measures. The model indicated a strong main effect of method on search time, F ( 2 , 12 ) = 45.25 , p = 2.58 × 10 6 (AIC = 248.0; 18 observations from six missions). The estimated mean search time for the ground-based method (model intercept) was 1179 s. Relative to Ground, the Assisted approach was faster by 245 s ( t ( 12 ) = 3.61 , p = 0.0036 ), and the Autonomous (Auto) approach was faster by 641 s ( t ( 12 ) = 9.43 , p = 6.75 × 10 7 ). Random-effects estimates showed between-mission variability (SD = 233 s) in addition to residual variability (SD = 118 s). Both piloted UAV assistance and autonomous UAV operation shorten search times versus conventional ground teams, with the autonomous mode yielding the largest improvement. The analysis was performed using R version 4.2.2. Taken together, these findings provide strong empirical evidence that the autonomous system consistently outperforms traditional and piloted SAR strategies in terms of search time efficiency, highlighting the potential of swarm autonomy to substantially accelerate victim localization in time-critical rescue missions.
Figure 6 illustrates a representative episode from scenario four. In this example, the agents successfully establish a visual line of sight to the target and proceed to cover the target area (the valley) in a coordinated and efficient manner.

5. Discussion

The experimental results demonstrate that autonomous UAV swarms can substantially outperform both conventional ground-based and UAV-assisted SAR operations in terms of target localization time and coverage efficiency. Across all evaluated scenarios, the decentralized multi-agent RL framework enables coordinated exploration without centralized control, resulting in faster victim detection and more uniform spatial coverage in complex three-dimensional terrain. These findings provide empirical evidence that decentralized learning is a viable and effective strategy for time-critical SAR missions where centralized supervision is impractical.
A key strength of the proposed approach lies in its decentralized decision-making structure. Each UAV operates based on local observations within a partially observable environment, while coordination emerges implicitly through shared reward optimization rather than explicit task assignment. This design reduces reliance on continuous global communication and mitigates single points of failure commonly associated with centralized planners. The consistent performance gains observed across different terrain configurations and repeated simulation runs indicate that the learned policies generalize beyond individual mission instances and are not overly tuned to specific starting conditions or trajectories.
The robustness of the autonomous swarm is further supported by the architectural choices made in the framework design. The Dec-POMDP formulation, stochastic policy execution, and randomized initialization of agent positions and energy budgets expose the agents to variability during training and evaluation. As a result, the learned policies demonstrate stable cooperative behavior under changing mission conditions, maintaining effective coverage even in the presence of partial observability and limited information exchange. Importantly, decentralized execution ensures that local decision-making remains functional even when shared information is delayed or temporarily unavailable.
From an operational perspective, the reduction in target localization time has direct implications for real-world SAR effectiveness. Faster aerial localization shortens the treatment-free interval, a critical determinant of survival in emergency medicine, and enables rescue teams to plan safer and more efficient access routes before physical deployment. By reducing the time spent searching in hazardous terrain, autonomous swarms have the potential to improve both victim outcomes and responder safety. The integration of intelligent swarm behavior into existing SAR workflows therefore represents a meaningful advancement beyond incremental improvements offered by manually piloted UAV support.
While the autonomous swarm was evaluated in a simulation environment, the experimental setup was grounded in realistic mission parameters. The terrain model, UAV kinematics, and visual detection constraints were derived from prior field trials and established sensor specifications, ensuring that the simulated scenarios reflect plausible operational conditions. The strong agreement between simulated mission outcomes and empirically observed performance trends in ground-based and UAV-assisted SAR further reinforces the practical relevance of the proposed benchmarking framework.
Overall, the results indicate that decentralized multi-agent RL can serve as a powerful enabler for scalable, efficient, and adaptive SAR operations. By quantitatively benchmarking autonomous swarms against human-only and UAV-assisted strategies, this work provides a structured and evidence-based foundation for advancing UAV autonomy toward real-world deployment.

6. Limitations and Future Work

The proposed framework demonstrates substantial performance gains and provides a rigorous benchmarking comparison across SAR paradigms; however, several limitations remain that define important directions for future research. This study is intentionally positioned as a benchmarking and evaluation framework rather than a full operational deployment, and the autonomous swarm component was evaluated exclusively in simulation to ensure safety, repeatability, and controlled experimentation.
A primary limitation concerns the explicit assessment of robustness under degraded sensing and communication conditions. Although robustness is partially embedded in the framework through decentralized execution, partial observability, stochastic policy learning, and randomized mission initialization, current experiments do not include targeted sensitivity analyzes. Factors such as localization noise, sensor uncertainty, intermittent communication loss, or delayed state updates were not systematically injected into the simulation. As a result, robustness is inferred indirectly through consistent performance across repeated runs and varying terrain configurations rather than quantified through controlled perturbation experiments. Future work will therefore incorporate simple but informative sensitivity analyzes—such as bounded Gaussian localization noise or probabilistic loss of communication packets loss—to evaluate performance degradation trends and identify operational margins [8,47]. Similar approaches have been widely adopted to assess resilience in autonomous and multi-agent UAV systems operating under uncertainty.
A second limitation relates to the treatment of failure modes and recovery mechanisms. The framework includes a Dijkstra-based fallback planner to ensure safe navigation and emergency recovery when the learned policy becomes unreliable, for example, due to communication interruptions or navigation conflicts. However, this fallback mechanism was not explicitly tested for stress against concrete failure scenarios such as sustained localization drift, GPS degradation, or perception errors that affect target visibility [57]. Previous research has shown that hybrid autonomy strategies—combining learning-based control with classical planning, redundancy in sensing, or intermittent human supervision—can significantly enhance resilience in safety-critical UAV applications [11,35]. Extending the current framework to systematically evaluate fallback behavior and integrate additional tolerance strategies represents an important step toward operational robustness.
A further limitation concerns simulation-to-reality validation. Although the simulation environment is physics-based and grounded in realistic terrain geometry, UAV kinematics, and sensor specifications, the calibration of perception-related parameters remains indirect. Visual detection constraints, including line-of-sight conditions and pixel-based visibility thresholds, were derived from established pedestrian detection literature and camera models, but the study does not include a direct quantitative comparison between simulated detection times and those observed during field trials. A dedicated validation step that aligns the simulated and real-world detection performance would strengthen credibility and reduce the simulation-to-reality gap, which remains a recognized challenge in RL-based robotic systems [12,47].
Finally, the scope of the current evaluation is limited in terms of scalability and mission complexity. Experiments were conducted with three UAV agents, fixed planning horizons, and stationary targets. Larger swarm sizes, non-stationary victims, longer-duration missions, and dynamic mission updates introduce additional coordination, communication, and energy-management challenges that were not addressed in this study. Moreover, the framework prioritizes search efficiency rather than fully autonomous perception, with victim detection abstracted through geometric visibility constraints. Future extensions will integrate multimodal onboard sensing, adaptive energy management, and regulatory considerations to support the safe, scalable, and end-to-end autonomous deployment of UAV swarms in real-world SAR operations.

7. Conclusions

The presented work introduced a unified benchmarking framework for evaluating SAR missions across three operational paradigms: human-only, UAV-assisted, and fully autonomous swarm configurations. By integrating controlled field experiments with physics-based simulations, the framework established a rigorous foundation for quantifying the operational impact of UAV autonomy in time-critical rescue missions. The decentralized multi-agent Reinforcement Learning approach, implemented through PPO, enabled cooperative exploration and dynamic adaptation to complex, partially observable environments. Experimental results demonstrated that the autonomous swarm achieved markedly faster victim localization and more efficient coverage than both conventional and semi-autonomous approaches, confirming the effectiveness of decentralized learning for scalable multi-agent coordination. Beyond computational performance, these results translate to tangible operational advantages, including shorter treatment-free intervals, reduced risk to human rescuers, and improved mission adaptability. Overall, the framework advances the understanding of how autonomous UAV swarms can enhance the efficiency, safety, and resilience of modern SAR operations, marking a significant step toward their broader real-world deployment.

Author Contributions

Conceptualization, J.B. and M.D.; Methodology, J.B. and M.R.M.; Software, J.B. and R.K.; Validation, M.J.v.V.; Formal analysis, J.B. and M.J.v.V.; Investigation, J.B., A.M.-A. and M.R.M.; Data curation, J.B.; Writing—original draft preparation, M.R.M. and J.B.; Writing—review and editing, J.B., M.R.M., A.M.-A. and R.K.; Visualization, J.B. and M.J.v.V.; Supervision, M.D.; Project administration, M.D.; Funding acquisition, M.D. All authors have read and agreed to the published version of the manuscript.

Funding

The research leading to these results has received funding from the European Regional Development Fund under the Cooperation Program Interreg VI-A Italia Austria 2021–2027, ITAT11-008, START Living Lab and from the European Union’s Horizon Europe research and innovation program under grant agreement No 101168017 (HURRICANE); and from the Austrian Federal Ministry of Labor and Economy, the National Foundation for Research, Technology, and Development, and the Christian Doppler Research Association (Josef Ressel Center for multimedia analysis in the mobility domain).

Data Availability Statement

The data presented in this study are available on reasonable request from the corresponding author due to logistical considerations.

Acknowledgments

During the preparation of this manuscript, the authors used ChatGPT (GPT-5) for assistance with spelling and minor sentence revisions. The authors have reviewed and edited the output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
UAVUnmanned Aerial Vehicle
SARSearch and Rescue
RLReinforcement Learning
PPOProximal Policy Optimization
UASUnmanned Aerial Systems
GPSGlobal Positioning System
Dec-POMDPDecentralized Partially Observable Markov Decision Process
SLZStart and Landing Zone
TATarget Area

References

  1. Rauch, S.; Dal Cappello, T.; Strapazzon, G.; Palma, M.; Bonsante, F.; Gruber, E.; Ströhle, M.; Mair, P.; Brugger, H.; Group, I.A.T.R.S.; et al. Pre-hospital times and clinical characteristics of severe trauma patients: A comparison between mountain and urban/suburban areas. Am. J. Emerg. Med. 2018, 36, 1749–1753. [Google Scholar] [CrossRef]
  2. Martinez-Alpiste, I.; Golcarenarenji, G.; Wang, Q.; Alcaraz-Calero, J.M. Search and rescue operation using UAVs: A case study. Expert Syst. Appl. 2021, 178, 114937. [Google Scholar] [CrossRef]
  3. Roveri, G.; Crespi, A.; Eisendle, F.; Rauch, S.; Corradini, P.; Steger, S.; Zebisch, M.; Strapazzon, G. Climate change and human health in Alpine environments: An interdisciplinary impact chain approach understanding today’s risks to address tomorrow’s challenges. BMJ Glob. Health 2024, 8, e014431. [Google Scholar] [CrossRef] [PubMed]
  4. Mohebbi, M.R.; Sena, E.W.; Döller, M.; Klinger, J. Wildfire Spread Prediction Through Remote Sensing and UAV Imagery-Driven Machine Learning Models. In Proceedings of the 2024 18th International Conference on Control, Automation, Robotics and Vision (ICARCV), Dubai, United Arab Emirates, 12–15 December 2024; pp. 827–834. [Google Scholar]
  5. Karaca, Y.; Cicek, M.; Tatli, O.; Sahin, A.; Pasli, S.; Beser, M.F.; Turedi, S. The potential use of unmanned aircraft systems (drones) in mountain search and rescue operations. Am. J. Emerg. Med. 2018, 36, 583–588. [Google Scholar] [CrossRef] [PubMed]
  6. Khan, A.; Gupta, S.; Gupta, S. Emerging UAV technology for disaster detection, mitigation, response, and preparedness. J. Field Robot. 2022, 39, 905–955. [Google Scholar] [CrossRef]
  7. Elmokadem, T.; Savkin, A. Towards fully autonomous UAVs: A survey. Sensors 2021, 21, 6223. [Google Scholar] [CrossRef]
  8. Javaid, S.; Saeed, N.; Qadir, Z.; Fahim, H.; He, B.; Song, H.; Bilal, M. Communication and control in collaborative UAVs: Recent advances and future trends. IEEE Trans. Intell. Transp. Syst. 2023, 24, 5719–5739. [Google Scholar] [CrossRef]
  9. Lyu, M.; Zhao, Y.; Huang, C.; Huang, H. Unmanned aerial vehicles for search and rescue: A survey. Remote Sens. 2023, 15, 3266. [Google Scholar] [CrossRef]
  10. Rudol, P.; Doherty, P. Human body detection and geolocalization for UAV search and rescue missions using color and thermal imagery. In Proceedings of the 2008 IEEE Aerospace Conference, Big Sky, MT, USA, 1–8 March 2008; pp. 1–8. [Google Scholar]
  11. Queralta, J.; Taipalmaa, J.; Pullinen, B.; Sarker, V.; Gia, T.; Tenhunen, H.; Gabbouj, M.; Raitoharju, J.; Westerlund, T. Collaborative multi-robot search and rescue: Planning, coordination, perception, and active vision. IEEE Access 2020, 8, 191617–191643. [Google Scholar] [CrossRef]
  12. Phadke, A.; Medrano, F.; Sekharan, C.; Chu, T. An analysis of trends in UAV swarm implementations in current research: Simulation versus hardware. Drone Syst. Appl. 2024, 12, 1–10. [Google Scholar] [CrossRef]
  13. Bogue, R. Disaster relief, and search and rescue robots: The way forward. Ind. Robot. Int. J. Robot. Res. Appl. 2019, 46, 181–187. [Google Scholar] [CrossRef]
  14. Callender, N.; Ellerton, J.; Macdonald, J.H. Physiological demands of mountain rescue work. Emerg. Med. J. 2012, 29, 753–757. [Google Scholar] [CrossRef] [PubMed]
  15. Cooper, D.C. Fundamentals of Search and Rescue; Jones & Bartlett Learning: Burlington, MA, USA, 2005. [Google Scholar]
  16. Sharma, K.; Doriya, R.; Pandey, S.; Kumar, A.; Sinha, G.; Dadheech, P. Real-time survivor detection system in SaR missions using robots. Drones 2022, 6, 219. [Google Scholar] [CrossRef]
  17. Milani, M.; Roveri, G.; Falla, M.; Dal Cappello, T.; Strapazzon, G. Occupational accidents among search and rescue providers during mountain rescue operations and training events. Ann. Emerg. Med. 2023, 81, 699–705. [Google Scholar] [CrossRef]
  18. AlAli, Z.; Alabady, S. A survey of disaster management and SAR operations using sensors and supporting techniques. Int. J. Disaster Risk Reduct. 2022, 82, 103295. [Google Scholar] [CrossRef]
  19. Mohebbi, M.R.; Kafash, E.; Döller, M. Multi-Agent Trajectory Prediction for Urban Environments with UAV Data Using Enhanced Temporal Kolmogorov-Arnold Networks with Particle Swarm Optimization. In Proceedings of the 17th International Conference on Agents and Artificial Intelligence, Porto, Portugal, 23–25 February 2025; Volume 2, pp. 586–597. [Google Scholar] [CrossRef]
  20. Mohebbi, M.R.; Klinger, J.; Mohebbi Najm Abad, J.; Döller, M.; Tavasoli, M. Advanced Driving Behavior Analysis through Kolmogorov-Arnold Network and UAV Traffic Data. In Proceedings of the 17th International Conference on Machine Learning and Computing (ICMLC 2025), Guangzhou, China, 14–17 February 2025; Volume 1475, pp. 305–322. [Google Scholar]
  21. Seguin, C.; Blaquière, G.; Loundou, A.; Michelet, P.; Markarian, T. Unmanned aerial vehicles (drones) to prevent drowning. Resuscitation 2018, 127, 63–67. [Google Scholar] [CrossRef]
  22. van Veelen, M.J.; Vinetti, G.; Dal Cappello, T.; Eisendle, F.; Mejia-Aguilar, A.; Parin, R.; Oberhammer, R.; Falla, M.; Strapazzon, G. Drones reduce the time to defibrillation in a highly visited non-urban area: A randomized simulation-based trial. Am. J. Emerg. Med. 2024, 86, 5–10. [Google Scholar] [CrossRef]
  23. Valsan, A.; Parvathy, B.; GH, V.; Unnikrishnan, R.; Reddy, P.; Vivek, A. Unmanned aerial vehicle for search and rescue mission. In Proceedings of the 2020 4th International Conference on Trends in Electronics and Informatics (ICOEI)(48184), Tirunelveli, India, 15–17 June 2020; pp. 684–687. [Google Scholar]
  24. Telli, K.; Kraa, O.; Himeur, Y.; Ouamane, A.; Boumehraz, M.; Atalla, S.; Mansoor, W. A comprehensive review of recent research trends on unmanned aerial vehicles (UAVs). Systems 2023, 11, 400. [Google Scholar] [CrossRef]
  25. MahmoudZadeh, S.; Yazdani, A.; Kalantari, Y.; Ciftler, B.; Aidarus, F.; Al Kadri, M. Holistic review of UAV-centric situational awareness: Applications, limitations, and algorithmic challenges. Robotics 2024, 13, 117. [Google Scholar] [CrossRef]
  26. Chang, Y.; Cheng, Y.; Manzoor, U.; Murray, J. A review of UAV autonomous navigation in GPS-denied environments. Robot. Auton. Syst. 2023, 170, 104533. [Google Scholar] [CrossRef]
  27. Gupta, A.; Fernando, X. Simultaneous localization and mapping (SLAM) and data fusion in unmanned aerial vehicles: Recent advances and challenges. Drones 2022, 6, 85. [Google Scholar] [CrossRef]
  28. Amer, K.; Samy, M.; Shaker, M.; ElHelw, M. Deep convolutional neural network based autonomous drone navigation. In Proceedings of the Thirteenth International Conference on Machine Vision, Rome, Italy, 2–6 November 2020; Volume 11605, pp. 16–24. [Google Scholar]
  29. Ekechi, C.; Elfouly, T.; Alouani, A.; Khattab, T. A Survey on UAV Control with Multi-Agent Reinforcement Learning. Drones 2025, 9, 484. [Google Scholar] [CrossRef]
  30. Mohsan, S.; Khan, M.; Noor, F.; Ullah, I.; Alsharif, M. Towards the unmanned aerial vehicles (UAVs): A comprehensive review. Drones 2022, 6, 147. [Google Scholar] [CrossRef]
  31. Ingale, K.; Deshmukh, A.; Deshpande, A.; Deshmukh, S.; Deshmukh, M.; Bhise, S. Multi-agent swarm robotics for accurate position detection in disaster scenarios. In Proceedings of the 2023 International Conference on Inventive Computation Technologies (ICICT), Lalitpur, Nepal, 26–28 April 2023; pp. 1454–1460. [Google Scholar]
  32. Pang, J.; He, J.; Mohamed, N.; Lin, C.; Zhang, Z.; Hao, X. A hierarchical reinforcement learning framework for multi-UAV combat using leader–follower strategy. Knowl.-Based Syst. 2025, 316, 113387. [Google Scholar] [CrossRef]
  33. Wang, X.; Yang, D.; Chen, S. Particle swarm optimization based leader-follower cooperative control in multi-agent systems. Appl. Soft Comput. 2024, 151, 111130. [Google Scholar] [CrossRef]
  34. Virágh, C.; Nagy, M.; Gershenson, C.; Vásárhelyi, G. Self-organized UAV traffic in realistic environments. In Proceedings of the 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Daejeon, Republic of Korea, 9–14 October 2016; pp. 1645–1652. [Google Scholar]
  35. Adoni, W.; Fareedh, J.; Lorenz, S.; Gloaguen, R.; Madriz, Y.; Singh, A.; Kühne, T. Intelligent Swarm: Concept, Design and Validation of Self-Organized UAVs Based on Leader–Followers Paradigm for Autonomous Mission Planning. Drones 2024, 8, 575. [Google Scholar] [CrossRef]
  36. Bai, Y.; Zhao, H.; Zhang, X.; Chang, Z.; Jäntti, R.; Yang, K. Toward autonomous multi-UAV wireless network: A survey of reinforcement learning-based approaches. IEEE Commun. Surv. Tutorials 2023, 25, 3038–3067. [Google Scholar] [CrossRef]
  37. Ewers, J.; Anderson, D.; Thomson, D. Deep reinforcement learning for time-critical wilderness search and rescue using drones. Front. Robot. AI 2025, 11, 1527095. [Google Scholar] [CrossRef]
  38. Huang, X.; Wang, W.; Ji, Z.; Cheng, B. Representation Enhancement-Based Proximal Policy Optimization for UAV Path Planning and Obstacle Avoidance. Int. J. Aerosp. Eng. 2023, 2023, 6654130. [Google Scholar] [CrossRef]
  39. Bouhamed, O.; Ghazzai, H.; Besbes, H.; Massoud, Y. Autonomous UAV navigation: A DDPG-based deep reinforcement learning approach. In Proceedings of the 2020 IEEE International Symposium on Circuits and Systems (ISCAS), Virtual, 10–21 October 2020; pp. 1–5. [Google Scholar]
  40. Wenhong, Z.; Jie, L.; Zhihong, L.; Lincheng, S. Improving multi-target cooperative tracking guidance for UAV swarms using multi-agent reinforcement learning. Chin. J. Aeronaut. 2022, 35, 100–112. [Google Scholar]
  41. Yang, M.; Liu, G.; Zhou, Z.; Wang, J. Partially observable mean field multi-agent reinforcement learning based on graph attention network for UAV swarms. Drones 2023, 7, 476. [Google Scholar] [CrossRef]
  42. Zeng, Q.; Nait-Abdesselam, F. Multi-agent reinforcement learning-based extended boid modeling for drone swarms. In Proceedings of the ICC 2024–IEEE International Conference on Communications, Denver, CO, USA, 9–13 June 2024; pp. 1551–1556. [Google Scholar]
  43. Lowe, R.; Wu, Y.I.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-agent actor-critic for mixed cooperative-competitive environments. In Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
  44. Rashid, T.; Samvelyan, M.; De Witt, C.S.; Farquhar, G.; Foerster, J.; Whiteson, S. Monotonic value function factorisation for deep multi-agent reinforcement learning. J. Mach. Learn. Res. 2020, 21, 1–51. [Google Scholar]
  45. Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The surprising effectiveness of PPO in cooperative multi-agent games. In Proceedings of the 36th International Conference on Neural Information Processing Systems, New Orleans, LA, USA, 28 November–9 December 2022; Volume 35, pp. 24611–24624. [Google Scholar]
  46. Christianos, F.; Schäfer, L.; Albrecht, S. Shared experience actor-critic for multi-agent reinforcement learning. In Proceedings of the Advances in Neural Information Processing Systems, NeurIPS 2020, Online, 6–12 December 2020; Volume 33, pp. 10707–10717. [Google Scholar]
  47. Zhao, W.; Queralta, J.; Westerlund, T. Sim-to-real transfer in deep reinforcement learning for robotics: A survey. In Proceedings of the 2020 IEEE Symposium Series on Computational Intelligence (SSCI), Canberra, Australia, 1–4 December 2020; pp. 737–744. [Google Scholar]
  48. Boubin, J.; Burley, C.; Han, P.; Li, B.; Porter, B.; Stewart, C. Programming and deployment of autonomous swarms using multi-agent reinforcement learning. arXiv 2021, arXiv:2105.10605. [Google Scholar] [CrossRef]
  49. Su, K.; Qian, F. Multi-UAV Cooperative Searching and Tracking for Moving Targets Based on Multi-Agent Reinforcement Learning. Appl. Sci. 2023, 13, 11905. [Google Scholar] [CrossRef]
  50. Liao, G.; Wang, J.; Yang, D.; Yang, J. Multi-UAV Escape Target Search: A Multi-Agent Reinforcement Learning Method. Sensors 2024, 24, 6859. [Google Scholar] [CrossRef]
  51. Van Veelen, M.J.; Roveri, G.; Voegele, A.; Dal Cappello, T.; Masè, M.; Falla, M.; Regli, I.B.; Mejia-Aguilar, A.; Mayrgündter, S.; Strapazzon, G. Drones reduce the treatment-free interval in search and rescue operations with telemedical support–A randomized controlled trial. Am. J. Emerg. Med. 2023, 66, 40–44. [Google Scholar] [CrossRef]
  52. Kovanic, L.; Topitzer, B.; Petovsky, P.; Blistan, P.; Gergelova, M.B.; Blistanova, M. Review of Photogrammetric and Lidar Applications of UAV. Appl. Sci. 2023, 13, 6732. [Google Scholar] [CrossRef]
  53. Google Maps. Blätterbachklamm. 2025. Available online: https://www.google.com/maps/search/bl%C3%A4tterbachklamm/@46.3670343,11.3659854,1260a,35y,87.19h,77.88t/data=!3m1!1e3?entry=ttu&g_ep=EgoyMDI1MDkxNS4wIKXMDSoASAFQAw%3D%3D (accessed on 17 September 2025).
  54. Dollar, P.; Wojek, C.; Schiele, B.; Perona, P. Pedestrian detection: An evaluation of the state of the art. IEEE Trans. Pattern Anal. Mach. Intell. 2011, 34, 743–761. [Google Scholar] [CrossRef]
  55. Bialas, J.; Doeller, M.; Walch, S.; van Veelen, M.J.; Mejia-Aguilar, A. Optimizing Multi-Agent Coverage Path Planning UAV Search and Rescue Missions with Prioritizing Deep Reinforcement Learning. In Proceedings of the 2024 IEEE International Conference on Robotics and Biomimetics (ROBIO), Bangkok, Thailand, 10–14 December 2024; pp. 85–90. [Google Scholar]
  56. Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef]
  57. Alghamdi, S.; Alahmari, S.; Yonbawi, S.; Alsaleem, K.; Ateeq, F.; Almushir, F. Autonomous Navigation Systems in GPS-Denied Environments: A Review of Techniques and Applications. In Proceedings of the 2025 11th International Conference on Automation, Robotics, and Applications (ICARA), Zagreb, Croatia, 12–14 February 2025; pp. 290–299. [Google Scholar]
Figure 1. Geologic surroundings of the mission. (a) Reconstructed mission scenario showing the headquarters and designated target locations within the Bletterbach canyon. (b) Satellite Image of the Bletterbach canyon [53] (Map data: Google). ((a): The surrounding exported as a mesh, (b): the satellite image).
Figure 1. Geologic surroundings of the mission. (a) Reconstructed mission scenario showing the headquarters and designated target locations within the Bletterbach canyon. (b) Satellite Image of the Bletterbach canyon [53] (Map data: Google). ((a): The surrounding exported as a mesh, (b): the satellite image).
Drones 10 00079 g001
Figure 2. Model architecture of the actor network.
Figure 2. Model architecture of the actor network.
Drones 10 00079 g002
Figure 3. (a) User Interface of the mission planner. (b) Flowchart of the user interactions. (c) Flowchart of the server-side processes. System architecture illustrating (a) The mission configuration interface, (b) the user workflow, and (c) the server-side spatial processing and task allocation logic.
Figure 3. (a) User Interface of the mission planner. (b) Flowchart of the user interactions. (c) Flowchart of the server-side processes. System architecture illustrating (a) The mission configuration interface, (b) the user workflow, and (c) the server-side spatial processing and task allocation logic.
Drones 10 00079 g003aDrones 10 00079 g003b
Figure 4. (a) Ant Colony Optimization (ACO). (b) Particle Swarm Optimization (PSO). (c) Genetic Algorithm (GA). Comparison of the non-finetuned model (dashed line) against three hours of optimizing with ACO, PSO, and GA.
Figure 4. (a) Ant Colony Optimization (ACO). (b) Particle Swarm Optimization (PSO). (c) Genetic Algorithm (GA). Comparison of the non-finetuned model (dashed line) against three hours of optimizing with ACO, PSO, and GA.
Drones 10 00079 g004
Figure 5. First 500 episodes of the training process for the RL backend.
Figure 5. First 500 episodes of the training process for the RL backend.
Drones 10 00079 g005
Figure 6. Example trajectories for scenario four.
Figure 6. Example trajectories for scenario four.
Drones 10 00079 g006
Table 1. Parameters used for the Framework.
Table 1. Parameters used for the Framework.
ParameterValue
w , d , h 64, 64, 15
ϵ 0.2
γ 0.99
c c 5
c m 0.2
Table 2. Comparison of mission durations (hh:mm:ss) with its standard deviation in seconds across three SAR strategies for six locations. Ground and drone-assisted teams show two repeated trials each, the autonomous swarm results the mean of repeated simulations.
Table 2. Comparison of mission durations (hh:mm:ss) with its standard deviation in seconds across three SAR strategies for six locations. Ground and drone-assisted teams show two repeated trials each, the autonomous swarm results the mean of repeated simulations.
ScenarioGround TeamDrone-AssistedAutonomous
L100:20:15 ± 21 s00:20:10 ± 70 s00:14:18 ± 23 s
L200:23:31 ± 472 s00:19:56 ± 228 s00:13:45 ± 26 s
L300:23:30 ± 475 s00:19:37 ± 87 s00:12:43 ± 14 s
L400:17:53 ± 567 s00:12:25 ± 13 s00:05:01 ± 6 s
L500:20:49 ± 9 s00:10:08 ± 38 s00:04:41 ± 7 s
L600:11:55 ± 146 s00:11:05 ± 42 s00:03:20 ± 7 s
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Bialas, J.; Mohebbi, M.R.; van Veelen, M.J.; Mejia-Aguilar, A.; Kathrein, R.; Döller, M. From Human Teams to Autonomous Swarms: A Reinforcement Learning-Based Benchmarking Framework for Unmanned Aerial Vehicle Search and Rescue Missions. Drones 2026, 10, 79. https://doi.org/10.3390/drones10020079

AMA Style

Bialas J, Mohebbi MR, van Veelen MJ, Mejia-Aguilar A, Kathrein R, Döller M. From Human Teams to Autonomous Swarms: A Reinforcement Learning-Based Benchmarking Framework for Unmanned Aerial Vehicle Search and Rescue Missions. Drones. 2026; 10(2):79. https://doi.org/10.3390/drones10020079

Chicago/Turabian Style

Bialas, Julian, Mohammad Reza Mohebbi, Michiel J. van Veelen, Abraham Mejia-Aguilar, Robert Kathrein, and Mario Döller. 2026. "From Human Teams to Autonomous Swarms: A Reinforcement Learning-Based Benchmarking Framework for Unmanned Aerial Vehicle Search and Rescue Missions" Drones 10, no. 2: 79. https://doi.org/10.3390/drones10020079

APA Style

Bialas, J., Mohebbi, M. R., van Veelen, M. J., Mejia-Aguilar, A., Kathrein, R., & Döller, M. (2026). From Human Teams to Autonomous Swarms: A Reinforcement Learning-Based Benchmarking Framework for Unmanned Aerial Vehicle Search and Rescue Missions. Drones, 10(2), 79. https://doi.org/10.3390/drones10020079

Article Metrics

Back to TopTop