1. Introduction
Under the combined pressures of intensifying global climate change, continuous population growth, and increasingly scarce arable land, climate-smart agriculture (CSA) [
1,
2] has become a core paradigm for driving the sustainable transformation of agriculture. Technological innovation, represented by precision agriculture, is widely regarded as an important pathway toward achieving CSA goals. Within the technical framework of precision agriculture, the automatic navigation of unmanned ground vehicle (UGV) (e.g., agricultural machinery and spraying vehicles) is recognized as a core supporting component, playing an irreplaceable role in improving production efficiency and operation quality while reducing resource consumption. In particular, high-precision and low-latency positioning of unmanned vehicles is the key to enabling automatic navigation. At present, positioning based on the global navigation satellite system (GNSS) is the primary approach [
3,
4], such as GPS [
5]. Conventional GNSS-assisted positioning methods typically rely on densely deployed infrastructure, such as base stations (BSs) [
6] or access points (APs) [
7], to achieve high-precision positioning through real-time kinematic (RTK) [
8] techniques. However, in large-scale farmland scenarios, the autonomous operation of unmanned vehicles faces two practical limitations. First, BSs and APs are sparse. The sparse infrastructure prevents agricultural machinery from receiving stable RTK correction information, causing the positioning accuracy to degrade sharply—typically from the centimeter level to the meter level. Second, 4G/5G network coverage is poor. In large-scale agricultural scenarios, weak signal strength, limited backhaul links, and narrow channel bandwidth lead to significant transmission delays of positioning correction information, which can easily result in short-term loss of positioning data.
To ensure precise and real-time positioning of unmanned vehicles when the GNSS signal is briefly lost, several assisted positioning architectures have been proposed to compensate for the missing positioning information.
Multi-sensor fusion-based assisted positioning architectures can combine the advantages of different sensors for joint vehicle positioning, achieving higher accuracy and stability than standalone positioning methods [
9,
10,
11]. However, to attain the required accuracy, these methods rely on redundant sensor installation, as each intelligent device must be equipped with multiple sensors, which substantially increases the system cost.
With the development of unmanned aerial vehicle (UAV) technology, assisted positioning architectures that employ multiple UAVs as “mobile point base stations” have been proposed as an alternative to fixed and dense ground infrastructure. Such architectures offer more flexible and convenient deployment and have been widely adopted. An assisted location analysis framework for vehicle positioning in smart-city industrial zones has been proposed [
9], which consists of distributed mobile edge computing and a UAV-deployed Internet of Things to enable flexible positioning. A location-based robust beamforming algorithm has been proposed [
10], which fuses UAV position and mobility information to correct direction-of-arrival (DOA) estimation errors and optimize communication and positioning performance. A source localization method for multi-UAV networks based on a deep neural network and Cramér–Rao bound (CRB) weighting has been proposed [
11], which improves positioning accuracy and computational efficiency through the DeepSSF network and a weighted fusion strategy. A cooperative positioning system employing multiple UAVs equipped with monostatic multiple-input multiple-output (MIMO) radar has been proposed [
12], which achieves high-precision maritime target localization in a space-air-ground integrated network (SAGIN) environment based on weighted subspace fitting. A three-dimensional reconstruction strategy based on multi-angle observations from UAV-swarm synthetic aperture radar (UAV-SAR) has been proposed [
13] for tracking vehicle targets. A multi-UAV cooperative positioning architecture based on an interactive multiple-model maneuvering-target tracking algorithm has been proposed [
14], which is applicable to UAV platforms with bearing-only measurements.
However, Problem 1: existing frameworks all assume that UAVs remain stationary during the positioning process. In practical positioning scenarios, environmental factors such as strong winds cause local positional deviations (oscillations) of the UAVs, which severely degrades the target positioning accuracy.
In addition, the design of the positioning algorithm is another important aspect of the assisted positioning architecture. In multi-UAV-assisted positioning architectures, positioning algorithms fall into four main categories. First, time-of-arrival/time-difference-of-arrival (TOA/TDOA)-based algorithms [
15,
16] require strict time synchronization between the UAVs and ground devices and are highly sensitive to clock offsets—a requirement that is difficult to satisfy in farmland with poor 4G/5G coverage. Second, received-signal-strength-indicator (RSSI)-based algorithms [
17] suffer severe accuracy loss from multipath fading in terrain-obstructed environments, typically achieving only meter-level accuracy. Third, fingerprint-based algorithms [
18] rely on extensive offline field surveys, which is impractical for vast farmland. Fourth, DOA based algorithms [
19] require only a compact uniform linear array (ULA) at the receiver, do not need strict time synchronization, and can achieve sub-degree angular accuracy with only a moderate number of antenna elements. Considering the practical constraints of agricultural deployment, DOA estimation algorithms strike the best balance among accuracy, hardware cost, and deployment feasibility. Therefore, this algorithm is adopted in this paper.
Generally, DOA estimation algorithms can be divided into two categories according to their sampling frequency. The first category comprises algorithms that follow the Nyquist sampling rate. The beamforming method [
19] and Capon’s method [
20] serve as fundamental algorithms for DOA estimation; however, their estimation accuracy is limited and cannot meet the accuracy requirements of practical systems. To address this problem, subspace decomposition algorithms such as MUSIC [
21,
22,
23] and ESPRIT [
24] have been proposed for DOA estimation. These algorithms are widely used in practical systems because they offer acceptable estimation accuracy and low computational complexity under high SNR and sufficient snapshots. However, MUSIC and ESPRIT have poor capability in handling coherent signals and must rely on preprocessing techniques [
23] to assist estimation. Therefore, the FBSS-ESPRIT algorithm [
25] has been proposed, which can process incoherent signals with good real-time performance but suffers from poor estimation accuracy. To overcome this limitation, subspace fitting algorithms such as stochastic maximum likelihood (SML) [
26,
27] and weighted subspace fitting (WSF) [
28] have been proposed. These algorithms achieve higher accuracy than subspace decomposition algorithms and can also handle coherent signals simultaneously. However, they involve a multidimensional nonlinear optimization problem, resulting in extremely high computational complexity that often prevents their direct application to practical systems.
To address this issue, various optimization algorithms based on subspace fitting have been proposed. A robust block-sparse reconstruction DOA estimation algorithm based on WSF (MUC-WSF) has been proposed [
12]; by parameterizing the mutual coupling manifold matrix and constructing an optimal weighted subspace fitting model, it achieves sparse reconstruction without aperture loss under unknown mutual coupling and improves the accuracy and robustness of multi-UAV cooperative maritime target positioning in SAGIN environments. A sieve maximum likelihood method (SiML) has been proposed [
29]; by extending the SML criterion to an infinite-dimensional functional data model and using a sieve to constrain the subspace dimension, it effectively overcomes the high computational cost of conventional SML and its limitation to point-source models, achieving efficient, unbiased, and asymptotically efficient estimation of arbitrarily shaped sources. A stochastic maximum likelihood estimation algorithm based on the alternating direction method of multipliers (MESA) has been proposed [
30]; by reformulating the SML problem as a rank-constrained Toeplitz covariance estimation and solving it with a sequential alternating direction method of multipliers (ADMM), it enables the localization of more uncorrelated sources than sensors in sparse linear arrays and provides robust estimation for highly correlated and coherent source scenarios, while theoretically proving the robustness of this SML method to source correlation. Although these algorithms reduce the computational complexity, the problem of real-time performance still remains.
With the continuous development of DOA estimation technology, the signals received by the array carry an increasing amount of information, and signal processing requires an ever-wider bandwidth, which in turn drives the Nyquist sampling rate ever higher. Consequently, traditional DOA estimation algorithms based on the Nyquist sampling theorem impose increasingly demanding requirements in practical applications. Therefore, how to reduce the sampling frequency—or find an algorithm that operates at a rate far below the Nyquist frequency while still recovering the original information—has attracted widespread attention from researchers.
Building on the above, the second category of algorithms achieves DOA estimation by sampling the signal at a rate far below the Nyquist frequency. Algorithms that exploit the sparse representation of signals are known as compressed sensing (CS) algorithms [
31,
32,
33,
34,
35]. In practical farmland positioning scenarios, the limited cost budget for signal-reception hardware restricts the number of signal samples (i.e., snapshots), resulting in insufficient information about the original signal. Therefore, this paper adopts the CS algorithm as the assisted positioning algorithm. The CS algorithm involves three key aspects: sparse representation, compressive sampling, and signal reconstruction.
At present, according to the completeness of the dictionary, the sparse representation stage can be classified into three types: on-grid [
31], off-grid [
32], and gridless [
33]. The off-grid and gridless methods avoid grid constraints, but they introduce additional floating-point operations into the overall DOA estimation system, leading to excessively high computational complexity and poor positioning timeliness. Therefore, this paper focuses on the on-grid case. On-grid algorithms sparsify the signal by constructing a sparse basis matrix from an uniform overcomplete dictionary (UOD). An algorithm based on iterative dictionary learning has been proposed, which alleviates the grid-mismatch problem in DOA estimation by alternately updating the dictionary and performing sparse recovery [
31]. A two-stage fast matching pursuit (TSFMP) algorithm has been proposed [
34] to reduce computational complexity and improve the positioning accuracy of multi-target DOA estimation. A generalized spatio-temporal coprime sampling (GSTCS) method has been proposed [
35], aiming to maximize the degrees of freedom and jointly estimate the DOA and Doppler frequency through compressed sensing.
Problem 2: The method of constructing the sparse basis matrix from a UOD performs sparsification over the entire angular domain. Because the angles of incident waves are random, DOA estimation must traverse the entire angular domain. Consequently, higher positioning accuracy requires a larger dictionary size, which adversely affects the real-time performance of the positioning system.
Compressive sampling stage: The signal is first sparsified in the sparse representation stage, and the sparsified signal is then passed through the measurement matrix in the compressive sampling stage to obtain the signal measurements. To guarantee the uniqueness of the reconstruction result, the measurement matrix and the sparse basis matrix must be mutually incoherent, that is, they must satisfy the restricted isometry property (RIP) [
36,
37].
In the signal reconstruction stage, the orthogonal matching pursuit (OMP) algorithm [
38] is commonly used for signal recovery. However, the OMP algorithm suffers from low angular resolution and an inability to distinguish highly correlated grids, resulting in poor DOA estimation accuracy. To overcome these limitations, the focused OMP (FOMP) [
39] and generalized OMP (GOMP) [
40] algorithms have been developed, which partially address the DOA estimation accuracy problem. In addition, the MSCM-OMP algorithm, based on the signal covariance matrix and multiple sub-covariance matrices, has been proposed [
41]; it significantly improves angular estimation accuracy at low SNR while reducing computational complexity. For real-time super-resolution direction-of-arrival estimation in automotive radar sensors, the maximum likelihood real-time super-resolution (MARS) algorithm has been proposed [
42]; by exploiting the data correlation between adjacent time slots to shrink the search space and employing the OMP algorithm in place of exhaustive search to efficiently solve the maximum likelihood objective function, it achieves millisecond-level real-time and high-accuracy DOA estimation.
Problem 3: The OMP algorithm and its variants encounter a source-correlation problem when resolving multiple signal sources. Specifically, the reconstruction accuracy of each signal source depends on the reconstruction result of the preceding source. If a reconstruction error occurs, the entire DOA estimation system may fall into a local optimum.
To address the above problems, this paper makes four main contributions:
- 1.
A DOA-estimation-based multi-UAV-assisted positioning architecture is proposed for large-scale smart-farming cultivation scenarios. For the first time, the architecture takes UAV oscillation into account, deriving an assisted-positioning error model together with the associated positioning constraints, thereby providing systematic theoretical support for practical positioning deployment.
- 2.
A displacement-aware array signal reception model that integrates the spatio-temporal information of UAVs is constructed. In this paper, the UAV position and transmission instant are encapsulated into the data payload to eliminate the transmit–receive geometric mismatch, and the instantaneous position is decomposed into a nominal hovering component and a Gaussian random displacement component. For the first time, within a unified framework, we explicitly reveal the mechanism by which UAV jitter affects positioning accuracy through two coupled paths—the delay domain and the angular domain—thereby establishing an analytically tractable physical modeling foundation for the subsequent DOA estimation algorithm.
- 3.
An adaptive local overcomplete dictionary (LOD) model is proposed to improve the real-time performance of positioning. This paper adopts a two-stage “coarse-estimation–local-refinement” dictionary construction strategy: the closed-form solution of FBSS-ESPRIT is first used as the prior center, after which a fine grid is generated only within its neighborhood. Innovatively taking the RMSE as the basic grid unit, we reveal—through Monte Carlo experiments—the phase-transition behavior of the dictionary-width coefficient as a function of SNR, and propose a constrained, SNR-adaptive piecewise-constant value criterion. Based on this criterion, the adaptive LOD model is constructed.
- 4.
A multi-source decoupled signal reconstruction algorithm based on particle swarm optimization (PSO) is proposed. Departing from the traditional serial matching-pursuit paradigm, this paper reformulates the joint DOA estimation of multiple sources as a swarm-intelligence optimization problem in which the angle combination of all sources serves as a single particle state vector. This method transforms the originally serially coupled coherent-source estimation into parallel and independent optimization at the particle level, and, by leveraging the continuous search capability of the PSO algorithm within the narrow LOD domain, breaks through the grid-quantization limitation to achieve high-accuracy DOA estimation. The final results demonstrate that the proposed joint optimization algorithm offers lower computational complexity and higher DOA estimation accuracy, and that within the assisted positioning architecture it achieves a lower positioning error than existing algorithms.
The remainder of this paper is organized as follows. The proposed multi-UAV-assisted positioning architecture and positioning error model are introduced in
Section 2. The proposed signal reception model based on the local position offset of UAVs is presented in
Section 3. The proposed compressed-sensing DOA estimation algorithm optimized by the LOD-PSO model is described in
Section 4. The simulation results are provided in
Section 5. The experimental results are discussed in
Section 6. The conclusions are drawn in
Section 7.
2. System Architecture
Based on the capability of UAVs to assist target positioning and navigation in SAGIN scenarios, a multi-UAV-assisted positioning architecture based on DOA estimation is constructed for farmland positioning scenarios in this paper, as illustrated in
Figure 1.
denotes the working area of the agricultural equipment, and the position of the agricultural equipment
is located within
. Each UAV serves as a signal transmission node, with positions denoted as
, and
, respectively. Each piece of agricultural equipment follows an independent working trajectory and is equipped with a ULA for signal reception [
43,
44]. It is worth noting that the ULA only needs to operate in the sub-6 GHz band, and the half-wavelength inter-element spacing results in an aperture of only a few centimeters. Each array element can be implemented using a standard patch antenna. In addition, the receiving chain of the ULA requires only a conventional Software Defined Radio (SDR) level RF front end, including a low-noise amplifier, mixer, filter, and ADC, without multi-sensor integration or any additional precision components on the ground equipment. Therefore, the unit hardware cost of the ULA is significantly lower than that of a survey-grade RTK-GNSS receiver or a light detection and ranging–inertial measurement unit (LiDAR–IMU) fusion module.
The UGV and unmanned aerial vehicles (UAVs) communicate via a local area network. The yellow lines represent information exchange between the UAVs and the UGV, including the real-time UAV position, battery level, and timestamp. The UAVs only need to transmit radio signals, as indicated by the red lines. The UGV is equipped with a local data-processing system and can process the signals transmitted by the UAVs locally, without forwarding them to the mobile edge computing (MEC) server. This avoids localization-information transmission latency caused by poor network quality and limited signal bandwidth. It should be noted that no signal interference exists between the two communication links of the UGV and UAVs because the communication bands are isolated in frequency.
In
Figure 1, the blue region denotes the coverage area of the base station. When the UGV enters the base-station coverage area, the RTK-GNSS positioning method is adopted to perform conventional unmanned operations. When the UGV moves outside the base-station coverage area, the accuracy of conventional positioning methods decreases substantially. Therefore, a multi-UAV-based DOA estimation-assisted localization algorithm is adopted for cooperative localization. When the assisted localization algorithm is used to localize the UGV, each UAV fuses multi-sensor data, including RTK-GNSS and IMU data [
45], and the resulting localization information can typically be output at a frequency of
[
46]. However, after the electromagnetic signal is transmitted to the agricultural machinery and processed by the local data-processing system, the localization frequency of the agricultural machinery is typically lower than
.
The accuracy target adopted here stems from concrete tillage and seeding operations. Row following in seeding and inter-row mechanical weeding typically demand a lateral accuracy of 2–
, whereas broadcast spraying, fertilization, and coarse tillage can tolerate 10–
. During brief RTK-GNSS outages, the assisted localization layer proposed in this paper is intended not to replace RTK but to constrain drift, keeping the implement within the agronomic tolerance of the operation and thereby preventing the skips and overlaps that waste seed, chemicals, and fuel. Since a UGV traveling at a typical operating speed of 1–
moves
–
per
frame, the sub-meter errors reported in
Section 5.4, delivered at the equipment rate of below
, are sufficient to keep medium-to-coarse operations within tolerance during such outages.
In practical localization scenarios, environmental factors such as strong winds may cause local position deviations of the UAVs. After a UAV transmits an electromagnetic signal to the agricultural machinery, the agricultural machinery processes the received signal and calculates the DOA estimate. Owing to the frequency mismatch between the UAV transmission rate and the signal-processing rate of the agricultural machinery, when the DOA value is calculated and the RTK-GNSS position of the UAV is required to derive the coordinates of the agricultural machinery, the UAV position may already have shifted. Consequently, the transmitted RTK-GNSS position no longer corresponds to the original position from which the electromagnetic signal was transmitted, resulting in localization errors. The specific architectural schematic is shown in
Figure 2.
In a multi-UAV cooperative direction-finding localization system, each platform can obtain only the DOA of the target, so the two-dimensional position of the target must be determined by the spatial intersection of several lines of sight (LOS). The localization accuracy, however, depends not only on the direction-finding algorithm itself but also on the relative geometric configuration of the UAVs, the angle-measurement errors, and the hovering jitter of the platforms. In this work we first establish the linear intersection model derived from the azimuth observations together with the least-squares position estimate; we then develop, on this basis, a first-order propagation model that maps the angle-measurement error and platform jitter to the final localization error; and finally we summarize the geometric feasibility constraints that the relative positions of the three UAVs must satisfy.
To facilitate the subsequent derivations, the principal notation is defined as follows. The position of the k-th UAV is denoted with ; without loss of generality UAV 2 is taken as the coordinate origin, i.e., , and the coordinates of UAV 1 and UAV 3 are written and , respectively. The true position of the target is , and its least-squares estimate is . The azimuth measured by the k-th UAV is , its noisy observation is , and the range from the k-th UAV to the target is .
Each UAV
obtains the azimuth
of the target through direction finding. In the two-dimensional plane, this azimuth defines a line of sight (a ray) originating at
and pointing toward the target; any point
on the LOS satisfies the direction-consistency condition
where
are the coordinates of a running point on the LOS,
and
are the abscissa and ordinate of the
k-th UAV, and
is its measured azimuth. Equation (
1) contains the tangent function and becomes singular as
. To remove this numerical singularity and obtain a linear form valid for all azimuths, we cross-multiply and rearrange it into
with symbols as in (
1). Expanding and moving the constant term to the right-hand side yields the standard linear equation in the target coordinates,
in which the left-hand side is a linear combination of the target coordinates and the right-hand side is a constant determined solely by the position and azimuth of the
k-th UAV.
To simplify the exposition, define the unit normal vector of the
k-th LOS,
where
is perpendicular to the direction of the
k-th LOS and satisfies
. With this definition, (
3) can be written compactly as
where
is the true target position and
is the constant term of the
k-th LOS equation, determined jointly by the UAV position and its azimuth. Stacking the LOS equations (
5) of the three UAVs gives an over-determined linear system
where
is the observation matrix, each row of which is the LOS normal
, and
is the constant vector with components
as defined in (
5).
In the noise-free case (
6) holds exactly and the three lines of sight are concurrent; in practice, however, the azimuths are perturbed to
, the three LOS no longer intersect at a single point, and a least-squares criterion must be adopted. Defining the residual vector
and minimizing its squared Euclidean norm
, the normal equations yield the least-squares estimate of the target position,
where
is the estimated target position and
is the information matrix. This estimate exists and is unique if and only if
has full column rank, i.e., the information matrix is nonsingular, whose geometric meaning is that the three UAVs are not collinear. To quantify the influence of the geometric configuration on localization accuracy, and under the simplifying assumption that the LOS observation errors are independent and identically distributed, we define the geometric dilution of precision (GDOP),
where
denotes the matrix trace and
is the direction difference between the
i-th and
j-th lines of sight. A smaller GDOP indicates a better geometric configuration; as the three UAVs tend to collinearity,
, the GDOP diverges, and the localization error is amplified without bound. The least-squares estimate (
7) and the geometric-accuracy measure (
8) together form the basis for the error-propagation analysis that follows.
Building on the observation model and the least-squares solution, we next analyze how two classes of error sources—angle-measurement error and platform jitter—propagate to the final localization error. Collecting the azimuth errors of the three UAVs into the vector
, where
is the angle-measurement error of the
k-th UAV. Consider first the case in which the UAV positions retain their nominal values while only the azimuths are perturbed. Taking the first-order total differential of the LOS constraint function
(whose form derives from (
5)) and enforcing that it remains identically zero, we obtain
where
is the localization offset induced by the angle-measurement error and
is the range from the
k-th UAV to the target. The coefficient
arises from
: differentiating the normal vector
with respect to the azimuth gives exactly the LOS direction unit vector
, which, combined with the geometric relation
, yields the result. Stacking the three equations (
9) into the matrix form
and taking the least-squares solution gives
where
is the diagonal matrix of the three ranges and
is the angular Jacobian matrix. The corresponding expected mean-square localization error is
where
is the angle-error covariance matrix and
is the angle-measurement variance of the
k-th UAV. The
k-th diagonal element of the inner factor
equals
, showing that the localization error is proportional to the product of the squared range and the angle-measurement variance, while the outer factor
is governed by the GDOP of (
8). This term depends only on the direction-finding accuracy and the geometric configuration, and is independent of time.
Within the total latency
from direction finding to reporting, each UAV undergoes a displacement
due to hovering jitter, where
is the actual UAV position after the latency. From the partial derivative of the constraint function with respect to the UAV position,
, an analogous stacking yields the localization offset due to jitter,
where
is the localization offset induced by jitter and
is the vector of jitter projections onto the LOS normals. This shows that only the jitter component perpendicular to the LOS injects localization error, acting equivalently as an angular perturbation of
. Assuming the jitter is isotropic and mutually independent across UAVs, with normal-projection variance
, the expected mean-square localization error is
where
is the velocity variance of the
k-th UAV’s jitter and
is the total latency. This term is proportional to the square of the latency, indicating that any processing and communication delay is amplified by the platform jitter into an irreducible localization error, which is precisely why the system imposes a hard real-time constraint. The detailed derivation is provided in
Appendix A.
The final localization error is the vector sum of the two terms above,
. Since the angle-measurement noise and the platform jitter are statistically independent, the expected cross-covariance term vanishes, and the expected total mean-square error decouples into two independent components,
where
is the Cramér–Rao lower bound of the
k-th UAV’s angle measurement,
is the efficiency factor of the direction-finding algorithm, and
is the
k-th column of the Jacobian
. Equation (
14) decomposes the localization error rigorously into two indispensable sources: the accuracy term is determined jointly by the direction-finding efficiency and the geometric configuration (the GDOP), and it approaches its theoretical lower limit when
and the deployment is optimal; the real-time term is proportional to the square of the total latency, so that even with perfectly accurate angle measurement, any delay still injects error through (
12). The two constitute an intrinsic trade-off—a more complex algorithm may raise the efficiency
but increases the processing latency
—and therefore an optimal cooperative-localization algorithm must jointly minimize both.
The error expressions (
11) and (
14) remain bounded only under the premise that the three-UAV configuration is well conditioned. Accordingly, we summarize the geometric feasibility constraints that the relative positions of the three UAVs must satisfy; together, these constraints delimit the feasible deployment region over which the least-squares localization of (
7) is solvable, stable, and of controllable accuracy. Setting aside the physical-feasibility constraints related to detection and communication, the purely geometric constraints can be summarized in three interrelated conclusions.
Constraints:
- 1.
The three UAVs must not lie on the same straight line, i.e., the area of the triangle they span must be strictly positive,
where
is the area of the triangle formed by the three UAVs,
are the coordinates of UAV 1 relative to the origin (taken at UAV 2), and
are those of UAV 3.
- 2.
The observation matrix
must have full column rank, and the condition number of its information matrix must not exceed a prescribed threshold,
where
and
are the largest and smallest eigenvalues of the information matrix and
is the maximum admissible condition number.
- 3.
The intersection angle of any two UAV lines of sight must avoid the ill-conditioned region near 0 or
and fall within a safe interval,
where
is the minimum admissible intersection angle and
is the maximum admissible geometric amplification factor.
The threshold
introduced in (C2) is the key parameter determining the size of the feasible region. Regarding its value, the choice of
is essentially a compromise between localization accuracy and deployment flexibility, and in practice it is usually determined by back-solving from the tolerable error-amplification factor
. Common values fall roughly into three tiers. In high-accuracy scenarios with stringent requirements,
is often taken between 4 and 10, corresponding to an intersection angle no smaller than about
–
and a geometric amplification of about 2–3 times, for which the error ellipse is fairly close to a circle. In general engineering scenarios it may be relaxed to
–100, corresponding to an intersection angle no smaller than about
–
, in exchange for deployment flexibility. In practical systems,
is often adopted as a robust yet not overly conservative default threshold, with the final value still to be calibrated in light of the angle-measurement accuracy
, the target ranges
(cf. (
11)), and the mission’s accuracy specification.
3. Signal Reception Model Based on UAV State Variations
Assume that the number of array elements is M. In theory, the array manifold can be arbitrary. However, in practical DOA estimation, the array model is typically configured as a specific array geometry, such as a ULA. Furthermore, it is assumed that K far-field narrowband signals impinge on the sensor array from distinct directions relative to the reference point, and the center frequencies of these signals are known. The signals are modeled as far-field narrowband signals, and .
The baseband complex signal transmitted by the
k-th UAV at time instant
is expressed as follows:
where
denotes the baseband complex signal transmitted by the
k-th UAV.
represents the known pilot/ranging sequence, which is used for signal synchronization, time delay estimation, and DOA estimation.
represents the transmission time instant of the
k-th signal.
denotes the duration of the signal frame.
denotes the data modulation phase, which carries the encoded data packet. The
is defined as follows:
where
denotes the unique identifier of the
k-th UAV.
represent the three-dimensional position of the UAV at the transmission time instant
.
denotes the Cyclic Redundancy Check code, which is employed to ensure data integrity.
The signal arrives at the agricultural equipment array after propagating through the wireless channel. The steering vector for the signal direction is given by the following:
where
is the steering vector corresponding to the direction
.
denotes the relative azimuth angle of the
k-th signal with respect to the heading direction of the agricultural equipment.
denotes the absolute azimuth angle of the
k-th signal in the global coordinate system.
denotes the yaw angle (heading angle) of the agricultural equipment.
d is the inter-element spacing of the array, which is typically set to half a wavelength, i.e.,
.
denotes the carrier wavelength of the signal.
At the reception time instant
, the baseband signal received by the
array element is the superposition of all
k-th signals and noise, which is expressed as follows:
where
denotes the baseband complex signal received by the
m-th array element at time instant
.
denotes the complex amplitude of the
k-th signal, incorporating the path loss and the random channel phase.
denotes the effective time argument of the received baseband signal after compensating for the transmission timestamp and the propagation delay. The variable
denotes the signal reception time at the agricultural equipment array,
denotes the transmission time of the signal emitted by the
k-th UAV, and
denotes the propagation delay from the
k-th UAV to the receiving array.
denotes the position of the UAV at the transmission time instant.
denotes the position of the agricultural equipment at the transmission time instant, which is retrospectively obtained from the historical positioning data of the agricultural equipment. Since the propagation delay is much smaller than the motion timescale of the agricultural equipment,
is used as an approximation of the receiver position during signal propagation.
c denotes the speed of light.
denotes the additive noise at the
m-th array element, which is typically modeled as complex white Gaussian noise.
To make the physical mechanism of UAV oscillation mathematically explicit in Equation (
18), the time-varying UAV position
is further decomposed into a nominal hovering component and a stochastic local displacement component:
where
denotes the nominal hovering position reported by the UAV’s onboard GPS/IMU fusion module at
, and
represents the wind/vibration-induced local positional displacement of the UAV. The displacement is statistically modeled as a zero-mean Gaussian process
where the displacement intensity
characterizes the severity of environmental disturbance. In farmland scenarios,
typically ranges from a few centimeters under calm conditions to tens of centimeters under strong-wind conditions. The case
reduces to the conventional stationary-UAV assumption adopted by existing works, whereas
corresponds to the practical scenario considered in this paper.
Substituting Equation (
19) into the propagation delay defined in Equation (
18), the displacement-aware delay is obtained as follows:
Equation (
21) is mathematically consistent with the original definition of
in Equation (
18); it merely makes the previously implicit displacement component explicit. Because the UAV displacement also perturbs the geometric direction from the UAV to the agricultural equipment, the relative azimuth
acquires a small perturbation
. A first-order Taylor expansion gives the following:
so that the perturbed steering vector can be expressed as follows:
The
represents the gradient of
with respect to the UAV position vector
, indicating how the angle
changes with small variations in the UAV position. The
represents the higher-order terms whose magnitude is on the order of
or smaller, which are usually neglected under small-angle perturbation assumptions. Equations (
19)–(
23) constitute the displacement-aware refinement of the reception model. They explicitly reveal two coupled error pathways through which UAV oscillation degrades positioning accuracy: a delay-domain error introduced via Equation (
21), and an angular-domain error introduced via Equations (
22) and (
23). Both pathways are absent from the conventional stationary-UAV formulation, which directly motivates the displacement-aware reception model proposed in this paper.
By stacking the signals from all
M array elements into a vector, the compact array received signal model is obtained as follows:
where
denotes the effective time-argument vector of the received UAV signals after compensating for the transmission timestamps and propagation delays. Specifically, the
k-th element of
is given by
. This notation indicates that different UAV signals generally have different transmission timestamps and propagation delays.
denotes the array manifold matrix.
represents the relative azimuth angle vector.
denotes the diagonal matrix with
as its diagonal entries.
represents the propagation delay vector.
According to the continuous-time array received signal model in Equation (
24b), let
denote the sampling frequency and
denote the sampling period. To avoid spectral aliasing, the sampling frequency should satisfy the Nyquist criterion, i.e.,
where
B is the bandwidth of the received signal. Let
,
, denote the
q-th sampling instant. Then, the discrete-time array received signal can be expressed as
By stacking
Q snapshots column-wise, the array received data matrix is obtained as
where
denotes the source signal snapshot matrix incorporating the transmission timestamps and propagation delays, and
denotes the array noise snapshot matrix.
Based on the discrete snapshots, the theoretical covariance matrix of the array received signal is defined as
where
denotes the Hermitian transpose. Substituting the discrete received signal model into the above expression yields
where
is the covariance matrix of the element-wise delayed source signal vector, and
is the noise covariance matrix. If the noise is modeled as zero-mean additive white Gaussian noise, then
. In practical implementation, the theoretical covariance matrix is usually approximated by the sample covariance matrix as
where
The proposed array received signal model explicitly incorporates the transmission timestamps and propagation delays, thereby reducing the model mismatch caused by processing latency in conventional methods. Moreover, the model can be incorporated into a variety of DOA estimation algorithms.
Furthermore, in practical farmland positioning scenarios, the CS algorithm is an effective approach for DOA estimation under the condition of limited snapshots. The improved CS algorithm model is presented in detail in the next chapter.
5. Results
In this section, the proposed LOD-PSO algorithm is compared with MSCM-OMP [
41], FOMP [
39], FBSS-ESPRIT [
25], MESA [
30], MUC-WSF [
12], ss-MUSIC [
23], and MARS [
42] through multiple sets of experiments. Among these methods, ss-MUSIC [
23] and FBSS-ESPRIT [
25] are subspace-decomposition-based algorithms, MESA [
30] and MUC-WSF [
12] are subspace-fitting-based algorithms, and MSCM-OMP [
41], FOMP [
39], and MARS [
42] are CS-based algorithms. By comparing the LOD-PSO algorithm with different types of algorithms, the general applicability of the proposed algorithm is demonstrated.
A uniform linear array is considered. The parameters are set as
,
, and
. In addition, the distance
d between two adjacent sensors is set to one half of the signal wavelength
, i.e.,
. The angular search step size of all algorithms is set to
. All results are obtained from 100 Monte Carlo (MC) trials. The RMSE is used to evaluate the estimation accuracy and is defined as in (
33). All experiments are conducted utilizing MATLAB R2023b on a personal computer with an Intel R7-4800H CPU and 16 GB RAM, without GPU acceleration. The per-frame processing latency of all methods is measured on this identical platform using the same input data. The timing covers the complete pipeline from the array received data to the output localization estimate, and, for LOD-PSO, it includes its FBSS-ESPRIT coarse-estimation stage. A warm-up run is performed in advance to eliminate the first-call overhead, thereby ensuring a fair comparison among all algorithms under identical conditions.
The parameters of the PSO stage are configured as follows: the swarm size is set to and the maximum number of iterations to , while the inertia weight decreases linearly from to . Since the particle swarm is confined to a LOD centered on the FBSS-ESPRIT coarse estimate, such a small swarm size and iteration number are sufficient for stable convergence, which is one of the reasons for the low latency of LOD-PSO. Under the hovering condition, the position jitter of each UAV is modeled as a zero-mean, isotropic, temporally white Gaussian process. Given a hovering oscillation velocity of and a per-frame coherent processing time of as the characteristic integration interval of the jitter, the displacement standard deviation of each UAV is , and the corresponding jitter variance is (, identical for all three UAVs).
5.1. DOA Estimation Accuracy
5.1.1. Angular Distribution of DOA Estimates
Figure 7 shows the angular distributions of DOA estimates obtained by different algorithms under incoherent signal conditions. The true DOAs are randomly generated as
, and 100 Monte Carlo (MC) trials are conducted.
Figure 7a presents the angular distributions of different algorithms when the SNR is
dB, whereas
Figure 7b presents the corresponding results when the SNR is 2 dB.
As shown in
Figure 7a, when the SNR is
dB, the DOA estimates obtained by MSCM-OMP [
41], FOMP [
39], FBSS-ESPRIT [
25], MESA [
30], MUC-WSF [
12], ss-MUSIC [
23], and MARS [
42] exhibit relatively large angular fluctuations around the two true DOAs. In particular, the estimated angles of several comparison algorithms are widely scattered, indicating that their estimation performance is significantly affected by noise in low-SNR scenarios. This phenomenon is especially evident near the two target angles, where the estimated DOAs deviate noticeably from the ground-truth values and the distribution of the estimation points becomes dispersed. The large dispersion reflects the reduced stability and robustness of these algorithms under severe noise interference.
In contrast, the proposed LOD-PSO algorithm yields DOA estimates that are tightly concentrated around the true DOAs even at dB. The estimated angle distributions corresponding to the two sources show narrow spreads and remain close to and , respectively. This demonstrates that the proposed algorithm can effectively suppress the influence of noise and maintain high estimation accuracy under low-SNR conditions. Compared with the other algorithms, LOD-PSO exhibits better robustness and stronger global search capability, thereby reducing the probability of large angular estimation errors.
When the SNR increases to 2 dB, as shown in
Figure 7b, the angular distributions of all algorithms become more concentrated around the true DOAs, indicating that the improvement in SNR enhances the DOA estimation accuracy. However, differences among the algorithms can still be observed. Although the estimation fluctuations of the comparison algorithms are reduced compared with those at
dB, several methods still present a certain degree of angular dispersion. By contrast, the proposed LOD-PSO algorithm produces the most compact distributions around the two true DOAs, with almost all estimated angles located near the reference lines. This indicates that LOD-PSO not only performs well in low-SNR environments but also maintains superior estimation stability when the signal quality improves.
5.1.2. RMSE vs. SNR
Figure 8 presents the comparison of RMSE versus SNR for different algorithms under incoherent signal conditions, as shown in
Figure 8a, and coherent signal conditions, as shown in
Figure 8b. It can be observed that the RMSE values of all algorithms decrease with increasing SNR, indicating that the DOA estimation accuracy improves as the noise level decreases. However, the proposed LOD-PSO algorithm consistently achieves the lowest RMSE over the entire SNR range in both scenarios.
For incoherent signals, as shown in
Figure 8a, MSCM-OMP [
41], FOMP [
39], and FBSS-ESPRIT [
25] exhibit relatively large DOA estimation errors in the low-SNR region. In particular, when
dB, their RMSE values are
,
, and
, respectively, whereas the RMSE of the proposed LOD-PSO algorithm is only
. Thus, the RMSE of LOD-PSO is
,
, and
of that achieved by MSCM-OMP [
41], FOMP [
39], and FBSS-ESPRIT [
25], respectively. This indicates that the proposed algorithm can effectively suppress noise-induced estimation errors and provide more accurate DOA estimates under low-SNR conditions.
For coherent signals, as shown in
Figure 8b, the estimation accuracy of most algorithms deteriorates due to the correlation among sources. This degradation is particularly significant for OMP-derivative algorithms and conventional subspace-based methods, since source coherence may cause covariance matrix rank deficiency and inaccurate direction selection. At
dB, the RMSE values of MSCM-OMP [
41], FOMP [
39], and FBSS-ESPRIT [
25] are
,
, and
, respectively, while LOD-PSO achieves an RMSE of
. Therefore, the RMSE of LOD-PSO is
,
, and
of that obtained by MSCM-OMP [
41], FOMP [
39], and FBSS-ESPRIT [
25], respectively. These results demonstrate that LOD-PSO can effectively alleviate the negative effects of source correlation and noise.
Moreover, the enlarged regions in
Figure 8 further show that LOD-PSO maintains a clear advantage in the medium- and high-SNR ranges. Although the RMSE values of all algorithms gradually decrease as the SNR increases, LOD-PSO still provides the smallest estimation errors. Overall, the proposed LOD-PSO algorithm consistently outperforms the existing state-of-the-art algorithms in terms of DOA estimation performance. The improvement is especially significant under low-SNR and coherent-source conditions, demonstrating the robustness, accuracy, and adaptability of the proposed method.
5.2. Real-Time Performance Comparison of Different Algorithms
Table 1 presents the computational complexity comparison of the eight algorithms considered in this article, and
Table 2 summarizes the parameters involved in the analysis. Throughout the following discussion, we treat the array size
M, the grid size
L, the dictionary size
, and the snapshot number
Q as the growing dimensions, whereas the source number
K, the swarm size
, and the iteration counts (
,
,
) are regarded as bounded constants, as is standard in asymptotic complexity analysis. Under the experimental configuration
,
, and
, we have
, so that the eigendecomposition term
dominates the covariance-formation term
. The computational cost of each algorithm can be decomposed into two parts:
covariance/subspace construction and
parameter search/iteration.
5.2.1. Covariance/Subspace Construction
The sample covariance matrix is shared by ss-MUSIC, MSCM-OMP, FBSS-ESPRIT, MESA, MUC-WSF, the proposed LOD-PSO, and its formation costs . Among these, ss-MUSIC, FBSS-ESPRIT, and MUC-WSF must additionally perform an eigendecomposition of (or of the forward–backward smoothed covariance) at a cost of to extract the signal/noise subspaces and , giving a construction overhead of . MSCM-OMP and MESA operate directly on the covariance and thus incur only , whereas FOMP works on the raw snapshots and MARS processes a single-snapshot statistic, so their construction cost is negligible, i.e., . Crucially, the proposed LOD-PSO evaluates the deterministic maximum-likelihood (SML) criterion directly on and never forms nor inverts an matrix; hence its construction overhead is only , strictly cheaper than the eigendecomposition-based subspace methods.
5.2.2. Parameter Search/Iteration
For ss-MUSIC, the dominant cost is the exhaustive traversal of the spatial spectrum, , which grows linearly with the grid size L and becomes the bottleneck once the angular resolution is tightened to . MSCM-OMP and FOMP rely on greedy sparse reconstruction over an overcomplete dictionary, costing and , respectively, whose dominant factor is the dictionary size rather than M; to guarantee angular resolution, must be kept large, which keeps their complexity high. FBSS-ESPRIT avoids any grid search and only solves the rotational-invariance equation, at a cost of ; its overall complexity is therefore governed by the eigendecomposition. MESA is dominated by the sequential ADMM outer loop, , which cannot be parallelized. MUC-WSF performs a nonlinear weighted-subspace-fitting search, . MARS employs OMP over a temporally correlated reduced search space of size , giving ; although far cheaper than the full-grid ss-MUSIC, it still scales with the reduced grid size .
For the proposed LOD-PSO, the parameter-search cost is that of the particle swarm confined to the limited optimization domain (LOD). Each fitness evaluation forms the array manifold and the orthogonal projector , and it then evaluates ; since the only inversion is of the matrix , the per-evaluation cost is , with no term. With particles evolved over iterations, the parameter-search complexity is . Because the LOD, extracted from a low-cost coarse localization pass, confines the swarm to only a small fraction of the full field of view , the seeding of the particles takes only scatter operations and the product behaves as a bounded small constant; consequently, for any practically useful angular resolution.
5.2.3. Overall Comparison
As can be seen from
Table 1, the proposed LOD-PSO attains a highly competitive computational complexity among all the compared algorithms and in particular delivers the lowest
measured single-frame latency among all
accurate super-resolution estimators. This is due to two reasons: (i) the limited optimization domain restricts the parameter search to a very small feasible region, so that
acts as a bounded small constant; and (ii) LOD-PSO evaluates the SML criterion directly on the sample covariance, so that its refinement stage scales only as
and never invokes a full-resolution grid or dictionary traversal. Asymptotically, treating
K,
, and
as constants, the refinement cost of LOD-PSO reduces to
; however, as detailed below, its
overall complexity still inherits the
term of the coarse-initialization stage.
For completeness, the overall complexity of LOD-PSO must include its coarse-initialization stage. LOD-PSO first runs FBSS-ESPRIT to obtain the dictionary centers at a cost of , and it subsequently performs the swarm refinement at ; its total cost is therefore . It follows that LOD-PSO shares the same eigendecomposition lower bound as the subspace-based methods, rather than lying strictly below it. Its practical advantage does not stem from avoiding the eigendecomposition, but from replacing the expensive fine-resolution search with a swarm search confined to the narrow LOD, for which behaves as a bounded small constant. This is precisely why, at the resolution adopted throughout, the measured single-frame latency of LOD-PSO is significantly lower than that of the grid/dictionary-heavy methods.
In contrast, ss-MUSIC, FBSS-ESPRIT, and MUC-WSF all retain the eigendecomposition (which dominates since ); MESA carries the sequential ADMM cost; MSCM-OMP, FOMP, and MARS remain governed by the dictionary/grid factors or ; and ss-MUSIC additionally suffers from the spectral traversal. Grouping these observations by dominating order, the sparse-recovery methods MARS, MSCM-OMP, and FOMP scale with the reduced grid or dictionary size ( or ), whereas the eigendecomposition-based methods LOD-PSO, FBSS-ESPRIT, MESA, and MUC-WSF, together with the full-grid ss-MUSIC, are all bounded below by ; among the latter, LOD-PSO and FBSS-ESPRIT remain the cheapest because their post-eigendecomposition stages contribute only a bounded constant, while MESA, MUC-WSF, and ss-MUSIC pay heavy iterative or full-grid overheads on top of the same term. Where FBSS-ESPRIT is the fastest owing to its closed-form estimator on the small array , LOD-PSO follows immediately because its swarm refinement adds only a bounded small constant on top of the same initialization, the placement of MARS ahead of MSCM-OMP follows from the typical regime , and ss-MUSIC ranks slowest because its full-grid spectral traversal explodes at the resolution on top of the eigendecomposition. LOD-PSO is therefore the only algorithm that pairs a subspace-grade lower bound with a bounded-constant refinement cost, while still delivering maximum-likelihood-grade super-resolution accuracy.
5.3. Ablation Study
To attribute the observed performance gains to each individual contribution, we conduct an ablation study that isolates the effect of the three interrelated components proposed in this paper: the displacement-aware reception model, the LOD, and the PSO-based joint reconstruction. Starting from a common baseline—uniform overcomplete dictionary OMP reconstruction built on a stationary-UAV reception model (denoted UOD + OMP)—we activate the three contributions one at a time and evaluate the resulting DOA root-mean-square error (RMSE) and the per-frame processing latency. To make the ablation consistent with the main claims of this paper, we consider two representative cases. The first case corresponds to non-coherent sources at , which evaluates the individual contribution of each component under a moderate-SNR non-coherent condition. The second case corresponds to coherent sources at , which is a more challenging low-SNR multipath scenario and directly reflects the regime where the proposed LOD-PSO method is expected to provide its main advantage. In both cases, every other array, channel, snapshot, and jitter parameter is held identical to the common baseline so that any change in performance can be ascribed solely to the newly activated component.
The ablation reveals that the three contributions are complementary and address distinct error mechanisms rather than overlapping in function. For the non-coherent case in
Table 3, progressing from B0 to B1, the introduction of the displacement-aware reception model reduces the DOA RMSE from
to
while leaving the latency almost unchanged (from
to
). This is expected: the displacement-aware model does not alter the estimator or its search space but instead removes the transceiver geometric mismatch introduced by UAV hovering displacement, thereby suppressing the motion-induced component of the estimation error. In the notation of our error-propagation framework, it acts primarily on
.
Progressing from B1 to B2, activating the LOD—while still employing the greedy OMP solver—reduces the per-frame latency by roughly an order of magnitude, from to . This dramatic reduction arises because the LOD confines the reconstruction to a compact, statistically bounded search window centered on the FBSS-ESPRIT estimate, rather than scanning the full uniform grid. Notably, the DOA RMSE also improves from to , indicating that pruning implausible grid atoms not only accelerates the solver but also removes spurious candidates that would otherwise degrade the estimate. The LOD therefore acts principally on the latency term, and thereby—through the algorithmic delay —indirectly on .
Finally, progressing from B2 to B3, replacing OMP with the joint PSO reconstruction yields a further improvement in DOA accuracy under non-coherent sources, driving the DOA RMSE down from to , with a further slight reduction in latency (from to ). This confirms that the PSO reconstruction refines the sparse solution beyond the reach of the greedy OMP by jointly optimizing the on-grid atoms, thereby mitigating the residual grid-mismatch and multipath-induced errors that persist even in the non-coherent regime, operating directly on the reconstruction-induced error term .
The additional low-SNR coherent-source ablation in
Table 4 further confirms the necessity of the proposed components in the most challenging operating regime. Compared with the non-coherent case, the baseline C0 suffers a much larger RMSE of
, because source coherence causes covariance-rank degradation and also aggravates the sequential error-propagation problem of OMP. After introducing the displacement-aware reception model, the RMSE decreases from
to
, while the latency changes only slightly from
to
. This indicates that the displacement-aware model remains effective under coherent low-SNR conditions, but its role is mainly to correct the UAV-motion-induced geometric mismatch rather than to resolve the source-correlation problem.
When the LOD is further activated, the latency is reduced from to , demonstrating that the LOD is the principal contributor to real-time performance even under low-SNR coherent-source conditions. Meanwhile, the RMSE decreases from to . This improvement is more pronounced than in the non-coherent case because the global UOD contains many highly correlated and noise-sensitive atoms under low-SNR coherent conditions. By restricting the reconstruction to an SNR-adaptive local window centered on the FBSS-ESPRIT estimate, the LOD removes a large number of implausible atoms and therefore reduces the probability that OMP selects spurious directions.
Finally, replacing OMP with PSO in the LOD gives a further RMSE reduction from to , while the latency is further reduced from to . This additional gain is particularly important for coherent sources. Unlike OMP, which estimates the sources sequentially and therefore suffers from residual-coupling and error-cascade effects, the PSO-based reconstruction treats the multi-source DOA vector as a single particle state and jointly optimizes all source directions. Therefore, it directly mitigates the coherent-source coupling that remains after the displacement-aware modeling and LOD pruning stages. This result explains why the full LOD-PSO method achieves its largest relative advantage in the low-SNR coherent-source regime.
In summary, the ablation study demonstrates that each contribution plays a distinct and mutually reinforcing role: the displacement-aware reception model targets the geometric mismatch term , the LOD targets the latency and, through , further reduces , and the PSO reconstruction targets the residual reconstruction error . The newly added coherent-source ablation at further verifies that this division of function still holds under the low-SNR coherent condition emphasized in this paper. In particular, the PSO reconstruction contributes a larger relative accuracy gain in the coherent case than in the non-coherent case, confirming that the proposed joint reconstruction is essential for breaking the source-correlation cascade of greedy OMP. It is their combination in the full LOD-PSO method that attains simultaneously the lowest DOA RMSE and the lowest latency in both tested ablation scenarios.
5.4. Positioning Error
The O agricultural machinery units are randomly positioned within the UAV coverage region, forming the set . In the UAV reference frame, the positions of UAVs , , and are assumed to be , , and , respectively, and at least one agricultural machinery unit within the UAV coverage region has available GPS positioning information during the localization process.
In the experiment, the signal sources are coherent,
, and the SNR is set to
and
. The proposed LOD-PSO algorithm is compared against seven representative baselines spanning the three mainstream technical routes: the subspace-decomposition methods ss-MUSIC [
23] and FBSS-ESPRIT [
25], the subspace-fitting methods MESA [
30] and MUC-WSF [
12], and the compressed-sensing methods MSCM-OMP [
41], FOMP [
39], and MARS [
42]. Each algorithm first performs DOA estimation for the
O machinery positions to obtain the azimuth information
, from which the two-dimensional position
is recovered by the least-squares cross-fix of Equation (
7). Utilizing these angular estimates, the total average positioning error
over
O trials is computed as
where
and
denote the true
x- and
y-coordinate positions of the
jth agricultural machinery in the UAV coordinate system, and
and
represent the corresponding estimated positions of the
jth machinery in the
ith trial.
Consistent with the mean-square error decomposition of Equation (
14), the empirical total positioning error is the superposition of two mutually independent components,
. Here,
is the
DOA-estimation error corresponding to the accuracy term of Equation (
11), which is governed by the angular RMSE together with the geometric configuration (the GDOP of Equation (
8)); it is measured under the ideal stationary-UAV setting and after the proposed displacement-aware reception model (Equations (
19)–(
23)) has removed the platform-drift contribution. In contrast,
is the UAV-oscillation error corresponding to the real-time term of Equation (
13), which grows quadratically with the reporting latency
t (i.e.,
) and reflects the platform jitter accumulated during processing.
Table 5 reports, for each algorithm, the DOA-estimation error
, the single-frame processing latency
t, the isolated oscillation-induced error
, and the resulting total error
.
DOA-estimation error . The proposed LOD-PSO consistently attains the smallest DOA-induced error at both SNR levels (
at
and
at
). This directly follows from the accuracy term of Equation (
11): by refining the angular resolution within the local overcomplete dictionary and jointly optimizing the multi-source angle vector through the particle swarm, LOD-PSO delivers the lowest RMSE among all methods, so its efficiency factor approaches the ideal limit
. By contrast, the subspace-decomposition methods FBSS-ESPRIT and ss-MUSIC suffer from covariance rank deficiency under coherent sources and yield the largest values (
and
at
), and the greedy FOMP is degraded by the source-correlation cascade; MESA, MUC-WSF, and MARS occupy an intermediate tier, with MSCM-OMP being the strongest baseline yet still inferior to LOD-PSO.
UAV-oscillation error and latency t. During the interval between the direction-finding instant and the reporting instant, each UAV drifts at an approximately constant operating speed
, so the oscillation-induced displacement error is governed solely by the algorithm’s processing latency,
(Equations (
13) and (
14)). Consequently,
increases
monotonically with
t: the shorter the latency, the smaller the drift. The closed-form FBSS-ESPRIT, being the fastest estimator for the small array
, therefore incurs the
smallest oscillation error (
/
), immediately followed by LOD-PSO (
/
), whose complexity is the lowest among all
accurate super-resolution estimators (Equation (
40)). At the opposite extreme, the dictionary/grid/iteration-heavy methods pay a heavy price: the full-grid ss-MUSIC (
), the nonlinear-search MUC-WSF (
), and the ADMM-based MESA (
) allow the UAVs to drift appreciably, producing the largest oscillation errors (
,
, and
at
); the large-dictionary FOMP (
) follows the same trend. This behavior is now fully consistent with the model: latency and oscillation error rise and fall together for every method, without exception.
Total error . The total positioning error combines the two orthogonal contributions through their root-mean-square,
, rather than a plain summation, so that the angular-estimation error and the UAV-oscillation error are aggregated in the mean-square sense consistent with their statistically independent origins. Under this metric, LOD-PSO achieves the lowest total positioning error in every configuration (
at
and
at
). At
, its total error is
,
,
,
,
,
, and
of that of MSCM-OMP, FOMP, FBSS-ESPRIT, ss-MUSIC, MESA, MUC-WSF, and MARS, respectively; the corresponding ratios at
are
,
,
,
,
,
, and
, showing that the advantage is most pronounced at low SNR, where the DOA error is most severe. Because LOD-PSO provides the smallest total positioning error while maintaining a short latency, it achieves a favorable balance between localization accuracy and real-time efficiency in Equation (
14). It should also be noted that the practical applicability of the obtained accuracy depends on both the SNR condition and the required precision of the agricultural operation. For example, at
, the total error of LOD-PSO is
, which is more suitable for agricultural UAV tasks with moderate positioning requirements, such as field inspection, crop-growth monitoring, wide-area mapping, and general navigation assistance. In contrast, under a more favorable signal condition such as
, the total error is reduced to
, making the method more applicable to higher-precision tasks such as seeding, plant protection, and precision spraying. Therefore, rather than claiming that all sub-meter errors are sufficient for operations requiring 10–
accuracy, the revised interpretation qualifies the applicability of LOD-PSO according to the operating SNR and the target agricultural operation. Combined with a latency compatible with the real-time response requirement of farm equipment, these results indicate that LOD-PSO provides the best overall accuracy, robustness, and real-time performance for cooperative multi-UAV localization in smart-tillage scenarios.
All results reported in this section are obtained from
independent Monte Carlo trials at each SNR point (SNR = signal-to-noise ratio, in dB). In each trial, the source angles (the true directions-of-arrival of the UAV emitter), the noise realizations (the independently generated measurement-noise samples), and the platform-jitter displacements (the small random offsets of the agricultural machinery platform) are redrawn independently so that the
N trials are mutually independent. To make the statistical uncertainty of these estimates explicit, we additionally characterize their trial-to-trial variability. For a second-order statistic such as the RMSE, the estimator
, where
is the estimated root-mean-square error,
is the estimated position vector in trial
i,
is the true position vector, and
is the squared localization error of trial
i, has, by the delta method (a first-order error-propagation technique), a coefficient of variation
, where
is the relative standard deviation of the RMSE estimate,
and
v denote the mean and variance of the per-trial squared error
, and
N is the trial count. Under the standard zero-mean Gaussian-error approximation (
, i.e., the squared error follows an approximate chi-square distribution so that its variance equals twice the square of its mean), this reduces to the conservative bound
so that the corresponding
confidence interval (the interval covering the true value with
probability, using the normal critical value
) of every reported RMSE value has a relative half-width of
(i.e., each RMSE should be read as
). Because the performance separations in
Table 5 are far larger than this interval—for instance, the
gap between LOD-PSO (
, the proposed method’s total error at
) and the next-best MARS (
, the smallest error among competing methods at the same
), whose
intervals
(computed as
) and
(computed as
) do not overlap—the confidence intervals of the compared methods are mutually disjoint (non-overlapping) and the method ranking is statistically unambiguous (clearly distinguishable) at the present sample size (the current
). Increasing
N would narrow these intervals further but would neither alter the ranking nor affect any conclusion of this work.
6. Discussion
The simulation results reported in
Section 5 can be coherently interpreted only when they are read against the error-decomposition theory established in
Section 2. Equation (
14) predicts that the expected total mean-square localization error separates into a configuration- and accuracy-governed term
and a latency-governed term
, and that these two terms constitute an intrinsic trade-off rather than two independently optimizable quantities. The central empirical finding of this work is that LOD-PSO is the only method among the eight compared that suppresses both terms simultaneously: it attains the smallest DOA-induced error
(
Table 5,
at
), thereby driving the efficiency factor
toward its ideal limit, while its grid/dictionary-free complexity
keeps its latency
t among the lowest of all methods (
at
, second only to FBSS-ESPRIT) and hence keeps its oscillation-induced error
correspondingly small (
). No other method holds both terms small at once: the fastest estimator, FBSS-ESPRIT, attains the smallest oscillation error (
) only at the price of by far the largest accuracy error, whereas the more accurate spectral and sparse baselines incur latencies that inflate their oscillation term, so that LOD-PSO alone reaches the smallest total error
(
).
A closer reading of
Table 5 exposes a failure mode that a purely latency-oriented analysis would overlook, and that only the decomposition of Equation (
14) renders visible. FBSS-ESPRIT records the lowest wall-clock latency of all methods (
at
) and, as a direct consequence, the
smallest oscillation-induced error (
); yet it produces by far the
largest DOA-induced error (
) so that its total error (
) is dominated almost entirely by the accuracy term and ranks among the worst of all eight methods. This is not a contradiction but a direct manifestation of Equation (
14): the raw subspace fix returns the fastest solution precisely because it performs no super-resolution refinement, but the large angular bias it leaves behind, propagated through the accuracy term, dwarfs whatever is saved on the real-time term. The two contributions therefore act as genuinely separate levers—minimizing latency protects only the oscillation term and leaves the accuracy term unbounded. This is the deeper reason why merely minimizing computational latency, as is common in the real-time DOA literature, is insufficient for cooperative localization, and why an estimator must be simultaneously accurate and fast: LOD-PSO is the only compared method that satisfies both requirements, trading a negligible latency increment over FBSS-ESPRIT (
versus
, i.e., an oscillation-error increase of only about
) for a more-than-tenfold reduction in accuracy error (
versus
).
Situated against prior multi-UAV DOA-assisted localization frameworks, which uniformly assume stationary platforms during the direction-finding-to-report interval, the present contribution is best understood as making the previously hidden real-time term of Equation (
14) both explicit and controllable. The displacement-aware reception model of Equations (
19)–(
23) is what allows the platform-drift contribution to be absorbed at the level of the observation model rather than propagated into the world-coordinate solution, and the isolated column
in
Table 5 quantifies, for the first time in this setting, the magnitude of the error that the stationary-UAV assumption silently incurs—on the order of
to
depending on the estimator and SNR, since this term scales directly with the algorithmic latency of each method. From a methodological standpoint, the coarse-to-fine architecture of the LOD is likewise a departure from the conventional uniform-overcomplete-dictionary paradigm: by confining the sparse recovery to an SNR-adaptive neighborhood of the FBSS-ESPRIT initialization, the multi-source search space contracts by a factor exceeding
, which is the structural mechanism that makes maximum-likelihood-grade super-resolution feasible on embedded ARM-class processors at the sub-
rate required by field equipment, without GPU acceleration. Equally important, the PSO reformulation converts the sequential, residual-coupled atom selection of OMP-type recovery into a joint evaluation of the complete
K-source angle vector, which is the key to breaking the source-correlation cascade that degrades greedy methods under coherent signals—a condition that is the norm, not the exception, in farmland multipath environments.
Beyond the specific algorithm, the results carry a broader implication for the deployment of SAGIN-based positioning in precision agriculture. The error budget of Equation (
14), together with the geometric feasibility constraints (
C1)–(
C3), provides a principled design tool that decouples the two levers available to a system integrator: the accuracy term can be reduced by improving the estimator (raising
), whereas the real-time term can be reduced either by shortening
or, geometrically, by maintaining a well-conditioned UAV configuration that bounds
. Because
enters quadratically while the geometric amplification enters through the condition number, the analysis suggests that, in high-jitter conditions, investing in either faster onboard computation or tighter formation control yields a larger marginal return than further refining angular accuracy—a quantitative guideline that the raw RMSE-versus-SNR curves alone cannot provide.
From the viewpoint of practical feasibility, the above real-time and accuracy results must also be considered together with UAV autonomy, endurance, and deployment cost. In typical farmland scenarios, the assisting UAVs are not required to perform high-speed trajectory tracking; rather, they mainly need to maintain stable hovering or slow cooperative flight while providing DOA measurements for the UGV. Therefore, the proposed framework is compatible with commercial multi-rotor UAV platforms equipped with lightweight communication, antenna-array, and embedded-computing modules. For short- to medium-duration operations, the available flight time of such UAVs can support continuous localization assistance over a single operation window. For longer smart-tillage tasks, continuous service can be maintained by UAV rotation, scheduled battery replacement, or field-side charging, so that one subset of UAVs performs localization while another subset is recharged or replaced. This deployment strategy prevents UAV endurance from becoming a fundamental limitation of the proposed localization architecture.
The overall system cost is also quantifiable at the level of the main hardware components. A practical implementation mainly consists of multiple commercial UAV platforms, lightweight antenna or receiving arrays, communication modules, and an embedded processing unit. Since the proposed LOD-PSO algorithm achieves the required update rate without GPU acceleration in the simulated setting, the processing unit can potentially be implemented on an ARM-class embedded platform rather than a high-power workstation, which may reduce both cost and energy consumption. In a typical small-scale deployment using three to four assisting UAVs, the dominant cost would come from the UAV platforms and their onboard sensing/communication modules, while the additional algorithmic cost would be limited to the embedded processor and software implementation. Consequently, the proposed system would be more expensive than a single GNSS-only receiver, but it may provide a complementary positioning mechanism for GNSS-degraded or temporarily unavailable conditions. This cost–benefit trade-off may be more acceptable for high-value smart-tillage operations where localization performance can influence operation quality and continuity.