1. Introduction
Unmanned aerial vehicles (UAVs) are increasingly used in low-altitude sensing, intelligent inspection, and emergency response. In these applications, the UAV continuously captures visual scenes and reports task-relevant information to the ground side. However, raw image transmission is difficult to sustain under the bandwidth, transmit-power, and onboard computing constraints of UAV platforms. Semantic communication alleviates this bottleneck by shifting the focus from bit-level fidelity to task-oriented meaning. Recent surveys and reviews have further clarified its theoretical basis, task-oriented objective, and future-Internet relevance in emerging wireless and AI-native systems [
1,
2,
3]. Recent implementation-oriented discussions also address semantic transmission compatibility with digital wireless systems and cloud-edge-device collaborative settings [
4,
5]. For image-oriented transmission, recent semantic coding studies also show that useful task semantics can be preserved under constrained bandwidth and dynamic channels through task-aware representation learning and lightweight semantic reconstruction [
6,
7]. UAV-based visual sensing is an important application domain for such systems: robust aerial semantic segmentation under resource constraints has been studied in [
8], and edge perception frameworks for next-generation networks with semantic awareness are discussed in [
9]. Recent high-quality studies further investigate distortion-resilient goal-oriented transmission and privacy-aware image JSCC [
10,
11], while recent implementations continue to explore adaptive image transmission under resource constraints [
12].
However, in UAV semantic communication, the key challenge is no longer efficient transmission alone, but privacy-aware control under coupled resource constraints. On the visual side, intermediate semantic features may still preserve contours, scene structure, and category cues, leaving the system vulnerable to semantic inference, feature analysis, model inversion, or model-theft-assisted eavesdropping [
13,
14,
15]. Broader security reviews and recent secure semantic transmission studies likewise emphasize adaptive and context-aware protection as a growing theme in semantic secure communications [
15,
16]. On the spatial side, bounding-box centers and derived location descriptors can reveal target positions, mission areas, and mobility traces, which remain high-risk structured spatial data in recent edge offloading privacy, trajectory privacy publishing, and privacy-preserving UAV location-authentication studies [
17,
18,
19]. UAV-specific collaborative sensing studies further show that distributed localization and fusion can also leak geometric state information unless privacy-preserving masking is incorporated into the fusion stage [
20]. Recent mobility-oriented semantic communication surveys likewise emphasize that semantic messages may leak trajectories and contextual behavior unless privacy and security are designed together [
21]. Related UAV privacy studies also show that privacy-preserving offloading and geofence-aware authentication are becoming practical requirements for resource-constrained aerial systems [
19,
22]. Meanwhile, UAV platforms must operate under limited transmit power, fluctuating air–ground channels, and constrained onboard computation. Therefore, the core problem addressed in this paper is how to jointly suppress visual and location privacy leakage while preserving legitimate task recovery in a resource-constrained UAV semantic communication link.
Existing studies remain insufficient in four closely related respects. First, semantic communication foundations remain task-oriented rather than dual-privacy-coupled. Recent reviews and representative studies consolidate the semantic communication paradigm and task-oriented transmission principles [
1,
2,
6]. However, they do not address coordinated visual and location privacy in UAV sensing scenarios. Second, semantic transmission adaptation and edge-inference optimization still emphasize utility maximization rather than joint privacy control. Existing studies improve task utility and resource efficiency under dynamic constraints [
4,
23]. However, they do not coordinate visual privacy, location privacy, and transmit-power budgets within one framework. Third, secure semantic communication studies mainly focus on visual leakage, model security, or representation inversion. They clarify semantic-feature privacy risks and the need for representation-level protection [
11,
13,
14]. However, they usually omit location leakage carried by object descriptors. Fourth, location privacy studies provide useful spatial protection tools, but rarely integrate them into semantic communication utility optimization. Edge offloading privacy, trajectory publishing, and UAV location authentication address important facets of spatial disclosure [
17,
18,
19]. UAV-specific distributed localization privacy is also emerging [
20]. However, these strands are seldom coupled with semantic decoding quality and air–ground resource allocation in a unified UAV setting. Therefore, a unified setting that controls legitimate-receiver task utility, attacker-side recoverability, and communication-resource budgets is still lacking.
Accordingly, this paper develops a privacy-enhanced UAV semantic communication framework in which dedicated perturbation mechanisms are inserted into both the visual semantic stream and the location-descriptor stream before transmission. The main contributions are summarized as follows:
Dual-Privacy UAV Semantic Communication Framework: We establish a unified system model that couples dual-privacy modeling—visual semantic leakage and location-descriptor leakage—with communication-resource constraints in a single UAV air–ground link. Unlike prior work that either protects only visual features [
11,
13] or handles location privacy in isolation [
17,
18], the proposed framework explicitly coordinates both protection streams and their interaction with transmit power, establishing a principled privacy–utility operating point for the legitimate receiver while suppressing attacker-side recoverability.
Differential-Privacy-based Visual and Location Protection Mechanisms: For visual data, we design a region-aware DP mechanism that constructs a class-conditioned, cell-wise privacy-budget map from the global budget , applying stronger Gaussian noise to sensitive semantic regions (e.g., pedestrians, vehicles) while preserving utility in non-critical areas. The semantic grounding of the perturbation map—derived from task-level object detections and the feature-grid coordinate transform—directly links the protection intensity to the task-oriented nature of the semantic representation. For location data, we propose a scenario-adaptive strategy that selects between randomized DP on a discrete grid and planar Laplace geo-indistinguishability according to the spatial granularity required by the mission, both with formal DP guarantees.
Joint Optimization of Privacy Budgets and Transmit Power: We formulate a scenario-dependent joint utility maximization over the visual privacy budget , location privacy budget , and transmit power , capturing the coupled influence of privacy noise, channel quality, and resource constraints on both the legitimate receiver’s task performance and the attacker’s recoverability. A BCD-based algorithm is developed to solve this non-convex problem, converging within 6–8 iterations in all tested scenarios.
Experimental Validation: Simulation results on the VisDrone dataset demonstrate stable convergence, differentiated scenario-adaptive behavior, and a superior privacy–utility trade-off relative to uniform DP. In the fixed-budget surveillance setting, region-aware DP improves semantic class accuracy by 7.1%, reduces sensitive-region SSIM by 60.3%, and lowers sensitive-feature similarity by 18.2% compared with uniform DP, while maintaining comparable whole-image attack suppression.
The rest of this paper is organized as follows.
Section 2 presents the UAV semantic communication system model, the receiver/eavesdropper observation boundaries, and the privacy-enhanced transmission model.
Section 3 describes the differential-privacy-based visual and location protection mechanisms together with their theoretical guarantees.
Section 4 formulates the joint optimization problem of transmit power and privacy budgets and presents the proposed algorithm.
Section 5 gives the simulation results and analysis.
Section 6 concludes the paper.
Table 1 summarises the principal symbols used in this paper.
2. System Model
2.1. UAV Semantic Communications Under Third-Party Eavesdropping
Following the encoder–channel–decoder narrative widely adopted in semantic image transmission studies, such as [
6], we consider a UAV-assisted air–ground semantic communication system composed of a UAV transmitter, a legitimate ground station, and a third-party passive eavesdropper.
Figure 1 illustrates the overall scene, including the UAV sensing side, the legitimate receiver, the passive eavesdropper, and the coupled pressures of privacy leakage and resource constraints.
The UAV semantic communication system with a third-party eavesdropper comprises four stages: transmitter-side semantic encoding, air-to-ground transmission, legitimate receiver decoding, and eavesdropper observation. Specifically, the UAV captures scene images and converts them into task-oriented semantic messages before transmission. The legitimate ground station receives the protected semantic messages and reconstructs a task-oriented output. The passive eavesdropper can intercept the same transmitted variables, but does not belong to the authorized processing chain.
2.2. Transmitter-Side Semantic Encoding Model
Let
denote the input UAV image. The transmitter-side semantic encoder and detector are modelled by
where
is the visual semantic feature tensor and
is the detected object set.
For each detected object
, the transmitter records its bounding box
and category label
, where
is the normalized object center and
are the normalized box width and height. The target-related descriptor organizer
is
where
denotes the normalized center coordinate of object
j.
The tensor is the main carrier of visual task semantics, whereas preserves coarse object layout and category metadata. The latter is not introduced as an independent task head; rather, it is a structured semantic description accompanying the visual tensor.
2.3. Air–Ground Transmission Model for Visual and Location Messages
At the transmission stage, the UAV places two coupled semantic messages on the air–ground link: a visual semantic tensor and a compact target-related descriptor set. To keep the channel model independent of the concrete protection realization, these transmitted messages are denoted by and , respectively.
To capture the multipath propagation characteristics of realistic UAV air–ground links, the wireless channel is modelled as a Rayleigh block-fading channel with LoS-dominated large-scale path loss. The effective channel power gain comprises a distance-dependent path-loss term
and a Rayleigh-distributed fading coefficient
with
(unit mean), so that
. The receiver-side thermal noise is represented by additive white Gaussian noise with variance
. Under this model, the transmitted visual tensor and descriptor set are jointly written as
where
and
denote the received visual and descriptor messages after air–ground transmission.
Within one transmission block,
h is assumed to remain constant across blocks. It varies with the UAV-ground geometry and propagation environment. The received SNR at the legitimate receiver is therefore
The same block-level channel state is used for both message components at the system-model level, so the coupled influence of
,
h, and
on semantic recovery is summarized through
for the later analysis.
2.4. Legitimate Receiver Semantic Decoding Model
At the legitimate ground station, the received visual semantic tensor is decoded together with detector-derived prior cues reconstructed from the received descriptor set . Let denote the reconstructed prior cues. These cues encode category-presence information and coarse spatial supports derived from the received object descriptors.
The receiver-side semantic recovery process is written as
where
denotes the legitimate-side semantic decoder and
is the recovered semantic mask or task-oriented output.
At the receiver, the visual tensor enters the semantic decoder as the main information carrier. The received descriptor set is reorganized into rather than passed through a separate location decoder, and then serves as a refinement prior indicating which object categories are present and where coarse object regions are likely to appear.
2.5. Threat Model and Observation at the Eavesdropper
In this paper, we consider two coupled privacy risks. The first is visual privacy leakage, in which the intercepted semantic feature tensor still reveals sensitive scene content. The second is location privacy leakage, in which intercepted object descriptors reveal target positions or coarse mission geometry. If either side remains weakly protected, the attacker can fuse the two information sources and improve its recovery ability.
The passive eavesdropper can intercept the same transmitted variables as the legitimate receiver, but it cannot access raw onboard images, modify the transmission protocol, or obtain authorized side information [
13,
15]. To conservatively model the attacker, we assume that it knows the overall system architecture (a white-box/grey-box assumption) and may employ representation inversion, semantic inference, or joint descriptor-assisted recovery. This paper employs two attacker formulations that probe different capability levels. The gradient-based optimization attacker iteratively minimizes the reconstruction distance between a candidate image and the intercepted protected feature at inference time, without any prior training on the data distribution [
13], thereby providing a data-free lower bound on attacker capability. The trained inversion network pre-trains a lightweight convolutional inverter
on in-distribution training images and applies it at inference time to reconstruct the original visual content from the intercepted feature map, without relying on iterative gradient optimisation [
14]; as a strictly stronger adversary exploiting in-distribution training data, it validates that the protection conclusions hold under a more capable, realistic threat. Both attackers share the same knowledge boundary: they know the encoder architecture and feature dimensionality but cannot access raw UAV images, DP noise realizations, or authorized side information. The privacy model is defined at the semantic-representation level rather than through physical-layer secrecy assumptions.
At the representation level, the attacker-side visual recovery attempt is abstracted as
where
denotes a generic inference or attack operator. In addition to visual inversion, the eavesdropper may directly exploit
or its derived spatial statistics to infer sensitive target positions, activity trajectories, or mission areas.
5. Simulation Results and Analysis
5.1. Dataset and Parameter Settings Across Different Scenarios
The experiments use the VisDrone dataset and are conducted on the official validation split [
27], which is the standard protocol used in the VisDrone detection and tracking benchmark. To ensure a consistent statistical basis across all reported comparisons, each scenario-level or privacy–utility summary is aggregated over 500 validation samples drawn from this split. For the attacker-side reconstruction evaluation, the results are reported as mean ± 95% confidence interval computed over 100 attack cases per DP mode (20 images × 5 seeds), providing statistically reliable evidence of the privacy protection effectiveness across diverse image content.
Three scenario policies are studied. The surveillance scenario uses and W, emphasizes balanced sensing utility under a high-risk environment, and adopts the DP grid-based location mechanism. The precision-oriented scenario uses a tighter privacy budget together with a larger power limit W and a planar Laplace branch, which reflects tasks requiring higher semantic fidelity and link quality. The resource-limited scenario sets and W, and adopts the DP grid-based location mechanism to emphasize conservative resource usage.
5.2. Baseline Schemes Design
Three complementary experimental protocols are considered. The first protocol evaluates the proposed region-aware pipeline under the surveillance, precision-oriented, and resource-limited policies, and is used for the cross-scenario comparison. The second protocol fixes the location privacy budget at and sweeps the visual privacy budget over to compare uniform differential privacy with region-aware differential privacy. The third protocol fixes and in the surveillance scenario and compares no protection, uniform differential privacy, and region-aware differential privacy.
The adopted baselines serve different purposes. No protection provides an upper-bound reference for legitimate semantic utility while exposing the privacy risk of transmitting unprotected features. Uniform differential privacy serves as a non-selective privacy baseline and tests whether privacy can be improved without exploiting semantic-region sensitivity. Region-aware differential privacy is then compared against these references to isolate the benefit of spatially selective perturbation. All compared methods share the same feature extractor, the same semantic decoder, the same channel model, the same task setting, and the same attack-side evaluation protocol. Therefore, the reported differences can be attributed to the protection mechanism or the scenario policy rather than to inconsistent budgets, links, or datasets.
5.3. Evaluation Metrics
Overall utility measures the final system-level privacy–utility trade-off achieved by the joint optimization. Semantic category accuracy quantifies semantic usability at the legitimate receiver after privacy protection and air–ground transmission. Sensitive-region SSIM and sensitive-feature cosine similarity quantify attacker-side visual recoverability at the structural and feature levels, respectively. Location error measures the cost introduced by the location-protection mechanism, and received SNR characterizes the communication condition supporting semantic decoding. Attack-side PSNR and SSIM are also reported to quantify reconstruction suppression under different visual privacy mechanisms [
13].
5.4. Hyperparameter Selection
The key hyperparameters are set as follows. The grid size for the discrete location DP branch is chosen to provide a spatial resolution of approximately of the normalized coordinate range, which balances grid granularity with the sensitivity of the exponential-mechanism normalization constant. Finer grids (larger N) provide higher utility at the cost of a smaller effective privacy amplification from randomization. The coordinate scaling factor that maps continuous normalized coordinates to the grid is absorbed into the effective privacy parameter and does not change the mechanism class. For the visual DP branch, the context-ring scale and background scale are set so that the context ring receives approximately half the noise of the most sensitive object cells (), and the background receives slightly weaker noise than the global budget ( yields , corresponding to of the global-budget noise standard deviation); these values were selected to maintain task utility while suppressing contextual leakage around sensitive regions and ensuring that background noise does not fall to zero. The clipping norm is set to bound the sensitivity of each feature vector, and follows the standard choice for Gaussian mechanism analysis in machine-learning settings.
5.5. Implementation Details and Statistical Protocol
All experiments are conducted in a unified Python 3.10/PyTorch 2.1.0 simulation environment and can run on either CPU or CUDA devices. The same semantic feature extraction, privacy protection, air–ground transmission, and semantic decoding pipeline is used across all compared methods, and the compared settings differ only in the privacy mechanism or the scenario policy. This task configuration is consistent with recent visual semantic communication studies that evaluate task-level semantic output quality under shared transmission and task heads rather than pixel-level fidelity [
4,
6,
10]. This unified configuration ensures fair comparison across baselines and across scenario policies.
For optimization, the BCD-based solver uses a stopping tolerance of , a maximum of 50 iterations, a minimum of 6 iterations before the stopping criterion is activated, and a damping factor of . For visual privacy, three protection modes are considered, namely no protection, uniform differential privacy, and region-aware differential privacy, and the Gaussian mechanism uses . On the location side, the protection branch is selected by the scenario policy and takes the form of either the DP grid-based location mechanism or the planar Laplace branch. The scenario-dependent privacy budgets and transmit-power budgets follow the scenario definitions stated above.
The fixed-budget evaluation reported in
Section 5.8 uses
and a gradient-based reconstruction reference. At
, the region-aware DP applies noise with
to the most sensitive cells (pedestrians), under which the trained inversion network produces near-random reconstructions across all DP modes and cannot discriminate between mechanism designs; the gradient-based attacker is therefore used in that evaluation as a controlled ablation indicator. Section Adversarial Evaluation: Gradient-Based and Trained Inversion Attackers uses the trained inversion network at
, a regime in which the attacker achieves partial reconstruction under no-protection and can meaningfully distinguish the selective protection behavior of region-aware DP from uniform DP.
5.6. Privacy–Utility Trade-Off Under Different Visual Privacy Budgets
Figure 3 shows the privacy–utility frontier under a fixed location privacy budget
and a visual privacy sweep over
. Both uniform DP and region-aware DP improve semantic category accuracy as the visual privacy budget becomes looser. However, the region-aware mechanism remains consistently better positioned on the frontier.
The numerical results further clarify this gap. At , region-aware DP achieves a semantic category accuracy of 0.802, compared with 0.748 under uniform DP, corresponding to a 7.2% relative improvement, while the sensitive-region SSIM is reduced from 0.038 to 0.015, corresponding to a 60.5% relative reduction. At , region-aware DP still preserves a higher semantic category accuracy of 0.819 versus 0.764, corresponding to a 7.2% relative improvement, and meanwhile keeps the sensitive-region SSIM much lower at 0.018 versus 0.047, corresponding to a 61.7% relative reduction. The same trend is observed in sensitive-feature similarity, indicating that the proposed mechanism preserves more task-relevant semantics while suppressing attacker-side recovery more effectively under the same budget sweep.
5.7. Cross-Scenario Performance Under Heterogeneous Mission Priorities
Figure 4 compares cross-scenario behavior under heterogeneous mission priorities. The same region-aware pipeline is evaluated under the surveillance, precision-oriented, and resource-limited policies. The surveillance scenario reaches the highest overall utility, with a mean value of 0.6733, while the precision-oriented and resource-limited scenarios obtain 0.3418 and 0.2365, respectively. This result shows that the proposed framework converges to different operating points according to scenario requirements rather than applying a fixed privacy-budget or transmit-power setting to all tasks.
The scenario-level means further explain these differences. The precision-oriented setting allocates a larger average transmit power of 0.2596 W and achieves the highest mean received SNR of 8.216 dB, which is consistent with its emphasis on semantic fidelity. The resource-limited setting reduces the average transmit power to 0.0094 W and correspondingly yields the lowest mean SNR of dB, reflecting strict resource control. Meanwhile, the surveillance scenario keeps the lowest mean location error of 0.0248, compared with 0.0312 in the precision-oriented case and 0.0584 in the resource-limited case. These results confirm that the proposed joint optimization framework coordinates privacy budgets, communication quality, and location distortion according to mission priority rather than optimizing a single metric in isolation.
5.8. Fixed-Budget DP Ablation and Mechanism Interpretation
Figure 5,
Figure 6 and
Figure 7 report the fixed-budget comparison under the surveillance scenario with
and
, where no protection, uniform DP, and region-aware DP are evaluated under the same setting. Because the privacy budgets and scenario policy are fixed while the visual protection strategy changes, this comparison directly reveals the contribution of the visual privacy mechanism.
The numerical results show a clear three-way trade-off. Without protection, the semantic category accuracy is 0.843, but the sensitive-region SSIM reaches 0.201 and the sensitive-feature similarity remains 1.0, indicating almost complete recoverability on the attacker side. Uniform DP suppresses the sensitive-region SSIM to 0.0383 and reduces the sensitive-feature similarity to 0.22, but it also lowers the semantic category accuracy to 0.748. Region-aware DP improves the semantic category accuracy to 0.802 while further reducing the sensitive-region SSIM to 0.0152 and the sensitive-feature similarity to 0.18. Relative to uniform DP, this means a 7.2% improvement in semantic category accuracy, a 60.3% reduction in sensitive-region SSIM, and an 18.2% reduction in sensitive-feature similarity. At the whole-image attack level, the attack SSIM under region-aware DP is 0.01687, which is comparable to the 0.01708 of uniform DP, but is achieved with noticeably better legitimate-task performance.
These comparisons show that region-aware perturbation is the key visual-protection mechanism of the proposed framework. Relative to uniform DP, it preserves more legitimate-task semantics while further suppressing sensitive-region leakage, with simultaneous gains in task accuracy and reductions in both structural and feature-level recoverability. Relative to no protection, it sharply reduces attacker-side recoverability while maintaining a substantially stronger task-side response than uniform perturbation under the same fixed-budget setting.
Adversarial Evaluation: Gradient-Based and Trained Inversion Attackers
This section evaluates attacker-side recoverability under two complementary threat models introduced in
Section 2.5. The gradient-based optimization attacker serves as a lower-bound baseline: it iteratively minimizes the L1 distance between the feature map of a candidate image and the intercepted protected feature, starting from a random initialization, without any training data (300 optimization steps per image). The trained inversion network is a strictly stronger adversary that pre-trains a lightweight convolutional inverter
on 450 training images drawn from the same distribution, then applies it at inference time to reconstruct the original image from intercepted semantic features (no channel noise, direct feature interception). Two metrics are reported: Attack SSIM ↓, the full-image structural similarity of the attacker’s reconstruction relative to the original; and Sensitive-region PSNR ↓, the peak signal-to-noise ratio computed exclusively on pixels inside ground-truth sensitive bounding boxes. All results cover 100 independent runs per condition (20 test images × 5 random seeds) and are reported as mean ± 95% CI via the
t-distribution (
Table 3).
Table 3 presents results for both attacker formulations. The gradient-based optimizer achieves Attack SSIM
under no protection, confirming that even training-free iterative optimization can extract limited structural information from unprotected features. The trained inversion network, with access to 450 in-distribution training images, achieves Attack SSIM
under the same condition—nearly
higher—establishing it as the dominant reconstruction threat. Both DP mechanisms substantially degrade both attackers. For the trained inverter at
, Uniform DP reduces Attack SSIM by 74% (
; 95% CIs non-overlapping); for the gradient attacker, region-aware DP achieves a 77% reduction (
). At the sensitive-region level, the region-aware mechanism provides additional selective protection: at
, sensitive-region PSNR reaches
dB (trained inverter) and
dB (gradient attacker), both substantially lower than Uniform DP (
dB and
dB respectively; 95% CIs non-overlapping). At the full-image level, region-aware DP yields a higher Attack SSIM than Uniform DP for the trained inverter at both
values (
vs.
at
;
vs.
at
). This reflects the inherent selectivity trade-off: by concentrating noise on sensitive regions while applying weaker perturbation to background areas (
), region-aware DP intentionally allows the attacker to recover background structure more accurately in exchange for substantially stronger protection of privacy-sensitive content. The protection advantage of region-aware DP is therefore selective rather than global, targeting the semantic content that carries the greatest privacy risk.
5.9. Convergence Behavior and Optimization Overhead
Figure 8 reports the BCD-based joint optimization dynamics. Because the optimizer uses a minimum of 6 iterations, a convergence tolerance of
, and an upper bound of 50 iterations, the recorded histories can be directly interpreted as optimization-overhead evidence. In the precision-oriented scenario, the mean utility increases from 0.2649 at iteration 1 to 0.2817 at iteration 6. In the resource-limited scenario, the mean utility rises from 0.2839 to 0.2986 over the same interval. In the surveillance scenario, the mean utility grows from 0.6325 at iteration 1 to 0.6733 at iteration 6.
Most cases already stabilize within six to seven iterations. Only a very small subset of surveillance samples continues to an eighth iteration, where the mean utility reaches 0.6801 for the remaining four hard cases. These observations indicate that the proposed joint optimization reaches stable operating points with a short iteration horizon and moderate iteration-level overhead.
5.10. Qualitative Semantic Output Visualization, Privacy Visualization, and Limitations
Figure 9 compares the original image with a task-level semantic output visualization rendered from the decoder mask at the legitimate receiver. The right panel is not a pixel-level reconstructed image; it is a visualization generated from the decoder’s semantic output to highlight object categories and coarse spatial layout. It shows that the protected features preserve the relative positions and semantic distribution of major objects such as vehicles, pedestrians, bicycles, and coarse scene structure after air–ground transmission. Because this visualization is derived from task-level semantic masks rather than photorealistic rendering, fine-grained textures, colors, and detailed visual appearances are intentionally absent. This observation is consistent with the design goal of semantic communication, namely preserving task-oriented semantics rather than directly recognizable visual details.
Figure 10 provides qualitative comparisons of multiple scenes under no protection, uniform differential privacy, and region-aware differential privacy. Under no protection, vehicle contours, parking layouts, and road textures remain relatively clear, indicating that the attacker may still recover sensitive visual content with identifying value. Uniform differential privacy significantly perturbs the whole image and weakens the visibility of sensitive content, but it also destroys a considerable amount of non-sensitive background structure. By contrast, region-aware differential privacy introduces stronger perturbations over sensitive target regions such as vehicles while preserving more global scene outlines such as roads, vegetation, and the overall layout. This qualitative evidence visually supports the quantitative observations in
Figure 5,
Figure 6 and
Figure 7.
5.11. Computational Complexity
Table 4 reports the measured runtime and peak memory overhead of the three core modules.
The DP protection modules add negligible per-frame overhead relative to the semantic encoding and decoding pipeline. The BCD solver converges within 6–8 iterations in all reported scenarios (
Figure 8), with each iteration performing a one-dimensional search over a compact interval for each sub-problem.
5.12. Simulation-to-Deployment Gap
The experiments are conducted in a controlled simulation environment; we explicitly characterise the key gaps relative to practical UAV deployment. (i) Detection accuracy: the simulation assumes a YOLO detector with performance matched to the VisDrone benchmark; in practice, detection accuracy may vary with altitude, lighting, and motion blur, which would affect the sensitive-region mask used by the visual DP mechanism. (ii) Channel realism: the Rayleigh block-fading model captures small-scale multipath fading but does not account for interference, blockage, or Doppler spread under high UAV mobility; extensions to time-varying and interference-limited channel models remain future directions. (iii) Onboard computation: the complexity measurements in
Section 5.11 indicate that the DP and BCD modules add negligible latency, but end-to-end deployment on constrained UAV hardware requires hardware-level profiling beyond the simulation environment. Practical UAV deployment and hardware validation are listed as priority future directions.
Taken together, the quantitative and qualitative results indicate that the proposed framework achieves a favorable privacy–utility trade-off and differentiated operating points across task scenarios. Across trade-off evaluation, cross-scenario adaptation, fixed-budget mechanism comparison, convergence behavior, and representative visual cases, the method consistently preserves more task-relevant semantics while reducing attacker-side recoverability. Robustness analyses over channel quality, DP noise scaling, scene density, and attack preprocessing conditions would further broaden the evaluation scope.