In this section, we present the datasets, evaluation metrics, and implementation details used in our experiments, followed by comprehensive quantitative and qualitative comparisons. In addition, we conduct ablation studies to verify the effectiveness of the proposed components.
4.1. Datasets and Evaluation Metrics
Datasets. In UAV image restoration experiments, we adopt the UAVIR-5D [
27] dataset to verify the performance of our method. UAVIR-5D is a multi-degradation benchmark tailored for UAV viewpoints, containing 4520 paired training images and 500 paired test images, with five typical degradation types including blur, noise, haze, low-light, and raindrops. UAVIR-5D covers various common UAV aerial scenarios such as urban roads, building complexes, and campus environments, providing a unified evaluation benchmark for multi-degradation UAV image restoration.
To further evaluate the generalization capability of our model, we also conduct experiments on the MDRS-Landsat [
52] remote sensing dataset proposed in Ada4DIR. MDRS-Landsat is constructed using 5500 high-quality Landsat-8 images from the RSHaze dataset, and synthetic degradations including blur, noise, haze, and low-light are systematically applied. This dataset contains 5130 training samples, 100 validation samples, and 270 test samples, covering diverse large-scale land surfaces and man-made structures, providing a more realistic benchmark for cross-domain generalization in multi-degradation restoration.
Evaluation Metrics. For evaluation, consistent with most existing image restoration works, we report Peak Signal-to-Noise Ratio (PSNR) [
53] and Structural Similarity Index Measure (SSIM) [
54] for each degradation task as well as overall performance. PSNR measures pixel-wise reconstruction fidelity between restored and reference images, where a higher value indicates lower reconstruction error. SSIM assesses perceptual similarity in terms of luminance, contrast, and structure, focusing more on structural information and human visual consistency.
4.2. Implementation Details
For implementation, our model is developed based on the PyTorch 2.0.1 framework and trained using an environment configured with Anaconda3, and CUDA 11.8. All experiments are carried out on a single NVIDIA RTX 4090D GPU. To ensure reproducibility, we specify the key architectural parameters: the initial channel dimension
C is set to 48, and the encoder follows a four-stage hierarchical design with channel dimensions of
. The prompt pool size is set to
, with components initialized using a standard Gaussian distribution with a standard deviation of
. In the AFPB, we employ a fixed frequency separation strategy, where the low-frequency component is defined as the central rectangular region occupying 50% of the spectrum area, while the remaining area is treated as high-frequency. Additionally, the number of Prompt Components is set to 5. We adopt the Adam [
55] optimizer with
and
, setting an initial learning rate of
and employ a cosine annealing strategy to gradually decay the learning rate to
for more stable convergence. The batch size is set to 4, and input images are randomly cropped into
patches during training. The total number of training iterations is fixed to 300 K. To improve the robustness and generalization of our model, standard data augmentation strategies such as random horizontal and vertical flips are applied, while strictly maintaining alignment between degraded images and their corresponding ground truth. During inference, the best checkpoint is selected to generate the final results.
4.3. Comparisons with the State of the Art
In this section, we evaluate the overall performance of different models across various restoration tasks on the UAVIR-5D dataset, including deblurring, denoising, dehazing, de-raindropping, and low-light enhancement. Both quantitative and qualitative results are provided to comprehensively assess the effectiveness of each method. Furthermore, to verify the generalization capability of our approach, we additionally conduct experiments on the publicly available MDRS-Landsat dataset for multi-degradation remote sensing image restoration.
Evaluations on UAVIR-5D datasets for all-in-one UAV image restoration. To validate the effectiveness of the proposed method in handling multiple UAV degradation tasks, we conduct comprehensive comparisons on the UAVIR-5D dataset with a series of representative all-in-one image restoration approaches, including AirNet [
24], TransWeather [
50], DGUNet [
47], IDR [
11], IR-SDE [
46], PromptIR [
28], MambaIR [
35], AutoDIR [
26], DiffUIR [
25], InstructIR [
32], PG-WSSM [
27], and FoundIR [
56].
As shown in
Table 2, the proposed method achieves superior performance on most degradation types, except for a few cases where the leading results are not obtained. In terms of overall restoration quality, our approach reaches an average PSNR of 31.64 dB across the five degradation tasks, outperforming the second-best method PG-WSSM by 0.43 dB, and thus achieving the best overall results. The notable improvement clearly demonstrates the effectiveness of our frequency prompts mechanism in mitigating degradation interference and enhancing robust multi-degradation restoration performance.
In terms of individual restoration tasks, our method achieves different levels of improvement under various degradations. For the deblurring task, the performance gain over the second-best approach is relatively small, which may be attributed to the fact that motion blur typically causes more severe structural information loss, thereby increasing the difficulty of accurate recovery. Regarding the denoising task, we observe a slight performance gap compared with the best-performing method PromptIR. This trade-off is deeply rooted in the distinct mechanisms of frequency-domain versus spatial-domain prompting. As illustrated in the frequency spectrum visualization in
Figure 1, Gaussian noise manifests as a uniform distribution across high-frequency bands, lacking the specific directional or structural patterns observed in motion blur or rain streaks. Existing spatial-based methods like PromptIR excel at pixel-wise correction, enabling them to effectively suppress such stochastic, non-structural noise through local feature modulation.
In contrast, our proposed frequency-domain prompts are designed to capture and disentangle global degradation patterns based on their spectral signatures. Since random noise presents as a global disturbance without a compact spectral representation, our frequency-aware filters are slightly less precise in distinguishing noise from high-frequency image textures compared to localized spatial attention. However, this marginal trade-off in pure denoising allows our model to gain significant robustness in handling complex, spatially coupled degradations. For instance, in the dehazing task, our approach surpasses the second-best method by 0.69 dB, benefiting from the proposed frequency prompt that compensates for low-frequency attenuation and effectively improves the overall contrast and visibility of the restored images. For the de-raindrop task and low-light enhancement, our method achieves remarkable improvements of 0.81 dB and 0.53 dB, respectively. This demonstrates that while slightly less sensitive to random noise, the frequency prompts provide superior guidance in distinguishing occlusion artifacts (rain) from background textures and ensuring robust luminance recovery.
To further assess the restoration capability of the proposed method in UAV imaging scenarios, we provide visual comparisons across various degradation tasks.
Figure 3,
Figure 4,
Figure 5 and
Figure 6 illustrate the results of our approach and several representative methods under different degradation conditions. Overall, compared with existing advanced approaches, our method preserves better global visual consistency, enhances contrast and dynamic range, and substantially improves texture reconstruction quality, achieving superior structural fidelity and color restoration.
Figure 3 presents the deblurring results. As observed, our method effectively suppresses streak-like motion artifacts and significantly improves object boundaries and overall clarity. In contrast, TransWeather and AutoDIR still retain noticeable blur, while MambaIR and DiffUIR generate overly smoothed results, leading to the loss of fine details.
Figure 4 shows dehazing performance under challenging non-uniform haze conditions, where strong local scattering causes spatially varying degradation. Our method better captures low-frequency haze components and reduces scattering halos, restoring natural color appearances for buildings and vegetation. In comparison, PromptIR and InstructIR fail to fully remove haze and produce residual artifacts, while AutoDIR and DiffUIR partially suppress haze but suffer from luminance imbalance and color shifts.
Figure 5 illustrates the de-raindrop results. AutoDIR and InstructIR still exhibit prominent residual rain droplets or blur artifacts, whereas DGUNet and PromptIR mistakenly treat large rain streaks as background textures, leaving streak remnants. By contrast, our method effectively removes raindrop occlusions and streak artifacts without compromising background details.
Figure 6 presents the low-light enhancement results. Our approach achieves a more natural illumination enhancement by simultaneously improving brightness, suppressing noise, and maintaining chromatic consistency. DGUNet, although increasing brightness, fails to recover fine-grained textures, while IR-SDE suffers from insufficient contrast enhancement or abnormal color saturation, degrading the overall perceptual quality.
Evaluations on MDRS-Landsat for Remote Sensing Image Restoration. To further validate the cross-scene generalization capability of the proposed method, we conduct experiments on the MDRS-Landsat dataset and compare our approach with several representative image restoration methods, including NAFNet [
48], Restormer [
45], DGUNet [
47], TransWeather [
50], AirNet [
24], PromptIR [
28], IDR [
11], SrResNet-AP [
57], Restormer-AP [
57], Uformer-AP [
57], and Ada4DIR-s [
52]. All compared methods follow the official train/test split and unified experimental configuration of MDRS-Landsat. Specifically, we randomly select 2048 samples from the dataset as the training set, with 512 images for each degradation type, while the entire validation and test sets are used for evaluation under their respective tasks.
As shown in
Table 3, our method achieves an average PSNR of 38.39 dB over four degradation types, outperforming the second-best approach, Ada4DIR-s, by 0.92 dB. This significant improvement verifies the strong effectiveness and generalization capability of combining frequency prompts with state space modeling for unified restoration in multi-degradation remote sensing scenarios.
Regarding specific restoration tasks, our method delivers consistent and notable performance gains across most degradation types. For the deblurring task, our approach surpasses IDR by 1.89 dB, demonstrating its superior ability to recover structural information. For the dehazing task, we obtain an improvement of 0.07 dB, indicating robustness in compensating low-frequency contrast attenuation. In the denoising task, our method achieves 36.85 dB, outperforming PromptIR by 1.86 dB, highlighting a better balance between noise suppression and detail preservation. For the low-light enhancement task, although our model achieves the second-best PSNR, it still shows clear advantages over most competing methods. Overall, even when trained with only 2048 samples, our method maintains excellent performance across all degradation categories in the remote sensing domain, demonstrating strong generalization ability under limited data conditions.
Evaluation of Model Complexity. To further assess the computational efficiency of our model, we compare its complexity with recent state-of-the-art approaches in terms of floating-point operations (FLOPs) and the number of learnable parameters. As reported in
Table 4, our method achieves an effective balance between computational cost and parameter capacity, demonstrating competitive efficiency while maintaining strong restoration performance.
4.4. Ablation Study
In this section, we validate the effectiveness of the key components integrated into our network through comprehensive ablation experiments conducted on the UAVIR-5D dataset. We report the average PSNR results in
Table 3. To ensure fairness and reliability, all experiments follow the same training settings described in
Section 4.2.
Effectiveness of Key Components. We first analyze the contribution of the two major components proposed in our architecture, namely the PGMB and the AFPB. In the experiments, SS2D denotes the Mamba block equipped with a 2D selective scanning strategy, and APB refers to an adaptive prompt block operating purely in the spatial domain. As shown in
Table 5, each component contributes noticeably to the overall performance, indicating that both modules play indispensable roles in multi-degradation restoration. Moreover, the visual comparisons in
Figure 7 further demonstrate the qualitative benefits introduced by each component.
Effectiveness from Structural and Parametric Perspectives. Following standard ablation study protocols, we extended our analysis to investigate the optimal configuration of the prompt modules from both structural and parametric perspectives. First, regarding the structural placement,
Table 6 demonstrates that integrating the AFPB into the skip connections is crucial for performance enhancement. This highlights that a rational structural design is as critical as the presence of the module itself. Second, to provide a detailed parameter comparison, we investigated the impact of the number of Prompt Components (N) on restoration performance. As reported in
Table 7, the results exhibit a clear trend: the performance peaks at N = 5 with a PSNR of 31.64 dB. Specifically, at N = 3, the restricted representational capacity limits the restoration quality. Conversely, increasing N beyond 5 to 7 and 9 results in performance saturation or slight degradation rather than improvement. This non-monotonic trend implies that excessive components may introduce semantic redundancy or feature interference. Consequently, N = 5 is adopted as the optimal balance between representation capability and model compactness.
Beyond these quantitative analyses, we further examine the internal working mechanism of the proposed modules. The AFPB functions effectively because different degradation types exhibit distinct spectral signatures—for instance, haze concentrates in low-frequency bands, whereas noise dominates high-frequency regions. By explicitly separating these components in the Fourier domain, the AFPB enables the model to apply targeted suppression or enhancement strategies, thereby effectively addressing the challenge of decoupling distinct degradation components, which remains difficult for spatial-domain methods (such as standard convolution). This rationale explains why our method outperforms existing approaches in complex coupled tasks such as dehazing and raindrop removal, as verified by the superior PSNR results.
4.5. Limitations and Discussion
Although the proposed Degradation-Aware Frequency Prompt State Space Model achieves state-of-the-art performance across multiple tasks, we have identified several limitations and potential failure cases that merit further discussion.
First, regarding extreme coupled degradation scenarios, the model’s reliance on implicit frequency-domain encoding faces challenges when multiple severe degradations co-occur. As raised in our experimental analysis, restoration performance depends heavily on the accuracy of self-generated frequency prompts. Under extremely challenging conditions—such as the simultaneous presence of dense haze and low illumination—the signal-to-noise ratio drops critically. In these cases, the spectral characteristics of haze and low-light become highly entangled in the low-frequency band. The implicit degradation encoding may fail to correctly disentangle these dominant degradation patterns from the inherent image structure, potentially leading to failure cases characterized by suboptimal brightness recovery or residual color casts.
Second, there is an inherent performance trade-off regarding denoising. As indicated in
Table 1, our method performs slightly worse (−0.51 dB) than the spatial-domain method PromptIR in pure denoising tasks. This is a consequence of our architectural design choice: our frequency prompts prioritize capturing global degradation patterns (like streaks or haze distribution) to maximize all-in-one generalization. However, Gaussian noise manifests as a uniform distribution across high frequencies without specific structural patterns. Consequently, our frequency-aware filters are less sensitive to such stochastic pixel-level noise compared to spatially localized modulation methods. We acknowledge this as a necessary trade-off to achieve superior robustness in handling complex, spatially variant degradations.
Third, the domain gap between synthetic and real-world data remains a challenge. Our model is trained on the UAVIR-5D dataset, which is constructed using physics-based synthesis. However, real-world UAV imagery often suffers from more complex, non-uniform artifacts that are difficult to simulate perfectly, such as the coupling of video compression blocking, sensor jitter caused by high-speed flight, and atmospheric turbulence. When applied to real-world wild data, the model may exhibit minor artifacts or reduced restoration quality due to these unseen degradation combinations, limiting its flexibility compared to multi-expert systems trained on diverse real data.
Finally, we must consider computational constraints for practical deployment. While the Mamba-based architecture offers linear complexity and is more efficient than standard Transformers, the proposed AFPB introduces Fast Fourier Transform (FFT) operations. FFT involves complex number calculations that can introduce additional inference latency, especially on edge computing devices with limited hardware acceleration for spectral operations. This could potentially act as a bottleneck for time-sensitive UAV navigation missions requiring high-FPS processing.
In future work, we plan to address these limitations by exploring controllable prompt mechanisms with user intervention to handle extreme cases, and by investigating lightweight spectral approximations to reduce the computational overhead on UAV edge devices.