Next Article in Journal
High-Resolution 3D Structural Documentation of the Saqqara Pyramids, Egypt, Using Terrestrial Laser Scanning and Integrated Geomatics Techniques for Heritage Preservation
Previous Article in Journal
Where the Hills Slide Slowly: A LiDAR-Based Morphometric Framework for Landslide Instability Regimes in Soft-Rock Terrains
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

SDTformer: Scale-Adaptive Differential Transformer Network for Remote Sensing Image Dehazing

1
Department of Mathematics, Imperial College of London, London SW7 2AZ, UK
2
School of Remote Sensing and Information Engineering, Wuhan University, Wuhan 430079, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(8), 1136; https://doi.org/10.3390/rs18081136
Submission received: 12 March 2026 / Revised: 8 April 2026 / Accepted: 9 April 2026 / Published: 11 April 2026

Highlights

What are the main findings?
  • A scale-adaptive differential Transformer is proposed to suppress attention noise by modeling differential attention across multiple spatial scales for remote sensing image dehazing.
  • The proposed SDTformer improves reconstruction fidelity and achieves favorable performance compared with state-of-the-art methods on benchmark datasets.
What are the implications of the main findings?
  • The scale-adaptive differential attention mechanism enables more effective modeling of haze features in complex remote sensing scenes.
  • The proposed framework enhances feature representation and improves the reliability of Transformer-based remote sensing image dehazing.

Abstract

In Transformer-based image restoration models, the self-attention mechanism often introduces attention noise from irrelevant contextual feature, hindering the recovery of underlying clear content. Although many methods have been proposed to suppress attention noise, we note that most existing approaches are often developed for general vision tasks and fail to generalize across remote sensing image dehazing, where large-scale spatial structures pose additional challenges for attention modeling. How to effectively model scale-aware attention to suppress redundant activations becomes crucial for remote sensing image dehazing. In this paper, we propose a scale-adaptive differential Transformer (SDTformer), an architecture designed to suppress attention noise through a differential attention mechanism, thereby improving reconstruction fidelity. Specifically, the model incorporates a scale-adaptive differential self-attention module, which models contextual dependencies across different spatial scales and reduces redundant contextual interference by computing differential attention maps. Additionally, a dynamic differential feed-forward network is proposed to adaptively select informative spatial features, strengthening feature aggregation. To further enhance feature representation, a gated fusion module is introduced to aggregate multi-scale features generated by different encoder blocks, which facilitates the learning process of each decoder block and improves the final reconstruction performance. Extensive experimental results on the commonly used benchmarks show that our method achieves favorable performance against state-of-the-art approaches.

1. Introduction

In modern Earth observation [1], atmospheric degradation has become one of the main bottlenecks affecting the quality of remote sensing images, especially under complex weather conditions, such as fog, haze, and clouds. Suspended aerosols and fine particles absorb and scatter incident light, leading to reduced image contrast and deviation in color representation from true distributions, and then cause details and textures to be obscured or blurred. This degradation not only weakens the informational content of remote sensing images but also further propagates to subsequent applications, resulting in a significant decline in model performance and decision-making accuracy in critical tasks such as environmental monitoring [2], land cover classification [3], disaster assessment [4], and urban planning [5,6]. Therefore, it is crucial to develop an effective algorithm to address weather-induced distortions in remote sensing images.
The rapid advancement of deep learning has greatly promoted remote sensing image restoration. Early studies widely adopted convolutional neural networks (CNNs) [7,8,9,10,11] to learn end-to-end mappings from degraded images to clear images, achieving notable progress in image dehazing [12,13]. However, due to their inherently local receptive fields, CNNs struggle to capture long-range dependencies. In remote sensing scenarios, atmospheric degradation often exhibits global characteristics and cross-region consistency, which limits the effectiveness of CNN-based modeling. To address these limitations, researchers have introduced attention mechanisms [14,15,16,17] and Transformer architectures [18,19,20] into remote sensing image restoration. Benefiting from the self-attention paradigm, Transformers enable global context interaction across the feature space, allowing the model to better capture atmospheric degradation patterns while preserving fine-grained structural details. Although these approaches achieve better performance than most CNN-based methods, they typically rely on the standard softmax-based self-attention mechanism to model long-range dependencies, where the attention value denotes the weighted aggregation of features, while the non-negligible responses assigned to irrelevant regions introduce undesired attention noise during feature aggregation.
To this end, several approaches introduce Top-k sparse self-attention [19] into deep neural networks to suppress redundant responses by selecting only the most significant attention scores for feature interaction, as shown in Figure 1. While such strategies alleviate the influence of irrelevant regions to some extent, they rely on a hard selection mechanism that may discard potentially useful contextual information, thereby limiting the model’s ability to capture complex degradation patterns in remote sensing scenes. To alleviate such limitations, traditional differential self-attention [21] has been introduced to enhance the discriminative capability of attention by modeling the difference between attention responses, thereby suppressing irrelevant contextual interactions during feature aggregation. However, such mechanisms are mainly designed for general vision tasks and do not explicitly consider the characteristics of high-resolution remote sensing imagery. In particular, the large spatial scale and complex structural distribution in remote sensing scenes make it difficult for conventional differential attention to effectively distinguish haze-related regions from diverse background textures, limiting its effectiveness for remote sensing dehazing tasks. Thus, it is of great interest to design a scale-adaptive differential attention mechanism that can effectively suppress attention noise while capturing haze-relevant representations across multiple spatial scales for high-resolution remote sensing image dehazing.
To address the above challenges, we propose a scale-adaptive differential Transformer for remote sensing image dehazing. The proposed framework aims to suppress attention noise while enhancing the representation of haze-relevant features under complex high-resolution scenes. Specifically, we introduce a scale-adaptive differential self-attention (SDSA) mechanism that models contextual dependencies across different spatial scales and reduces redundant contextual interference by constructing differential attention maps. In addition, a multi-scale differential feed-forward network (MDFN) is designed to adaptively select informative spatial features, thereby strengthening feature aggregation and improving representation capability. Furthermore, a gated fusion module (GFM) is incorporated to effectively integrate multi-scale features produced by different encoder stages, which facilitates the learning of the decoder and enhances reconstruction fidelity. Extensive experiments on several widely used benchmarks demonstrate that the proposed method achieves favorable performance compared with existing state-of-the-art approaches.
The main contributions are summarized as follows:
  • We propose a scale-adaptive differential Transformer to generate high-quality remote sensing dehazing images and achieve more precise restoration of details and textures.
  • We develop a scale-adaptive differential self-attention that mitigates attention noise, thereby encouraging the model to focus more on critical regions and informative features.
  • We design a multi-scale differential feed-forward network, which enhances feature representation while effectively reducing interference from redundant information.
  • Extensive experimental results on various benchmarks demonstrate that our method achieves favorable performance against SOTA approaches.

2. Related Work

2.1. Remote Sensing Image Dehazing

Recovering accurate surface radiance from images degraded by atmospheric scattering remains a complex and mathematically indeterminate problem in remote sensing and computer vision [22,23]. To mitigate these effects, early methodologies relied heavily on empirical statistical assumptions, known as handcrafted priors. For example, the dark channel prior (DCP) [24] is based on the premise that local patches in haze-free images contain at least one color channel with very low intensity. Similarly, the color attenuation prior [25] estimates haze density by exploiting the linear correlation between scene depth and the saturation–brightness difference. By utilizing these priors to determine transmission maps and atmospheric light, the physical imaging model can be inverted to retrieve clear images [24,25,26]. Nevertheless, while effective in controlled settings, these prior-based techniques often struggle with the heterogeneous and complex conditions found in real-world remote sensing data.
The advent of deep learning [27,28,29,30] has introduced powerful new paradigms for image restoration, allowing models to learn non-linear mappings from hazy to clear inputs automatically. This has led to a proliferation of convolutional neural network (CNN) architectures [11,12,31,32,33,34]. As a seminal work, DehazeNet [11] utilized Maxout units to extract haze-relevant features and estimate transmission maps end-to-end. Subsequently, AOD-Net [35] optimized the physical scattering model by unifying the transmission and atmospheric light parameters into a single variable for direct estimation. Addressing network depth and scale, GridDehazeNet [32] employed a grid-structured network with residual dense connections to learn direct dehazing mappings. Furthermore, acknowledging that haze is rarely uniform, FFA-Net [36] integrated pixel and channel attention mechanisms, enabling the model to prioritize high-frequency textures and heavily obscured regions adaptively.
Although originally developed for natural language processing (NLP), Transformer architectures have increasingly replaced CNNs in computer vision tasks due to their superior ability to model long-range dependencies [37,38]. For instance, Restormer [39] introduced multi-dconv head transposed attention, shifting the computational focus from spatial to channel dimensions, alongside a gated feed-forward network. To balance global context with local detail, Guo et al. [40] developed a hybrid architecture employing a Feature Modulation Mechanism to fuse CNN and Transformer strengths. In parallel, DehazeFormer [37] improved performance through a RescaleNorm layer, preserving low-frequency information, and a window attention mechanism tailored for uneven haze. More recently, the lightweight dual-attention network by Hua et al. [41] utilized efficient residual learning to capture spatial-channel features with minimal parameters. However, a persistent challenge in Transformer-based dehazing is the accumulation of attention noise, where the aggregation of irrelevant background features and haze interference degrades the quality of the generated attention maps.

2.2. Vision Transformer

Initially engineered for sequence modeling in NLP, the Transformer has garnered widespread acclaim across disciplines for its exceptional proficiency in modeling long-range dependencies. The advent of the Vision Transformer (ViT) [37,38] represented a paradigm shift, proving that pure self-attention architectures could transcend the local receptive field limitations of conventional CNNs to deliver superior visual representations. Nevertheless, despite the improved global fusion afforded by self-attention, the architecture confronts a formidable dichotomy of challenges in image processing, computational inefficiency and compromised feature purity.
Primarily, the quadratic complexity of standard self-attention relative to token count creates a prohibitive bottleneck for high-resolution visual tasks. To mitigate this, strategies such as shifted-window self-attention [20] have been introduced to linearize complexity by confining computations to non-overlapping local windows. Alternatively, channel self-attention [39] reorients the focus towards inter-channel covariance, aggregating global context while sidestepping the heavy computational burden of spatial dimensions. Furthermore, a pervasive yet frequently underestimated defect in Transformers is the long-tail distribution of attention maps, which are prone to pollution by a multitude of low-weight, irrelevant background features. This noise not only incurs computational waste but also dilutes the saliency of critical features, impairing the model’s ability to distinguish between haze and background in complex remote sensing scenarios. Consequently, sparse attention mechanisms have arisen as a structured solution to refine attention focus. Notable examples include DRSformer [19], which utilizes a Top-k channel selection strategy to retain highly responsive features, and SparseViT [42], which leverages window pruning to filter out non-semantic background regions.
However, the pursuit of sparsity in current methods often entails significant trade-offs. Reliance on hard thresholding or coarse block-level pruning frequently disrupts feature integrity, causing the irreversible loss of valuable context. Moreover, the sensitivity of these threshold-dependent methods to hyperparameters often induces training instability. To overcome these impediments, we propose a multi-scale differential Transformer. Diverging from rigid hard sparsification, our approach exploits differential operations to amplify feature contrast. This mechanism adaptively suppresses insignificant background noise and meticulously extracts discriminative high-frequency details, all while preserving the topological integrity of the underlying features.

3. Proposed Method

3.1. Overall Network

To address the non-uniform distribution and continuous spatial diffusion characteristics of haze in remote sensing imagery, this paper proposes a scale-adaptive differential Transformer network, as shown in Figure 2. The primary objective is to progressively decouple haze degradation features across multiple resolution levels. Operating with a pyramid structure of remote sensing hazy images as multi-path inputs, the network initially employs 3 × 3 convolutional layers to extract shallow semantic features. Building upon this foundation, N stacked differential Transformer blocks are utilized to deeply extract and fuse multi-source haze information within each internal scale. Specifically, the differential Transformer block integrates a scale-adaptive differential self-attention mechanism with a dynamic differential feed-forward neural network. This architecture is designed to precisely capture the most discriminative visual cues at each pyramid scale. Furthermore, to achieve the complementarity and enhancement of cross-scale features, a gated fusion module is introduced. The GFM adaptively selects discriminative features and optimizes the fusion strategy for scale-adaptive representations, thereby maximizing feature expressiveness.
In terms of implementation, the scale-adaptive feature learning network employs a Gaussian kernel to construct a Gaussian pyramid [30]. The original image is progressively downsampled to 1/2 and 1/4 resolutions, denoted as S 2 and S 3 respectively (with S 1 representing the original resolution), establishing a sequence of different scales. Subsequently, images at each scale are processed by differential Transformer-based U-Net. Both input and output mappings are executed via 3 × 3 convolutions, while residual connections are incorporated to enhance gradient propagation and feature preservation. The network adopts a cascaded upsampling strategy, where output features from smaller scales are sequentially upsampled and concatenated with the input of the subsequent level, thereby constructing a complete coarse-to-fine feature learning flow. This process is formally defined as follows:
f ( x ) = x + O m D T U I m ( x ) ,
where D T U ( · ) represents the data flow of the differential Transformer-based U-Net, and I m and O m denote the input and output feature mapping operations, respectively. Accordingly, the single-level data flow is represented as O i c = f I i c , where O l c and I l c correspond to the intermediate output and input features at the l-th level of the multi-resolution branch, and C signifies the coarse-to-fine learning process. To facilitate information interaction between levels, the output O l c from the l-th level is upsampled and added to the input of the ( l + 1 ) -th level. The connection mechanism is defined as follows:
I ^ l + 1 c = I l + 1 c + U P O l c ,
where U P ( · ) denotes the upsampling operation.

3.2. Differential Transformer Block (DTB)

Traditional Transformers [39] often struggle with the repetitive textures and pervasive haze distribution inherent in remote sensing imagery. In these architectures, conventional channel-wise self-attention mechanisms tend to perform an indiscriminate global aggregation of all tokens in the feature space during the reconstruction of query and key projections. Consequently, this approach assimilates substantial background noise from non-informative regions, which significantly hampers both computational efficiency and dehazing precision. To mitigate these limitations, we propose the differential Transformer block. The core philosophy of the DTB involves incorporating a differential operator that acts as a filter to suppress redundant low-frequency information while reserving discriminative high-frequency components. This compels the model to concentrate on high-frequency variations in the feature space, ensuring the precise capture and preservation of highly discriminative haze degradation features, even within sparse, blurred, or heavily occluded complex scenarios. Structurally, each DTB comprises two pivotal components: the scale-adaptive differential self-attention mechanism and the multi-scale differential feed-forward network. Taking the l-th module as an instance, the feature propagation and evolution process is formally expressed as follows:
Y l = X l 1 + SDSA L N X l 1 , X l = Y l + MDFN L N Y l ,
where L N ( · ) denotes the layer normalization operation. X l 1 represents the output feature from the preceding module. It yields the intermediate feature Y l after being processed by the SDSA and augmented with a residual connection. Subsequently, Y l undergoes further refinement via the MDFN to generate the final output feature X l of the current module.

3.2.1. Scale-Adaptive Differential Self-Attention (SDSA)

Starting with a normalized input feature map X R H × W × C , the proposed SDSA operates on multiple spatial scales to capture haze structures of different sizes. Specifically, three parallel attention branches are constructed, each corresponding to a different scale s { 1 , 2 , 3 } . For each scale, the query ( Q s ), key ( K s ), and value ( V s ) projections are generated by sequencing a 1 × 1 pointwise convolution followed by a 3 × 3 depthwise convolution with scale-specific parameters. Once generated, the Q s and K s tensors are reshaped into matrices of size R H W × C and R C × H W , respectively, and then split along the channel axis to form two separate groups: { Q 1 s , Q 2 s } and { K 1 s , K 2 s } . Pixel-wise similarity is then computed to produce two transposed attention maps A 1 s , A 2 s R C × C . A differential attention operation is further applied to emphasize discriminative structural cues, where the difference between the two attention maps is modulated by a learnable coefficient matrix λ s . The SDSA computation at scale s can be formulated as follows:
  Q s = W d ( W p ( X ) ) , K s = W d ( W p ( X ) ) , V s = W d ( W p ( X ) ) , s { 1 , 2 , 3 } Q 1 s , Q 2 s = Q s , K 1 s , K 2 s = K s , A 1 s = softmax K 1 s Q 1 s α , A 2 s = softmax K 2 s Q 2 s α , X ^ s = A 1 s λ s · A 2 s V s ,
where X corresponds to the input feature, and X ^ s corresponds to the output feature of different scales. W p ( · ) and W d ( · ) denote the 1 × 1 pointwise and 3 × 3 depthwise convolutions, respectively. The terms Q 1 s , Q 2 s R H W × C and K 1 s , K 2 s R C × H W represent the grouped queries and keys, and V s R H W × C is obtained by flattening the spatial dimensions of the original tensor from R H × W × C . α is an optional temperature factor defined by α = d .
In conventional differential approaches, the straightforward application of a uniform λ value often results in the simultaneous suppression of feature representations in both low- and high-response regions. Consequently, this limitation hinders the achievement of substantial performance gains in dehazing tasks. To mitigate the issue where high query–key matching scores are excessively suppressed, we introduce a learnable parameter matrix λ l e a r n s R C × C for different scales. This matrix is designed to compensate for deviations in the feature extraction process by addressing dimensional mismatches in intermediate representations. The refined parameter λ s is defined as follows:
λ s = λ l e a r n s + λ i n i t , s { 1 , 2 , 3 } ,
where λ s denotes the modulation coefficient for differential attention, parameterized as λ s = λ l e a r n s + λ i n i t , where λ l e a r n s R C × C aligns with the attention matrix, and λ i n i t = 0.8 0.6 × exp ( 0.3 · ( i + 1 ) ) with i [ 1 , L ] .

3.2.2. Multi-Scale Differential Feed-Forward Network (MDFN)

While conventional feed-forward networks play a pivotal role in enriching feature semantics, they often fail to effectively decouple redundant information within the spatial dimension when processing spatially correlated image data. To surmount this limitation, we propose an advanced architecture termed the multi-scale differential feed-forward network. The fundamental principle of this architecture is the application of spatial differential mechanisms to suppress irrelevant background interference, thereby ensuring the extraction of durable and discriminative feature representations. Operationally, the MDFN begins by dividing the input tensor X R H × W × C and subjecting the components to parallel convolutional layers with varying kernel sizes (specifically 5 × 5 and 3 × 3). This multi-branch strategy allows the model to effectively encapsulate contextual information across different scales. Subsequently, these features are aggregated to generate the intermediate variable X m i d , which is further partitioned into two subsets of features for weighted differential computation. This mathematical process is formally expressed as follows:
X 1 , X 2 = X , X 1 = Conv 5 × 5 ( ReLU X 1 ) , X 2 = Conv 3 × 3 ( ReLU X 2 ) , X m i d = Conv 3 × 3 ( Concat X 1 , X 2 ) , X m i d 1 , X m i d 2 = X m i d , X f i n a l = L N X m i d 1 λ × X m i d 2 , X ^ = Concat X f i n a l X 1 , X f i n a l X 2 ,
where ⊙ denotes the Hadamard product, and L N ( · ) represents layer normalization. Through this design, the resulting X f i n a l effectively highlights regions exhibiting high-frequency variations. Notably, when the MDFN is integrated with the differential attention module, the overall Transformer architecture achieves dual spatial–channel sparsity. This capability allows for the simultaneous elimination of redundant features across both channel and spatial domains, thereby significantly enhancing the dehazing efficiency of the model.

3.3. Gated Fusion Module (GFM)

In multi-scale architectures, effectively managing discrepancies in spatial resolution is critical to maximizing the utility of complementary features from different levels. While standard approaches typically utilize basic interpolation for resizing, this often results in aliasing artifacts and the erosion of high-frequency texture details, ultimately reducing image quality. To address these issues, we introduce the gated fusion module. The GFM is engineered to dynamically modulate and merge semantic content with detailed information across varying scales using an adaptive gating technique. Fundamentally, this module substitutes straightforward concatenation with a gating strategy to facilitate the intelligent selection and fusion of cross-scale data. Structurally, the GFM handles inputs from three distinct scales ( S 3 , S 2 , S 1 ) through a nested configuration, creating a dual-level filtering system. The mathematical formulation of this operation is given by:
G F M = Conv 1 × 1 Gate Conv 1 × 1 Gate ( S 3 , S 2 ) ,   S 1 .
In this equation, Conv 1 × 1 is employed for channel dimensionality adjustment. Feature selection is executed by the gating function Gate ( X , Y ) , which selectively highlights or suppresses areas of the primary feature X guided by the auxiliary feature Y:
Gate ( X , Y ) = σ Conv 3 × 3 Conv 1 × 1 ( Y ) Conv 3 × 3 Conv 1 × 1 ( X ) ,
where σ represents the tanh function, serving to provide non-linear gain modulation. Following processing, the fused features are forwarded to the respective decoding stage. This approach effectively reduces information loss associated with scale changes and ensures model robustness throughout the multi-scale learning procedure.

3.4. Loss Function

To ensure the restored images maintain both high fidelity and clear details, we implement a composite objective function that optimizes the network simultaneously across the pixel, gradient, and frequency domains. The total loss is defined as a weighted sum of three distinct components:
L t o t a l = λ 1 L c h a r + λ 2 L e d g e + λ 3 L f r e q .
The details of each component are described below. First, we utilize the Charbonnier loss ( L c h a r ) to measure the pixel-level difference between the predicted output x and the ground truth (GT) y. Unlike Mean Squared Error (MSE), Charbonnier loss is more robust to outliers and approximates the sparsity of the L 1 norm while remaining differentiable everywhere. Its mathematical form is:
L c h a r ( x , y ) = ( x y ) 2 + ε 2 ,
where the constant ε is added to ensure numerical stability and smooth gradients.
Second, to address the common issue of over-smoothing in dehazing tasks, we incorporate an Edge loss ( L e d g e ). This component forces the model to prioritize high-frequency texture recovery by minimizing the L 1 distance between the first-order gradients of the images:
L e d g e = | x y | 1 .
Here, ∇ represents the gradient operator, implemented using Sobel kernels to capture edge information in both horizontal and vertical orientations, ensuring sharper boundaries in the restoration.
Finally, we introduce frequency loss ( L f r e q ) to recover global structural patterns and periodic textures that are frequently missed by spatial-only methods. By applying the Fast Fourier Transform (FFT), we map the images to the spectral domain and minimize the L 1 difference between their magnitude spectra:
L f r e q = | F ( x ) F ( y ) | 1 ,
where F ( · ) denotes the operation to extract frequency magnitude. In our experiments, the hyperparameters are empirically set to λ 1 = 1 , λ 2 = 0.05 , and λ 3 = 0.01 .

4. Experiments

4.1. Experimental Setup

4.1.1. Datasets and Metrics

To comprehensively evaluate the performance of our method on the remote sensing image dehazing task, we conduct experiments on two synthetic datasets, SateHaze1k [43] and RICE [44], and one real-world dataset, RRSD300 [45]. The StateHaze1k dataset is divided into three haze levels, namely thin, moderate, and thick, with each subset comprising 320 training image pairs and 45 testing image pairs. The RICE dataset is captured by switching cloud-layer visibility and then cropped into non-overlapping 512 × 512 patches. The dataset consists of 500 image pairs in total, including 425 training pairs and 75 test pairs. The real-world dataset RRSD300 consists of 300 test images covering diverse natural environments, including urban, rural, and mountainous scenes, thereby simulating various real-world scenarios.
We adopt Peak Signal-to-Noise Ratio (PSNR) [46] and Structural Similarity Index Measure (SSIM) [47] as evaluation metrics. PSNR measures pixel-level similarity between the restored image and its reference, where larger values correspond to smaller errors. SSIM evaluates visual similarity by jointly considering luminance, contrast, and structure, with a stronger focus on structural consistency and alignment with human visual perception. In addition, we use the no-reference metrics Multi-Scale Image Quality (MUSIQ) [48] and Integrated Local Natural Image Quality Evaluator (ILNIQE) [49] to quantitatively evaluate restoration results in real-world scenes. Higher MUSIQ values indicate better model performance, whereas lower ILNIQE values indicate better performance.

4.1.2. Comparison Methods

We evaluate our method against state-of-the-art approaches, including prior-based methods (DCP [24]), CNN-based methods (DehazeNet [11], AOD-Net [35], LD-Net [50], GCA-Net [33], GrideDehazeNet [32], FFA-Net [36], MSBDN [34], FCTF-Net [51], M2SCN [52]), Mamba-based methods (UVM-Net [53], MambaIR [54], MambaIRv2 [55]), and Transformer-based methods (Restormer [39], DeHamer [40], DehazeFormer [37], Uformer [56], AIDTransformer [57], RSDformer [58]). All methods are retrained and evaluated with their official implementations to guarantee fair evaluation.

4.1.3. Implementation Details

During training, our model is built on PyTorch (v2.1.0) and optimized using the Adam optimizer [59]. The learning rate is decayed from an initial value of 2 × 10−4 to 1 × 10−4 using a cosine annealing strategy. The total number of epochs is set to 500. The model is trained with a batch size of 4 on image patches of size 256 × 256. All experiments are performed on a single NVIDIA RTX 4090 GPU (24 GB). Following previous work [16], we employ a combination of the Charbonnier loss, edge loss, and frequency loss.

4.2. Evaluation on the SateHaze1k Dataset

As shown in Table 1, we conduct a systematic quantitative comparison on the StateHaze1k dataset and report PSNR and SSIM under three haze density levels, including thin haze, moderate haze, and thick haze, as well as the overall average performance. Overall, our method achieves 25.93 dB in average PSNR and 0.9086 in SSIM, outperforming all competing methods and achieving the best performance. Compared with the closest-performing method MambaIRv2, our approach improves the average PSNR by 1.12 dB, which validates its overall effectiveness for remote sensing image dehazing. On the more severely degraded thick haze subset, our method outperforms the Transformer-based RSDformer by 1.33 dB in PSNR, demonstrating strong restoration capability under complex degradation conditions.
We further present visual comparison results on the three subsets in Figure 3, Figure 4 and Figure 5. As shown in Figure 3, on the thin subset, our method restores fine texture details and accurate colors. In contrast, the CNN-based method FCTF-Net and the Mamba-based method MambaIRv2 leave visible haze residues, which obscure structural details and introduce color distortion. The results on the thick subset in Figure 5 further demonstrate that our method can still recover clear images under severe haze degradation, indicating the robustness of the proposed approach.

4.3. Evaluation on the RICE Dataset

As shown in Table 2, we conduct quantitative evaluation on the RICE dataset with several representative dehazing methods. Our method achieves 30.47 dB in PSNR and 0.9513 in SSIM, which are the best results among all compared methods. Compared with the closest competitor, DehazeFormer (30.26 dB/0.9448), our method improves PSNR by 0.21 dB and SSIM by 0.0065, indicating more accurate structural restoration and better perceptual quality. In addition, compared with the state space modeling-based method MambaIR (29.85 dB/0.9456), our method also shows consistent advantages, further validating the effectiveness of the proposed architecture in long-range dependency modeling and fine-grained feature recovery. Our method shows competitive effectiveness along with consistent stability.
Qualitative comparisons are provided in Figure 6. Our method provides superior visual quality, especially for images containing dense vegetation with complex textures and fine details. For example, the restoration results of MSBDN contain noticeable noise and fail to recover the original colors. In contrast, our method produces visual representations that closely match the details of mountain ridges and building textures, demonstrating the robustness of the proposed approach. To further analyze the qualitative performance, we visualize the pixel-level SSIM differences between the dehazed results of different methods and the ground-truth images on the RICE dataset, as shown in Figure 7. By overlaying the SSIM difference maps onto the restored images, local structural discrepancies are clearly highlighted. Regions with warmer colors correspond to larger structural deviations, allowing for intuitive comparison of how well different methods preserve fine details and local structures. This visualization provides additional insight beyond visual inspection and emphasizes the structural restoration capability of our approach.

4.4. Generalization Results on Real-World RRSD300 Dataset

To assess real-world generalization, we additionally evaluate our method on the RRSD300 dataset. We first train the model on the thin subset of StateHaze1k and then use the trained weights for generalization testing on RRSD300. The quantitative results of the no-reference evaluation are presented in Table 3. Our method achieves the best performance on both MUSIQ and ILNIQE, indicating that it not only effectively removes haze degradation, but also suppresses over-enhancement and structural distortion, thereby preserving the naturalness and authenticity of the image content. The qualitative results are shown in Figure 8. It can be observed that our method exhibits strong haze removal capability. Compared with other representative state-of-the-art methods, our approach not only removes haze more effectively but also better restores authentic background structures and color information. Other methods tend to introduce color distortion or shift, and may even lose image details. These real-world generalization results further validate the effectiveness of our method.

4.5. Evaluation of Model Complexity

To comprehensively evaluate our model, we further conduct a model complexity analysis. As reported in Table 4, we compare our method with several representative state-of-the-art approaches in terms of model complexity, measured by the number of parameters, and computational cost, measured by floating-point operations (FLOPs). For fair comparison, all efficiency evaluations are performed at a resolution of 256 × 256. The analysis shows that, compared with the Transformer-based AIDTransformer, which has 439.9G FLOPs and 27.1M parameters, our model is significantly more lightweight with 65.0G FLOPs and 6.2M parameters. Quantitative results on the StateHaze1k and RICE datasets indicate that our method effectively reconciles restoration quality with computational cost.

4.6. Ablation Studies

In the following, we conduct ablation studies on the challenging thin subset of the StateHaze1k dataset for analysis and discussion.

4.6.1. Effectiveness of the SDSA

The effectiveness of the scale-adaptive differential self-attention module is examined through ablation studies, where the attention mechanism in our network is replaced with several representative self-attention variants, including multi-dconv head transposed attention (MDTA) [39], Top-k sparse self-attention (TKSA) [19], and traditional differential self-attention (TDSA) [21]. According to Table 5, the baseline model equipped with MDTA attains 27.09 dB PSNR and 0.9234 SSIM. Replacing MDTA with TKSA slightly improves the performance to 27.16 dB/0.9241, indicating that sparse attention can reduce redundancy but provides limited benefit for remote sensing dehazing. In contrast, incorporating the proposed SDSA module yields the best performance, achieving 27.32 dB PSNR and 0.9247 SSIM, consistently outperforming all compared attention variants. This benefits from the design of our SDSA, which explicitly models complementary query–key components and regulates differentiated attention responses across layers. This enables more effective suppression of self-attention noise and better adaptation to spatially varying haze in remote sensing scenarios.

4.6.2. Effectiveness of the Numbers of Scale Layers

An ablation study is conducted to evaluate the scale-adaptive feature learning architecture by progressively adding different resolution levels. Quantitative results are summarized in Table 5. Specifically, the variant (d) that adopts only a single-scale architecture at the original resolution achieves a PSNR of 26.26 dB and an SSIM of 0.9184. This result indicates that relying solely on a single resolution limits the network’s ability to capture haze patterns with different spatial extents. When an additional half-scale branch is introduced, i.e., with two scale layers (variant (e)), the performance improves significantly to 26.93 dB PSNR and 0.9225 SSIM. This improvement demonstrates that incorporating multi-scale representations enables the model to better capture medium-scale haze structures and enhances feature robustness through cross-scale information complementarity. By further extending the architecture to adopt a three-layer multi-scale design, our method achieves the best performance, reaching a PSNR of 27.32 dB and an SSIM of 0.9247. Compared with the single-scale configuration, the complete scale-adaptive architecture improves PSNR and SSIM by 1.06 dB and 0.0063, respectively. These consistent improvements verify that progressively decoupling haze degradation across multiple resolution levels is effective, and the scale-adaptive learning strategy facilitates more accurate feature representation and restoration.

4.6.3. Effectiveness of MDFN

We conduct ablation experiments to investigate the contribution of the multi-scale differential feed-forward network by replacing the feed-forward module in our network with several representative feed-forward network variants, including the conventional feed-forward network [60], gated-dconv feed-forward network [39], and multi-scale feed-forward network [19]. Quantitative comparisons are reported in Table 6. The baseline model with the feed-forward network achieves 27.15 dB PSNR and 0.9242 SSIM. While a feed-forward network can enhance feature representation through channel expansion, its dense transformation inevitably introduces spatial redundancy, which limits its effectiveness in remote sensing image dehazing. Similarly, the multi-scale feed-forward network achieves 26.87 dB PSNR and 0.9198 SSIM, suggesting that multi-scale feature aggregation contributes positively to modeling haze with varying spatial extents. Our proposed MDFN first employs parallel multi-scale depthwise convolutions, enabling the model to capture spatial context under different receptive fields. Then, the learnable parameters explicitly enhance feature contrast, allowing the network to emphasize information-rich structures while suppressing redundant responses. Together, these designs strengthen feature representation and thereby improve dehazing performance.

4.6.4. Effectiveness of GFM

The effectiveness of the gated fusion module is examined in an ablation study through a comparison of multiple multi-scale feature fusion strategies. Quantitative results are shown in Table 7. Specifically, completely removing the feature fusion module (Variants (i)) leads to a PSNR of 26.96 dB and an SSIM of 0.9231, indicating that directly propagating multi-scale features without explicit fusion is insufficient to fully exploit cross-scale complementary information. When a simple feature concatenation strategy (Variants (j)) is adopted, the performance is slightly improved, reaching 27.08 dB PSNR and 0.9234 SSIM. This improvement suggests that aggregating features from different resolutions is beneficial; however, naive concatenation lacks adaptive selection and may introduce redundant or misaligned information. In contrast, our model equipped with the proposed GFM achieves the best performance, with PSNR and SSIM reaching 27.32 dB and 0.9247, respectively. These consistent improvements demonstrate that GFM effectively alleviates the adverse effects caused by spatial resolution mismatch and enables more discriminative integration of semantic and detailed features across scales, thereby significantly enhancing the performance of the proposed dehazing framework.

4.6.5. Effectiveness of Loss Function

We perform an ablation study to examine each component of the composite loss function, gradually incorporating different sub-objectives during training. Quantitative results are shown in Table 7. Specifically, although the Charbonnier loss provides stable pixel-level supervision, relying solely on pixel-domain optimization limits the model’s ability to recover fine structures and perceptual details. As a result, this setting achieves the lowest PSNR of 25.98 dB. By further incorporating edge loss and frequency loss, the performance consistently improves. The full model that integrates all three loss terms achieves the best performance, reaching 27.32 dB in PSNR and 0.9247 in SSIM. These results indicate that the three loss components are complementary: Charbonnier loss ensures image fidelity, edge loss enhances local details, and frequency loss promotes global consistency. Their joint optimization leads to more structurally accurate dehazing results.

4.7. Limitations and Discussion

Despite the strong performance achieved by the proposed SDTformer, several limitations and open issues remain and deserve further discussion, particularly with respect to its modeling assumptions and potential extensions to more complex remote sensing scenarios.
First, the proposed multi-scale differential attention mechanism is designed to suppress common and redundant attention responses by amplifying discriminative feature differences. This design implicitly assumes that meaningful structural or textural variations are present in the feature space to support effective contrast enhancement. In most remote sensing images, such variations naturally arise from man-made structures, terrain changes, or object boundaries, enabling the differential operation to function reliably. However, in extremely low-contrast or highly homogeneous regions, such as large water surfaces, desert areas, or snow-covered scenes under dense haze, the similarity distributions across spatial locations may become nearly uniform. Under these conditions, the differential operation has limited contrast to exploit, which may reduce its ability to highlight informative regions. As shown in Figure 9, a failure case in a desert areas scenario is presented. While the proposed model is effective in the majority of practical cases, further improvement in differential modeling in texture-deficient or visually unclear scenes warrants future investigation.
Second, the strength of the scale-adaptive differential attention in SDTformer is controlled through a combination of learnable parameters and layer-wise initialization. This strategy enables stable training and balanced attention modulation across network depths, and it proves effective across different datasets, as evidenced by the quantitative results reported in Table 1 and Table 2. Nevertheless, the initialization and adjustment of differential strength are still guided by empirical design choices. Remote sensing scenes vary widely in terms of land cover types, haze density, and spatial distribution, which implies that the optimal degree of attention suppression or enhancement may differ across scenes and even across regions within a single image. A more adaptive mechanism that dynamically regulates differential strength based on scene characteristics or feature statistics could further improve the robustness and flexibility of the proposed framework.
Finally, the GFM plays a critical role in selectively aggregating multi-scale features from different encoder stages. Table 7 indicates that simple concatenation of multi-scale features (Variant (b)) leads to only marginal performance gains, whereas the gated fusion mechanism consistently enhances restoration quality by enabling more selective information flow. This observation indicates that unstructured aggregation of multi-scale features may introduce redundant or even conflicting information, particularly in the presence of complex haze distributions. Nevertheless, the current gated fusion design relies on manually defined fusion stages and fixed scale granularity. While this design is effective and stable, it may not fully exploit complex cross-scale interactions in highly heterogeneous scenes. More flexible fusion strategies, such as hierarchical or content-adaptive fusion mechanisms, could potentially further enhance multi-scale representation learning.

5. Conclusions

In this paper, we propose SDTformer, a scale-adaptive differential Transformer framework for remote sensing image dehazing, which addresses the accumulation of redundant and noisy responses in conventional self-attention mechanisms by explicitly modeling attention differences to suppress irrelevant activations and enhance discriminative features. By integrating differential attention with a multi-scale architecture and a gated fusion module for selective feature aggregation, SDTformer effectively captures haze-related degradations across spatial scales, and extensive experiments on multiple remote sensing benchmarks demonstrate consistent improvements in both quantitative performance and visual quality, particularly in challenging scenes with complex haze distributions. These results indicate that the proposed framework highlights the importance of differential modeling and selective multi-scale fusion in Transformer-based remote sensing image restoration, and provides a solid foundation for future extensions to more complex atmospheric degradation scenarios.

Author Contributions

Conceptualization, B.L.; methodology, B.L.; software, B.L.; validation, B.L.; formal analysis, B.L. and Q.Z.; investigation, B.L.; data curation, B.L. and Q.Z.; writing—original draft preparation, B.L.; writing—review and editing, B.L. and Q.Z.; visualization, B.L.; supervision, Q.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

All online experimental datasets referenced in this work can be found at: https://github.com/MingTian99/RSDformer (accessed on 14 November 2025).

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Anderson, K.; Ryan, B.; Sonntag, W.; Kavvada, A.; Friedl, L. Earth observation in service of the 2030 Agenda for Sustainable Development. Geo-Spat. Inf. Sci. 2017, 20, 77–96. [Google Scholar] [CrossRef] [Scilit]
  2. Yuan, Q.; Shen, H.; Li, T.; Li, Z.; Li, S.; Jiang, Y.; Xu, H.; Tan, W.; Yang, Q.; Wang, J.; et al. Deep learning in environmental remote sensing: Achievements and challenges. Remote Sens. Environ. 2020, 241, 111716. [Google Scholar] [CrossRef] [Scilit]
  3. Phiri, D.; Morgenroth, J. Developments in Landsat land cover classification methods: A review. Remote Sens. 2017, 9, 967. [Google Scholar] [CrossRef] [Scilit]
  4. Tang, A.; Wen, A. An intelligent simulation system for earthquake disaster assessment. Comput. Geosci. 2009, 35, 871–879. [Google Scholar] [CrossRef] [Scilit]
  5. Levy, J.M.; Hirt, S.A.; Dawkins, C.J. Contemporary Urban Planning; Routledge: Oxfordshire, UK, 2009. [Google Scholar]
  6. Yang, Q.; Chen, X.; Li, P.; Guan, Q.; Jin, G.; Jin, J. Rethinking Rainy 3D Scene Reconstruction via Perspective Transforming and Brightness Tuning. arXiv 2025, arXiv:2511.06734. [Google Scholar] [CrossRef] [Scilit]
  7. O’shea, K.; Nash, R. An introduction to convolutional neural networks. arXiv 2015, arXiv:1511.08458. [Google Scholar] [CrossRef] [Scilit]
  8. Song, T.; Li, P.; Fan, S.; Jin, J.; Jin, G.; Fan, L. Exploring a context-gated network for effective image deraining. J. Vis. Commun. Image Represent. 2024, 98, 104060. [Google Scholar] [CrossRef] [Scilit]
  9. Li, P.; Shu, X.; Feng, C.M.; Feng, Y.; Zuo, W.; Tang, J. Surgical video workflow analysis via visual-language learning. Npj Health Syst. 2025, 2, 5. [Google Scholar] [CrossRef] [Scilit]
  10. Chen, X.; Pan, J.; Dong, J. Bidirectional multi-scale implicit neural representations for image deraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 25627–25636. [Google Scholar]
  11. Cai, B.; Xu, X.; Jia, K.; Qing, C.; Tao, D. Dehazenet: An end-to-end system for single image haze removal. IEEE Trans. Image Process. 2016, 25, 5187–5198. [Google Scholar] [CrossRef] [Scilit]
  12. Ren, W.; Liu, S.; Zhang, H.; Pan, J.; Cao, X.; Yang, M.H. Single image dehazing via multi-scale convolutional neural networks. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2016; pp. 154–169. [Google Scholar]
  13. Ren, W.; Pan, J.; Zhang, H.; Cao, X.; Yang, M.H. Single image dehazing via multi-scale convolutional neural networks with holistic edges. Int. J. Comput. Vis. 2020, 128, 240–259. [Google Scholar] [CrossRef] [Scilit]
  14. Niu, Z.; Zhong, G.; Yu, H. A review on the attention mechanism of deep learning. Neurocomputing 2021, 452, 48–62. [Google Scholar] [CrossRef] [Scilit]
  15. Fan, S.; Song, T.; Jin, G.; Jin, J.; Li, Q.; Xia, X. A lightweight cloud and cloud shadow detection transformer with prior-knowledge guidance. IEEE Geosci. Remote Sens. Lett. 2024, 21, 8003405. [Google Scholar] [CrossRef] [Scilit]
  16. Chen, H.; Chen, X.; Lu, J.; Li, Y. Rethinking Multi-Scale Representations in Deep Deraining Transformer. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Palo Alto, CA, USA, 2024; Volume 38, pp. 1046–1053. [Google Scholar]
  17. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
  18. Guan, Q.; Chen, X.; Jin, G.; Jin, J.; Fan, S.; Song, T.; Pan, J. Rethinking Nighttime Image Deraining via Learnable Color Space Transformation. arXiv 2025, arXiv:2510.17440. [Google Scholar] [CrossRef] [Scilit]
  19. Chen, X.; Li, H.; Li, M.; Pan, J. Learning a sparse transformer network for effective image deraining. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 5896–5905. [Google Scholar]
  20. Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; Timofte, R. Swinir: Image restoration using swin transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 1833–1844. [Google Scholar]
  21. Wortsman, M.; Lee, J.; Gilmer, J.; Kornblith, S. Replacing softmax with relu in vision transformers. arXiv 2023, arXiv:2309.08586. [Google Scholar] [CrossRef] [Scilit]
  22. Fattal, R. Single image dehazing. ACM Trans. Graph. 2008, 27, 72. [Google Scholar] [CrossRef] [Scilit]
  23. Liu, Y.; Pan, J.; Ren, J.; Su, Z. Learning deep priors for image dehazing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 2492–2500. [Google Scholar]
  24. He, K.; Sun, J.; Tang, X. Single image haze removal using dark channel prior. IEEE Trans. Pattern Anal. Mach. Intell. 2010, 33, 2341–2353. [Google Scholar] [CrossRef] [Scilit]
  25. Zhu, Q.; Mai, J.; Shao, L. A fast single image haze removal algorithm using color attenuation prior. IEEE Trans. Image Process. 2015, 24, 3522–3533. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Li, B.; Ren, W.; Fu, D.; Tao, D.; Feng, D.; Zeng, W.; Wang, Z. Benchmarking single-image dehazing and beyond. IEEE Trans. Image Process. 2018, 28, 492–505. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. LeCun, Y.; Bengio, Y.; Hinton, G. Deep learning. Nature 2015, 521, 436–444. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Goodfellow, I. Deep Learning; MIT Press: Cambridge, MA, USA, 2016. [Google Scholar]
  29. Wu, X.; Chen, H.; Chen, X.; Xu, G. Multi-scale transformer with conditioned prompt for image deraining. Digit. Signal Process. 2025, 156, 104847. [Google Scholar] [CrossRef] [Scilit]
  30. Wu, X.; Xiao, Z.; He, J.; Lei, J.; Zeng, X.; Xu, G. Multi-weather unmanned aerial vehicle remote sensing image restoration via scale-aware Trident Mamba. J. Appl. Remote Sens. 2025, 19, 046507. [Google Scholar] [CrossRef] [Scilit]
  31. Zhang, H.; Patel, V.M. Densely connected pyramid dehazing network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 3194–3203. [Google Scholar]
  32. Liu, X.; Ma, Y.; Shi, Z.; Chen, J. Griddehazenet: Attention-based multi-scale network for image dehazing. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 7314–7323. [Google Scholar]
  33. Chen, D.; He, M.; Fan, Q.; Liao, J.; Zhang, L.; Hou, D.; Yuan, L.; Hua, G. Gated context aggregation network for image dehazing and deraining. In Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vvision (WACV); IEEE: New York, NY, USA, 2019; pp. 1375–1383. [Google Scholar]
  34. Dong, H.; Pan, J.; Xiang, L.; Hu, Z.; Zhang, X.; Wang, F.; Yang, M.H. Multi-scale boosted dehazing network with dense feature fusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 2157–2167. [Google Scholar]
  35. Li, B.; Peng, X.; Wang, Z.; Xu, J.; Feng, D. Aod-net: All-in-one dehazing network. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 4770–4778. [Google Scholar]
  36. Qin, X.; Wang, Z.; Bai, Y.; Xie, X.; Jia, H. FFA-Net: Feature fusion attention network for single image dehazing. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Palo Alto, CA, USA, 2020; Volume 34, pp. 11908–11915. [Google Scholar]
  37. Song, Y.; He, Z.; Qian, H.; Du, X. Vision transformers for single image dehazing. IEEE Trans. Image Process. 2023, 32, 1927–1941. [Google Scholar] [CrossRef] [Scilit]
  38. Han, K.; Wang, Y.; Chen, H.; Chen, X.; Guo, J.; Liu, Z.; Tang, Y.; Xiao, A.; Xu, C.; Xu, Y.; et al. A survey on vision transformer. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 87–110. [Google Scholar] [CrossRef] [Scilit]
  39. Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.H. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 5728–5739. [Google Scholar]
  40. Guo, C.L.; Yan, Q.; Anwar, S.; Cong, R.; Ren, W.; Li, C. Image dehazing transformer with transmission-aware 3d position embedding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 5812–5820. [Google Scholar]
  41. Hua, Z.; Hua, Z.; Li, J. LWDA-Net: A lightweight dual-attention network for single image dehazing. In Proceedings of the 2024 4th International Conference on Neural Networks, Information and Communication; IEEE: New York, NY, USA, 2024; pp. 415–420. [Google Scholar]
  42. Chen, X.; Liu, Z.; Tang, H.; Yi, L.; Zhao, H.; Han, S. Sparsevit: Revisiting activation sparsity for efficient high-resolution vision transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 2061–2070. [Google Scholar]
  43. Huang, B.; Zhi, L.; Yang, C.; Sun, F.; Song, Y. Single satellite optical imagery dehazing using SAR image prior based on conditional generative adversarial networks. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Snowmass Village, CO, USA, 1–5 March 2020; pp. 1806–1813. [Google Scholar]
  44. Lin, D.; Xu, G.; Wang, X.; Wang, Y.; Sun, X.; Fu, K. A remote sensing image dataset for cloud removal. arXiv 2019, arXiv:1901.00600. [Google Scholar] [CrossRef] [Scilit]
  45. Wen, Y.; Gao, T.; Li, Z.; Zhang, J.; Chen, T. Encoder-minimal and decoder-minimal framework for remote sensing image dehazing. In Proceedings of the ICASSP 2024—2024 IEEE International Conference on Acoustics, Speech and Signal Processing; IEEE: New York, NY, USA, 2024; pp. 36–40. [Google Scholar]
  46. Huynh-Thu, Q.; Ghanbari, M. Scope of validity of PSNR in image/video quality assessment. Electron. Lett. 2008, 44, 800–801. [Google Scholar] [CrossRef] [Scilit]
  47. Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image quality assessment: From error visibility to structural similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [Scilit]
  48. Ke, J.; Wang, Q.; Wang, Y.; Milanfar, P.; Yang, F. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 10–17 October 2021; pp. 5148–5157. [Google Scholar]
  49. Zhang, L.; Zhang, L.; Bovik, A.C. A feature-enriched completely blind image quality evaluator. IEEE Trans. Image Process. 2015, 24, 2579–2591. [Google Scholar] [CrossRef] [Scilit]
  50. Ullah, H.; Muhammad, K.; Irfan, M.; Anwar, S.; Sajjad, M.; Imran, A.S.; de Albuquerque, V.H.C. Light-DehazeNet: A novel lightweight CNN architecture for single image dehazing. IEEE Trans. Image Process. 2021, 30, 8968–8982. [Google Scholar] [CrossRef] [Scilit]
  51. Li, Y.; Chen, X. A coarse-to-fine two-stage attentive network for haze removal of remote sensing images. IEEE Geosci. Remote Sens. Lett. 2020, 18, 1751–1755. [Google Scholar] [CrossRef] [Scilit]
  52. Li, S.; Zhou, Y.; Xiang, W. M2SCN: Multi-model self-correcting network for satellite remote sensing single-image dehazing. IEEE Geosci. Remote Sens. Lett. 2022, 20, 6000605. [Google Scholar] [CrossRef] [Scilit]
  53. Zheng, Z.; Wu, C. U-shaped vision mamba for single image dehazing. arXiv 2024, arXiv:2402.04139. [Google Scholar] [CrossRef] [Scilit]
  54. Guo, H.; Li, J.; Dai, T.; Ouyang, Z.; Ren, X.; Xia, S.T. Mambair: A simple baseline for image restoration with state-space model. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2024. [Google Scholar]
  55. Guo, H.; Guo, Y.; Zha, Y.; Zhang, Y.; Li, W.; Dai, T.; Xia, S.T.; Li, Y. Mambairv2: Attentive state space restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025. [Google Scholar]
  56. Wang, Z.; Cun, X.; Bao, J.; Zhou, W.; Liu, J.; Li, H. Uformer: A general u-shaped transformer for image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022. [Google Scholar]
  57. Kulkarni, A.; Murala, S. Aerial image dehazing with attentive deformable transformers. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Waikoloa, HI, USA, 2–7 January 2023; pp. 6305–6314. [Google Scholar]
  58. Song, T.; Fan, S.; Li, P.; Jin, J.; Jin, G.; Fan, L. Learning an effective transformer for remote sensing satellite image dehazing. IEEE Geosci. Remote Sens. Lett. 2023, 20, 8002305. [Google Scholar] [CrossRef] [Scilit]
  59. Diederik, P.K. Adam: A method for stochastic optimization. arXiv 2015, arXiv:1412.6980. [Google Scholar] [CrossRef] [Scilit]
  60. Dosovitskiy, A. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Illustration of different attention mechanisms. (a) standard self-attention, (b) sparse self-attention, (c) differential self-attention, and (d) the proposed scale-adaptive differential self-attention, which models multi-scale contextual dependencies to better suppress attention noise in remote sensing scenes.
Figure 1. Illustration of different attention mechanisms. (a) standard self-attention, (b) sparse self-attention, (c) differential self-attention, and (d) the proposed scale-adaptive differential self-attention, which models multi-scale contextual dependencies to better suppress attention noise in remote sensing scenes.
Remotesensing 18 01136 g001
Figure 2. The overall architecture of the proposed SDTformer for remote sensing image dehazing.
Figure 2. The overall architecture of the proposed SDTformer for remote sensing image dehazing.
Remotesensing 18 01136 g002
Figure 3. Compare the haze removal performance on the thin subset of the StateHaze1k dataset with existing SOTA methods.
Figure 3. Compare the haze removal performance on the thin subset of the StateHaze1k dataset with existing SOTA methods.
Remotesensing 18 01136 g003
Figure 4. Comparison of the haze removal performance on the moderate subset of the StateHaze1k dataset with existing SOTA methods.
Figure 4. Comparison of the haze removal performance on the moderate subset of the StateHaze1k dataset with existing SOTA methods.
Remotesensing 18 01136 g004
Figure 5. Comparison of the haze removal performance on the thick subset of the StateHaze1k dataset with existing SOTA methods.
Figure 5. Comparison of the haze removal performance on the thick subset of the StateHaze1k dataset with existing SOTA methods.
Remotesensing 18 01136 g005
Figure 6. Comparison of the haze removal performance on the RICE dataset with existing SOTA methods.
Figure 6. Comparison of the haze removal performance on the RICE dataset with existing SOTA methods.
Remotesensing 18 01136 g006
Figure 7. Visualization of pixel-level SSIM differences between the dehazed results of different methods and the ground-truth images on the RICE dataset, highlighting local structural discrepancies. The SSIM difference maps are overlaid on the restored images, where warmer colors (red) indicate larger structural deviations.
Figure 7. Visualization of pixel-level SSIM differences between the dehazed results of different methods and the ground-truth images on the RICE dataset, highlighting local structural discrepancies. The SSIM difference maps are overlaid on the restored images, where warmer colors (red) indicate larger structural deviations.
Remotesensing 18 01136 g007
Figure 8. Comparison of the haze removal performance on the real-world RRSD300 dataset.
Figure 8. Comparison of the haze removal performance on the real-world RRSD300 dataset.
Remotesensing 18 01136 g008
Figure 9. Failure case visualization.
Figure 9. Failure case visualization.
Remotesensing 18 01136 g009
Table 1. The performance of the proposed method is compared with several existing state-of-the-art methods on the StateHaze1k dataset. The best and second-best entries are denoted by bold and underlining, respectively.
Table 1. The performance of the proposed method is compared with several existing state-of-the-art methods on the StateHaze1k dataset. The best and second-best entries are denoted by bold and underlining, respectively.
MethodsThin HazeModerate HazeThick HazeAverage
PSNRSSIMPSNRSSIMPSNRSSIMPSNRSSIM
DCP [24]13.450.70159.780.591610.900.572011.380.6217
DehazeNet [11]16.570.488716.930.299215.440.368916.320.3856
AOD-Net [35]18.740.858417.690.796913.420.652316.620.7692
LD-Net [50]17.830.852119.800.903316.600.764718.070.8400
GCA-Net [33]22.270.903024.890.932720.510.830722.560.8888
GrideDehazeNet [32]20.040.861420.960.907118.670.790319.890.8529
FFA-Net [36]22.300.901325.460.935620.840.834822.860.8906
MSBDN [34]18.020.735120.760.800616.780.547118.520.6943
FCTF-Net [51]20.060.876923.430.927218.680.794320.720.8661
M2SCN [52]25.210.917526.110.941621.330.828924.220.8960
UVM-Net [53]24.500.918326.140.942122.150.836824.260.8991
MambaIR [54]24.740.919325.960.940822.300.850924.330.9037
MambaIRv2 [55]25.370.917326.120.940422.950.848024.810.9019
Restormer [39]24.970.918626.770.942221.280.824324.340.8951
DeHamer [40]20.940.864922.890.869119.800.797921.210.8440
DehazeFormer [37]23.920.905625.940.942322.030.826823.970.8916
Uformer [56]21.680.888521.140.832119.880.806220.900.8423
AIDTransformer [57]23.120.898225.080.913620.560.821722.920.8778
RSDformer [58]24.050.911825.970.936122.870.854324.290.9007
Ours27.320.924726.260.941524.200.859625.930.9086
Table 2. The performance of the proposed method is compared with several existing state-of-the-art methods on the RICE dataset. The best and second-best entries are denoted by bold and underlining, respectively.
Table 2. The performance of the proposed method is compared with several existing state-of-the-art methods on the RICE dataset. The best and second-best entries are denoted by bold and underlining, respectively.
DatasetRICE
MethodsDCP [24]AOD-Net [35]LD-Net [50]FFA-Net [36]
PSNR17.4823.7728.8828.54
SSIM0.78410.87310.93360.9396
MethodsMSBDN [34]DehazeFormer [37]MambaIR [54]Ours
PSNR29.9630.2629.8530.47
SSIM0.91050.94480.94560.9513
Table 3. The performance of the proposed method is compared with several existing state-of-the-art methods on the real-world RRSD300 dataset. The best and second-best entries are denoted by bold and underlining, respectively.
Table 3. The performance of the proposed method is compared with several existing state-of-the-art methods on the real-world RRSD300 dataset. The best and second-best entries are denoted by bold and underlining, respectively.
MethodsAOD-Net
[35]
LD-Net
[50]
FFA-Net
[36]
MSBDN
[34]
FCTF-Net
[51]
Restormer
[39]
DeHamer
[40]
RSDformer
[58]
Ours
MUSIQ45.903045.880545.075843.647144.483945.589345.915144.357046.3567
ILNIQE34.572333.410824.100124.767325.853024.072325.130624.531923.8228
Table 4. Model complexity comparison with a test image size of 256 × 256 pixels.
Table 4. Model complexity comparison with a test image size of 256 × 256 pixels.
MethodsFFA-Net
[36]
MSBDN
[34]
Restormer
[39]
DehazeFormer
[37]
Uformer
[56]
AIDTransformer
[57]
Ours
Parameters (M)4.631.326.19.720.627.16.2
FLOPs (G)287.541.5140.989.841.1439.965.0
Table 5. Ablation study on different self-attention modules and numbers of scale layers. “✔” indicates that the component is enabled, while “✗” indicates that the component is disabled. The best is indicated in bold.
Table 5. Ablation study on different self-attention modules and numbers of scale layers. “✔” indicates that the component is enabled, while “✗” indicates that the component is disabled. The best is indicated in bold.
VariantsSelf-Attention ModulesNumber of Scale LayersPSNRSSIM
MDTATKSATDSASDSA123
(a)27.090.9234
(b)27.160.9241
(c)26.830.9224
(d)26.260.9184
(e)26.930.9225
Ours27.320.9247
Table 6. Ablation study of the different feed-forward networks. “✔” indicates that the component is enabled, while “✗” indicates that the component is disabled. The best is indicated in bold.
Table 6. Ablation study of the different feed-forward networks. “✔” indicates that the component is enabled, while “✗” indicates that the component is disabled. The best is indicated in bold.
VariantsFeed-Forward NetworkGated-Dconv
Feed-Forward Network
Multi-Scale
Feed-Forward Network
MDFNPSNRSSIM
(f)27.150.9242
(g)27.210.9236
(h)26.870.9198
Ours27.320.9247
Table 7. Ablation studies on the GFM module and the loss function. “✔” indicates that the component is enabled, while “✗” indicates that the component is disabled. The best is indicated in bold.
Table 7. Ablation studies on the GFM module and the loss function. “✔” indicates that the component is enabled, while “✗” indicates that the component is disabled. The best is indicated in bold.
VariantsFeature Fusion ModuleLoss FunctionPSNRSSIM
Concat GFM Charbonnier Loss Edge Loss Frequency Loss
(i)26.960.9231
(j)27.080.9234
(k)25.980.9043
(l)26.640.9206
(m)26.480.9198
Ours27.320.9247
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Liu, B.; Zhang, Q. SDTformer: Scale-Adaptive Differential Transformer Network for Remote Sensing Image Dehazing. Remote Sens. 2026, 18, 1136. https://doi.org/10.3390/rs18081136

AMA Style

Liu B, Zhang Q. SDTformer: Scale-Adaptive Differential Transformer Network for Remote Sensing Image Dehazing. Remote Sensing. 2026; 18(8):1136. https://doi.org/10.3390/rs18081136

Chicago/Turabian Style

Liu, Boyu, and Qi Zhang. 2026. "SDTformer: Scale-Adaptive Differential Transformer Network for Remote Sensing Image Dehazing" Remote Sensing 18, no. 8: 1136. https://doi.org/10.3390/rs18081136

APA Style

Liu, B., & Zhang, Q. (2026). SDTformer: Scale-Adaptive Differential Transformer Network for Remote Sensing Image Dehazing. Remote Sensing, 18(8), 1136. https://doi.org/10.3390/rs18081136

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop