Next Article in Journal
Geospatial Foundation Models Improve Atoll Island Ecosystem Mapping: A Case Study Using AlphaEarth Embeddings
Previous Article in Journal
Geometry-Guided Semi-Supervised Multimodal Segmentation for UAV-Based Rice-Lodging Mapping
Previous Article in Special Issue
Absolute Calibration of Weather Radars Using Metal Spheres Based on Sector Scanning
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

TFCRNet: Dual-Discriminator SAR-to-Optical Translation and Region-Gated Cross-Attention Fusion for Thick-Cloud Removal

1
College of Computer Science and Software Engineering, Hohai University, Nanjing 211100, China
2
Information Center, Ministry of Water Resources, Beijing 100053, China
3
The Key Laboratory of Water Big Data Technology of Ministry of Water Resources, Hohai University, Nanjing 211100, China
4
Information Center, Yellow River Conservancy Commission (YRCC), Zhengzhou 450004, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(17), 2962; https://doi.org/10.3390/rs18172962
Submission received: 3 June 2026 / Revised: 27 July 2026 / Accepted: 28 August 2026 / Published: 2 September 2026

Highlights

What are the main findings?
  • This study proposes TFCRNet, a two-stage SAR–optical thick cloud removal framework that explicitly combines SAR-to-optical translation and cloud-aware multimodal fusion.
  • A dual-discriminator design separately constrains spectral fidelity and structural integrity during SAR-to-optical translation.
  • A region-gated cross-attention fusion module uses cloud masks to adaptively balance translated SAR-derived cues and reliable optical information in different regions.
What are the implications of the main findings?
  • The results suggest that reducing cross-modal discrepancy before fusion and applying region-aware interaction during fusion are both important for SAR–optical thick-cloud removal.
  • The proposed framework improves reconstruction quality while reducing cloud-region artifacts and degradation in cloud-free regions.

Abstract

Thick-cloud contamination severely limits the usability of optical remote sensing imagery because cloud-covered regions may suffer from complete loss of surface information. Synthetic aperture radar (SAR) imagery provides complementary structural cues due to its cloud-penetrating capability, but the substantial cross-modal discrepancy between SAR and optical images makes high-fidelity SAR–optical fusion challenging. Existing methods usually either directly fuse heterogeneous SAR and optical features or use SAR-to-optical translation with insufficient spectral and structural constraints, which may lead to spectral distortion, structural artifacts, or degradation of cloud-free regions. To address these issues, we propose TFCRNet, a two-stage translation-and-fusion network for SAR–optical thick-cloud removal. In the translation stage, a Multi-Scale Feature Fusion Generator (MSFFG) transforms SAR imagery into optical-like images, while a Spectral Discriminator (SpeD) and a Structural Discriminator (StrD) separately constrain spectral fidelity and structural integrity. In the fusion stage, a Region-Gated Cross-Attention Fusion (RGCAF) module performs cloud-aware feature interaction between the translated optical image and the cloudy optical image. Using an externally supplied cloud mask, RGCAF emphasizes translated SAR-derived cues in cloud-covered regions while retaining reliable optical information in cloud-free regions. TFCRNet therefore requires a cloud mask during inference. Experiments on the SEN12MS-CR and SMILE-CR datasets show that TFCRNet achieves the best overall performance among the baseline methods reproduced under the unified experimental protocol adopted in this study. TFCRNet obtains 33.01/30.53 dB PSNR and 0.91/0.88 SSIM on SEN12MS-CR and SMILE-CR, respectively. These controlled results should be distinguished from literature-reported values obtained under different experimental settings, several of which are higher on selected metrics. Fine-grained ablations demonstrate that SpeD and StrD provide differentiated spectral and structural supervision, while RGCAF improves multimodal reconstruction through cloud-mask-guided regional information routing rather than spatially uniform feature fusion.

1. Introduction

Optical remote sensing imagery is one of the most important data sources for Earth observation, because it provides rich spatial and spectral information for land-cover mapping, environmental monitoring, agricultural assessment, disaster response, and urban analysis. However, optical sensors are highly sensitive to atmospheric conditions. Cloud contamination frequently reduces the availability and reliability of satellite observations, and cloud/shadow screening has therefore become a fundamental preprocessing issue in optical remote sensing workflows [1,2]. Once cloudy pixels enter downstream interpretation tasks, they can cause biased classification, inaccurate change detection, and unreliable surface-parameter retrieval [3]. Therefore, recovering cloud-free optical imagery from contaminated observations remains an important and practical problem.
Cloud removal methods differ substantially according to the severity of cloud degradation. For thin clouds or haze-like contamination, part of the surface signal can still be observed through the atmosphere, and restoration can exploit residual spectral responses, spatial smoothness, or physical cloud-transmission priors. Classical methods have therefore used inpainting, filtering, physical degradation modeling, and sparse/low-rank constraints to recover degraded pixels [4,5,6]. Deep learning methods further improve the nonlinear reconstruction ability of single-image cloud removal by learning spatial, spectral, and contextual representations from data [7,8,9,10]. Nevertheless, thick-cloud removal is more challenging because thick clouds may fully block the optical reflectance of the underlying surface. In this case, the contaminated optical image alone contains insufficient information for reliable reconstruction.
To compensate for the missing surface information, recent studies have introduced auxiliary observations. Multi-temporal optical methods recover cloudy regions by borrowing information from cloud-free observations of the same area acquired at other times [11,12,13,14]. These methods are effective when suitable temporal references are available, but their performance can be affected by land-cover changes, seasonal variation, illumination differences, and registration errors. Synthetic aperture radar (SAR) provides another valuable auxiliary modality because microwave signals can penetrate clouds and are less dependent on solar illumination. Early and deep learning-based studies have shown that SAR observations can provide structural and textural cues for optically obscured regions, making SAR–optical fusion a promising route for thick-cloud removal [15,16,17].
Despite this advantage, SAR–optical thick-cloud removal is still difficult because SAR and optical images are produced by fundamentally different imaging mechanisms. SAR images record microwave backscattering responses related to geometry, roughness, moisture, and acquisition configuration, whereas optical images mainly describe surface spectral reflectance in visible and infrared bands. Directly fusing these heterogeneous signals may introduce SAR-induced artifacts, spectral distortion, or structural inconsistency, especially when the network does not explicitly distinguish cloudy and cloud-free regions [18,19]. Cloud-covered areas require stronger SAR-derived compensation, while clear areas should preserve reliable optical observations. Therefore, an effective thick-cloud removal network should reduce the SAR–optical modality gap before fusion and should perform region-aware information selection during fusion.
SAR-to-optical translation provides a natural way to reduce this modality gap. Instead of directly injecting SAR features into optical reconstruction, translation-based methods first synthesize an optical-like representation from SAR and then use it as a more compatible cue for cloud removal [20,21]. However, SAR-to-optical translation is intrinsically ill-posed: a translated image may preserve geometric structures but produce unrealistic spectral relationships, or it may show plausible color appearance while blurring edges and contours. In addition, many fusion modules treat the whole image with a global strategy and therefore fail to adaptively balance translated SAR-derived information and original optical information according to cloud-region reliability. Accordingly, the distinction of TFCRNet does not arise from translation, adversarial learning, or attention in isolation, but from coupling spectral–structural adversarial supervision with cloud-mask-guided bidirectional cross-attention within a unified two-stage framework.
To address these issues, we propose TFCRNet, a two-stage SAR–optical thick-cloud removal framework that combines SAR-to-optical translation and cloud-aware fusion. In the first stage, a Multi-Scale Feature Fusion Generator (MSFFG) converts SAR imagery into an optical-like representation. To better constrain the translated image, we design a dual-discriminator adversarial structure composed of a Spectral Discriminator (SpeD) and a Structural Discriminator (StrD), which respectively emphasize spectral fidelity and structural integrity. In the second stage, a Region-Gated Cross-Attention Fusion (RGCAF) module uses the cloud mask as a spatial reliability prior to adaptively fuse the cloudy optical image and the translated optical image. In this way, TFCRNet jointly addresses two limitations in SAR–optical translation and fusion: the entanglement of spectral and structural errors during translation and the insufficient differentiation of information use between cloud-covered and cloud-free regions.
The main contributions of this work are summarized as follows:
  • We propose TFCRNet, a two-stage SAR–optical thick-cloud removal framework that couples SAR-to-optical modality alignment with cloud-aware optical reconstruction. Rather than relying on translation or fusion alone, TFCRNet integrates targeted translation constraints and region-adaptive information selection within a unified framework.
  • We design a Multi-Scale Feature Fusion Generator with a dual-discriminator adversarial constraint. SpeD models local and global spectral relationships, whereas StrD emphasizes gradient-defined edges and contours, providing complementary supervision for SAR-to-optical translation.
  • We introduce a Region-Gated Cross-Attention Fusion module that uses cloud masks to perform region-adaptive cross-modal interaction, thereby enhancing reconstruction in cloud-covered areas while reducing disturbance in cloud-free areas.
  • Experiments on SEN12MS-CR and SMILE-CR show that TFCRNet achieves competitive performance and outperforms the reproduced baselines under the unified protocol. Ablation results further verify the complementary effects of SpeD, StrD, and RGCAF.

2. Related Work

2.1. Optical Remote Sensing Image Cloud Removal

Optical remote sensing image cloud removal has been studied from traditional model-based restoration to recent deep learning reconstruction. Classical methods usually assume that the clean surface signal can be estimated from spatial continuity, spectral redundancy, physical transmission models, or low-rank structures. These assumptions are useful for thin clouds, small missing regions, or scenes with strong spatial regularity, but they become less reliable when thick clouds fully obscure the land surface. Deep networks alleviate this limitation by learning data-driven mappings from contaminated observations to clean targets, but single-image cloud removal remains inherently under-constrained when the target region contains little observable information. Related hyperspectral image analysis also shows that locally adaptive irregular filtering can suppress complex background interference while retaining informative spectral deviations [22].
Recent image restoration research provides useful technical foundations for cloud removal. Diffusion-based and low-rank tensor models have been explored to improve the reconstruction of missing or uncertain regions in remote sensing imagery [23,24,25]. More general restoration and generation studies also show that perceptual losses, adversarial learning, denoising priors, and contextual inpainting can improve the visual and structural quality of restored images [26,27,28,29]. Frequency-aware remote sensing networks further demonstrate that explicit frequency decoupling and frequency-guided denoising can preserve high-frequency semantic details while suppressing representation noise [30,31]. Edge- and structure-aware inpainting methods further indicate that explicit boundary guidance is helpful for recovering missing content [32,33]. Transformer- and efficient CNN-based restoration architectures enhance long-range dependency modeling and local reconstruction ability, which are both important for large-area cloud removal [34,35,36,37]. These studies motivate stronger representation modeling, but they do not fully solve the information-missing problem caused by optically thick clouds.

2.2. SAR–Optical Fusion for Thick-Cloud Removal

SAR–optical fusion has become an important solution for thick-cloud removal because SAR can provide cloud-penetrating structural information. Existing methods improve multimodal reconstruction from different perspectives. Graph-based feature aggregation, transformer modeling, two-flow encoding, and unified spatial–spectral residual learning have been used to enhance SAR–optical interaction [38,39,40,41]. Other studies emphasize global–local fusion, hierarchical spectral–structural preservation, and heterogeneous parallel encoding to better balance optical spectral details and SAR structural cues [42]. More broadly, dual-domain decoupled fusion and frequency-domain-enhanced spectral–spatial fusion have recently been used to strengthen complementary representation interaction in remote sensing imagery [43,44]. These methods demonstrate the effectiveness of SAR assistance when optical observations are severely contaminated.
Recent studies have further improved cloud removal through diffusion modeling, efficient Transformer learning, and auxiliary-data-driven codebook pretraining. EMRDM [45] formulates cloud removal as a mean-reverting diffusion process between cloudy and cloud-free images and improves the denoising and sampling procedures within an elucidated diffusion design space. ECRformer [46] introduces semantic-decoupled feature learning and efficient attention modules to separately optimize structural recovery and texture rendering in multimodal cloud removal. TCF-VQGAN [47] learns a discrete multimodal codebook through two-stage training and incorporates additional MODIS observations and unpaired auxiliary data to alleviate limited paired supervision. These methods primarily improve reconstruction through stronger generative priors, semantic feature decoupling, or auxiliary-data pretraining. In contrast, TFCRNet follows an explicit modality-alignment and region-aware reconstruction strategy. It first maps SAR observations into an optical-like domain, uses SpeD and StrD to separately constrain spectral and structural translation errors, and then employs RGCAF to regulate two direction-specific cross-attention responses according to the spatial reliability indicated by the cloud mask.
However, direct SAR–optical fusion still has two limitations. First, the large modality gap makes it difficult for a unified network to decide which SAR patterns are useful for optical reconstruction and which patterns are modality-specific noise. Second, the reliability of the optical input is spatially nonuniform. Cloudy regions should rely more on SAR-derived complementary information, whereas cloud-free regions should mainly retain the original optical observation. If the fusion strategy does not explicitly consider this regional difference, it may either underuse SAR information in cloud-covered areas or overuse it in clear areas. This motivates a translation-before-fusion design and a cloud-region-aware fusion mechanism.

2.3. SAR-to-Optical Image Translation and Image-to-Image Generation

SAR-to-optical image translation aims to map SAR observations into an optical-like domain. The development of GANs and image-to-image translation frameworks provides the methodological basis for this problem [48,49,50]. In remote sensing, attention-guided and GAN-based SAR-assisted translation methods have shown that SAR-derived structures can be converted into optical-like cues for cloud removal [51,52]. Cloud-guided translation-and-fusion methods further indicate that translated SAR information can be more useful when cloud masks are used to guide the final reconstruction process [53].
Although translation-before-fusion reduces the modality gap, SAR-to-optical synthesis remains an ill-posed task because SAR backscatter does not uniquely determine optical spectra. Existing translation networks often use a single adversarial discriminator or a unified reconstruction constraint, which may mix spectral and structural error modes during optimization. For thick-cloud removal, this is problematic because spectral distortion and structural artifacts have different effects on the final reconstruction. Therefore, TFCRNet separates adversarial supervision into spectral and structural discrimination, allowing the generator to receive more targeted guidance during SAR-to-optical translation.

2.4. Attention and Region-Aware Fusion in Image Restoration

Attention mechanisms have been widely used to improve feature selection and dependency modeling. Channel attention, spatial attention, non-local attention, and transformer self-attention allow networks to recalibrate informative responses and capture long-range context [54,55,56,57]. Multi-stage restoration networks and efficient transformer designs further show that combining local detail recovery with global context modeling is beneficial for image restoration [58,59]. Geometric-prior-guided refinement, Euclidean-affinity-augmented hyperbolic modeling, and position-aware differential denoising also illustrate complementary strategies for improving contextual reasoning and suppressing ambiguous responses in high-resolution remote sensing representations [60,61,62]. Modern vision backbones such as ViT, Swin Transformer, and ConvNeXt also demonstrate the importance of hierarchical representation learning and global–local feature interaction [63,64,65]. A recent dual-scale transformer for neural video compression similarly shows that jointly capturing global structure and local texture can improve compact visual representation learning [66].
For cloud removal, attention should not only model feature dependencies but also respect regional reliability. Cloud-covered regions, cloud-shadow regions, and cloud-free regions provide different levels of trustworthy optical information. A cloud mask can therefore serve as an explicit spatial prior for deciding where SAR-derived information should be emphasized and where original optical information should be preserved. In TFCRNet, the cloud mask is embedded into the computation of the region gate rather than being used only as an additional input channel or a reconstruction loss weight. The resulting gate weights two direction-specific cross-attention responses, enabling translated SAR-derived information and original optical information to be emphasized in different regions.

3. Materials and Methods

3.1. Problem Formulation

Given a cloudy optical image I c , a corresponding SAR image I s , and a cloud mask M, the goal of SAR–optical thick-cloud removal is to reconstruct a cloud-free optical image I o that is close to the ground-truth cloud-free image I g t . The cloud mask M is an externally supplied binary input that indicates cloud-contaminated regions and serves as a spatial reliability prior during fusion. It is assumed to be available during both training and inference, while cloud-mask prediction itself is outside the scope of TFCRNet. The overall mapping can be formulated as
I o = F ( I c , I s , M ) ,
where F ( · ) denotes the proposed cloud removal network.
TFCRNet decomposes this mapping into two stages. In the translation stage, the SAR image is converted into an optical-like image:
I t = G ( I s ) ,
where G ( · ) denotes the SAR-to-optical translation generator and I t denotes the translated optical image. In the fusion stage, the translated optical image, cloudy optical image, and cloud mask are jointly used to produce the final cloud-free result:
I o = R ( I c , I t , M ) ,
where R ( · ) denotes the cloud-aware fusion network.
This two-stage formulation is designed to first reduce the cross-modal discrepancy between SAR and optical imagery and then perform region-adaptive reconstruction. The translated optical image provides SAR-derived structural priors in an optical-like space, while the cloudy optical image preserves reliable spectral and spatial details in cloud-free regions.

3.2. Overall Architecture of TFCRNet

The overall architecture of TFCRNet is shown in Figure 1. The network consists of two main stages: SAR-to-optical translation and fusion-based cloud removal. Its encoder–decoder and attention components are inspired by representative deep image modeling architectures, including U-Net and ResNet [67,68], as well as Transformer-style and modern ConvNet backbones. The restoration-oriented design is also related to progressive and attention-based image restoration networks.
In the translation stage, the SAR image is fed into the proposed Multi-Scale Feature Fusion Generator (MSFFG). The generator extracts SAR representations at multiple spatial scales and adaptively fuses them to synthesize an initial optical-like image. To improve the quality of the translated optical image, two discriminators are introduced during adversarial learning. The Spectral Discriminator (SpeD) focuses on spectral consistency, while the Structural Discriminator (StrD) focuses on edge and geometric integrity. Both discriminators are used only to provide adversarial supervision during training and are removed from the inference path. Therefore, the deployed TFCRNet consists of the translation generator and the fusion-based cloud-removal network, without the additional inference cost of SpeD and StrD.
In the fusion stage, the translated optical image and the cloudy optical image are separately encoded by ConvNeXt-based feature extractors. Multi-level features are then fused by the proposed Region-Gated Cross-Attention Fusion (RGCAF) module. The cloud mask is used as a spatial reliability prior to guide different fusion preferences in cloudy and cloud-free regions. Finally, a decoder progressively restores the spatial resolution and outputs the final cloud-free optical image.

3.3. Multi-Scale Feature Fusion Generator

The SAR-to-optical translation stage aims to generate an optical-like image from the SAR input. Since SAR imagery contains structural information at different spatial scales, a single-scale generator may be insufficient to capture both global layout and fine local textures. Therefore, we design the Multi-Scale Feature Fusion Generator (MSFFG), as shown in Figure 2.
Given the SAR image I s , we first construct multi-scale inputs through downsampling and obtain three scales of SAR representations. Each scale is processed by convolution, normalization, and nonlinear activation to extract shallow features. These features are then sent to the Multi-Scale Feature Fusion (MSFF) module for adaptive integration. After feature fusion, several residual blocks are used to extract deeper representations. Finally, three parallel decoding paths recover the translated optical image at the original resolution. The multi-branch decoder helps preserve global structure and reconstruct fine details simultaneously.

Multi-Scale Feature Fusion Module

The structure of the MSFF module is shown in Figure 3. The MSFF module is designed to adaptively integrate features from different scales. Let F 1 , F 2 , and F 3 denote the feature maps extracted from three scales. Before fusion, all feature maps are aligned to the same spatial resolution through projection or resampling operations. The concatenated feature is obtained as
F c = Concat P 1 ( F 1 ) , P 2 ( F 2 ) , P 3 ( F 3 ) ,
where P 1 ( · ) , P 2 ( · ) , and P 3 ( · ) denote resolution- and channel-alignment operations.
To predict scale- and channel-adaptive fusion weights, F c is processed by a lightweight gating branch composed of 1 × 1 convolutions, a depthwise separable 3 × 3 convolution, and a sigmoid activation:
W m = σ Conv 1 × 1 DWConv 3 × 3 Conv 1 × 1 Conv 1 × 1 ( F c ) ,
where DWConv 3 × 3 ( · ) denotes depthwise separable convolution and σ ( · ) denotes the sigmoid function. The output of MSFF is then computed through element-wise reweighting and residual aggregation:
F m s f f = F c + F c W m .
This design avoids directly summing features from different scales and enables the generator to emphasize informative structural and textural responses. In this way, high-level features provide semantic and structural guidance, while low-level features contribute spatial positions and fine details.

3.4. Dual-Discriminator Adversarial Constraint

Conventional adversarial translation networks usually use a single discriminator to distinguish real and generated images. However, SAR-to-optical translation involves two different quality dimensions: spectral fidelity and structural integrity. A generated image may have plausible spectral appearance but distorted structures, or clear structures but unrealistic spectral relationships. To address this issue, we design a dual-discriminator structure (Dual-D) consisting of a Spectral Discriminator (SpeD) and a Structural Discriminator (StrD).
Both discriminators are built upon PatchGAN [49], which performs local patch-level discrimination and outputs a patch-level authenticity map rather than a single scalar. However, SpeD and StrD operate on different feature representations before PatchGAN discrimination. SpeD processes features enhanced through local spectral aggregation and channel-oriented linear attention, whereas StrD processes Sobel-gradient-enhanced structural features. Therefore, SpeD emphasizes cross-band radiometric relationships, while StrD emphasizes edges, contours, and local geometric continuity. The general PatchGAN architecture used in TFCRNet is shown in Figure 4.

3.4.1. Spectral Discriminator

The Spectral Discriminator (SpeD) is designed to constrain the spectral consistency of the translated optical image. Both real and translated samples are supplied to SpeD using all multispectral bands retained under the corresponding dataset configuration; no RGB-only band selection is applied before spectral feature extraction. Since SAR backscatter does not uniquely determine optical reflectance, a translated image may exhibit plausible spatial content while containing unrealistic relationships among optical bands. To address this problem, the input image is processed by a spectral feature extraction module before PatchGAN discrimination, as shown in Figure 5. The module contains two complementary branches. The local branch captures local spectral responses through pooling and convolution, while the global branch models long-range channel dependencies through linear attention. This design is related to channel and spatial attention mechanisms that have been widely used for adaptive feature recalibration and dependency modeling [54,55,56].
Given an input feature F, layer normalization is first applied [69]:
F n = Norm ( F ) .
The local spectral branch is defined as
F l = Conv ϕ Conv Pooling ( F n ) ,
where ϕ ( · ) denotes the activation function. If pooling changes the spatial resolution, the output is resized to match the resolution of F n before aggregation. The global spectral dependency branch is formulated as
F a = LinearAtt ( F n ) ,
where LinearAtt ( · ) denotes linear attention along the channel dimension. The two branch outputs are aggregated with the normalized input to obtain the intermediate spectral feature:
F m i d = F n + F l + F a .
A second residual block further enhances the feature:
F s p e = F m i d + ϕ Norm Conv ( F m i d ) .
The enhanced spectral feature F s p e is then sent to PatchGAN for adversarial discrimination. The local branch models neighborhood-level spectral responses, while the linear-attention branch captures longer-range dependencies among optical bands. Their combination allows SpeD to preserve the spatial layout of the input while emphasizing cross-band consistency, thereby penalizing unrealistic spectral relationships and radiometric distortion rather than explicitly focusing on edge information.

3.4.2. Structural Discriminator

As shown in Figure 5, the Structural Discriminator (StrD) focuses on edge, contour, and geometric structure. Unlike SpeD, StrD does not use spectral pooling or channel-oriented attention. Instead, the structural feature extraction module applies fixed Sobel operators along the horizontal and vertical directions to compute image gradients. For multi-channel inputs, the Sobel operation is applied channel-wise, and the resulting gradient responses are aggregated or projected to form an edge-aware attention map. The Sobel kernels are defined as
S x = 1 0 1 2 0 2 1 0 1 , S y = 1 2 1 0 0 0 1 2 1 .
Given an input feature F, the horizontal and vertical gradient responses are computed as
G x = S x F , G y = S y F ,
where ∗ denotes convolution. The edge-aware attention map is then obtained as
A e d g e = σ Conv 1 × 1 G x G x + G y G y + ϵ g ,
where ϵ g is a small constant for numerical stability, and the 1 × 1 convolution adjusts the channel dimension of the gradient response. The structural feature is obtained by applying the edge-aware attention map to the original feature and then aggregating it with convolution:
F s t r = Conv F A e d g e .
The structural feature F s t r is then input into PatchGAN. Because the discriminator operates on gradient-enhanced features, StrD mainly penalizes blurred boundaries, contour discontinuities, and local geometric deformation. This structural supervision is complementary to the radiometric supervision provided by SpeD.

3.5. Region-Gated Cross-Attention Fusion Module

After SAR-to-optical translation, the translated optical image I t provides useful SAR-derived structural information for cloud-covered regions, while the original cloudy optical image I c still contains reliable observations in cloud-free regions. Therefore, the fusion stage should treat cloudy and cloud-free regions differently. To this end, we propose the Region-Gated Cross-Attention Fusion (RGCAF) module, as shown in Figure 6. The module uses the cloud mask to generate a cloud-region gating map and then performs bidirectional cross-attention between cloudy optical features and translated optical features.
Given the externally supplied binary cloud mask M, we resize it to the spatial resolution of the current feature level using nearest-neighbor interpolation and denote the resized mask as M l . Nearest-neighbor interpolation is adopted to preserve the discrete cloud and cloud-free labels and to avoid introducing intermediate values during feature-level resizing. Therefore, M l remains a binary mask. RGCAF subsequently transforms M l through two convolution–batch-normalization–activation units [70] and a sigmoid function to generate the learnable soft cloud-region gating map:
W c = σ Conv 1 × 1 ϕ BN Conv 1 × 1 ( M l ) ,
where W c is a learned soft gating map with values between zero and one, whereas the resized input mask M l remains binary. The complementary gate 1 W c is used to weight the cloud-free-region response.
Let F c l o u d y and F t r a n s denote the cloudy optical feature and the translated optical feature, respectively. After layer normalization, query, key, and value projections are computed as
Q c , K c , V c = Proj LN ( F c l o u d y ) , Q t , K t , V t = Proj LN ( F t r a n s ) ,
where Proj ( · ) denotes linear or convolutional projection followed by reshaping to attention heads.
The two cross-attention responses are computed in a bidirectional manner:
A c t = Softmax Q c K t T d V t ,
A t c = Softmax Q t K c T d V c ,
where d is the feature dimension. Here, A c t denotes the response that uses the cloudy optical feature as query and the translated optical feature as key/value. It injects translated SAR-derived information into the cloudy optical branch and is therefore more suitable for cloud-covered regions. In contrast, A t c uses the translated optical feature as query and the cloudy optical feature as key/value. It retrieves information from the original optical branch and is therefore more suitable for preserving reliable cloud-free regions.
Finally, the region gating map is used to integrate the two attention responses:
F o u t = W c A c t + ( 1 W c ) A t c .
This formulation ensures that cloud-covered regions rely more strongly on translated SAR-derived cues, while cloud-free regions preserve information from the original optical observation. In implementation, W c is broadcast along the channel dimension to match the shape of the attention responses.

3.6. Loss Functions

The optimization of TFCRNet contains two types of objectives: the generator-side objective used to update the SAR-to-optical generator and the fusion network, and the discriminator-side objective used to update the two adversarial discriminators. The generator-side objective is defined as
L G = λ t L t r a n s + λ f L f u s i o n ,
where λ t and λ f are set to 0.5 and 0.5 , respectively.

3.6.1. Translation-Stage Loss

The translation-stage loss is used to guide SAR-to-optical synthesis and consists of a generator adversarial loss, a spectral angle loss, and a structural similarity loss:
L t r a n s = L G , a d v + λ s a m L S A M + λ s s i m L S S I M .
The adversarial supervision is provided by SpeD and StrD. For a discriminator D j { D s p e , D s t r } , the discriminator loss is written as
L D j = E I g t p d a t a log D j ( I g t ) E I s p s log 1 D j ( G ( I s ) ) ,
and the total discriminator loss is
L D = λ D s p e L D s p e + λ D s t r L D s t r ,
where λ D s p e and λ D s t r control the contributions of the Spectral and Structural Discriminator losses, respectively. Both weights are set to 1.0 in our implementation. For each discriminator D j { D s p e , D s t r } , the corresponding generator adversarial loss is defined using the non-saturating objective:
L G , a d v j = E I s p s log D j G ( I s ) .
The total generator adversarial loss is then formulated as
L G , a d v = λ G s p e L G , a d v s p e + λ G s t r L G , a d v s t r ,
where λ G s p e and λ G s t r control the contributions of the spectral and structural adversarial signals received by the generator, respectively. Both weights are set to 1.0 in the default configuration. This formulation explicitly separates the generator optimization objective from the discriminator objective while allowing the two adversarial branches to be independently weighted.
In our implementation, the loss weights are set as λ s a m = 1.0 and λ s s i m = 1.0 . The SAM loss measures the spectral angle between the translated optical image and the ground truth:
L S A M = 1 H W i = 1 H W arccos clip I t ( i ) · I g t ( i ) I t ( i ) 2 I g t ( i ) 2 + ϵ s , 1 + δ , 1 δ ,
where ϵ s and δ are small constants for numerical stability. The clipping operation prevents invalid values caused by floating-point errors.
The SSIM loss is used to enhance structural similarity:
L S S I M = 1 SSIM ( I t , I g t ) .

3.6.2. Fusion-Stage Loss

For the fusion stage, we use a cloud-mask-guided weighted Charbonnier loss. Let e i = I o ( i ) I g t ( i ) denote the reconstruction error at pixel i. For multi-band optical images, e i is a spectral vector. The fusion-stage loss is defined as
L f u s i o n = 1 H W i = 1 H W 1 + λ m M ( i ) e i 2 2 + ϵ 2 ,
where ϵ is set to 10 3 and λ m controls the relative importance of cloud-covered regions. In our implementation, λ m is set to 1.0 , which assigns larger weights to cloud-covered pixels while still constraining cloud-free regions. This loss improves cloudy-region reconstruction and helps preserve clear-region fidelity.

4. Results

4.1. Datasets

4.1.1. SEN12MS-CR

SEN12MS-CR is a large-scale global all-season SAR–optical cloud removal dataset [17]. It contains triplets composed of cloudy Sentinel-2 multispectral images, corresponding cloud-free Sentinel-2 reference images, and Sentinel-1 SAR images. The dataset contains globally distributed, non-overlapping regions of interest across different continents and seasons, and each region is cropped into image patches for deep learning. We follow the released ROI-level partition rather than randomly splitting the cropped patches. Patches derived from the same ROI are assigned to only one subset and therefore do not appear across the training, validation, and testing sets. After patch extraction and preprocessing, the adopted split contains 107,143 training samples, 7176 validation samples, and 7899 testing samples. Because the cloudy and cloud-free Sentinel-2 observations may originate from different acquisition dates, their pixel-wise agreement may also be affected by land-cover change, phenological variation, illumination differences, and residual registration errors.

4.1.2. SMILE-CR

SMILE-CR is a multimodal cloud removal benchmark built for Landsat-8 optical imagery and Sentinel-1 SAR imagery [71]. It contains 1400 paired samples collected from globally distributed regions with diverse land-cover types. Each sample consists of a cloudy Landsat-8 multispectral image, a corresponding Sentinel-1 SAR image, a cloud-free Landsat-8 reference image, and a cloud mask. The cloudy observation is generated by transferring realistic cloud and shadow effects to cloud-free Landsat-8 imagery, which reduces the temporal land-cover discrepancy between the contaminated observation and its reference. We use a fixed paired-sample-level split at a ratio of 5:1:1, resulting in 1000 training, 200 validation, and 200 testing samples. This protocol is treated as a controlled same-dataset evaluation and is not presented as a geographically independent generalization test.
The cloud-free references of the two datasets should be interpreted according to their construction protocols. SEN12MS-CR may contain temporal, seasonal, illumination, and registration differences between the cloudy observation and the cloud-free reference, whereas SMILE-CR reduces temporal-reference inconsistency by generating cloud contamination from cloud-free optical imagery. Consequently, the reported PSNR, SSIM, SAM, and MAE values quantify agreement with the dataset-provided reference rather than exact recovery of an unknown simultaneous surface state.

4.2. Implementation Details

Experiments are implemented using PyTorch 1.13.1 on Ubuntu 20.04 LTS with an NVIDIA A40 GPU (NVIDIA Corporation, Santa Clara, CA, USA). The Adam optimizer is used for training, with β 1 = 0.5 and β 2 = 0.999 [72]. The batch size is set to 4. The initial learning rate of the generator is set to 1 × 10 4 , and the discriminator learning rate is set to 1 × 10 5 . The weight decay is set to 1 × 10 4 . The input image size is 256 × 256 , and the maximum number of training epochs is 30. For controlled comparison, all reproduced methods are evaluated using the dataset partitions described above and the same implementations of PSNR, SSIM, SAM, and MAE. For methods with publicly available official implementations, we retain their recommended optimizer, learning-rate schedule, loss configuration, and training strategy, while adapting only the input and output interfaces required by SEN12MS-CR and SMILE-CR. The checkpoint with the highest validation PSNR is selected for final testing, and the reported metrics are calculated on the corresponding test split.

4.3. Evaluation Metrics

We use four commonly adopted metrics to evaluate cloud removal quality: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Spectral Angle Mapper (SAM), and Mean Absolute Error (MAE). PSNR measures pixel-level fidelity, SSIM evaluates structural similarity following the standard structural similarity formulation [73], SAM measures spectral consistency, and MAE reflects the absolute reconstruction error. Higher PSNR and SSIM values indicate better performance, whereas lower SAM and MAE values indicate better reconstruction quality. All quantitative metrics reported in this study are calculated over the complete image. Therefore, PSNR, SSIM, SAM, and MAE should not alone be interpreted as separate quantitative measurements of reconstruction in cloud-covered regions and preservation in cloud-free regions. The region-dependent behavior of TFCRNet is assessed jointly through the qualitative comparisons, the RGCAF ablation, and the cloud-region-weighted loss analysis.

4.4. Comparison Methods

We compare TFCRNet with five representative SAR–optical cloud removal methods. SAR2OPT uses conditional GANs to synthesize optical images from SAR inputs [20]. DSen2-CR uses a deep residual neural network for Sentinel-2 cloud removal with SAR–optical data fusion [16]. GLF-CR performs SAR-enhanced cloud removal through global–local fusion [18]. HS2P introduces hierarchical spectral and structure-preserving fusion for multimodal cloud and shadow removal [19]. HPN-CR designs a heterogeneous parallel network for SAR–optical data fusion cloud removal [42]. All reproduced baselines are evaluated using unified dataset partitions, test samples, and metric implementations. Because the compared methods differ in architecture and original optimization objectives, their recommended method-specific optimizers, loss functions, and training schedules are retained rather than being replaced by a single training configuration.

4.5. Quantitative Comparison on SEN12MS-CR

Table 1 reports the controlled quantitative comparison on SEN12MS-CR. Among the baseline methods reproduced using the unified dataset partition, test samples, and metric implementations adopted in this study, TFCRNet achieves the strongest performance across all four metrics, reaching 33.0115 dB PSNR, 0.9085 SSIM, 6.2795 SAM, and 0.0125 MAE. Compared with the reproduced SAR2OPT baseline, TFCRNet substantially improves both structural and spectral reconstruction quality, indicating that SAR-only translation is insufficient for high-fidelity optical reconstruction. TFCRNet also outperforms the reproduced DSen2-CR, GLF-CR, HS2P, and HPN-CR baselines under the same controlled evaluation protocol. These results establish the relative effectiveness of TFCRNet within the reproduced comparison set and should not be interpreted as a claim of universal state-of-the-art performance across results reported under different experimental settings.
The visual comparison on SEN12MS-CR is shown in Figure 7. SAR2OPT tends to generate optical images whose structures are close to SAR backscatter patterns but whose spectral appearance deviates from the real optical image. DSen2-CR improves structural reconstruction but may introduce redundant artifacts under dense cloud coverage. GLF-CR reduces some artifacts but still suffers from local blurring. HS2P improves multimodal feature selection but may produce color distortion in large cloud-covered regions. HPN-CR performs well among the baselines, but its global feature interaction may still disturb cloud-free regions. In contrast, TFCRNet removes thick clouds more completely and preserves the spectral appearance of cloud-free regions more effectively.

4.6. Quantitative Comparison on SMILE-CR

Table 2 reports the controlled quantitative comparison on SMILE-CR. Among the baseline methods reproduced under the unified evaluation protocol adopted in this study, TFCRNet obtains the strongest results across all four metrics, with 30.5291 dB PSNR, 0.8813 SSIM, 6.9864 SAM, and 0.0217 MAE. Compared with the reproduced SAR2OPT baseline, TFCRNet improves PSNR by 12.1707 dB and reduces SAM by 20.9835, demonstrating the importance of combining translated SAR-derived cues with the available optical observations. Compared with the reproduced HPN-CR baseline, TFCRNet improves PSNR by 1.3846 dB and further reduces MAE and SAM. These comparisons are restricted to the methods reproduced within our experimental framework and are not intended as a direct ranking against literature-reported results obtained under different dataset and implementation settings.
The visual comparison on SMILE-CR is presented in Figure 8. SAR2OPT suffers from severe color deviation because it only uses SAR information. DSen2-CR and GLF-CR introduce SAR–optical fusion but still generate redundant artifacts or local blurring under complex cloud conditions. HS2P may leave cloud-shadow-like residuals in some regions. HPN-CR improves feature interaction but lacks explicit region-aware reliability modeling. TFCRNet produces fewer artifacts and better preserves land-cover boundaries, showing stronger generalization on another sensor and dataset setting.

4.7. Ablation Study

To verify the effectiveness of the proposed components, we first conduct a component-level ablation of MSFF, Dual-D, and RGCAF on SEN12MS-CR, as reported in Table 3. Removing MSFF reduces PSNR from 33.0115 dB to 31.7293 dB because the generator cannot effectively integrate multi-scale structural and semantic information. Removing Dual-D causes the largest degradation in PSNR, SSIM, and SAM, indicating the importance of adversarial spectral and structural constraints. Removing RGCAF increases MAE from 0.0125 to 0.0242, showing that region-adaptive fusion is particularly important for reducing reconstruction errors. To further identify the source of these improvements, we separately analyze SpeD and StrD, cloud-mask guidance, and the individual loss terms in the following experiments.
Figure 9 shows the qualitative ablation results. Without MSFF, the reconstructed image tends to show blurred boundaries and weakened fine details. Without the dual-discriminator design, the translated optical image quality decreases, leading to spectral distortion and structural artifacts in the final result. Without RGCAF, the network cannot explicitly distinguish cloudy and cloud-free regions during fusion, which may cause incomplete reconstruction in cloud-covered areas and unnecessary disturbance in clear regions. The complete TFCRNet achieves the best visual quality, with clearer structures, fewer artifacts, and better spectral consistency.
To examine the individual contributions of the two discriminators, we compare four adversarial configurations: a conventional single PatchGAN, SpeD only, StrD only, and the complete SpeD + StrD design. In the single-PatchGAN setting, the spectral and structural feature extraction modules are removed, and the translated optical image is directly evaluated by one PatchGAN discriminator. In the SpeD-only and StrD-only settings, only the corresponding discriminator branch is retained. The generator, fusion network, dataset split, remaining loss terms, and training settings are kept unchanged to ensure a controlled comparison.
As shown in Table 4, both specialized single-discriminator configurations outperform the conventional single PatchGAN, indicating that the task-specific spectral and structural feature transformations provide useful adversarial signals. SpeD only obtains a lower SAM than StrD only, suggesting that its spectral representation provides a stronger constraint on cross-band radiometric consistency, whereas StrD only achieves a higher SSIM, suggesting that gradient-enhanced discrimination contributes more directly to structural preservation. The complete SpeD + StrD configuration achieves the best performance among the four evaluated discriminator settings, improving PSNR by 0.9241 dB and SSIM by 0.0088 while reducing SAM by 2.3328 and MAE by 0.0064 compared with the single PatchGAN. These results support the differentiated and complementary behavior of SpeD and StrD. However, the complete configuration contains two specialized adversarial branches and therefore has greater adversarial capacity than the single-PatchGAN baseline. Because the current experiment does not include a parameter-matched two-PatchGAN or widened single-PatchGAN control, it cannot completely separate the effect of task-specific specialization from the effect of increased adversarial capacity.
We further evaluate the contributions of cloud-mask guidance and the individual training objectives. To remove cloud-mask guidance while retaining the bidirectional cross-attention structure, the learned region gate is replaced by a constant gate W c = 0.5 , such that the two cross-attention responses contribute equally at every spatial position. To isolate the effect of cloud-region reconstruction weighting, λ m is set to zero while the cloud mask is still used by RGCAF. The remaining settings independently remove L G , a d v , L S A M , and L S S I M while keeping the network architecture and other training objectives unchanged.
As shown in Table 5, replacing the learned region gate with W c = 0.5 reduces PSNR by 0.6251 dB and increases MAE by 0.0054, demonstrating that the effectiveness of RGCAF depends not only on bidirectional cross-attention but also on cloud-mask-guided spatial routing. Removing cloud-region weighting causes a smaller but consistent degradation, indicating that the weighted reconstruction term provides an additional optimization emphasis on contaminated regions. Among the individual loss terms, removing L G , a d v produces the largest overall degradation, removing L S A M causes the most pronounced increase in SAM, and removing L S S I M leads to the largest reduction in SSIM. The complete objective therefore provides the best balance among reconstruction accuracy, spectral consistency, and structural preservation.

5. Discussion

5.1. Effectiveness of Translation-Before-Fusion

The experimental results show that directly relying on SAR-to-optical translation alone is insufficient for high-fidelity thick-cloud removal, as demonstrated by the relatively poor performance of SAR2OPT. Nevertheless, translation remains valuable when used as an intermediate alignment step. SAR images contain cloud-penetrating structural information, but their radiometric properties differ significantly from optical images. By first translating SAR imagery into an optical-like representation, TFCRNet reduces the cross-modal discrepancy and provides a more compatible structural prior for the subsequent fusion stage. This explains why the proposed two-stage framework substantially outperforms both SAR-only translation and direct fusion baselines. From the perspective of stage design, SAR2OPT provides a translation-only reference, whereas DSen2-CR, GLF-CR, HS2P, and HPN-CR represent direct SAR–optical fusion without the proposed translation-before-fusion pathway. Although these external methods are not exact internal architectural variants, they provide complementary comparisons for evaluating the motivation of reducing the modality discrepancy before multimodal fusion.

5.2. Effectiveness of Spectral–Structural Dual Discrimination

The fine-grained ablation results provide evidence for the differentiated roles of SpeD and StrD. The SpeD-only configuration obtains a lower SAM than the StrD-only configuration, suggesting that spectral feature enhancement provides a stronger constraint on cross-band radiometric consistency. In contrast, StrD only achieves a higher SSIM than SpeD only, suggesting that gradient-enhanced discrimination contributes more directly to structural preservation. The SpeD + StrD configuration achieves the best PSNR, SSIM, SAM, and MAE among the evaluated discriminator settings, supporting the use of differentiated spectral and structural adversarial pathways. Nevertheless, SpeD + StrD contains two specialized discriminator branches and is not parameter-matched to the conventional single-PatchGAN baseline. Therefore, the current results support, but do not conclusively prove, that the complete improvement arises exclusively from task-specific discriminator specialization.

5.3. Effectiveness of Region-Gated Cross-Attention Fusion

The proposed RGCAF module is motivated by the different reliability of cloudy and cloud-free regions. The original component-level ablation shows that removing RGCAF increases MAE from 0.0125 to 0.0242. The finer-grained results in Table 5 further show that retaining bidirectional cross-attention but replacing the learned mask-derived gate with W c = 0.5 reduces PSNR by 0.6251 dB and increases MAE by 0.0054. Therefore, the improvement does not arise only from cross-attention; the cloud-mask-derived gate is also necessary for assigning translated SAR-derived information and reliable optical information to different regions. Removing only the cloud-region weighting produces a smaller degradation, indicating that mask-guided feature routing is the primary regional mechanism, while the weighted reconstruction loss provides an additional optimization constraint.

5.4. Limitations and Future Work

Although TFCRNet achieves strong performance on SEN12MS-CR and SMILE-CR, several limitations remain. First, TFCRNet requires an externally supplied cloud mask during inference, and the current experiments use the masks associated with the adopted datasets. Robustness to mask erosion, dilation, spatial shifts, omission and commission errors, or independently predicted masks has not been systematically evaluated. Future work will investigate controlled mask perturbations and integrate cloud detection and cloud removal into a unified framework. Second, the reported PSNR, SSIM, SAM, and MAE values are full-image metrics and do not separately quantify reconstruction in cloud-covered regions and preservation in cloud-free regions. Region-specific evaluation will therefore be included in future studies. Third, although SpeD and StrD are removed during inference, the two-stage procedure and dual-discriminator optimization increase training complexity. Lightweight translation modules or knowledge distillation could be explored to improve efficiency. Finally, SAR-to-optical translation remains an ill-posed problem, particularly in regions where SAR backscatter provides limited spectral cues. Incorporating temporal observations or uncertainty modeling may further improve reconstruction reliability.

6. Conclusions

In this paper, we proposed TFCRNet, a two-stage SAR–optical translation-and-fusion network for thick-cloud removal in remote sensing images. The method first translates SAR imagery into an optical-like representation using a Multi-Scale Feature Fusion Generator and a dual-discriminator adversarial constraint. The Spectral Discriminator and Structural Discriminator separately guide spectral fidelity and structural integrity. In the fusion stage, a Region-Gated Cross-Attention Fusion module uses cloud masks to perform region-adaptive cross-modal interaction, enabling the network to enhance translated SAR-derived cues in cloudy regions while preserving original optical details in cloud-free regions. Experiments on SEN12MS-CR and SMILE-CR show that TFCRNet outperforms representative SAR–optical cloud removal methods in PSNR, SSIM, SAM, and MAE. Ablation studies further verify the effectiveness of MSFF, the dual-discriminator design, and RGCAF. These results demonstrate that combining translation-before-fusion with cloud-aware region-gated interaction is effective for high-quality SAR–optical thick-cloud removal.

Author Contributions

Conceptualization, S.G., X.L. (Xin Lyu), W.X. and X.L. (Xin Li); methodology, S.G., C.S., C.W. and C.Z.; software, S.G., C.S., Z.X. and X.L. (Xin Li); validation, S.G., W.X., H.H. and D.A.; formal analysis, S.G., C.W., C.Z. and Z.X.; investigation, S.G., W.X., H.H., D.A. and X.L. (Xin Li); resources, X.L. (Xin Lyu), W.X., H.H. and D.A.; data curation, S.G., C.S., C.W. and Z.X.; writing—original draft preparation, S.G., C.S., X.L. (Xin Li) and Z.X.; writing—review and editing, W.X., X.L. (Xin Lyu), H.H., D.A. and C.W.; visualization, S.G., C.Z., C.S. and X.L. (Xin Li); supervision, X.L. (Xin Lyu), W.X., H.H. and D.A.; project administration, X.L. (Xin Lyu), W.X., C.Z. and C.W.; funding acquisition, X.L. (Xin Lyu), W.X., H.H. and D.A. All authors have read and agreed to the published version of the manuscript.

Funding

This work was funded in part by the National Key Research and Development Program of China under Grant No. 2024YFC3210800, National Natural Science Foundation of China (Grant No. 62401196), Postgraduate Research & Practice Innovation Program of Jiangsu Province (26CXJY1535), and Natural Science Foundation of Jiangsu Province (Grant No. BK20241508).

Informed Consent Statement

Not applicable.

Data Availability Statement

The SEN12MS-CR dataset is publicly available at https://patricktum.github.io/cloud_removal/sen12mscr/ (accessed on 8 May 2026). The SMILE-CR dataset is publicly available at https://www.kaggle.com/datasets/yuxiawhu/smile-cr (accessed on 8 May 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. King, M.D.; Platnick, S.; Menzel, W.P.; Ackerman, S.A.; Hubanks, P.A. Spatial and Temporal Distribution of Clouds Observed by MODIS Onboard the Terra and Aqua Satellites. IEEE Trans. Geosci. Remote Sens. 2013, 51, 3826–3852. [Google Scholar] [CrossRef] [Scilit]
  2. Zhu, Z.; Woodcock, C.E. Object-Based Cloud and Cloud Shadow Detection in Landsat Imagery. Remote Sens. Environ. 2012, 118, 83–94. [Google Scholar] [CrossRef] [Scilit]
  3. Qiu, S.; Zhu, Z.; He, B. Fmask 4.0: Improved Cloud and Cloud Shadow Detection in Landsats 4–8 and Sentinel-2 Imagery. Remote Sens. Environ. 2019, 231, 111205. [Google Scholar] [CrossRef] [Scilit]
  4. Maalouf, A.; Carré, P.; Augereau, B.; Fernandez-Maloigne, C. A Bandelet-Based Inpainting Technique for Clouds Removal from Remotely Sensed Images. IEEE Trans. Geosci. Remote Sens. 2009, 47, 2363–2371. [Google Scholar] [CrossRef] [Scilit]
  5. Shen, H.; Li, H.; Qian, Y.; Zhang, L.; Yuan, Q. An Effective Thin Cloud Removal Procedure for Visible Remote Sensing Images. ISPRS J. Photogramm. Remote Sens. 2014, 96, 224–235. [Google Scholar] [CrossRef] [Scilit]
  6. Li, J.; Wu, Z.; Hu, Z.; Zhang, J.; Li, M.; Mo, L.; Molinier, M. Thin Cloud Removal in Optical Remote Sensing Images Based on Generative Adversarial Networks and Physical Model of Cloud Distortion. ISPRS J. Photogramm. Remote Sens. 2020, 166, 373–389. [Google Scholar] [CrossRef] [Scilit]
  7. Zheng, J.; Liu, X.-Y.; Wang, X. Single Image Cloud Removal Using U-Net and Generative Adversarial Networks. IEEE Trans. Geosci. Remote Sens. 2021, 59, 6371–6385. [Google Scholar] [CrossRef] [Scilit]
  8. Tao, C.; Fu, S.; Qi, J.; Li, H. Thick Cloud Removal in Optical Remote Sensing Images Using a Texture Complexity Guided Self-Paced Learning Method. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5619612. [Google Scholar] [CrossRef] [Scilit]
  9. Xu, M.; Deng, F.; Jia, S.; Jia, X.; Plaza, A.J. Attention Mechanism-Based Generative Adversarial Networks for Cloud Removal in Landsat Images. Remote Sens. Environ. 2022, 271, 112902. [Google Scholar] [CrossRef] [Scilit]
  10. Yu, W.; Zhang, X.; Pun, M.-O. Cloud Removal in Optical Remote Sensing Imagery Using Multiscale Distortion-Aware Networks. IEEE Geosci. Remote Sens. Lett. 2022, 19, 5512605. [Google Scholar] [CrossRef] [Scilit]
  11. Sarukkai, V.; Jain, A.; Uzkent, B.; Ermon, S. Cloud Removal from Satellite Images using Spatiotemporal Generator Networks. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Snowmass, CO, USA, 1–5 March 2020; pp. 1796–1805. [Google Scholar]
  12. Stucker, C.; Garnot, V.S.F.; Schindler, K. U-TILISE: A Sequence-to-Sequence Model for Cloud Removal in Optical Satellite Time Series. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5408716. [Google Scholar] [CrossRef] [Scilit]
  13. Ebel, P.; Xu, Y.; Schmitt, M.; Zhu, X.X. SEN12MS-CR-TS: A Remote-Sensing Data Set for Multi-Modal Multi-Temporal Cloud Removal. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5222414. [Google Scholar] [CrossRef] [Scilit]
  14. Zheng, W.-J.; Zhao, X.-L.; Zheng, Y.-B.; Lin, J.; Zhuang, L.; Huang, T.-Z. Spatial-Spectral-Temporal Connective Tensor Network Decomposition for Thick Cloud Removal. ISPRS J. Photogramm. Remote Sens. 2023, 199, 182–194. [Google Scholar] [CrossRef] [Scilit]
  15. Eckardt, R.; Berger, C.; Thiel, C.; Schmullius, C. Removal of Optically Thick Clouds from Multi-Spectral Satellite Images Using Multi-Frequency SAR Data. Remote Sens. 2013, 5, 2973–3006. [Google Scholar] [CrossRef] [Scilit]
  16. Meraner, A.; Ebel, P.; Zhu, X.X.; Schmitt, M. Cloud Removal in Sentinel-2 Imagery Using a Deep Residual Neural Network and SAR-Optical Data Fusion. ISPRS J. Photogramm. Remote Sens. 2020, 166, 333–346. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  17. Ebel, P.; Meraner, A.; Schmitt, M.; Zhu, X.X. Multisensor Data Fusion for Cloud Removal in Global and All-Season Sentinel-2 Imagery. IEEE Trans. Geosci. Remote Sens. 2021, 59, 5866–5878. [Google Scholar] [CrossRef] [Scilit]
  18. Xu, F.; Shi, Y.; Ebel, P.; Zhu, X.X.; Schmitt, M. GLF-CR: SAR-Enhanced Cloud Removal with Global–Local Fusion. ISPRS J. Photogramm. Remote Sens. 2022, 192, 268–278. [Google Scholar] [CrossRef] [Scilit]
  19. Li, Y.; Wei, F.; Zhang, Y.; Chen, W.; Ma, J. HS2P: Hierarchical Spectral and Structure-Preserving Fusion Network for Multimodal Remote Sensing Image Cloud and Shadow Removal. Inf. Fusion 2023, 94, 215–228. [Google Scholar] [CrossRef] [Scilit]
  20. Bermudez, J.D.; Happ, P.N.; Oliveira, D.A.B.; Feitosa, R.Q. SAR to Optical Image Synthesis for Cloud Removal with Generative Adversarial Networks. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2018, IV-1, 5–11. [Google Scholar] [CrossRef] [Scilit]
  21. Liu, R.; Meng, S.; Peng, Y.; Tian, X. TransFusion-CR: Two-Phase SAR-to-Optical Translation and Deep Feature Fusion for Cloud Removal. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5635911. [Google Scholar] [CrossRef] [Scilit]
  22. Xiang, P.; Zhang, J.; Qi, S.; Jung, S.K.; Zhou, H.; Zhao, D. Hyperspectral Anomaly Detection Using Taylor Expansion and Weighted Irregular Block Filter. Infrared Phys. Technol. 2025, 150, 105942. [Google Scholar] [CrossRef] [Scilit]
  23. Zou, X.; Li, K.; Xing, J.; Zhang, Y.; Wang, S.; Jin, L.; Tao, P. DiffCR: A Fast Conditional Diffusion Framework for Cloud Removal from Optical Satellite Images. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5612014. [Google Scholar] [CrossRef] [Scilit]
  24. Peng, H.; Huang, T.-Z.; Zhao, X.-L.; Lin, J.; Wu, W.-H.; Li, L.-Y. Deep Domain Fidelity and Low-Rank Tensor Ring Regularization for Thick Cloud Removal of Multitemporal Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5409314. [Google Scholar] [CrossRef] [Scilit]
  25. Xiang, P.; Qi, S.; Jung, S.K.; Qian, K.; Cheng, K.; Prasad, S.; Zhou, H.; Zhao, D. High-Low-Frequency Feature-Guided Diffusion Model for Hyperspectral Anomaly Detection. Opt. Laser Technol. 2026, 203, 115561. [Google Scholar] [CrossRef] [Scilit]
  26. Johnson, J.; Alahi, A.; Fei-Fei, L. Perceptual Losses for Real-Time Style Transfer and Super-Resolution. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 11–14 October 2016; pp. 694–711. [Google Scholar]
  27. Ledig, C.; Theis, L.; Huszár, F.; Caballero, J.; Cunningham, A.; Acosta, A.; Aitken, A.P.; Tejani, A.; Totz, J.; Wang, Z.; et al. Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 4681–4690. [Google Scholar]
  28. Zhang, K.; Zuo, W.; Chen, Y.; Meng, D.; Zhang, L. Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising. IEEE Trans. Image Process. 2017, 26, 3142–3155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Pathak, D.; Krähenbühl, P.; Donahue, J.; Darrell, T.; Efros, A.A. Context Encoders: Feature Learning by Inpainting. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 26 June–1 July 2016; pp. 2536–2544. [Google Scholar]
  30. Li, X.; Xu, F.; Yu, A.; Lyu, X.; Gao, H.; Zhou, J. A Frequency Decoupling Network for Semantic Segmentation of Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5607921. [Google Scholar] [CrossRef] [Scilit]
  31. Li, X.; Xu, F.; Zhang, J.; Zhang, H.; Lyu, X.; Liu, F.; Gao, H.; Kaup, A. Frequency-Guided Denoising Network for Semantic Segmentation of Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5400217. [Google Scholar] [CrossRef] [Scilit]
  32. Yu, J.; Lin, Z.; Yang, J.; Shen, X.; Lu, X.; Huang, T.S. Generative Image Inpainting with Contextual Attention. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 5505–5514. [Google Scholar]
  33. Nazeri, K.; Ng, E.; Joseph, T.; Qureshi, F.Z.; Ebrahimi, M. EdgeConnect: Generative Image Inpainting with Adversarial Edge Learning. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 3265–3274. [Google Scholar]
  34. Chen, H.; Wang, Y.; Guo, T.; Xu, C.; Deng, Y.; Liu, Z.; Ma, S.; Xu, C.; Xu, C.; Gao, W. Pre-Trained Image Processing Transformer. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 12299–12310. [Google Scholar]
  35. Liang, J.; Cao, J.; Sun, G.; Zhang, K.; Van Gool, L.; Timofte, R. SwinIR: Image Restoration Using Swin Transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, Montreal, QC, Canada, 11–17 October 2021; pp. 1833–1844. [Google Scholar]
  36. Wang, Z.; Cun, X.; Bao, J.; Zhou, W.; Liu, J.; Li, H. Uformer: A General U-Shaped Transformer for Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 17683–17693. [Google Scholar]
  37. Chen, L.; Chu, X.; Zhang, X.; Sun, J. Simple Baselines for Image Restoration. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 17–33. [Google Scholar]
  38. Chen, S.; Zhang, W.; Li, Z.; Wang, Y.; Zhang, B. Cloud Removal with SAR-Optical Data Fusion and Graph-Based Feature Aggregation Network. Remote Sens. 2022, 14, 3374. [Google Scholar] [CrossRef] [Scilit]
  39. Han, S.; Wang, J.; Zhang, S. Former-CR: A Transformer-Based Thick Cloud Removal Method with Optical and SAR Imagery. Remote Sens. 2023, 15, 1196. [Google Scholar] [CrossRef] [Scilit]
  40. Mao, R.; Li, H.; Ren, G.; Yin, Z. Cloud Removal Based on SAR-Optical Remote Sensing Data Fusion via a Two-Flow Network. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 7677–7686. [Google Scholar] [CrossRef] [Scilit]
  41. Wang, Y.; Zhang, B.; Zhang, W.; Hong, D.; Zhao, B.; Li, Z. Cloud Removal with SAR-Optical Data Fusion Using a Unified Spatial–Spectral Residual Network. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5600820. [Google Scholar] [CrossRef] [Scilit]
  42. Gu, P.; Liu, W.; Feng, S.; Wei, T.; Wang, J.; Chen, H. HPN-CR: Heterogeneous Parallel Network for SAR-Optical Data Fusion Cloud Removal. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5402115. [Google Scholar] [CrossRef] [Scilit]
  43. Li, X.; Xu, F.; Zhang, J.; Yu, A.; Lyu, X.; Gao, H.; Zhou, J. Dual-Domain Decoupled Fusion Network for Semantic Segmentation of Remote Sensing Images. Inf. Fusion 2025, 124, 103359. [Google Scholar] [CrossRef] [Scilit]
  44. Li, X.; Xu, F.; Li, J.; Su, Y.; Li, L.; Lyu, X.; Xu, Z.; Kaup, A. Frequency Domain-Enhanced Spectral-Spatial Fusion Transformer for Semantic Segmentation of Remote Sensing Images. Inf. Fusion 2026, 132, 104248. [Google Scholar] [CrossRef] [Scilit]
  45. Liu, Y.; Li, W.; Guan, J.; Zhou, S.; Zhang, Y. Effective Cloud Removal for Remote Sensing Images by an Improved Mean-Reverting Denoising Model with Elucidated Design Space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 11–15 June 2025; pp. 17851–17861. [Google Scholar]
  46. Zhang, Z.; Li, J.; Liang, Y.; Yan, J.; Xiao, Y.; Su, X.; Yuan, Q. ECRformer: An Efficient Cloud Removal Transformer with Semantic-Decoupled Learning for Multimodal Satellite Imagery. ISPRS J. Photogramm. Remote Sens. 2026, 237, 323–338. [Google Scholar] [CrossRef] [Scilit]
  47. Wang, C.; Feng, H.; Zheng, Y.; Yang, W.; Zhang, X.; Wang, G.; Wang, Y. TCF-VQGAN: Two-Stage Codebook Fusion Vector-Quantized GAN for Multimodal Remote Sensing Image Cloud Removal. Remote Sens. 2026, 18, 1643. [Google Scholar] [CrossRef] [Scilit]
  48. Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. In Proceedings of the 28th Annual Conference on Neural Information Processing Systems 2014, Montreal, QC, Canada, 8–13 December 2014; pp. 2672–2680. [Google Scholar]
  49. Isola, P.; Zhu, J.-Y.; Zhou, T.; Efros, A.A. Image-to-Image Translation with Conditional Adversarial Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 5967–5976. [Google Scholar]
  50. Zhu, J.-Y.; Park, T.; Isola, P.; Efros, A.A. Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 27–29 October 2017; pp. 2223–2232. [Google Scholar]
  51. Zhang, S.; Li, X.; Zhou, X.; Wang, Y.; Hu, Y. Cloud Removal Using SAR and Optical Images via Attention Mechanism-Based GAN. Pattern Recognit. Lett. 2023, 175, 8–15. [Google Scholar] [CrossRef] [Scilit]
  52. Darbaghshahi, F.N.; Mohammadi, M.R.; Soryani, M. Cloud Removal in Remote Sensing Images Using Generative Adversarial Networks and SAR-to-Optical Image Translation. IEEE Trans. Geosci. Remote Sens. 2022, 60, 4105309. [Google Scholar] [CrossRef] [Scilit]
  53. Xiang, X.; Tan, Y.; Yan, L. Cloud-Guided Fusion with SAR-to-Optical Translation for Thick Cloud Removal. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5633715. [Google Scholar] [CrossRef] [Scilit]
  54. Hu, J.; Shen, L.; Sun, G. Squeeze-and-Excitation Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 7132–7141. [Google Scholar]
  55. Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision, Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
  56. Wang, X.; Girshick, R.; Gupta, A.; He, K. Non-Local Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–22 June 2018; pp. 7794–7803. [Google Scholar]
  57. Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention Is All You Need. In Proceedings of the 31st Annual Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 4–9 December 2017; pp. 5998–6008. [Google Scholar]
  58. Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.-H. Multi-Stage Progressive Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 14821–14831. [Google Scholar]
  59. Zamir, S.W.; Arora, A.; Khan, S.; Hayat, M.; Khan, F.S.; Yang, M.-H. Restormer: Efficient Transformer for High-Resolution Image Restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 5728–5739. [Google Scholar]
  60. Li, X.; Xu, F.; Liu, F.; Tong, Y.; Lyu, X.; Zhou, J. Semantic Segmentation of Remote Sensing Images by Interactive Representation Refinement and Geometric Prior-Guided Inference. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5400318. [Google Scholar] [CrossRef] [Scilit]
  61. Li, X.; Xu, F.; Liu, F.; Lyu, X.; Gao, H.; Zhou, J.; Kaup, A. A Euclidean Affinity-Augmented Hyperbolic Neural Network for Semantic Segmentation of Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5636718. [Google Scholar] [CrossRef] [Scilit]
  62. Li, X.; Shi, C.; Xu, N.; Su, Y.; Kaup, A.; Liu, D.; Li, X. Position-Aware Differential Denoising Transformer for Semantic Segmentation of Remote Sensing Images. IEEE Geosci. Remote Sens. Lett. 2026, 23, 5000405. [Google Scholar] [CrossRef] [Scilit]
  63. Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale. In Proceedings of the International Conference on Learning Representations, Virtual Event, 3–7 May 2021. [Google Scholar]
  64. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Montreal, QC, Canada, 11–17 October 2021; pp. 10012–10022. [Google Scholar]
  65. Liu, Z.; Mao, H.; Wu, C.-Y.; Feichtenhofer, C.; Darrell, T.; Xie, S. A ConvNet for the 2020s. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 11976–11986. [Google Scholar]
  66. Wang, Y.; Wu, Y.; Zhang, Z.; Huang, Q.; Tang, B.; Yang, Z.; Zhang, K.; Zhang, L. Dual-Scale Transformer with Variable Bitrate Synchronization for Neural Video Compression. ACM Trans. Multimed. Comput. Commun. Appl. 2026, 22, 139. [Google Scholar] [CrossRef] [Scilit]
  67. Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the 18th International Conference on Medical Image Computing and Computer-Assisted Intervention, Munich, Germany, 5–9 October 2015; pp. 234–241. [Google Scholar]
  68. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
  69. Ba, J.L.; Kiros, J.R.; Hinton, G.E. Layer Normalization. arXiv 2016, arXiv:1607.06450. [Google Scholar]
  70. Ioffe, S.; Szegedy, C. Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift. In Proceedings of the International Conference on Machine Learning, Lille, France, 6–11 July 2015; pp. 448–456. [Google Scholar]
  71. Xia, Y.; He, W.; Huang, Q.; Yin, G.; Liu, W.; Zhang, H. CRformer: Multi-Modal Data Fusion to Reconstruct Cloud-Free Optical Imagery. Int. J. Appl. Earth Obs. Geoinf. 2024, 128, 103793. [Google Scholar] [CrossRef] [Scilit]
  72. Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. In Proceedings of the International Conference on Learning Representations, San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
  73. Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. Overall architecture of the proposed TFCRNet. The framework consists of a SAR-to-optical translation stage and a cloud-aware fusion stage. The translation stage uses MSFFG with dual-discriminator guidance, while the fusion stage uses RGCAF to perform region-adaptive cross-modal interaction.
Figure 1. Overall architecture of the proposed TFCRNet. The framework consists of a SAR-to-optical translation stage and a cloud-aware fusion stage. The translation stage uses MSFFG with dual-discriminator guidance, while the fusion stage uses RGCAF to perform region-adaptive cross-modal interaction.
Remotesensing 18 02962 g001
Figure 2. Architecture of the Multi-Scale Feature Fusion Generator (MSFFG). MSFFG extracts SAR representations at multiple scales, fuses them through MSFF, and reconstructs the translated optical image through parallel decoding paths.
Figure 2. Architecture of the Multi-Scale Feature Fusion Generator (MSFFG). MSFFG extracts SAR representations at multiple scales, fuses them through MSFF, and reconstructs the translated optical image through parallel decoding paths.
Remotesensing 18 02962 g002
Figure 3. Structure of the Multi-Scale Feature Fusion (MSFF) module. Different colors are used to visually distinguish the processing blocks and feature-flow components in the module.
Figure 3. Structure of the Multi-Scale Feature Fusion (MSFF) module. Different colors are used to visually distinguish the processing blocks and feature-flow components in the module.
Remotesensing 18 02962 g003
Figure 4. PatchGAN discriminator architecture used in the dual-discriminator design.
Figure 4. PatchGAN discriminator architecture used in the dual-discriminator design.
Remotesensing 18 02962 g004
Figure 5. Spectral and structural feature extraction modules used in the dual-discriminator design.
Figure 5. Spectral and structural feature extraction modules used in the dual-discriminator design.
Remotesensing 18 02962 g005
Figure 6. Structure of the Region-Gated Cross-Attention Fusion (RGCAF) module. The cloud-region gate assigns larger weights to translated SAR-derived responses in cloud-covered regions and larger weights to optical responses in cloud-free regions.
Figure 6. Structure of the Region-Gated Cross-Attention Fusion (RGCAF) module. The cloud-region gate assigns larger weights to translated SAR-derived responses in cloud-covered regions and larger weights to optical responses in cloud-free regions.
Remotesensing 18 02962 g006
Figure 7. Visual comparison on the SEN12MS-CR dataset. (a) Cloudy optical image; (b) SAR image; (c) SAR2OPT result; (d) GLF-CR result; (e) DSen2-CR result; (f) HS2P result; (g) HPN-CR result; (h) TFCRNet result; (i) cloud-free reference image. The red boxes highlight representative regions selected for detailed visual comparison.
Figure 7. Visual comparison on the SEN12MS-CR dataset. (a) Cloudy optical image; (b) SAR image; (c) SAR2OPT result; (d) GLF-CR result; (e) DSen2-CR result; (f) HS2P result; (g) HPN-CR result; (h) TFCRNet result; (i) cloud-free reference image. The red boxes highlight representative regions selected for detailed visual comparison.
Remotesensing 18 02962 g007
Figure 8. Visual comparison on the SMILE-CR dataset. (a) Cloudy optical image; (b) SAR image; (c) SAR2OPT result; (d) GLF-CR result; (e) DSen2-CR result; (f) HS2P result; (g) HPN-CR result; (h) TFCRNet result; (i) cloud-free reference image. The yellow boxes highlight representative regions selected for detailed visual comparison.
Figure 8. Visual comparison on the SMILE-CR dataset. (a) Cloudy optical image; (b) SAR image; (c) SAR2OPT result; (d) GLF-CR result; (e) DSen2-CR result; (f) HS2P result; (g) HPN-CR result; (h) TFCRNet result; (i) cloud-free reference image. The yellow boxes highlight representative regions selected for detailed visual comparison.
Remotesensing 18 02962 g008
Figure 9. Ablation visualization on the SEN12MS-CR dataset. (a) Cloudy optical image; (b) SAR image; (c) cloud-free reference image; (d) full TFCRNet; (e) TFCRNet without RGCAF; (f) TFCRNet without the dual-discriminator design; (g) TFCRNet without MSFF. The red boxes highlight representative regions selected for detailed visual comparison.
Figure 9. Ablation visualization on the SEN12MS-CR dataset. (a) Cloudy optical image; (b) SAR image; (c) cloud-free reference image; (d) full TFCRNet; (e) TFCRNet without RGCAF; (f) TFCRNet without the dual-discriminator design; (g) TFCRNet without MSFF. The red boxes highlight representative regions selected for detailed visual comparison.
Remotesensing 18 02962 g009
Table 1. Controlled quantitative comparison among the baseline methods reproduced on the SEN12MS-CR dataset under the unified evaluation protocol adopted in this study. The best reproduced result is shown in bold, and the second-best reproduced result is underlined.
Table 1. Controlled quantitative comparison among the baseline methods reproduced on the SEN12MS-CR dataset under the unified evaluation protocol adopted in this study. The best reproduced result is shown in bold, and the second-best reproduced result is underlined.
MethodInputPSNRSSIMSAMMAE
SAR2OPT [20SAR only24.61400.779119.66840.0411
GLF-CR [18SAR + Optical28.35740.85629.10870.0288
DSen2-CR [16SAR + Optical27.06560.876813.73210.0294
HS2P [19SAR + Optical28.26650.857710.36130.0296
HPN-CR [42SAR + Optical28.87610.898011.83790.0239
TFCRNetSAR + Optical33.01150.90856.27950.0125
Table 2. Controlled quantitative comparison among the baseline methods reproduced on the SMILE-CR dataset under the unified evaluation protocol adopted in this study. The best reproduced result is shown in bold, and the second-best reproduced result is underlined.
Table 2. Controlled quantitative comparison among the baseline methods reproduced on the SMILE-CR dataset under the unified evaluation protocol adopted in this study. The best reproduced result is shown in bold, and the second-best reproduced result is underlined.
MethodInputPSNRSSIMSAMMAE
SAR2OPT [20SAR only18.35840.610827.96990.0926
GLF-CR [18SAR + Optical28.69430.86388.25040.0256
DSen2-CR [16SAR + Optical28.40310.84848.64070.0251
HS2P [19SAR + Optical28.76320.84997.43020.0358
HPN-CR [42SAR + Optical29.14450.87497.98050.0226
TFCRNetSAR + Optical30.52910.88136.98640.0217
Table 3. Component-level ablation study on the SEN12MS-CR dataset. MSFF denotes the Multi-Scale Feature Fusion module, Dual-D denotes the complete dual-discriminator design composed of SpeD and StrD, and RGCAF denotes the Region-Gated Cross-Attention Fusion module. A checkmark indicates that the corresponding component is included. The best results are shown in bold.
Table 3. Component-level ablation study on the SEN12MS-CR dataset. MSFF denotes the Multi-Scale Feature Fusion module, Dual-D denotes the complete dual-discriminator design composed of SpeD and StrD, and RGCAF denotes the Region-Gated Cross-Attention Fusion module. A checkmark indicates that the corresponding component is included. The best results are shown in bold.
MSFFDual-DRGCAFPSNRSSIMSAMMAE
31.72930.89569.76020.0204
31.37880.894810.12080.0216
31.96720.90798.92270.0242
33.01150.90856.27950.0125
Table 4. Fine-grained ablation of SpeD and StrD on the SEN12MS-CR dataset. The best results are shown in bold, and the second-best results are underlined.
Table 4. Fine-grained ablation of SpeD and StrD on the SEN12MS-CR dataset. The best results are shown in bold, and the second-best results are underlined.
Discriminator SettingPSNRSSIMSAMMAE
Single PatchGAN32.08740.89978.61230.0189
SpeD only32.46280.90347.18460.0161
StrD only32.35170.90527.84290.0165
SpeD + StrD33.01150.90856.27950.0125
Table 5. Ablation of cloud-mask guidance and individual loss terms on the SEN12MS-CR dataset. The best results are shown in bold, and the second-best results are underlined.
Table 5. Ablation of cloud-mask guidance and individual loss terms on the SEN12MS-CR dataset. The best results are shown in bold, and the second-best results are underlined.
ConfigurationPSNRSSIMSAMMAE
w/o cloud-mask guidance32.38640.90467.53180.0179
w/o cloud-region weighting32.67290.90696.84670.0148
w/o L G , a d v 31.37880.894810.12080.0216
w/o L S A M 32.10860.90219.38450.0175
w/o L S S I M 32.27630.89767.21490.0169
Full TFCRNet33.01150.90856.27950.0125
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Gao, S.; Xie, W.; Lyu, X.; He, H.; An, D.; Li, X.; Shi, C.; Wu, C.; Zhang, C.; Xu, Z. TFCRNet: Dual-Discriminator SAR-to-Optical Translation and Region-Gated Cross-Attention Fusion for Thick-Cloud Removal. Remote Sens. 2026, 18, 2962. https://doi.org/10.3390/rs18172962

AMA Style

Gao S, Xie W, Lyu X, He H, An D, Li X, Shi C, Wu C, Zhang C, Xu Z. TFCRNet: Dual-Discriminator SAR-to-Optical Translation and Region-Gated Cross-Attention Fusion for Thick-Cloud Removal. Remote Sensing. 2026; 18(17):2962. https://doi.org/10.3390/rs18172962

Chicago/Turabian Style

Gao, Shengkai, Wenjun Xie, Xin Lyu, Houjun He, Dong An, Xin Li, Chengyi Shi, Caifeng Wu, Chengming Zhang, and Zhennan Xu. 2026. "TFCRNet: Dual-Discriminator SAR-to-Optical Translation and Region-Gated Cross-Attention Fusion for Thick-Cloud Removal" Remote Sensing 18, no. 17: 2962. https://doi.org/10.3390/rs18172962

APA Style

Gao, S., Xie, W., Lyu, X., He, H., An, D., Li, X., Shi, C., Wu, C., Zhang, C., & Xu, Z. (2026). TFCRNet: Dual-Discriminator SAR-to-Optical Translation and Region-Gated Cross-Attention Fusion for Thick-Cloud Removal. Remote Sensing, 18(17), 2962. https://doi.org/10.3390/rs18172962

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop