Next Article in Journal
Integrated Fault Estimation and Fault-Tolerant Control for Drive-by-Wire Steering Systems in Four-Wheel Independent Drive Vehicles
Previous Article in Journal
Online-Tuned Fuzzy Pre-Filtering with an Attention BiLSTM for Misbehavior Detection in Vehicular Named Data Networking
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

DualCDM: Dual-Domain Conditional Diffusion for SAR-to-Optical Translation with Spatial–Frequency Correlation and Adaptive Feature Recalibration

1
Institute of Space Science and Technology, Nanchang University, Nanchang 330031, China
2
Faculty of Geo-Information Science and Earth Observation, University of Twente, 7500 AE Enschede, The Netherlands
*
Author to whom correspondence should be addressed.
Sensors 2026, 26(13), 4183; https://doi.org/10.3390/s26134183
Submission received: 26 May 2026 / Revised: 23 June 2026 / Accepted: 29 June 2026 / Published: 2 July 2026
(This article belongs to the Section Remote Sensors)

Abstract

Translating Synthetic aperture radar (SAR) images into optical images is intrinsically ill-posed because microwave backscatter and optical reflectance describe different physical properties of the observed scene. Although frequency-domain modeling has been introduced into diffusion-based translation, existing methods mainly rely on independent weighting of individual Fourier coefficients and provide limited modeling of interactions among neighboring frequencies and feature channels. To address this limitation, we propose dualCDM, a conditional diffusion model that jointly exploits spatial- and frequency-domain representations. In the diffusion backbone, a spatial-frequency hybrid residual block (SFHRB) combines a spatial convolution branch with complex-valued convolution in the Fourier domain. The complex convolution aggregates neighboring Fourier coefficients across all input feature channels, enabling local cross-frequency and cross-channel modeling, while its response is modulated by the diffusion timestep. In the SAR conditional encoder, an adaptive frequency-domain feature recalibration block (AFFRB) predicts input-dependent real-valued gains from magnitude and trigonometric phase representations of intermediate GRD features. These gains adaptively recalibrate the complex frequency responses without introducing an additional phase shift, while the residual connection preserves the original conditional information. A dual-domain objective further constrains both the predicted diffusion noise and the one-step optical reconstruction in the spatial and frequency domains. We also construct the S1S2 dataset using 16-bit Sentinel-2 reflectance data, retaining the original 0–10,000 value range and including the near-infrared band. Experiments on SEN1-2 and S1S2 show that dualCDM improves radiometric accuracy, spectral consistency, and structural preservation over six representative methods. Paired statistical tests further confirm significant improvements over the strongest competing method across all six evaluation metrics on both datasets.

1. Introduction

Optical and synthetic aperture radar (SAR) sensors provide complementary observations of the Earth surface. Optical imagery records spectral reflectance and spatial appearance in an intuitive form, but its availability is often limited by clouds and atmospheric effects [1]. SAR provides all-weather and day-and-night observations, but its side-looking imaging geometry, coherent scattering mechanism, and speckle make visual interpretation difficult. SAR-to-optical image translation, therefore, offers a way to generate optical-like representations from SAR observations, supporting visual interpretation and providing auxiliary information when cloud-free optical acquisitions are unavailable [2]. However, this translation remains challenging because SAR backscatter and optical reflectance are governed by substantially different imaging mechanisms.
Early SAR-to-optical translation methods mainly relied on convolutional neural networks to extract and recombine spatial features. Fu and Zhang [3] formulated remote sensing image translation as a sequence of image understanding, object transformation, and representation. Generative adversarial networks (GANs) subsequently became widely used because adversarial learning can improve the visual realism of generated images [4,5,6,7,8]. More recent studies have introduced Transformer architectures [9,10] and diffusion models [2,11,12] to improve long-range modeling and generation stability. A comprehensive review of this rapidly developing field was provided by Xiong et al. [13].
Frequency-domain modeling has recently been introduced to improve the representation of textures, boundaries, and other fine-scale structures. For example, Qin et al. [14] proposed a spatial-frequency diffusion framework that jointly processes target features in the two domains. This direction is also motivated by the spectral bias of neural networks, which tend to learn low-frequency components earlier than high-frequency components [15].
Modeling the correlations among spectral intensities and channels in the frequency domain helps capture a more comprehensive data distribution, properties that has yet to be exploited by diffusion translation networks. Existing frequency-domain modules apply pointwise or channel-wise weighting to individual Fourier coefficients. Although such operations can recalibrate spectral responses, they provide limited interaction among neighboring frequency components and feature channels. This limitation is important because spatial structures are generally represented by groups of correlated Fourier coefficients rather than isolated spectral points. Edges, textures, and repeated patterns produce locally dependent responses within spectral neighborhoods [16]. Moreover, applying the Fourier transform independently to each feature channel does not eliminate inter-channel dependence, which remains encoded in the complex-valued spectral coefficients. Effective frequency-domain modeling should, therefore, capture both local cross-frequency interactions and cross-channel correlations while preserving the complex relationships among Fourier coefficients.
The quality of the SAR conditional features presents another challenge. Speckle-related fluctuations can interfere with the extraction of boundaries, regional patterns, and strong scattering structures that provide important guidance for optical image generation. A common strategy is to apply an external despeckling filter to the SAR input, such as probabilistic patch-based filtering [14]. However, this preprocessing is performed independently of the translation objective and follows a predetermined filtering rule, which may remove structurally useful details together with noise-related variations. These limitations leave open how to reduce noise-related interference without sacrificing the structural information required for SAR-to-optical translation.
To address these limitations, we propose dualCDM, a conditional diffusion framework that combines spatial-frequency feature modeling, adaptive frequency-domain recalibration of SAR conditional features, and dual-domain supervision. In the diffusion backbone, SFHRB uses complex-valued convolution to model local cross-frequency and cross-channel correlations, while a parallel spatial branch preserves local image context. In the SAR conditional encoder, AFFRB learns input-dependent gains to recalibrate intermediate frequency-domain feature responses. A spatial-frequency objective further constrains both noise prediction and one-step optical reconstruction. We also construct a native-scale multispectral dataset to evaluate SAR-to-optical translation beyond conventional 8-bit RGB settings.
The main contributions are summarized as follows:
(1)
Local cross-frequency and cross-channel modeling: We introduce SFHRB, which combines a spatial convolution branch with a complex-valued convolution in the Fourier domain. Unlike independent per-frequency reweighting, the complex-valued convolution jointly processes the real and imaginary components of neighboring Fourier coefficients while mixing information across all input channels.
(2)
Adaptive frequency-domain recalibration of SAR conditional features: We introduce AFFRB into the SAR conditional encoder to adaptively modulate the Fourier-domain responses of intermediate real-valued GRD features. The module predicts input-dependent real-valued gains from the spectral magnitude and trigonometric phase representations and applies them to the complex Fourier coefficients.
(3)
Dual-domain supervision: We constrain the predicted diffusion noise and the one-step optical estimate in the frequency domain in addition to the standard spatial-domain noise loss.
(4)
Native-scale multispectral benchmark: We construct the S1S2 dataset using Sentinel-2 reflectance values retained on the 0–10,000 scale and including the near-infrared (NIR) band.

2. Related Work

In the context of remote sensing image translation, some studies have demonstrated the effectiveness of GANs in translating SAR images into optical images. Enomoto et al. [4] implemented SAR-to-optical image translation using conditional GAN. Reyes et al. [5] conducted image translation experiments using cycle-consistent GAN on both high-resolution data (TerraSAR-X and ALOS PRISM) and low-resolution data (Sentinel-1 and Sentinel-2). To improve the image quality and training stability, several new GAN-based methods were proposed. Li et al. [6] developed a modified cGAN with strong constraints based on structural similarity and L1 norm. Zhang et al. [7] introduced a feature-guided method that effectively utilizes the diversity of structural and texture features. A reciprocal adversarial network scheme that utilizes cascaded residual connections and a hybrid L1-GAN loss function was proposed by Fu et al. [8].
In addition to GANs, Transformer has also been used for the translation. Kong et al. [9] proposed an encoder-decoder generator based on the SWIN Transformer. Wang et al. [10] proposed a hybrid conditional GAN that combines the advantages of convolutional neural network (CNN) and Vision Transformer, enhancing image fidelity by incorporating a classification loss.
Central to these studies, diffusion models [17] have recently demonstrated excellent capabilities in generation tasks. By learning specific data distributions, diffusion models generate high-quality images and outperform GANs [18]. For the SAR-to-optical translation task, diffusion models have been used in [2,11,12,14]. Furthermore, ref. [19] proposed to use the Brownian Bridge diffusion model by modelling image translation tasks as a stochastic Brownian Bridge process.

3. Methodology

3.1. Network Structures

The architecture of the proposed dual-domain conditional diffusion model is illustrated in Figure 1. It consists of a diffusion backbone and a SAR conditional encoder. The diffusion backbone follows a U-Net encoder–decoder architecture and predicts the noise added to the optical image at diffusion timestep t. In parallel, the conditional encoder extracts multi-scale structural and textural features from the SAR observation. These conditional features are injected into the corresponding resolution levels of the diffusion backbone through lateral connections, thereby providing scene-dependent guidance for noise prediction. Long skip connections between the encoder and decoder of the diffusion backbone preserve spatial information and facilitate the fusion of shallow structural details with deeper semantic features.
In the diffusion backbone, the noisy optical image O t is first projected into a 128-channel feature space using a 3 × 3 convolution (Conv3). The backbone contains four resolution levels indexed by l 0 , 1 , 2 , 3 , with spatial resolutions of H / 2 l × W / 2 l and channel dimensions of 128, 256, 384, and 512, respectively. Each encoder level contains two consecutively stacked residual blocks, denoted by × 2 in Figure 1, followed by downsampling except at the deepest level. The decoder follows the reverse multi-scale hierarchy and progressively upsamples the features while fusing them with the corresponding encoder features through skip connections. Each decoder level contains three consecutively stacked residual blocks, denoted by × 3 , because an additional block is used to refine the fused decoder and skip-connected features. The two shallow resolution levels use SFHRB, whereas the two deeper and lower-resolution levels use SFHRBAtt. SFHRBAtt is the attention-enhanced variant of SFHRB and consists of an SFHRB followed by an AttnBlock. Self-attention is, therefore, applied only at the lower-resolution levels, where long-range dependencies can be modeled with lower memory consumption and computational cost. Finally, GroupNorm, Swish activation, and a 3 × 3 convolution are applied to generate the predicted noise Z, which has the same spatial resolution and number of channels as O t .
The SAR conditional encoder adopts the same four-level resolution hierarchy and channel configuration as the diffusion encoder. The SAR image S is first projected into the feature space through a 3 × 3 convolution. At the two shallow resolution levels, the conditional features are extracted using AFFRB, whereas the two deeper levels employ AFFRBAtt. AFFRBAtt is the attention-enhanced variant of AFFRB and is composed of an AFFRB followed by an AttnBlock. Similar to the diffusion encoder, each resolution level contains two consecutively stacked blocks, denoted by × 2 . The resulting conditional features are injected into the corresponding encoder and decoder stages of the diffusion backbone through lateral connections. Throughout Figure 1, the notation × N indicates that N blocks of the specified type are consecutively stacked at the same spatial resolution and channel dimension.

3.1.1. SFHRB Module

To preserve local spatial textures while improving frequency-domain representation, we design SFHRB based on the standard diffusion residual block. As shown in Figure 2, it contains a spatial convolution branch operating in the spatial domain and a parallel complex-valued convolution branch operating in the Fourier domain. The frequency branch is further modulated by the diffusion timestep embedding. Unlike pointwise spectral weighting, which processes frequency locations independently [14,20], the complex-valued convolution aggregates neighboring frequencies across multiple channels, enabling local cross-frequency and cross-channel modeling.
Given the input feature map x R B × C in × H × W , where B denotes the batch size, C i n is the number of input channels, and H × W represents the spatial resolution, we further consider the timestep embedding vector t R B × C t and the SAR conditional feature c R B × C cond × H × W . The spatial-domain branch applies normalization, nonlinear activation, and a 3 × 3 convolution to capture local spatial context,
h s = Conv 3 ϕ GN ( x ) R B × C out × H × W ,
where GN · denotes group normalization (GroupNorm), ϕ · denotes the Swish activation function (Swish), and Conv 3 · denotes a 3 × 3 convolution.
In parallel, the frequency-domain branch first applies a fast Fourier transform (FFT) to the normalized input features to obtain a complex spectral representation,
H C in ω = F GN ( x ) C B × C in × H × W ,
where F · denotes FFT, and ω = ω x , ω y denotes the frequency-domain coordinates.
During the forward diffusion process, the intensity of injected random noise varies across timesteps, causing the statistical properties of features in the frequency domain to change accordingly. To adapt to the timestep-dependent spectral distribution and enable adaptive spectral recalibration, the timestep embedding vector t is linearly projected into a complex-valued modulation vector, which is used to modulate the input and output channels of the spectral features,
g i n C B × C in , g o u t C B × C out ,
where C out is the number of output channels.
First, complex-valued channel modulation is applied to the input spectral features, i.e., each channel of each sample is multiplied by a complex-valued scaling factor shared across spatial locations, thereby recalibrating the spectral response of each channel,
H ˜ C in ω = H C in ω g in ,
where ⊙ denotes element-wise product.
Then, a complex-valued convolution is performed in the frequency domain,
Y C out ω = C in H ˜ C in ω W C out , C in ω ,
where ω denotes the two-dimensional discrete convolution with respect to the frequency-domain coordinates ω , and W C out , C in is the complex-valued convolution kernel that connects the input channel C in to the output channel C out (with a kernel size of 3 × 3 in this work).
Subsequently, complex-valued channel modulation is applied again to the output spectral features,
Y ˜ C out ω = Y C out ω g out .
Finally, an inverse FFT (IFFT) is applied to return to the spatial domain, and the real part is taken as the output of the frequency-domain branch,
h f = F 1 Y ˜ R B × C out × H × W ,
where F 1 · denotes IFFT, and · denotes the real part.
Specifically, the timestep embedding t is processed through a multilayer perceptron (MLP) to predict a per-channel scaling factor γ R B × C out and bias β R B × C out :
[ γ , β ] = Split ( MLP ( t ) ) , h 1 = h s + h f 1 + γ + β .
In addition, the SAR image is used as a conditional feature and mapped to C out via a 3 × 3 convolution, and then injected into the main branch with a learnable coefficient α ,
h 2 = h 1 + α · Conv 3 ϕ c .
The fused feature h 2 is further refined through hybrid refinement, and then added to the original input feature x via a skip connection to produce the final output,
y = x + Conv 3 ϕ GN ( h 2 ) .

3.1.2. AFFRB Module

As shown in Figure 3, the Sentinel-1 GRD intensity image is decomposed into low-, middle-, and high-frequency components for visualization. Specifically, the centered Fourier spectrum is divided using two radial thresholds, R 1 and R 2 , and the corresponding components are reconstructed through inverse Fourier transform. The low-frequency component mainly retains the large-scale spatial layout, overall object outlines, and slowly varying background. In contrast, granular fluctuations and fine-scale textures become increasingly evident in the middle- and high-frequency components, indicating that these frequency ranges are more sensitive to speckle-related variations. Nevertheless, structurally relevant information, including sharp boundaries, strong gradients, and point-like scatterers, is also present in these frequency ranges. Therefore, directly suppressing the middle- or high-frequency components may remove both noise-related fluctuations and useful structural details. The fixed frequency decomposition shown in Figure 3 is used only for visualization and is not explicitly employed in the proposed network.
This observation motivates an adaptive frequency-domain feature recalibration strategy that selectively attenuates noise-dominated responses while preserving structurally informative frequency components. To implement this strategy, we propose a frequency-domain adaptive attention module, termed FreqAttn, and embed it into the residual blocks of the SAR conditional encoder to form AFFRB. Rather than dividing the feature spectrum into predefined frequency bands, FreqAttn predicts a continuous, input-dependent gain for each feature channel and frequency location. The detailed architecture of AFFRB is illustrated in Figure 4.
Given the intermediate feature h s a r R B × C × H × W obtained after the first convolution, FreqAttn first applies an FFT to it and performs spectrum centering,
H s = S F h s a r C B × C × H × W ,
where S ( · ) denotes the spectrum centering operation.
Then, the magnitude and phase are extracted from H s to construct a joint feature,
Z = concat M , cos Φ , sin Φ R B × 3 C × H × W ,
where M = H s denotes the magnitude and Φ = H s denotes the phase. To avoid representation ambiguity and numerical instability caused by the phase discontinuity at π , π , we apply a trigonometric encoding to the phase and use cos ( Φ ) and sin ( Φ ) as its representation.
The resulting joint feature Z is then fed into a bottleneck network composed of two consecutive 1 × 1 convolutions, which maps it to frequency-domain gain with the same number of channels as the original spectrum features H s ,
L = σ Conv 1 ( 2 ) ϕ Conv 1 ( 1 ) Z R B × C × H × W ,
where σ · denotes the sigmoid function. To avoid amplifying noise and reduce the risk of numerical instability, we constrain the frequency-domain gain L to the range 0 , 1 , so that it only attenuates rather than amplifies.
The gain L is used to multiplicatively modulate the spectral features on a per-channel and per-frequency basis, adjusting only the magnitude distribution while keeping the phase unchanged,
H s _ a t t = H s L .
Finally, the filtered spectrum is mapped back to the spatial domain via inverse centering and an IFFT, and its real part is taken as the output of FreqAttn,
h o u t = F 1 S 1 H s _ a t t R B × C × H × W .

3.1.3. Auxiliary Modules

The other basic modules in the diffusion model are illustrated in Figure 5. The residual block (ResBlock in Figure 5a) consists of two convolutional layers, each followed by group normalization (GroupNorm) and a Swish activation. The diffusion timestep t is projected by an MLP and injected into the block to modulate and integrate information across different diffusion stages. The last convolution is zero-initialized to improve convergence stability and accelerate convergence in large-scale training [21,22]. The attention block (AttenBlock in Figure 5b) adopts a self-attention mechanism, where the correlation weights between the query (q) and key (k) are computed via softmax and applied to the value (v) for adaptive feature aggregation, thereby enhancing the modeling of long-range dependencies and improving spatial representations.
At the deeper, lower-resolution stages of the diffusion backbone, SFHRBAtt, shown in Figure 5c, applies an AttnBlock after SFHRB to incorporate long-range dependencies into the spatial-frequency representation. At the bottleneck between the encoder and decoder, MidBlock, shown in Figure 5d, sequentially stacks a ResBlock, an AttnBlock, and another ResBlock to refine the feature representation and capture global context. Similarly, at the deeper stages of the SAR conditional encoder, AFFRBAtt, shown in Figure 5e, applies self-attention after AFFRB, allowing long-range structural relationships to be modeled on top of the adaptively recalibrated frequency-domain features.

3.2. Loss Functions

To effectively train the proposed dual-domain conditional diffusion model, we design a spatial-frequency dual-domain loss that jointly constrains network learning in both the spatial and frequency domains. An MSE loss is used in the spatial domain. In the frequency domain, additional constraints are introduced by applying an L 1 loss to supervise the spectral representations of both the noise and the target samples. The overall loss is defined as
L = L s + λ 1 L f ( 1 ) + λ 2 L f ( 2 ) ,
L s = ϵ Z 2 2 ,
L f ( 1 ) = ϵ FFT Z FFT 1 ,
L f ( 2 ) = O 0 _ FFT O 0 _ FFT 1 ,
where λ 1 and λ 2 are weighting factors that balance the contributions of the two frequency-domain loss terms, and both are set to 0.1 in this work. ϵ denotes the random Gaussian noise added during the forward process, Z denotes the noise image predicted by the network, and ϵ FFT , Z FFT , O 0 _ FFT and O 0 _ FFT are the frequency-domain representations obtained from ϵ , Z, O 0 and O 0 , respectively. O 0 denotes the target estimate obtained from the predicted noise Z via one-step prediction,
O 0 = O t 1 α ¯ t · Z α ¯ t .

4. Experimental Scheme

4.1. Datasets

Two sets of training data were used in this study. One is a public dataset, and the other is newly developed in this study.
  • SEN1-2 dataset
The SEN1-2 dataset [23] includes 282,384 image pairs from Sentinel-1 and Sentinel-2 satellites worldwide, each with dimensions of 256 × 256 pixels. These SEN1-2 images also encompass all four seasons of the year, ensuring a diverse and comprehensive sample collection. Sentinel-1 images were acquired in interferometric wide swath (IW) mode as ground-range-detected (GRD) products using VV polarization, while Sentinel-2 provides corresponding optical images in red, green, and blue bands.
  • S1S2 dataset
Considering that quantization accuracy provides greater distinction between image details, a new dataset is built for this work, named Sentinel-1 to Sentinel-2 (S1S2). S1S2 has 15 image pairs from different cities using Sentinel-1 and corresponding Sentinel-2 images. The values range from 0 to 10,000. The time difference between SAR and optical data was restricted to no more than one day to mitigate temporal discrepancies. Sentinel-1 images utilized the IW mode of GRD products in dual-polarization (VV and VH) channels, which can achieve better image quality compared to single polarization [24]. The Sentinel-2 images have the red, green, blue, and near-infrared (NIR) bands. Preprocessing of the Sentinel-1 data included orbit correction, thermal noise removal, radiometric calibration, speckle noise removal (with the Refined Lee filter), and terrain correction. The backscattering values of preprocessed SAR images are converted into decibels. The Sentinel-1 images were geographically registered towards Sentinel-2. The data sources for the newly constructed dataset are listed in Table 1. This dataset can be downloaded from https://github.com/isstncu/s1s2 (accessed on 28 June 2026).
Both datasets were divided into training and testing sets. From the spring subset of the SEN1-2 dataset, 6000 image pairs were randomly chosen for the training set, with an additional 500 pairs allocated for testing. Meanwhile, the newly created S1S2 dataset obtained a total of 7150 pairs of 256 × 256-sized image patches without overlapping. Among them, 6000 pairs were randomly selected for training, with 500 pairs designated for testing.

4.2. Methods for Comparison

The proposed method, dualCDM, was compared with six state-of-the-art image translation methods, including Attn-CycleGAN (ASGIT) [25], ICGAN [26], IUNet [27], SWIN-GAN (SWIN) [9], CDDPM [2], and SFDiff [14]. All compared methods were implemented using their default parameter settings. The code of the proposed method can be downloaded from https://github.com/isstncu/dualcdm (accessed on 28 June 2026).
The proposed method was trained over 600 epochs with a learning rate of 0.00002. The Adam approach was used as the optimizer with parameter β 1 = 0.9 , β 2 = 0.999 . The input size is 64 × 64. The proposed method randomly cropped each 256 × 256 image into sixteen 64 × 64 patches as the input. The batch size is 48. During inference, the proposed method employed deterministic accelerated sampling with 25 steps, following [28]. Training and testing were carried out on an NVIDIA Tesla A40.
Because GAN-based methods generate an output through a single forward pass, whereas diffusion models rely on iterative denoising, their computational profiles are not directly comparable. We, therefore, limit the complexity analysis to the three diffusion-based methods, namely CDDPM, SFDiff, and dualCDM. Table 2 reports the number of trainable parameters, training configurations, estimated memory consumption at a batch size of 1, and testing time for 500 images. DualCDM has the largest parameter count because of its additional spatial-frequency modeling and adaptive frequency-domain feature recalibration modules. Nevertheless, its estimated per-image memory requirement remains relatively low, mainly because the model processes 64 × 64 patches, which reduces the memory occupied by intermediate activations.

4.3. Metrics

The translated images were evaluated with six complementary metrics covering radiometric error, structural preservation, spectral consistency, and global synthesis quality: RMSE, RASE, SSIM, SAM, Q4 [29], and ERGAS [30]. SSIM was computed on grayscale images converted from the RGB channels. SSIM and Q4 are optimal at 1, whereas RMSE, RASE, SAM, and ERGAS are optimal at 0. Bold in comparison tables indicate the best scores.

5. Experimental Results

5.1. Qualitative Comparison on Two Datasets

To better illustrate the performance of our approach, we selected a group of images from the test data and presented the results visually, followed by a qualitative discussion. Figure 6 and Figure 7 show two groups of the selected images from the test data with different characteristics and study areas.
Figure 6 presents representative SAR-to-optical translation results over mountain, city, river, snowfield, desert, and cropland scenes, with all models trained on the SEN1-2 dataset. GAN-based methods produce visually plausible results in some regions, but noticeable artifacts, color shifts, and structural inconsistencies remain in challenging scenes. The diffusion-based methods, particularly CDDPM and SFDiff, generally provide more stable results. SFDiff achieves the strongest visual performance among the competing methods, with improved texture continuity and fewer artifacts than CDDPM in several cases. The proposed method further improves structural coherence and color consistency in most scenes, although some local artifacts remain in particularly ambiguous regions. Owing to space limitations, a detailed enlargement is provided only for the representative cropland scene. The enlarged region shows that the proposed method better preserves the boundaries between fields, roads, and built-up areas while producing a more coherent local color distribution.
Figure 7 presents additional results over cropland, forest, river, and city scenes from the S1S2 dataset. The models use two-channel Sentinel-1 GRD observations (VV and VH) together with multi-band optical references. The additional polarization and spectral information provides richer constraints for SAR-to-optical translation and generally improves the visual quality of the generated images. In urban areas, some GAN-based methods begin to recover clearer geometric patterns than those obtained from single-band inputs, but color distortion and unstable textures are still evident. Diffusion-based approaches produce more consistent color distributions and structural patterns across different land-cover types. Among them, SFDiff and the proposed method achieve comparable overall visual quality, while the proposed method generally provides improved color fidelity and structural coherence. A local enlargement of the representative city scene is included to make these differences more visible. The enlarged comparison shows that the proposed method retains the repetitive road and building patterns more clearly and produces fewer local color inconsistencies than most competing methods.

5.2. Quantitative Comparison on the SEN1-2 Dataset

To extend the analysis from qualitative visual inspection to the entire test set, we further conduct a comprehensive quantitative evaluation using multiple metrics computed over all test samples. Table 3 reports the quantitative comparison of different SAR-to-optical translation methods on the SEN1-2 dataset, where the results are averaged over 500 test images. As shown in the table, the proposed method consistently achieves the best performance across all evaluation metrics, including RMSE, SAM, RASE, ERGAS, SSIM, and Q4, indicating superior radiometric accuracy, structural fidelity, and spectral consistency. Among the GAN-based methods, more advanced architectures such as IUNet and SWIN show improved performance. However, their overall performance remains significantly inferior to that of diffusion-based methods, particularly in terms of spectral and structural consistency.
The SEN1-2 dataset contains multiple scene categories, which enables an evaluation of the proposed method under different land cover conditions. We perform scene specific tests on four representative categories, namely cropland, mountain, city, and river. Each category includes 10 test images. The quantitative results are reported in Table 4. The proposed method achieves the best performance in all four categories for most metrics, showing stable advantages in radiometric accuracy, structural fidelity, and spectral consistency. Performance in the city category is generally lower for all methods because the scene contains dense man made structures and complex spatial patterns that are difficult to infer from SAR observations alone. Even in this challenging setting, the proposed method remains the top performer across all metrics.

5.3. Quantitative Comparison on the S1S2 Dataset

Similar to the evaluation on the SEN1-2 dataset, we further report the quantitative comparison on the S1S2 dataset in Table 5. Benefiting from the higher radiometric resolution and multi-band optical targets in S1S2, the performance differences among competing methods become more clearly observable. As shown in the table, the proposed method achieves the best results across all metrics, including lower errors (RMSE, SAM, RASE, and ERGAS) and higher perceptual or structural scores (SSIM and Q4), indicating improved radiometric accuracy, spectral consistency, and structural fidelity. Moreover, diffusion-based approaches (CDDPM and SFDiff) consistently outperform GAN-based methods on this dataset, while the proposed method further advances the overall performance beyond the strongest diffusion baseline.

5.4. Test on Large Images

To further evaluate the effectiveness of the proposed method and its ability to produce spatially consistent predictions at a larger scale, we conducted an additional large image experiment on the S1S2 dataset. Specifically, we used the 12th image pair, where a 4120 × 5120 region was used for training and a disjoint 1000 × 1000 region was reserved for testing.
The qualitative and quantitative results are reported in Figure 8 and Table 6. As shown, GAN-based methods exhibit noticeable deviations from the ground truth, among which SWIN yields the best performance but still suffers from non-negligible radiometric and structural errors. In contrast, diffusion-based methods provide substantially improved results.
Notably, different from the previous two evaluations, the performance gap between SFDiff and CDDPM becomes smaller in this large image setting, while the proposed method still maintains a clear advantage across all metrics. A possible explanation is that this setting uses only a single large image for training, which substantially reduces sample diversity and limits the coverage of frequency patterns. As a result, SFDiff, whose frequency modeling mainly relies on point-wise modulation at each frequency, may have limited capacity to learn a reliable spectral prior and to generalize under such restricted training data. In contrast, the proposed method explicitly models cross-frequency and cross-channel correlations in the frequency domain, making it less sensitive to reduced data diversity and enabling more stable improvements in radiometric fidelity and structural consistency.

5.5. Statistical Significance Analysis

Although dualCDM generally improves the visual quality, it is not uniformly superior to the competing methods in every local region. For example, the mountain scene in Figure 6 still contains a bright granular artifact in the upper-left region, similar to that observed in the CDDPM result. In the river scenes in Figure 7, some areas retain SAR-derived spatial patterns and exhibit insufficient spectral transformation over water and shoreline regions. These cases indicate that challenging scattering conditions may still lead to local artifacts or incomplete modality conversion. Therefore, the qualitative comparisons alone are insufficient to establish a consistent advantage across the complete test set.
To determine whether the overall improvements are statistically reliable rather than being driven by a limited number of favorable samples, we performed a paired comparison between dualCDM and SFDiff, the strongest baseline in the quantitative evaluation. The six metrics were calculated separately for each test image, yielding paired observations for the two methods. A two-sided Wilcoxon signed-rank test was adopted because normality could not be assumed for the per-image metric differences. The resulting p-values were adjusted across the six metrics using the Holm procedure. We additionally report 95% bootstrap confidence intervals for the paired mean differences and rank-biserial correlations as effect sizes. A difference was considered statistically significant when the Holm-adjusted p-value was below 0.05.
The results are summarized in Table 7. On both SEN1-2 and S1S2, dualCDM significantly outperforms SFDiff across all six metrics. All adjusted p-values are below 0.05, and none of the 95% confidence intervals includes zero. The rank-biserial correlations range from 0.6031 to 0.9771 on SEN1-2 and from 0.6229 to 0.9975 on S1S2, indicating that the improvements are both statistically significant and broadly consistent across the test images. The strongest consistency is observed for SSIM, for which dualCDM outperforms SFDiff on 97.2% of the SEN1-2 samples and 99.6% of the S1S2 samples. These results show that, despite several local limitations, the overall performance gains are not determined by a small number of favorable cases.

5.6. Ablation Study

To validate the rationality and effectiveness of the proposed SFHRB and AFFRB modules, as well as the role of the two frequency-domain loss functions during training, we conduct a systematic and comprehensive ablation study on the S1S2 dataset. To expedite the ablation comparisons, all variants are trained for only 100 epochs and evaluated on the entire test set of S1S2. The results are summarized in Table 8.
First, by comparing Column 1 with Columns 2 and 3, we observe that incorporating frequency-domain learning substantially improves the overall quality of SAR-to-optical translation. The results further indicate that the time modulation mechanism is crucial for effective frequency-domain feature modeling in diffusion models. Second, by comparing Column 1 with Columns 4 and 5, we find that performing frequency-domain denoising on the conditional SAR input helps improve the translation results. Meanwhile, representing the phase using cos ( Φ ) and sin ( Φ ) effectively mitigates representation ambiguity and numerical instability caused by phase discontinuities. The results in Column 6 further demonstrate that jointly modeling the above two designs yields more pronounced improvements across multiple metrics. In addition, comparing Column 1 with Columns 7 and 8 shows that both the frequency-domain loss on noise and the frequency-domain loss on the one-step reconstruction can consistently enhance translation performance, where the latter brings a larger gain. Column 9 suggests that jointly optimizing the two frequency-domain losses leads to further improvements across all metrics. Finally, Column 10, which integrates the SFHRB module, the AFFRB module, and both frequency-domain losses, achieves the best performance on all evaluation metrics, thereby validating the synergistic effectiveness of the proposed modules and loss designs.

5.7. Visualization and Interpretation of SFHRB and AFFRB

Figure 9 visualizes the spatial branch, frequency branch, and their fused representations in SFHRB at three feature resolutions. At the 64 × 64 level, the spatial branch mainly responds to localized boundaries and fine-scale structures, whereas the frequency branch produces broader responses around the dominant diagonal structure. Their fusion retains local details while introducing more spatially continuous responses. At the 32 × 32 level, the frequency branch highlights several coherent and approximately parallel structures that are less evident in the spatial branch, and the fused representation integrates these responses with the locally extracted spatial features. At the 16 × 16 level, both branches mainly represent coarse regional patterns, while their combination provides a more complete description of the dominant large-scale structure. These observations indicate that the spatial and frequency branches learn complementary rather than redundant representations. The 8 × 8 level is omitted because its spatial resolution is too coarse for meaningful structural interpretation.
Figure 10 illustrates the frequency-dependent modulation learned by AFFRB and its effect on multi-scale SAR feature representations. The spectral gain maps are clearly non-uniform and vary across encoder levels, showing that AFFRB does not apply a fixed attenuation factor or a predefined frequency filter. At the 64 × 64 level, the response-change map shows a broad enhancement of fine-scale local variations. At the 32 × 32 level, the enhanced responses become more spatially selective and are concentrated around several linear and boundary-like patterns. At the 16 × 16 level, the response changes mainly correspond to broader regional structures, including the prominent diagonal feature in the lower-right region. These results indicate that the frequency-attention-guided AFFRB performs scale-dependent spectral recalibration and translates it into spatially varying refinement of SAR features.

6. Discussions: Generalization Performance

To further evaluate the generalization capability of the proposed method, we conduct additional experiments on test data that differ from the training conditions. Specifically, 150 test images are selected from the SEN1-2 dataset, covering three seasons other than spring, which is exclusively used for training. In this way, the evaluation reflects both seasonal variations and regional diversity that were not observed during training. The quantitative and qualitative results are summarized in Table 9 and Figure 11, respectively.
As expected, due to differences in geographic regions and acquisition times, the performance of all compared methods degrades to some extent. Such variations introduce changes in illumination conditions, vegetation states, and surface properties, which pose additional challenges for SAR-to-optical translation. Nevertheless, the proposed method consistently achieves the best performance across most evaluation metrics and produces visual results that are closest to the ground truth optical images. Both the quantitative scores and visual comparisons demonstrate that the proposed method exhibits stronger robustness to seasonal and regional changes, indicating superior generalization ability compared with competing methods.

7. Conclusions

In this study, we proposed dualCDM, a conditional diffusion framework for SAR-to-optical image translation that jointly models spatial- and frequency-domain information. In the diffusion backbone, SFHRB combines a spatial convolution branch with complex-valued convolution in the Fourier domain. By aggregating neighboring Fourier coefficients across input feature channels, the frequency branch captures local cross-frequency and cross-channel interactions, while timestep modulation adapts the learned responses to different diffusion stages. In the SAR conditional encoder, AFFRB predicts input-dependent real-valued gains from the magnitude and trigonometric phase representations of intermediate GRD features. This design provides adaptive frequency-domain recalibration without introducing an additional phase shift, while the residual connection preserves the original conditional information. The spatial-frequency dual-domain objective further constrains both noise prediction and one-step optical reconstruction, improving structural and frequency-domain consistency.
We also constructed the S1S2 dataset using 16-bit Sentinel-2 reflectance data, retaining the original 0–10,000 value range and including the near-infrared band. Experiments on SEN1-2 and S1S2 demonstrate that dualCDM improves radiometric accuracy, spectral consistency, and structural preservation over representative GAN- and diffusion-based methods. Paired statistical analyses further show significant improvements over SFDiff, the strongest competing method, across all six evaluation metrics on both datasets. Nevertheless, the qualitative results indicate that dualCDM is not uniformly superior in every local region, and artifacts or incomplete modality conversion can still occur in scenes with complex terrain, strong backscatter, or ambiguous water and shoreline responses. Future work will, therefore, focus on improving the robustness of the conditional representation under these challenging scattering conditions.

Author Contributions

Data curation, Y.M.; investigation, J.W.; software, Y.M.; writing—original draft, Y.M. and H.A.; writing—review and editing, L.C. and J.W.; resources, J.W. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported by the National Natural Science Foundation of China under grant 42267070 and the Natural Science Foundation of Jiangxi Province under grant 20252BAC240068.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Cao, R.; Chen, Y.; Chen, J.; Zhu, X.; Shen, M. Thick Cloud Removal in Landsat Images Based on Autoregression of Landsat Time-Series Data. Remote Sens. Environ. 2020, 249, 112001. [Google Scholar] [CrossRef] [Scilit]
  2. Bai, X.; Pu, X.; Xu, F. Conditional Diffusion for SAR-to-Optical Image Translation. IEEE Geosci. Remote Sens. Lett. 2023, 21, 4000605. [Google Scholar]
  3. Fu, Z.; Zhang, W. Research on Image Translation Between SAR and Optical Imagery. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2012, I-7, 273–278. [Google Scholar] [CrossRef] [Scilit]
  4. Enomoto, K.; Sakurada, K.; Wang, W.; Kawaguchi, N.; Matsuoka, M.; Nakamura, R. Image Translation Between SAR and Optical Imagery with Generative Adversarial Nets. In Proceedings of IEEE International Geoscience and Remote Sensing Symposium (IGARSS); IEEE: New York, NY, USA, 2018; pp. 1752–1755. [Google Scholar]
  5. Reyes, M.F.; Auer, S.; Merkle, N.; Henry, C.; Schmitt, M. SAR-to-Optical Image Translation Based on Conditional Generative Adversarial Networks: Optimization, Opportunities and Limits. Remote Sens. 2019, 11, 2067. [Google Scholar]
  6. Li, Y.; Fu, R.; Meng, X.; Jin, W.; Shao, F. A SAR-to-Optical Image Translation Method Based on Conditional Generative Adversarial Network (cGAN). IEEE Access 2020, 8, 60338–60343. [Google Scholar]
  7. Zhang, J.X.; Zhou, J.J.; Lu, X.W. Feature-Guided SAR-to-Optical Image Translation. IEEE Access 2020, 8, 70925–70937. [Google Scholar]
  8. Fu, S.; Xu, F.; Jin, Y.Q. Reciprocal Translation Between SAR and Optical Remote Sensing Images with Cascaded-Residual Adversarial Networks. Sci. China-Inf. Sci. 2021, 64, 122301. [Google Scholar] [CrossRef] [Scilit]
  9. Kong, Y.; Liu, S.; Peng, X. Multi-Scale Translation Method from SAR to Optical Remote Sensing Images Based on Conditional Generative Adversarial Network. Int. J. Remote Sens. 2022, 43, 2837–2860. [Google Scholar] [CrossRef] [Scilit]
  10. Wang, Z.W.; Ma, Y.; Zhang, Y. Hybrid cGAN: Coupling Global and Local Features for SAR-to-Optical Image Translation. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5236016. [Google Scholar]
  11. Seo, M.; Jung, J.; Choi, D.G. Improved Flood Insights: Diffusion-Based SAR-to-EO Image Translation. Remote Sens. 2025, 17, 2260. [Google Scholar]
  12. Shi, H.; Cui, Z.; Chen, L.; He, J.; Yang, J. A Brain-Inspired Approach for SAR-to-Optical Image Translation Based on Diffusion Models. Front. Neurosci. 2024, 18, 1352841. [Google Scholar] [PubMed]
  13. Xiong, Q.; Li, G.; Yao, X.; Zhang, X. SAR-to-Optical Image Translation and Cloud Removal Based on Conditional Generative Adversarial Networks: Literature Survey, Taxonomy, Evaluation Indicators, Limits and Future Directions. Remote Sens. 2023, 15, 1137. [Google Scholar]
  14. Qin, J.; Wang, K.; Zou, B.; Zhang, L.; van de Weijer, J. Conditional Diffusion Model with Spatial-Frequency Refinement for SAR-to-Optical Image Translation. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5226914. [Google Scholar]
  15. Xu, Z.Q.J.; Zhang, Y.Y.; Luo, T.; Xiao, Y.Y.; Ma, Z. Frequency Principle: Fourier Analysis Sheds Light on Deep Neural Networks. Commun. Comput. Phys. 2020, 28, 1746–1767. [Google Scholar] [CrossRef] [Scilit]
  16. Ruderman, D.L. The statistics of natural images. Netw. Comput. Neural Syst. 1994, 5, 517. [Google Scholar] [CrossRef] [Scilit]
  17. Ho, J.; Jain, A.; Abbeel, P. Denoising Diffusion Probabilistic Models. In Proceedings of the Advances in Neural Information Processing Systems; ACM: New York, NY, USA, 2020; Volume 33, pp. 6840–6851. [Google Scholar]
  18. Dhariwal, P.; Nichol, A. Diffusion Models Beat GANs on Image Synthesis. In Proceedings of the Advances in Neural Information Processing Systems; ACM: New York, NY, USA, 2021; Volume 34, pp. 8780–8794. [Google Scholar]
  19. Li, B.; Xue, K.; Liu, B.; Lai, Y. BBDM: Image-to-Image Translation with Brownian Bridge Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR); IEEE: New York, NY, USA, 2023; pp. 1952–1961. [Google Scholar]
  20. Rao, Y.; Zhao, W.; Zhu, Z.; Lu, J.; Zhou, J. Global Filter Networks for Image Classification. In Proceedings of the Advances in Neural Information Processing Systems; ACM: New York, NY, USA, 2021; Volume 34, pp. 980–993. [Google Scholar]
  21. Goyal, P.; Dollár, P.; Girshick, R.B.; Noordhuis, P.; Wesolowski, L.; Kyrola, A.; Tulloch, A.; Jia, Y.; He, K. Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour. arXiv 2017, arXiv:1706.02677. [Google Scholar]
  22. Nichol, A.; Dhariwal, P. Improved Denoising Diffusion Probabilistic Models. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 8162–8171. [Google Scholar]
  23. Schmitt, M.; Hughes, L.H.; Zhu, X.X. The SEN1-2 Dataset for Deep Learning in SAR-Optical Data Fusion. ISPRS Ann. Photogramm. Remote Sens. Spat. Inf. Sci. 2018, IV-1, 141–146. [Google Scholar] [CrossRef] [Scilit]
  24. Zhang, Q.; Liu, X.; Liu, M.; Zou, X.; Zhu, L.; Ruan, X. Comparative Analysis of Edge Information and Polarization on SAR-to-Optical Translation Based on Conditional Generative Adversarial Networks. Remote Sens. 2021, 13, 128. [Google Scholar]
  25. Lin, Y.; Wang, Y.; Li, Y.; Gao, Y.; Wang, Z.; Khan, L. Attention-Based Spatial Guidance for Image-to-Image Translation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV); IEEE: New York, NY, USA, 2021; pp. 816–825. [Google Scholar]
  26. Yang, X.; Zhao, J.Y.; Wei, Z.Y.; Wang, N.N.; Gao, X.B. SAR-to-Optical Image Translation Based on Improved cGAN. Pattern Recognit. 2022, 121, 108208. [Google Scholar]
  27. Sun, Y.; Jiang, W.; Yang, J.; Li, W. SAR Target Recognition Using cGAN-Based SAR-to-Optical Image Translation. Remote Sens. 2022, 14, 1793. [Google Scholar]
  28. Song, J.; Meng, C.; Ermon, S. Denoising Diffusion Implicit Models. In Proceedings of the International Conference on Learning Representations, Virtual, 3–7 May 2021. [Google Scholar]
  29. Alparone, L.; Baronti, S.; Garzelli, A.; Nencini, F. A Global Quality Measurement of Pan-Sharpened Multispectral Imagery. IEEE Geosci. Remote Sens. Lett. 2004, 1, 313–317. [Google Scholar] [CrossRef] [Scilit]
  30. Wald, L.; Ranchin, T.; Mangolini, M. Fusion of Satellite Images of Different Spatial Resolutions: Assessing the Quality of Resulting Images. Photogramm. Eng. Remote Sens. 1997, 63, 691–699. [Google Scholar]
Figure 1. Architecture of the dual-domain conditional diffusion models (dualCDM).
Figure 1. Architecture of the dual-domain conditional diffusion models (dualCDM).
Sensors 26 04183 g001
Figure 2. Structure of the spatial-frequency hybrid residual block (SFHRB).
Figure 2. Structure of the spatial-frequency hybrid residual block (SFHRB).
Sensors 26 04183 g002
Figure 3. Visual comparison of original Sentinel-1 GRD images, their frequency components (low, middle, high), and corresponding magnitude spectrums.
Figure 3. Visual comparison of original Sentinel-1 GRD images, their frequency components (low, middle, high), and corresponding magnitude spectrums.
Sensors 26 04183 g003
Figure 4. Structure of the adaptive frequency-domain feature recalibration block (AFFRB).
Figure 4. Structure of the adaptive frequency-domain feature recalibration block (AFFRB).
Sensors 26 04183 g004
Figure 5. Other fundamental modules in the proposed diffusion model.
Figure 5. Other fundamental modules in the proposed diffusion model.
Sensors 26 04183 g005
Figure 6. Visual comparison of SAR-to-optical translation results on the SEN1-2 dataset. The yellow boxes indicate the areas to be magnified.
Figure 6. Visual comparison of SAR-to-optical translation results on the SEN1-2 dataset. The yellow boxes indicate the areas to be magnified.
Sensors 26 04183 g006
Figure 7. Visual comparison of SAR-to-optical translation results on the S1S2 dataset. The first row represents the red, green, and blue bands of the image. The second row represents the NIR, red, and green bands of the image. The yellow boxes indicate the areas to be magnified.
Figure 7. Visual comparison of SAR-to-optical translation results on the S1S2 dataset. The first row represents the red, green, and blue bands of the image. The second row represents the NIR, red, and green bands of the image. The yellow boxes indicate the areas to be magnified.
Sensors 26 04183 g007
Figure 8. Visual comparison of large image SAR-to-optical translation results on the S1S2 dataset (the first row represents the red, green, and blue bands of the image. The second row represents the NIR, red, and green bands of the image).
Figure 8. Visual comparison of large image SAR-to-optical translation results on the S1S2 dataset (the first row represents the red, green, and blue bands of the image. The second row represents the NIR, red, and green bands of the image).
Sensors 26 04183 g008
Figure 9. Multi-scale spatial-frequency feature decomposition and fusion in SFHRB.
Figure 9. Multi-scale spatial-frequency feature decomposition and fusion in SFHRB.
Sensors 26 04183 g009
Figure 10. Visualization of learned spectral gains and multi-scale feature-response changes induced by AFFRB.
Figure 10. Visualization of learned spectral gains and multi-scale feature-response changes induced by AFFRB.
Sensors 26 04183 g010
Figure 11. Visual comparison of cross-region and cross-season generalization on the SEN1-2 dataset.
Figure 11. Visual comparison of cross-region and cross-season generalization on the SEN1-2 dataset.
Sensors 26 04183 g011
Table 1. Data sources of the new dataset (all captured in 2022).
Table 1. Data sources of the new dataset (all captured in 2022).
DateCityLatLonImage Size
S1S2
07.1907.19Gniezeno (Poland)52.517.36400 × 6400
09.0709.07Alexandria (Romania)43.125.46400 × 6400
09.2109.22Paris (France)48.62.66400 × 6400
10.1910.18Dnipro (Ukraine)48.535.06400 × 6400
10.2810.28Chicago (United States)40.3−94.86400 × 6400
10.3111.01Walcourt (Belgium)50.34.36400 × 6400
11.0311.03Warsaw (Poland)52.621.53840 × 3840
11.0411.04Patras (Greece)39.322.12560 × 2560
11.1211.13Nijmegen (Netherlands)51.85.63840 × 3840
11.2511.25Birmingham (United Kingdom)52.7−2.25120 × 5120
11.2911.28New South Wales (Australia)−29.4149.85120 × 5120
12.1812.18Suqian (China)30.4113.45120 × 5120
12.1812.18Suceava (Romania)47.726.16400 × 6400
12.2212.21Lancaster (United States)40.1−76.36400 × 6400
12.3012.31Shumen (Bulgaria)43.427.35120 × 5120
Table 2. Model complexity and computational efficiency among diffusion-based methods.
Table 2. Model complexity and computational efficiency among diffusion-based methods.
CDDPMSFDiffDualCDM
Params (M)60.14116.72135.14
Est. memory (GB)4.757.302.97
Time/500 (s)8500999810,023
Time/image (s)17.00019.99620.046
Table 3. Quantitative evaluation of SAR-to-optical translation on the SEN1-2 dataset.
Table 3. Quantitative evaluation of SAR-to-optical translation on the SEN1-2 dataset.
ModelRMSESAMRASEERGASSSIMQ4
ASGIT69.20710.17601.01781.09970.16230.2377
ICGAN55.71270.17440.71050.73590.20880.1985
IUNet53.28920.16860.71190.74650.20970.2018
SWIN53.18380.16760.64690.66350.22880.1808
CDDPM50.76110.12710.63990.65680.24020.2409
SFDiff40.00030.11800.48720.49620.28280.3582
proposed37.74430.11160.46000.46860.32730.4014
Table 4. Quantitative evaluation across different land cover scenes on the SEN1-2 dataset.
Table 4. Quantitative evaluation across different land cover scenes on the SEN1-2 dataset.
ModelRMSESAMRASEERGASSSIMQ4
Corpland
ASGIT70.62320.34241.85121.97370.13710.2681
ICGAN54.31680.31041.32741.40360.22790.2316
IUNet47.06590.28581.16221.21500.25470.2015
SWIN45.49250.30671.10341.15780.30870.2292
CDDPM47.88100.23651.26151.33190.29170.3021
SFDiff34.48480.20750.85710.89360.35270.4610
proposed32.53840.21140.80250.83920.40000.4892
Mountain
ASGIT52.08510.16170.65810.66720.21840.3311
ICGAN48.84180.11910.58960.59030.17540.2440
IUNet54.52090.13360.68740.68730.19610.3190
SWIN45.88740.12170.56560.56820.19410.2489
CDDPM41.37230.08660.50430.50660.30240.4555
SFDiff32.78700.08010.39900.39940.33760.5870
proposed30.99190.07960.37850.37880.39710.6231
City
ASGIT76.01680.13251.13411.14230.06780.2420
ICGAN56.02540.14350.85120.86230.11770.1564
IUNet57.30620.12350.85850.86780.09450.1233
SWIN50.05400.13770.72510.73370.13540.1688
CDDPM54.11090.10280.76480.76920.12360.1639
SFDiff46.20640.11400.64590.64880.15640.2726
proposed43.40530.09910.60710.61010.18370.2907
River
ASGIT88.53260.23551.47091.54040.23300.3901
ICGAN61.96190.22611.01431.02990.31150.3365
IUNet56.54940.22270.91620.92680.37990.3016
SWIN55.08030.22120.92990.95230.38960.3440
CDDPM39.72510.12290.67970.69700.49640.5388
SFDiff31.83730.10380.52920.54320.57050.6573
proposed30.55160.10910.50500.51700.59090.6819
Table 5. Quantitative evaluation of SAR-to-optical translation on the S1S2 dataset.
Table 5. Quantitative evaluation of SAR-to-optical translation on the S1S2 dataset.
ModelRMSESAMRASEERGASSSIMQ4
ASGIT784.00770.18400.38050.30850.47680.2681
ICGAN523.42260.09880.25430.20850.52590.3994
IUNet519.86150.09870.25270.20850.52580.4085
SWIN534.74000.10560.25750.21620.50920.3920
CDDPM477.49490.08540.23060.19530.59850.4467
SFDiff418.23570.07540.20250.16660.62350.5141
proposed401.42970.07280.19430.15910.65070.5369
Table 6. Quantitative evaluation of large image testing on the S1S2 dataset.
Table 6. Quantitative evaluation of large image testing on the S1S2 dataset.
ModelRMSESAMRASEERGASSSIMQ4
ASGIT526.79460.11540.24940.20260.76590.2054
ICGAN486.67290.10560.23040.19150.74800.2227
IUNet515.68910.10870.24420.20550.70460.2264
SWIN453.88260.09920.21490.18010.74330.2398
CDDPM409.08490.08330.19370.16120.81410.3156
SFDiff405.89780.08340.19220.16030.80750.3311
proposed385.91930.08110.18270.15310.82050.3471
Table 7. Statistical significance analysis of dualCDM against SFDiff on the SEN1-2 and S1S2 datasets.
Table 7. Statistical significance analysis of dualCDM against SFDiff on the SEN1-2 and S1S2 datasets.
DatasetMetricSFDiffDualCDM Δ 95% CI of Δ Adjusted p r rb Win Rate (%)
SEN1-2RMSE40.000337.74432.2560[1.8984, 2.6084] 3.66 × 10 44 0.723582.2
SAM0.11800.11160.0063[0.0048, 0.0078] 1.51 × 10 31 0.603176.6
RASE0.48720.46000.0272[0.0213, 0.0323] 1.05 × 10 44 0.729982.2
ERGAS0.49620.46860.0276[0.0214, 0.0333] 1.05 × 10 44 0.729182.2
SSIM0.28280.32730.0450[0.0420, 0.0481] 3.78 × 10 79 0.977197.2
Q40.35820.40140.0432[0.0368, 0.0496] 1.28 × 10 34 0.636375.0
S1S2RMSE418.2357401.429716.806[14.9782, 18.6364] 1.17 × 10 57 0.829686.6
SAM0.07540.07280.0026[0.0022, 0.0029] 8.25 × 10 45 0.727579.6
RASE0.20250.19430.0082[0.0073, 0.0090] 6.23 × 10 58 0.832686.6
ERGAS0.16660.15910.0075[0.0068, 0.0083] 6.69 × 10 62 0.861988.6
SSIM0.62350.65070.0270[0.0259, 0.0281] 1.91 × 10 82 0.997599.6
Q40.51410.53690.0228[0.0192, 0.0264] 1.55 × 10 33 0.622974.8
Note: The paired difference Δ is defined such that a positive value consistently indicates better performance of dualCDM. The p-values are obtained using a two-sided Wilcoxon signed-rank test and adjusted across the six metrics within each dataset using the Holm procedure. r rb denotes the rank-biserial correlation, and the win rate is the percentage of test images for which dualCDM outperforms SFDiff.
Table 8. Ablation experiment.
Table 8. Ablation experiment.
SFHRB (without time modulation)
SFHRB
AFFRB (FreqAttn- concat M , Φ )
AFFRB (FreqAttn- concat M , cos ( Φ ) , sin ( Φ ) )
L f ( 1 )
L f ( 2 )
RMSE497.0738493.1099486.5744493.7522491.5883482.3787481.0528460.4320458.7144449.2712
SAM0.08610.08590.08500.08550.08540.08370.08360.08290.08200.0802
RASE0.24090.23890.23570.23930.23820.23360.23310.22300.22210.2175
ERGAS0.20390.20130.19770.20230.20110.19630.19420.18470.18370.1795
SSIM0.59120.59080.59830.59120.59320.60070.60780.60100.61070.6150
Q40.43180.43550.44500.43610.43660.44910.46660.47370.47970.4946
Table 9. Quantitative evaluation of cross-region and cross-season generalization on the SEN1-2 dataset.
Table 9. Quantitative evaluation of cross-region and cross-season generalization on the SEN1-2 dataset.
ModelRMSESAMRASEERGASSSIMQ4
ASGIT65.73330.20430.89010.95400.14200.1890
ICGAN60.39540.20930.79830.86340.19700.1363
IUNet60.97910.21090.79720.83520.19870.1332
SWIN57.52100.19890.76220.79830.22690.1336
CDDPM61.67160.19230.87040.91720.18750.1315
SFDiff56.77920.18960.76890.80560.21370.1340
proposed53.47200.18960.68840.71850.25310.1542
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Ma, Y.; Aghababaei, H.; Chang, L.; Wei, J. DualCDM: Dual-Domain Conditional Diffusion for SAR-to-Optical Translation with Spatial–Frequency Correlation and Adaptive Feature Recalibration. Sensors 2026, 26, 4183. https://doi.org/10.3390/s26134183

AMA Style

Ma Y, Aghababaei H, Chang L, Wei J. DualCDM: Dual-Domain Conditional Diffusion for SAR-to-Optical Translation with Spatial–Frequency Correlation and Adaptive Feature Recalibration. Sensors. 2026; 26(13):4183. https://doi.org/10.3390/s26134183

Chicago/Turabian Style

Ma, Yaobin, Hossein Aghababaei, Ling Chang, and Jingbo Wei. 2026. "DualCDM: Dual-Domain Conditional Diffusion for SAR-to-Optical Translation with Spatial–Frequency Correlation and Adaptive Feature Recalibration" Sensors 26, no. 13: 4183. https://doi.org/10.3390/s26134183

APA Style

Ma, Y., Aghababaei, H., Chang, L., & Wei, J. (2026). DualCDM: Dual-Domain Conditional Diffusion for SAR-to-Optical Translation with Spatial–Frequency Correlation and Adaptive Feature Recalibration. Sensors, 26(13), 4183. https://doi.org/10.3390/s26134183

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop