Figure 1.
Overall architecture of TCM-CR. The acquisition with the lowest scene-level cloud fraction is selected from the multi-temporal input as the reference image . An encoder–decoder—with cross-modal spatial modeling at each encoder level and cloud-aware temporal modeling on the deepest features—produces a residual, which is bounded by , scaled by , modulated by the cloud-driven soft gate , and superimposed on the reference image: .
Figure 1.
Overall architecture of TCM-CR. The acquisition with the lowest scene-level cloud fraction is selected from the multi-temporal input as the reference image . An encoder–decoder—with cross-modal spatial modeling at each encoder level and cloud-aware temporal modeling on the deepest features—produces a residual, which is bounded by , scaled by , modulated by the cloud-driven soft gate , and superimposed on the reference image: .
Figure 2.
Reference-image selection and soft-gate construction. (a) Comparison on one sample of the per-pixel least-cloudy composite (left), the least-cloudy acquisition (middle), and the ground truth (right), with the PSNR of the two selection schemes annotated in the panels: cloud-mask noise causes mis-selections in per-pixel compositing, whereas scene-level selection is more robust. (b) Construction of the soft gate: reference image (RGB), raw cloud mask, local smoothing, soft threshold, and the final gate (including the scene-fraction scaling factor); the gate confines the correction to cloud-covered regions.
Figure 2.
Reference-image selection and soft-gate construction. (a) Comparison on one sample of the per-pixel least-cloudy composite (left), the least-cloudy acquisition (middle), and the ground truth (right), with the PSNR of the two selection schemes annotated in the panels: cloud-mask noise causes mis-selections in per-pixel compositing, whereas scene-level selection is more robust. (b) Construction of the soft gate: reference image (RGB), raw cloud mask, local smoothing, soft threshold, and the final gate (including the scene-fraction scaling factor); the gate confines the correction to cloud-covered regions.
Figure 3.
Structure of the state-space modules. (a) Cross-modal spatial modeling: optical and SAR features of the same resolution are paired and aggregated bidirectionally along the horizontal and vertical directions, with SAR guiding the spatial propagation of the optical features. (b) Cloud-aware temporal modeling: multi-temporal features are integrated along both temporal directions, with each date’s update to the state attenuated by the factor according to its cloud fraction.
Figure 3.
Structure of the state-space modules. (a) Cross-modal spatial modeling: optical and SAR features of the same resolution are paired and aggregated bidirectionally along the horizontal and vertical directions, with SAR guiding the spatial propagation of the optical features. (b) Cloud-aware temporal modeling: multi-temporal features are integrated along both temporal directions, with each date’s update to the state attenuated by the factor according to its cloud fraction.
Figure 4.
Study regions and sample distribution. (a) Geographic locations of the four SEN12MS-CR-TS regions (asiaWest, europa, africa, america). (b) Distribution of samples by the scene-average cloud fraction of the reference image; vertical lines mark the grouping thresholds (5%, 20%, 50%) dividing the samples into the clear, light, moderate, and heavy groups.
Figure 4.
Study regions and sample distribution. (a) Geographic locations of the four SEN12MS-CR-TS regions (asiaWest, europa, africa, america). (b) Distribution of samples by the scene-average cloud fraction of the reference image; vertical lines mark the grouping thresholds (5%, 20%, 50%) dividing the samples into the clear, light, moderate, and heavy groups.
Figure 5.
Accuracy comparison grouped by cloud cover (four-region validation set, n = 1125, training regions included). (a) PSNR of each method across cloud-cover groups: TCM-CR nearly coincides with the reference image on the clear and light groups, whereas both UnCRtainTS versions fall about 9 dB below the reference image on the clear group. (b) Trade-off scatter of the accuracy change on clear samples (horizontal axis) against the gain on heavy samples (vertical axis); the shaded band marks the no-harm band dB, and TCM-CR lies at the end where accuracy is preserved and the gain remains high.
Figure 5.
Accuracy comparison grouped by cloud cover (four-region validation set, n = 1125, training regions included). (a) PSNR of each method across cloud-cover groups: TCM-CR nearly coincides with the reference image on the clear and light groups, whereas both UnCRtainTS versions fall about 9 dB below the reference image on the clear group. (b) Trade-off scatter of the accuracy change on clear samples (horizontal axis) against the gain on heavy samples (vertical axis); the shaded band marks the no-harm band dB, and TCM-CR lies at the end where accuracy is preserved and the gain remains high.
Figure 6.
Cross-region generalization and threshold robustness (the tested region is entirely excluded from training). (a) europa: PSNR by cloud-cover group; the heavy group improves from 12.88 dB to 20.15 dB (). On the same heavy samples, UnCRtainTS (L2) reaches 21.82 dB, but its public weights were trained on data covering europa, so the comparison favors it. (b) america: the clear and light groups are per-pixel identical to the reference image, and the moderate group improves by 0.32 dB. (c) Varying the soft-gate threshold over 0.01–0.50 leaves heavy-group PSNR at 20.15 dB throughout; only degrades the light group.
Figure 6.
Cross-region generalization and threshold robustness (the tested region is entirely excluded from training). (a) europa: PSNR by cloud-cover group; the heavy group improves from 12.88 dB to 20.15 dB (). On the same heavy samples, UnCRtainTS (L2) reaches 21.82 dB, but its public weights were trained on data covering europa, so the comparison favors it. (b) america: the clear and light groups are per-pixel identical to the reference image, and the moderate group improves by 0.32 dB. (c) Varying the soft-gate threshold over 0.01–0.50 leaves heavy-group PSNR at 20.15 dB throughout; only degrades the light group.
Figure 7.
Spatial distribution of the residual (one moderate and one heavy sample from the europa validation set; the model was not trained on this region). Columns: reference image (RGB), cloud mask, residual magnitude, and residual overlay; the color bar gives the residual magnitude in normalized reflectance units. Corrections concentrate on cloud-covered regions, while residuals at cloud-free pixels are strongly suppressed by the gate (not strictly zero; scene-wide zeroing occurs only when the cloud fraction of the reference image is below the threshold).
Figure 7.
Spatial distribution of the residual (one moderate and one heavy sample from the europa validation set; the model was not trained on this region). Columns: reference image (RGB), cloud mask, residual magnitude, and residual overlay; the color bar gives the residual magnitude in normalized reflectance units. Corrections concentrate on cloud-covered regions, while residuals at cloud-free pixels are strongly suppressed by the gate (not strictly zero; scene-wide zeroing occurs only when the cloud fraction of the reference image is below the threshold).
Figure 8.
Reconstruction examples and per-pixel absolute error at different cloud-cover levels. The four rows are, from top to bottom, a clear, light, moderate, and heavy sample (the heavy row is from europa, on which the model was not trained); the row label gives the group and the scene-average cloud fraction of the reference image. The six left columns show SAR (VV), the cloudiest acquisition, the reference image, UnCRtainTS, TCM-CR, and the ground truth (B4-B3-B2 true-color composites, annotated with PSNR, optimal values highlighted in green); the two right columns show the per-pixel absolute error of UnCRtainTS and TCM-CR (mean over 13 bands, reflectance scale, annotated with MAE), with the shared color bar capped at the 98th percentile of all error maps.
Figure 8.
Reconstruction examples and per-pixel absolute error at different cloud-cover levels. The four rows are, from top to bottom, a clear, light, moderate, and heavy sample (the heavy row is from europa, on which the model was not trained); the row label gives the group and the scene-average cloud fraction of the reference image. The six left columns show SAR (VV), the cloudiest acquisition, the reference image, UnCRtainTS, TCM-CR, and the ground truth (B4-B3-B2 true-color composites, annotated with PSNR, optimal values highlighted in green); the two right columns show the per-pixel absolute error of UnCRtainTS and TCM-CR (mean over 13 bands, reflectance scale, annotated with MAE), with the shared color bar capped at the 98th percentile of all error maps.
Table 1.
Overall accuracy on the asiaWest test set (n = 300).
Table 1.
Overall accuracy on the asiaWest test set (n = 300).
| Method | PSNR | SSIM | SAM | MAE | LPIPS |
|---|
| Reference image (least-cloudy acquisition) | 30.045 | 0.7830 | 0.2337 | 0.0544 | 0.2506 |
| Per-pixel least-cloudy composite | 20.384 | 0.6185 | 0.2959 | 0.0836 | 0.5570 |
| Temporal median composite | 19.280 | 0.6315 | 0.3495 | 0.1130 | 0.4556 |
| UnCRtainTS (L2) | 25.046 | 0.7943 | 0.2119 | 0.0457 | 0.2714 |
| UnCRtainTS (MGNLL) | 25.715 | 0.8079 | 0.1972 | 0.0487 | 0.2212 |
| TCM-CR (ours) | 30.015 | 0.7742 | 0.2376 | 0.0546 | 0.2646 |
Table 2.
PSNR (dB) on the four-region validation set, grouped by cloud cover.
Table 2.
PSNR (dB) on the four-region validation set, grouped by cloud cover.
| Cloud Cover | n | Reference Image | TCM-CR | UnCRtainTS (L2) | UnCRtainTS (MGNLL) |
|---|
| Clear (<5%) | 112 | 37.42 | 37.41 (−0.01) | 28.15 (−9.27) | 28.50 (−8.92) |
| Light (5–20%) | 539 | 25.46 | 25.45 (−0.01) | 25.12 (−0.34) | 24.13 (−1.33) |
| Moderate (20–50%) | 319 | 19.33 | 19.81 (+0.48) | 23.01 (+3.67) | 22.06 (+2.73) |
| Heavy (>50%) | 155 | 12.88 | 20.81 (+7.93) | 21.82 (+8.94) | 21.08 (+8.20) |
| Overall | 1125 | 23.18 | 24.40 (+1.22) | 24.37 (+1.19) | 23.56 (+0.38) |
Table 3.
PSNR (dB) on europa, grouped by cloud cover.
Table 3.
PSNR (dB) on europa, grouped by cloud cover.
| Cloud Cover | n | Reference Image | TCM-CR |
|---|
| Light (5–20%) | 51 | 18.47 | 18.47 (−0.00) |
| Moderate (20–50%) | 154 | 14.94 | 16.47 (+1.53) |
| Heavy (>50%) | 155 | 12.88 | 20.15 (+7.27) |
| Overall | 360 | 14.55 | 18.34 (+3.79) |
Table 4.
PSNR (dB) on america, grouped by cloud cover.
Table 4.
PSNR (dB) on america, grouped by cloud cover.
| Cloud Cover | n | Reference Image | TCM-CR |
|---|
| Clear (<5%) | 12 | 48.60 | 48.60 (−0.00) |
| Light (5–20%) | 174 | 30.31 | 30.31 (−0.00) |
| Moderate (20–50%) | 84 | 24.10 | 24.42 (+0.32) |
Table 5.
Accuracy on cloud-free versus cloud-covered pixels within moderate and heavy samples (four-region validation set).
Table 5.
Accuracy on cloud-free versus cloud-covered pixels within moderate and heavy samples (four-region validation set).
| Group | Pixel Subset | Reference PSNR | TCM-CR PSNR | Reference MAE | TCM-CR MAE |
|---|
| Moderate (n = 319) | Cloud-free | 19.94 | 20.09 | 0.2134 | 0.2008 |
| Moderate | Cloud-covered | 18.24 | 19.33 | 0.2380 | 0.2166 |
| Heavy (n = 155) | Cloud-free | 13.07 | 16.76 | 0.3900 | 0.2421 |
| Heavy | Cloud-covered | 12.86 | 21.67 | 0.4010 | 0.1377 |
Table 6.
Ablation results (asiaWest test set, n = 300).
Table 6.
Ablation results (asiaWest test set, n = 300).
| Configuration | PSNR | SSIM | SAM | MAE | LPIPS |
|---|
| Plain encoder–decoder | 23.52 | 0.750 | 0.292 | 0.057 | 0.318 |
| +reference-image residual, cross-modal spatial modeling | 30.16 | 0.784 | 0.231 | 0.053 | 0.259 |
| +temporal modeling | 30.11 | 0.786 | 0.227 | 0.054 | 0.256 |
| +cloud-aware weighting (full TCM-CR) | 30.10 | 0.782 | 0.233 | 0.054 | 0.257 |
| −reference-image residual | 22.31 | 0.574 | 0.458 | 0.076 | 0.399 |
| −SSIM loss | 30.12 | 0.787 | 0.226 | 0.054 | 0.257 |
Table 7.
Ablation on the four-region validation set, PSNR (dB), by cloud-cover group (heavy-group SSIM in the last column).
Table 7.
Ablation on the four-region validation set, PSNR (dB), by cloud-cover group (heavy-group SSIM in the last column).
| Configuration | Clear | Light | Moderate | Heavy | Heavy SSIM |
|---|
| Reference image | 37.42 | 25.46 | 19.33 | 12.88 | 0.539 |
| +reference-image residual, cross-modal spatial modeling | 37.41 | 25.45 | 19.87 | 20.09 | 0.701 |
| +temporal modeling | 37.41 | 25.45 | 19.95 | 20.65 | 0.698 |
| +cloud-aware weighting (full TCM-CR) | 37.41 | 25.45 | 19.81 | 20.81 | 0.726 |
Table 8.
Fusion-mechanism ablation, PSNR (dB) by cloud-cover group (four-region validation set).
Table 8.
Fusion-mechanism ablation, PSNR (dB) by cloud-cover group (four-region validation set).
| Configuration | Clear | Light | Moderate | Heavy |
|---|
| Reference image | 37.42 | 25.46 | 19.33 | 12.88 |
| Optical + SAR concatenation, no reference anchor (whole-scene) | 28.06 | 27.17 | 24.57 | 23.95 |
| Reference anchor + optical + SAR concatenation + temporal | 37.41 | 25.45 | 20.03 | 19.67 |
| Reference anchor + cross-modal spatial modeling + temporal (full) | 37.41 | 25.45 | 19.81 | 20.81 |
Table 9.
Effect of reference-image reliability on reconstruction, PSNR (dB); the four-region validation set is stratified into reliability quartiles.
Table 9.
Effect of reference-image reliability on reconstruction, PSNR (dB); the four-region validation set is stratified into reliability quartiles.
| Reference Reliability | n | Reference Image | TCM-CR | Offset |
|---|
| Low (most surface change) | 274 | 13.65 | 18.04 | +4.39 |
| Medium–low | 273 | 20.18 | 20.28 | +0.10 |
| Medium–high | 274 | 23.60 | 23.46 | −0.14 |
| High (reference reliable) | 274 | 36.60 | 36.44 | −0.16 |
Table 10.
Sensitivity to cloud-mask error at inference, PSNR (dB) by cloud-cover group (four-region validation set).
Table 10.
Sensitivity to cloud-mask error at inference, PSNR (dB) by cloud-cover group (four-region validation set).
| Mask Perturbation | Clear | Light | Moderate | Heavy | Overall |
|---|
| None (clean mask) | 37.41 | 25.45 | 19.81 | 20.81 | 24.40 |
| False positive 5% | 36.91 | 25.44 | 19.73 | 21.33 | 24.40 |
| False positive 10% | 36.89 | 25.39 | 18.82 | 21.26 | 24.10 |
| False positive 20% | 36.40 | 21.00 | 17.59 | 21.27 | 21.60 |
| False negative 5% | 32.96 | 25.46 | 19.73 | 20.72 | 23.93 |
| False negative 10% | 32.78 | 25.32 | 19.59 | 20.64 | 23.79 |
| False negative 20% | 32.78 | 24.57 | 19.33 | 20.45 | 23.33 |