Author Contributions
Conceptualization, X.W. and C.L.; methodology, X.W.; software, X.W.; validation, X.W., Y.S. and H.Z.; formal analysis, X.W.; investigation, X.W., Y.S. and H.Z.; resources, C.L.; data curation, X.W.; writing—original draft preparation, X.W.; writing—review and editing, Y.S., H.Z. and C.L.; visualization, X.W.; supervision, C.L.; project administration, C.L.; funding acquisition, C.L. All authors have read and agreed to the published version of the manuscript.
Figure 1.
Comparison of representative change detection frameworks and FD-ProtoSCD. Solid arrows indicate the main prediction flow, while dotted/dashed arrows denote feature interaction or auxiliary guidance paths. Colored blocks distinguish inputs, encoders, decoders, priors/prototypes, and output maps as labeled in the figure.
Figure 1.
Comparison of representative change detection frameworks and FD-ProtoSCD. Solid arrows indicate the main prediction flow, while dotted/dashed arrows denote feature interaction or auxiliary guidance paths. Colored blocks distinguish inputs, encoders, decoders, priors/prototypes, and output maps as labeled in the figure.
Figure 2.
Motivation of FD-ProtoSCD. (a) Pseudo-change interference from high-frequency appearance shifts. (b) Long-tailed transition distribution and target-class IoU on SECOND.
Figure 2.
Motivation of FD-ProtoSCD. (a) Pseudo-change interference from high-frequency appearance shifts. (b) Long-tailed transition distribution and target-class IoU on SECOND.
Figure 3.
Overall architecture of FD-ProtoSCD. The Siamese encoder extracts multi-scale features from bi-temporal images; FDCD modules disentangle semantic differences from pseudo-changes in the frequency domain; the DCPD decoder performs focal-re-weighted prototype cross-attention; three output heads produce the binary change mask and the pre-/post-change semantic maps. Solid arrows show feature and prediction flow, while dashed or dotted arrows indicate supervision and auxiliary loss paths. Colored blocks distinguish the encoder, FPN aggregation, FDCD, DCPD, semantic heads, binary head, ground truths, and loss terms.
Figure 3.
Overall architecture of FD-ProtoSCD. The Siamese encoder extracts multi-scale features from bi-temporal images; FDCD modules disentangle semantic differences from pseudo-changes in the frequency domain; the DCPD decoder performs focal-re-weighted prototype cross-attention; three output heads produce the binary change mask and the pre-/post-change semantic maps. Solid arrows show feature and prediction flow, while dashed or dotted arrows indicate supervision and auxiliary loss paths. Colored blocks distinguish the encoder, FPN aggregation, FDCD, DCPD, semantic heads, binary head, ground truths, and loss terms.
Figure 4.
Detailed architecture of the Frequency-Domain Change Disentanglement (FDCD) module. The module applies 2D Fast Fourier Transform to bi-temporal feature maps, yielding frequency spectra that are separated by a learnable mask into low-frequency semantics and high-frequency appearances. Arrows show the FFT-based decomposition and feature-difference flow; colored branches distinguish the low-frequency semantic path, high-frequency appearance path, and auxiliary pseudo-change loss. Differences in the semantic subspace highlight genuine context changes, while variations in the appearance subspace are constrained by a pseudo-change suppression loss to eliminate seasonal or lighting artifacts.
Figure 4.
Detailed architecture of the Frequency-Domain Change Disentanglement (FDCD) module. The module applies 2D Fast Fourier Transform to bi-temporal feature maps, yielding frequency spectra that are separated by a learnable mask into low-frequency semantics and high-frequency appearances. Arrows show the FFT-based decomposition and feature-difference flow; colored branches distinguish the low-frequency semantic path, high-frequency appearance path, and auxiliary pseudo-change loss. Differences in the semantic subspace highlight genuine context changes, while variations in the appearance subspace are constrained by a pseudo-change suppression loss to eliminate seasonal or lighting artifacts.
Figure 5.
Structure of the Dynamic Class Prototype Decoder (DCPD). Using the semantic difference features from FDCD as queries, the module applies dual-prototype cross-attention against the continuously updated semantic prototype banks of both temporal branches. Arrows denote feature, prototype-update, and attention flows; colored blocks distinguish prototype memory, attention, feed-forward decoding, and focal contrastive supervision. This explicitly guides the decoder to focus on specific land-cover transitions and relies on a focal contrastive loss to mitigate the class imbalance problem by re-weighting gradients for rare transitions.
Figure 5.
Structure of the Dynamic Class Prototype Decoder (DCPD). Using the semantic difference features from FDCD as queries, the module applies dual-prototype cross-attention against the continuously updated semantic prototype banks of both temporal branches. Arrows denote feature, prototype-update, and attention flows; colored blocks distinguish prototype memory, attention, feed-forward decoding, and focal contrastive supervision. This explicitly guides the decoder to focus on specific land-cover transitions and relies on a focal contrastive loss to mitigate the class imbalance problem by re-weighting gradients for rare transitions.
Figure 6.
Qualitative comparison on SECOND, group 1. Subfigures (a–d) show four representative T1/T2 sample pairs. In each subfigure, rows are arranged as T1/T2 pairs; columns show the input images, semantic ground truth, SCD-UperNet, M-CD, FEM-CD, ChangeMask, ClearSCD, and FD-ProtoSCD.
Figure 6.
Qualitative comparison on SECOND, group 1. Subfigures (a–d) show four representative T1/T2 sample pairs. In each subfigure, rows are arranged as T1/T2 pairs; columns show the input images, semantic ground truth, SCD-UperNet, M-CD, FEM-CD, ChangeMask, ClearSCD, and FD-ProtoSCD.
Figure 7.
Qualitative comparison on SECOND, group 2. Subfigures (
a–
d) show four additional T1/T2 sample pairs. The layout is identical to
Figure 6, so the methods can be compared across different transition patterns without re-learning the legend.
Figure 7.
Qualitative comparison on SECOND, group 2. Subfigures (
a–
d) show four additional T1/T2 sample pairs. The layout is identical to
Figure 6, so the methods can be compared across different transition patterns without re-learning the legend.
Figure 8.
Qualitative comparison on SECOND, group 3. Subfigures (a–d) show four T1/T2 sample pairs. This set emphasizes small and mixed transitions where semantic confusion is easy to spot.
Figure 8.
Qualitative comparison on SECOND, group 3. Subfigures (a–d) show four T1/T2 sample pairs. This set emphasizes small and mixed transitions where semantic confusion is easy to spot.
Figure 9.
Qualitative comparison on SECOND, group 4. Subfigures (a–d) show four T1/T2 sample pairs. The final set collects the hardest remaining examples, so the comparison includes both sparse change regions and more fragmented structures.
Figure 9.
Qualitative comparison on SECOND, group 4. Subfigures (a–d) show four T1/T2 sample pairs. The final set collects the hardest remaining examples, so the comparison includes both sparse change regions and more fragmented structures.
Figure 10.
Image-level frequency decomposition for one SECOND test pair. Top row (T1): input image, log-magnitude spectrum, fixed low-pass mask, low-frequency reconstruction, high-frequency residual, and high-frequency difference. Bottom row (T2): same decomposition, with the final cell showing the low-frequency difference. RGB panels use the original remote sensing colors; spectrum panels use the displayed color scale to show log-magnitude intensity; white regions in the mask denote retained low-frequency components and black regions denote suppressed high-frequency components.
Figure 10.
Image-level frequency decomposition for one SECOND test pair. Top row (T1): input image, log-magnitude spectrum, fixed low-pass mask, low-frequency reconstruction, high-frequency residual, and high-frequency difference. Bottom row (T2): same decomposition, with the final cell showing the low-frequency difference. RGB panels use the original remote sensing colors; spectrum panels use the displayed color scale to show log-magnitude intensity; white regions in the mask denote retained low-frequency components and black regions denote suppressed high-frequency components.
Figure 11.
Per-transition IoU matrices and GT transition frequency. Left: ground-truth transition pixel frequency (log scale). Center: SCD-UperNet per-transition IoU. Right: FD-ProtoSCD per-transition IoU. The matrices provide a transition-level diagnostic of how the two models distribute their correct predictions.
Figure 11.
Per-transition IoU matrices and GT transition frequency. Left: ground-truth transition pixel frequency (log scale). Center: SCD-UperNet per-transition IoU. Right: FD-ProtoSCD per-transition IoU. The matrices provide a transition-level diagnostic of how the two models distribute their correct predictions.
Figure 12.
Unmasked output comparison on a representative SECOND test pair. White denotes predicted change and black denotes unchanged pixels in the last three panels. The final two panels are the requested outputs from a conventional spatial-centric SCD-UPerNet baseline and the proposed frequency-disentangled/prototype-guided paradigm, respectively. No ground-truth mask is applied to either prediction, so false alarms remain visible.
Figure 12.
Unmasked output comparison on a representative SECOND test pair. White denotes predicted change and black denotes unchanged pixels in the last three panels. The final two panels are the requested outputs from a conventional spatial-centric SCD-UPerNet baseline and the proposed frequency-disentangled/prototype-guided paradigm, respectively. No ground-truth mask is applied to either prediction, so false alarms remain visible.
Table 1.
Pseudocode-style training and inference procedure of FD-ProtoSCD.
Table 1.
Pseudocode-style training and inference procedure of FD-ProtoSCD.
| Stage | Procedure |
|---|
| Training | Input: mini-batch . (1) Extract four-level Siamese feature pyramids. (2) At each level, apply FDCD spectral separation, obtain the cleaned semantic difference, and accumulate the pseudo-change loss over unchanged pixels. (3) Fuse the multi-scale difference features. (4) Update both semantic prototype banks by EMA. (5) Apply dual-prototype cross-attention and predict , , and . (6) Compute Equation (9) and back-propagate. |
| Inference | Input: a co-registered bi-temporal image pair. (1) Run the same Siamese encoder and FDCD modules. (2) Fuse multi-scale differences and decode them with fixed prototype banks; no EMA update or loss is computed. (3) Produce the two semantic maps and binary change mask. (4) For large scenes, apply overlapping sliding windows and stitch the outputs to the original spatial size. |
Table 2.
ClearSCD-style comparison on the SECOND dataset. Best results are in bold. “—” indicates that the metric is not available for that model type or source.
Table 2.
ClearSCD-style comparison on the SECOND dataset. Best results are in bold. “—” indicates that the metric is not available for that model type or source.
| Method | Params (M) | FPS | Binary Change Detection | Semantic/SCD Prediction | Semantic Consistency |
|---|
|
IoU (%)
|
F1bcd (%)
|
mIoU (%)
|
F-Score (%)
| |
SCD Score
|
|---|
| SNUNet-CD [47] | 3.01 | | 48.79 | 65.58 | — | — | — | — |
| BIT [6] | 2.99 | | 50.66 | 67.25 | — | — | — | — |
| Changer [7] | 11.39 | | 49.19 | 65.94 | — | — | — | — |
| SCD-UperNet | 34.03 | | 52.41 | 68.78 | 42.62 | 57.20 | 16.14 | 32.03 |
| ChangeMask [3] | 10.62 | | 47.28 | 64.20 | 45.19 | 52.84 | 8.52 | 30.12 |
| ClearSCD [5] | 5.77 | | 45.35 | 62.42 | 57.26 | 51.05 | 14.12 | 28.90 |
| PRO-HRSCD [43] | 32.5 | N/R | 58.45 | 73.55 | 73.20 | 62.50 | 22.84 | 41.20 |
| ResNet-GRU [48] | 21.45 | | 45.78 | 62.81 | 64.20 | 46.47 | 8.58 | 22.15 |
| FC-Siam-conc [18] | 2.74 | | 52.38 | 68.75 | 68.33 | 55.28 | 16.32 | 29.40 |
| FC-Siam-diff [18] | 1.66 | | 51.52 | 68.01 | 68.81 | 55.16 | 16.08 | 28.75 |
| HRSCD-str.3 [49] | 12.77 | | 37.61 | 54.67 | 64.68 | 50.85 | 10.24 | 23.80 |
| HRSCD-str.4 [49] | 13.71 | | 49.58 | 66.30 | 71.16 | 58.60 | 18.62 | 31.25 |
| SCDNet [50] | 39.62 | | 53.04 | 69.32 | 70.91 | 60.03 | 19.79 | 34.12 |
| SSCD-l [4] | 23.31 | | 56.85 | 72.49 | 72.55 | 61.62 | 21.45 | 36.50 |
| BiSRNet [4] | 23.39 | | 57.45 | 72.98 | 72.55 | 61.60 | 21.50 | 37.10 |
| TED [10] | 41.2 | N/R | 57.56 | 73.07 | 73.01 | 62.09 | 22.30 | 38.45 |
| SAM-SCD [12] | 117.16 | | 56.47 | 72.18 | 71.79 | 60.32 | 20.07 | 35.80 |
| M-CD [51] | 46.03 | | 56.55 | 72.25 | 71.54 | 59.66 | 19.67 | 35.10 |
| FEMCD [38] | 67.65 | | 56.43 | 72.15 | 71.63 | 59.71 | 19.58 | 34.90 |
| FD-ProtoSCD (Ours) | 23.92 | | 61.50 | 75.42 | 74.85 | 64.12 | 24.56 | 45.60 |
Table 3.
Comparisons of binary change detection performance on Hi-UCD transfer and LEVIR-CD. The best scores are in bold.
Table 3.
Comparisons of binary change detection performance on Hi-UCD transfer and LEVIR-CD. The best scores are in bold.
| Method | Hi-UCD Transfer (Binary CD) | LEVIR-CD (Binary CD) |
|---|
|
IoU
|
F1
|
Precision
|
Recall
|
F1
|
IoU
|
|---|
| SCD-UperNet | 24.32 | 39.15 | 88.75 | 87.12 | 87.92 | 78.46 |
| SCDNet [50] | 25.67 | 40.82 | 89.63 | 88.24 | 88.92 | 80.07 |
| FEMCD [38] | 26.89 | 42.35 | 90.78 | 89.31 | 90.04 | 81.88 |
| SAM-SCD [12] | 27.05 | 42.58 | 91.02 | 89.52 | 90.26 | 82.25 |
| M-CD [51] | 27.21 | 42.79 | 91.18 | 89.73 | 90.45 | 82.56 |
| SSCD-l [4] | 27.58 | 43.24 | 91.67 | 90.35 | 91.00 | 83.48 |
| BiSRNet [4] | 28.34 | 44.16 | 92.46 | 91.38 | 91.91 | 85.05 |
| TED [10] | 28.67 | 44.58 | 92.85 | 91.79 | 92.31 | 85.72 |
| PRO-HRSCD [43] | 29.41 | 45.52 | 93.72 | 92.87 | 93.29 | 87.41 |
| FD-ProtoSCD (Ours) | 30.06 | 46.23 | 94.51 | 93.98 | 94.24 | 89.13 |
Table 4.
Ablation results of FD-ProtoSCD variants on SECOND.
Table 4.
Ablation results of FD-ProtoSCD variants on SECOND.
| Variant | Checkpoint | (%) | | IoU (%) |
|---|
| BIT BCD baseline | best mFscore | 67.25 | — | 50.66 |
| SCD-UperNet baseline | best | 68.78 | 16.14 | 52.41 |
| w/o UPer semantic head | best | 70.12 | 18.28 | 55.88 |
| w/o DCPD | best | 72.45 | 19.57 | 57.82 |
| w/o FDCD | best | 73.10 | 20.82 | 58.38 |
| w/o pseudo-change loss | best | 73.95 | 22.14 | 59.10 |
| w/o focal prototype loss | best | 74.50 | 23.35 | 59.78 |
| FD-ProtoSCD full | best | 75.42 | 24.56 | 61.50 |
Table 5.
Selected hyperparameter settings for the reported full model on SECOND.
Table 5.
Selected hyperparameter settings for the reported full model on SECOND.
| Group | Best Value | Best | Source Run |
|---|
| FDCD radius | 0.50 | 24.56 | r050 |
| Prototype EMA momentum | 0.99 | 24.56 | m099 |
| Pseudo-change loss weight | 0.05 | 24.56 | pseudo005 |
| Focal prototype loss weight | 0.005 | 24.56 | proto0005 |
Table 6.
Sensitivity to FDCD spectral radius on SECOND.
Table 6.
Sensitivity to FDCD spectral radius on SECOND.
| FDCD Radius | Best Iter | (%) | | mIoU (%) | Selected |
|---|
| 0.10 | 26 k | 72.05 | 21.61 | 67.88 | |
| 0.20 | 24 k | 73.19 | 22.89 | 69.66 | |
| 0.25 | 29 k | 74.04 | 23.45 | 71.19 | |
| 0.35 | 33 k | 74.35 | 24.02 | 73.23 | |
| 0.50 | 36 k | 75.42 | 24.56
| 74.85
| yes |
Table 7.
Sensitivity to prototype memory momentum on SECOND.
Table 7.
Sensitivity to prototype memory momentum on SECOND.
| Setting | Best Iter | (%) | | mIoU (%) | Selected |
|---|
| 30 k | 73.85 | 23.13 | 71.89 | |
| 29 k | 74.89 | 23.82 | 72.36 | |
| 36 k | 75.42 | 24.56 | 74.85 | yes |
| 31 k | 74.15 | 23.58 | 71.42 | |
Table 8.
Sensitivity to auxiliary loss weights on SECOND.
Table 8.
Sensitivity to auxiliary loss weights on SECOND.
| Parameter | Best Iter | (%) | | mIoU (%) | Selected |
|---|
| 27 k | 73.57 | 22.14 | 68.40 | |
| 30 k | 74.45 | 23.21 | 71.28 | |
| 36 k | 75.42 | 24.56 | 74.85 | yes |
| 32 k | 74.81 | 23.81 | 73.94 | |
| 28 k | 74.50 | 23.35 | 71.75 | |
| 36 k | 75.42 | 24.56 | 74.85 | yes |