Next Article in Journal
Geometry-Guided Semi-Supervised Multimodal Segmentation for UAV-Based Rice-Lodging Mapping
Previous Article in Journal
Odometry and Mapping for Complex Environment Perception Under Partial-View Sensing
Previous Article in Special Issue
EGDNet: An Event-Guided Frequency-Aware Network with Cross-Modal Attention for Robust Weak Signature UAV Detection
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

MARC-Net: A Modality-Availability-Aware Robust Change Network for Missing-Optical Bi-Temporal Optical–SAR Change Detection of Reclaimed Cropland

1
Key Laboratory of National Geographical Census and Monitoring, Ministry of Natural Resources, Hangzhou 311121, China
2
Zhejiang Academy of Surveying and Mapping, Hangzhou 311121, China
3
Zhejiang Application Center of Nature Resources Satellite Technology, Hangzhou 311121, China
4
Key Laboratory for Information Science of Electromagnetic Waves (MoE), School of Information Science and Technology, Fudan University, Shanghai 200433, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(17), 2960; https://doi.org/10.3390/rs18172960
Submission received: 29 May 2026 / Revised: 27 July 2026 / Accepted: 18 August 2026 / Published: 2 September 2026

Highlights

What are the main findings?
  • This study formulates missing-modality bi-temporal optical–SAR change detection for reclaimed cropland and introduces MARC-Net with explicit modality-availability-aware input modeling.
  • MARC-Net reaches 0.7826 Full IoU, 0.7039 Missing-O IoU, and 0.7780 Missing-S IoU under the mixed-scene protocol.
What is the implication of the main findings?
  • Availability-aware optical–SAR fusion can support reclaimed-cropland monitoring when cloud-free optical observations are unavailable at one time point.
  • Combining mixed-scene diagnostics with scene-level tests provides a more complete assessment of missing-modality robustness.

Abstract

Reliable monitoring of reclaimed cropland is hindered when one optical acquisition is unavailable or degraded. We formulate missing-modality bi-temporal optical–SAR change detection and propose the Modality-Availability-Aware Robust Change Network (MARC-Net), a new architecture that combines explicit availability conditioning, condition-aware temporal proxy stabilization, a shared residual input adapter, multi-level signed temporal interaction, dilated context refinement, and hierarchical change decoding. A two-phase condition-balanced learning strategy jointly develops mixed-missing representations and optimizes Full, Missing-O, and Missing-S behavior without reconstructing the unavailable image. On four reclaimed-cropland scenes, the final model obtains IoUs of 0.7826, 0.7039, and 0.7780, respectively, with a three-condition mean of 0.7548. It exceeds the strongest evaluated external baseline mean (0.7381) while using one checkpoint and a fixed argmax decision rule. LOSO and controlled optical-degradation experiments further characterize robustness under geographic shift and progressive observation-quality degradation. These results demonstrate that availability-aware temporal stabilization and condition-balanced optimization provide an effective operating point for incomplete-input reclaimed-cropland monitoring while preserving complete-input performance.

1. Introduction

1.1. Background

Cropland protection, land consolidation, and post-disturbance land restoration are closely related to food security, agricultural production capacity, and sustainable land resource management. In such applications, reclaimed cropland should be understood not only as land that has been administratively restored, but also as land that exhibits an effective cropland state after reclamation. Reliable identification of effective reclaimed cropland can therefore support agricultural management, cultivated land protection, and the evaluation of reclamation policies. Remote sensing is well suited to this task because it provides repeated and spatially explicit observations over large areas. Recent reviews further highlight the growing role of artificial intelligence in extracting actionable information from such data [1]. Studies on abandoned cropland show that satellite observations can reveal cropland state transitions under complex seasonal conditions [2]. PolSAR-based abandoned land mapping further indicates that radar observations are valuable in cloudy and rainy agricultural regions [3]. Multi-source remote sensing has also been used to identify abandoned cropland and analyze its driving mechanisms, which highlights the importance of integrating diverse observation cues for cropland monitoring [4]. These studies motivate the use of remote sensing for reclaimed cropland monitoring, while the before–after nature of reclamation makes bi-temporal change detection a natural modeling perspective.
Optical and synthetic aperture radar (SAR) imagery provide complementary information for cropland monitoring. Optical imagery contains spectral, color, texture, and semantic cues that are useful for recognizing vegetation, bare soil, parcel boundaries, and surrounding land-cover context. SAR imagery, by contrast, records microwave backscatter related to structure, roughness, and moisture, and it can be acquired independently of solar illumination and under most weather conditions. Recent reviews have emphasized that optical–SAR deep feature fusion can improve remote sensing semantic segmentation by combining these complementary properties [5]. Optical–SAR image fusion and image translation have also been investigated for enhancing semantic segmentation under multimodal remote sensing settings [6]. Cross-fusion networks show that optical and SAR features can interact during land-cover interpretation [7]. Second-order attention-based bilinear fusion further suggests that channel relationships across modalities are useful for land-cover classification [8]. For effective reclaimed cropland detection, optical data can provide cropland semantics and boundary cues, while SAR can supply stable structural evidence when optical observations are degraded or missing.

1.2. Related Work

From a task perspective, effective reclaimed cropland detection is fundamentally a bi-temporal change detection problem. The target is not merely a static land-cover category at one date, but a land-surface transition between a pre-reclamation and a post-reclamation phase. Classical multisensor change detection has used post-classification comparison to compare SAR and optical observations [9]. Object-based hierarchical compound classification further shows that spatial units and temporal dependencies can reduce the instability caused by direct pixel-wise comparison in heterogeneous images [10]. With the development of deep learning, change detection and dense remote sensing interpretation have shifted toward representation learning from paired or pixel-level images. Fully convolutional networks provide an early foundation for dense prediction [11]. U-Net-like encoder–decoder structures are also important for recovering spatial details in segmentation maps [12]. DeepLab-style semantic segmentation demonstrates the value of contextual convolutional representations [13]. Residual learning provides a widely used backbone design for extracting robust hierarchical features [14]. Attention modules such as CBAM further show that channel and spatial recalibration can enhance convolutional representations [15]. Fully convolutional early-fusion and Siamese networks provide widely used change-detection baselines [16]. DTCDSCN further combines a shared SE-ResNet34 encoder, multi-level temporal differencing, and attention-guided decoding [17]. Transformer-based alternatives include the bitemporal image transformer (BIT) [18] and ChangeFormer [19]. Hierarchical shifted-window attention was introduced by Swin Transformer [20] and adapted to remote sensing change detection by SwinSUNet [21]. State-space alternatives include ChangeMamba’s spatiotemporal modeling [22] and CDMamba’s local–global change representation [23].
Bi-temporal optical change detection has achieved substantial progress because optical images provide intuitive and semantically rich observations. Semantic change detection networks explicitly model temporal semantic relationships and can improve the interpretation of land-cover transitions [24]. Temporal-transform designs further indicate that cross-temporal feature aggregation is useful for locating changed regions and understanding land-cover relationships [25]. More recent bi-temporal change detection models explore state-space modeling to improve global context modeling and computational efficiency [26]. Token interaction and bi-temporal spatial enhancement have also been used to improve fine-grained high-resolution change detection [27]. Knowledge distillation studies in semantic segmentation show that structured, relational, and texture information can improve dense prediction models [28,29]. Feature-augmented and remote-sensing-oriented distillation methods further indicate that model design and supervision can affect segmentation quality and efficiency [30,31]. However, most optical change detection methods implicitly rely on the availability of usable optical images at both time points, which is a fragile assumption in operational cropland monitoring.
Optical–SAR change detection offers a practical way to exploit complementary sensors, especially when optical observations are unreliable. Heterogeneous optical–SAR change detection studies have investigated homogeneous feature transformation to reduce the difference between optical and SAR representations [32]. Object-level optical–SAR change analysis also shows that heterogeneous images can provide useful change evidence even when their radiometric appearances are not directly comparable [10]. A particularly relevant setting uses pre-change optical–SAR data together with post-change SAR data, motivated by the fact that post-event optical imagery is often unavailable because of clouds, fog, smoke, or acquisition constraints [33]. Optical–SAR registration studies further remind us that geometric alignment is an important prerequisite for reliable multimodal fusion and comparison [34]. Urban surface change detection with SAR coherent scatterers also illustrates the ability of SAR data to capture spatio-temporal surface changes [35]. These works confirm the value of optical–SAR collaboration, but many of them either focus on heterogeneous cross-sensor comparison or assume a fixed complete input configuration rather than the missing-modality setting considered here.
Related progress in remote sensing classification and segmentation further supports the use of optical–SAR collaboration. Cross-modal feature learning has been shown to improve land-cover classification by aligning or fusing optical and SAR information [36]. Multilevel optical–SAR fusion networks suggest that features at different depths provide complementary cues for land-cover classification [37]. Multimodal feature fusion networks further demonstrate that different fusion levels can capture complementary land-cover evidence [38]. DCIFNet emphasizes correction and interaction between optical and SAR features for land-cover classification [39]. In semantic segmentation, CroFuseNet uses cross fusion for optical–SAR impervious surface extraction [40]. SOLSTM combines SAR–OPT matching attention with temporal modeling for multisource segmentation [41]. HAFNet highlights heterogeneous adaptive fusion for improved land-use classification [42]. Spectral-geometric iterative fusion further shows that multimodal segmentation must handle inconsistent spatial resolutions and geometric information [43]. Hybrid deep learning fusion models also indicate that feature extraction and spatial alignment are both important in multi-source remote sensing data fusion [44]. Although these studies are not designed specifically for reclaimed cropland change detection, they show that optical–SAR collaboration can enhance representation capacity and provide useful guidance for bi-temporal monitoring tasks.

1.3. Research Motivation

The main practical challenge addressed in this paper is missing-modality inference. Optical imagery is often affected by clouds, rainfall, haze, seasonal observation windows, and acquisition quality, so the optical image at either the first or the second time point may be absent or unusable. SAR imagery is usually more stable in this respect, making it an important source of information when one optical phase is missing. SAR-guided dehazing reflects the broader need to handle degraded optical observations in remote sensing workflows [45]. SAR-assisted cloud removal also shows that SAR can provide complementary information when optical imagery is contaminated or incomplete [46]. SAR water-body and flood mapping studies illustrate why all-weather SAR observations are operationally attractive when optical images are affected by clouds or rain [47,48]. Recent SAR flood mapping work further confirms the value of SAR for reliable mapping under adverse observation conditions [49]. Incomplete multimodal learning has recently attracted attention in remote sensing classification and data fusion [50]. Online knowledge distillation has also been explored for land-use/cover classification with full or missing modalities [51]. Earlier and recent missing-modality classification studies demonstrate that models trained only for complete multimodal input may degrade when a modality is absent at inference [52,53]. A missing-modality invariant cross-attention network has recently been proposed for robust land cover mapping under incomplete inputs [54]. Methods for arbitrary missing modalities and meta-modal representation further broaden this research direction [55,56]. Recent remote-sensing object detection, SAR detection, and multimodal fusion studies further suggest that robust interpretation benefits from multi-scale object perception, background-clutter suppression, and local–global feature coordination. These findings are relevant to our optical-missing setting because, when one optical phase is unavailable, SAR and multimodal features must carry more structural information for fragmented reclaimed parcels and complex agricultural backgrounds [57,58,59].
Nevertheless, these studies mostly address classification or general missing-modality learning. This paper instead focuses on the practically motivated missing-modality scenario in bi-temporal optical–SAR change detection, where SAR remains available and the objective is robust inference for effective reclaimed cropland detection.

1.4. Contributions and Paper Organization

To address this problem, this study develops MARC-Net, a modality-availability-aware robust change network for bi-temporal optical–SAR reclaimed-cropland monitoring. Its distinct missing-modality pathway comprises six-channel optical–SAR/availability packaging, an observation-only temporal proxy for an unavailable phase, shared residual input stabilization, and condition-balanced multi-state optimization. The proxy uses only the available observation from the other date and therefore does not reconstruct or access the removed image. A shared SE-ResNet34 encoder maps both dates into one feature space; signed differences at four scales are refined by sequential dilated convolution and decoded through hierarchical SCSE blocks. The study contributes both this integrated architecture and a traceable evaluation design covering complete input, missing optical, missing SAR, mixed-scene testing, LOSO generalization, partial optical degradation, and scene-scale deployment.
The remainder of this paper is organized as follows. Section 2 presents the materials and methods, including dataset construction, task formulation, the proposed MARC-Net, the missing-modality protocol, and optimization details. Section 3 reports the experimental results under full-input, missing-optical, and missing-SAR settings. Section 4 discusses the main findings, limitations, and implications for operational reclaimed cropland monitoring. Finally, Section 5 concludes this paper.

2. Materials and Methods

This section describes the data organization, task formulation, network architecture, missing-modality sample construction, and optimization strategy used for effective reclaimed cropland change detection. The method follows the practical assumption introduced in Section 1: optical and SAR observations are available at two time points during training, whereas at inference one modality at one time point may be missing. The primary operational scenario is single-phase optical missing with available SAR, but SAR-missing evaluation is also reported. The implementation is designed as a bi-temporal optical–SAR change detector with availability-aware input construction, rather than as a general arbitrary missing-modality framework.

2.1. Study Area, Data Sources, and Dataset Construction

2.1.1. Study Areas and Data Sources

The study area is located in the Hangzhou region, Zhejiang Province, China, a subtropical agricultural area where reclaimed cropland monitoring is challenged by frequent cloud cover and rainfall. The dataset comprises four independent geographic scenes acquired over distinct sub-areas within this region. Table 1 summarizes the scene-level properties. Each scene provides two co-registered bi-temporal optical–SAR pairs: Sentinel-2 optical images resampled to the working grid of 2 m resolution and corresponding SAR intensity images, together with a pixel-level binary label map that identifies effective reclaimed cropland change. The optical images at the first observation phase are acquired in June 2024 and the second-phase images in September 2024. For each sample, the four observations are denoted as O 1 , S 1 , O 2 , and S 2 , respectively. The processed dataset follows the directory structure used by the implementation, including optical_t1, sar_t1, optical_t2, sar_t2, and label folders for the training and validation subsets.
Figure 1 summarizes the dataset-construction pipeline, and Figure 2 shows the geographic locations of the four scenes.
Let the dataset be represented as
D = { ( O 1 ( i ) , S 1 ( i ) , O 2 ( i ) , S 2 ( i ) , Y ( i ) ) } i = 1 N ,
where i indexes a bi-temporal sample, N is the number of samples, and Y ( i ) { 0 , 1 } H × W is the binary effective reclaimed cropland change label. The value 1 denotes the target changed class, and 0 denotes the background or non-target class. In this study, effective reclaimed cropland denotes the within-year transition from absent or weak crop cover in the June observation to evident cultivated crop cover in the September observation; it does not include every land-cover change. The dataset provides one accepted binary mask for each scene, and these masks were checked against the paired optical images and the aligned SAR structure before patch extraction. Independent duplicate annotations and inter-annotator disagreement scores were not available in the dataset records, so annotation uncertainty is not claimed quantitatively and remains a dataset limitation.
The four scenes represent distinct agricultural landscapes and observation conditions, covering densely cultivated plains, riverine agricultural zones, hilly terraced cropland, and flat open-field areas. Scene 002, with the highest label density (41.7%) and favorable optical conditions, achieves the best model performance among the four scenes (see Section 3.1). Scene 003, by contrast, is the most challenging due to its steep terrain, fragmented parcels, and the lowest label density (3.9%), and the model performance on this scene degrades substantially. This scene-level diversity allows the evaluation to probe generalization across geographically separated and visually heterogeneous sub-regions, motivating the use of the leave-one-scene-out (LOSO) protocol.

2.1.2. Optical and SAR Preprocessing

Optical and SAR images are spatially organized into paired bi-temporal samples before model training. The images are read as raster arrays, and corresponding files in the optical, SAR, and label folders are matched consistently. The preprocessing workflow includes spatial alignment, channel arrangement, patch extraction, and numerical normalization. Since training operates on paired patches, geometric registration and spatial consistency between O t , S t , and Y are required before training.
For a time point t { 1 , 2 } , optical and SAR channels are concatenated after reading:
Z t = Concat ( O t , S t ) R C z × H × W ,
where C z = C o + C s = 5 , C o = 4 and C s = 1 are the numbers of optical and SAR channels, respectively. The availability map introduced in Section 2.4.1 is then appended to form the six-channel network input at each time point. After concatenation, the image array is normalized from the original image intensity range to a centered range:
Z ¯ t = 2 Z t 255 1 .
This operation is consistent with the implemented normalization function.
For reproducibility, the co-registration workflow is clarified as follows. Optical images were used as the spatial reference grid, and SAR images were geometrically corrected and resampled to the same map projection and pixel grid before patch generation. Tie-point inspection and overlay checks were performed at parcel boundaries, roads, rivers, and stable built-up edges. Because optical and SAR imagery have different imaging geometries, residual local misregistration may remain in hilly or fragmented areas; this limitation is considered when interpreting Scene 003 and the most challenging missing-optical cases. SAR intensities were normalized after spatial alignment, and the network receives the normalized SAR channels together with the optical channels rather than any generated optical substitute.

2.1.3. Label Construction for Effective Reclaimed Cropland

The task label is constructed as a binary change map. In the implementation, label values are remapped to { 0 , 1 } and then converted to a two-class target representation for training. The effective reclaimed cropland class corresponds to pixels where the land surface satisfies the task-specific change definition after reclamation. In this manuscript, the target is not a general land-cover change class, but an effective reclaimed cropland change class. Therefore, the label can be expressed as
Y ( p ) = 1 , p Ω erc , 0 , p Ω erc ,
where p denotes a pixel location and Ω erc denotes the set of pixels labeled as effective reclaimed cropland change in the reference masks. The dataset contains a single accepted binary mask per scene rather than independent annotations from multiple interpreters; consequently, this study evaluates model uncertainty but does not report an inter-annotator uncertainty statistic. The available records contain no separate ambiguity or ignore class, so potentially ambiguous pixels cannot be identified retrospectively after the accepted binary masks were produced.

2.1.4. Patch Generation and Data Split Strategy

The network is trained on fixed-size image patches. The default patch size in the implementation is 512 × 512 pixels, and the model uses the same spatial size to define the multi-scale feature resolutions. The mixed-scene split contains 1104 training pairs, 184 validation pairs, and 184 test pairs, for 1472 valid paired patches in total. Each pair contains exactly matched files in the optical_t1, sar_t1, optical_t2, sar_t2, and label folders. The LOSO experiments instead hold out all patches from one geographic scene and train on the other three scenes. Horizontal and vertical random flips are applied only to the training subset. Figure 3 illustrates the task definition, and Table 2 summarizes the dataset organization.

2.1.5. Single-Phase Missing-Modality Sample Construction

The primary missing-data setting is single-phase optical missing: O 1 or O 2 may be unavailable while both SAR observations remain present. Missing-S inference is retained as a secondary stress condition. For each incomplete sample, the pipeline first selects the unavailable modality and time point. The representation-learning phase uses zero replacement before normalization and appends a one-channel availability indicator; the condition-balanced phase incorporates the observation-only temporal proxy defined in Section 2.4 while preserving an explicit Missing-O flag.
For the primary optical-missing condition, let a t o { 0 , 1 } denote optical availability at time t. The masked optical input is defined as
O ˜ t = a t o O t .
For a single-phase optical-missing sample, the optical-availability pair satisfies
( a 1 o , a 2 o ) { ( 1 , 1 ) , ( 0 , 1 ) , ( 1 , 0 ) } ,
where ( 1 , 1 ) corresponds to complete optical input, ( 0 , 1 ) to missing O 1 , and ( 1 , 0 ) to missing O 2 . The analogous SAR-availability pair is used when SAR is selected. The final configuration uses train_missing_target=mixed, train_missing_prob=0.5, optical-target probability 0.6, and missing- t 2 probability 0.6 for optical-missing samples.

2.2. Problem Formulation

Given a bi-temporal optical–SAR sample, the objective is to estimate a binary effective reclaimed cropland change map. The complete-input formulation can be written as
Y ^ = f θ ( O 1 , S 1 , O 2 , S 2 ) ,
where f θ denotes the change detection network with trainable parameters θ . Under the single-phase optical-missing setting, the model receives O ˜ t , SAR observations, and availability indicators:
Y ^ = f θ ( O ˜ 1 , S 1 , a 1 , O ˜ 2 , S 2 , a 2 ) .
The learning problem is to obtain a model that performs well for ( a 1 , a 2 ) = ( 1 , 1 ) and remains robust for ( a 1 , a 2 ) = ( 0 , 1 ) or ( 1 , 0 ) . This formulation explicitly targets single-phase optical missing and does not assume arbitrary missing SAR inputs as the main operational scenario.

2.3. Overview of the Proposed Framework

MARC-Net contains five functional components: (1) an availability-conditioned optical–SAR representation, (2) condition-aware temporal proxy stabilization, (3) a residual input adapter shared between time points, (4) shared squeeze-and-excitation residual encoding with sequential dilated context refinement, and (5) a difference-driven hierarchical decoder. The change-reasoning stream adopts established principles of shared-weight temporal encoding and multi-level difference decoding [17], and reorganizes them within an availability-aware architecture. Specifically, MARC-Net integrates six-channel optical–SAR/availability packaging, an observation-only temporal proxy for an unavailable phase, residual input stabilization, mixed single-phase missing simulation, and condition-balanced multi-state optimization. These components enable one checkpoint to process complete, missing-optical, and missing-SAR inputs without reconstructing the unavailable image or using condition-specific parameter updates. Figure 4 presents the resulting end-to-end architecture.

2.4. Availability-Conditioned Optical–SAR Representation

2.4.1. Time-Point Input Construction

Let m t o , m t s { 0 , 1 } denote whether the optical and SAR observations at time t are usable. The primary operating condition has m t s = 1 and permits one optical phase to be absent, while missing-SAR inference is retained as a secondary stress test. Missing observations are first replaced by zeros in the raw intensity domain:
O ˜ t = m t o O t , S ˜ t = m t s S t .
The five image channels are concatenated and normalized to the centered range used by the data loader:
Z ¯ t = 2 Concat ( O ˜ t , S ˜ t ) 255 1 .
Consequently, a raw zero-filled missing channel becomes the constant value 1 after normalization. A one-channel joint-availability map explicitly distinguishes this sentinel pattern from valid low-intensity observations:
A t ( p ) = m t o m t s , p Ω , X t = Concat ( Z ¯ t , A t ) R 6 × H × W .
Only one modality is removed at a time in the reported experiments, so A t = 0 identifies an incomplete time point and the image channels identify which modality is absent.

2.4.2. Condition-Aware Temporal Proxy Stabilization

The condition-aware temporal proxy reduces the extreme distribution shift between a valid observation and the constant missing sentinel by using only information that remains available at inference. For modality q { o , s } at an unavailable time point t, let t ¯ denote the other date. The normalized proxy is
Z ¯ t q , = m t q Z ¯ t q + ( 1 m t q ) α q Z ¯ t ¯ q + ( 1 α q ) ( 1 ) ,
where α o = 0.25 for Missing-O and α s = 1 for Missing-S. Thus, an unavailable optical phase retains mostly the explicit sentinel while receiving a weak structural prior from the available optical date; an unavailable SAR phase uses the available SAR observation directly as a conservative temporal proxy. The corresponding stabilized network input is
X t = Concat ( Z ¯ t o , , Z ¯ t s , , A t ) , A t = 0 , Missing - O , 1 , Full or Missing - S .
The Missing-O flag continues to state explicitly that optical evidence is absent. For Missing-S, the copied SAR proxy is marked as stabilized input because the network receives a valid SAR-valued tensor rather than the missing sentinel. Crucially, Equation (12) never accesses the removed image, the test label, or a future prediction, so it does not introduce information leakage.

2.4.3. Residual Input Stabilization

Complete and incomplete tensors have different input distributions. MARC-Net therefore applies the same lightweight residual adapter to both time points:
X t = ReLU X t + BN 2 W 1 × 1 ReLU BN 1 ( W 3 × 3 X t ) .
The 3 × 3 convolution expands six channels to 16 hidden channels, and the 1 × 1 convolution projects them back to six channels. This module is conditioned on the concatenated image and availability values, but it performs neither image reconstruction nor cross-attention.

2.5. Shared Temporal Encoding and Difference Interaction

2.5.1. Shared SE-ResNet34 Encoder

The two adapted tensors pass through one shared encoder composed of squeeze-and-excitation residual blocks with the layer configuration [ 3 , 4 , 6 , 3 ] . The four output levels are
{ F t 1 , F t 2 , F t 3 , F t 4 } = E ψ ( X t ) , t { 1 , 2 } ,
with spatial strides 4 , 8 , 16 , and 32 and channel dimensions 64 , 128 , 256 , and 512, respectively. Within each residual block, channel recalibration is computed from global average pooled features:
s = sigmoid W 2 ReLU ( W 1 GAP ( U ) ) , U ^ = U s .
Weight sharing places the two dates in the same feature space, while channel recalibration suppresses weak or noisy responses before temporal comparison.

2.5.2. Multi-Level Signed Temporal Differences

MARC-Net uses signed subtraction rather than feature concatenation at the temporal interaction stage:
Δ l = F 1 l F 2 l , l { 1 , 2 , 3 , 4 } .
The operation preserves the direction of the learned feature transition and supplies difference evidence at both semantic and boundary-oriented resolutions.

2.5.3. Dilated Context Refinement

Only the deepest difference Δ 4 is passed through the context block. Four 3 × 3 convolutions with dilation rates d k { 1 , 2 , 4 , 8 } are applied sequentially:
Q 0 = Δ 4 , Q k = ReLU ( Conv 3 × 3 d k ( Q k 1 ) ) , k = 1 , , 4 ,
and their responses are accumulated by a residual sum:
C 4 = Q 0 + k = 1 4 Q k .
This block expands the effective receptive field without introducing Transformer layers, which is important to state explicitly because the implemented network is convolutional.

2.6. Hierarchical Difference Decoder and Decision Head

The decoder reconstructs spatial detail from C 4 and injects the shallower signed differences at their matching resolutions:
D 4 = D 4 ( C 4 ) + Δ 3 ,
D 3 = D 3 ( D 4 ) + Δ 2 ,
D 2 = D 2 ( D 3 ) + Δ 1 ,
D 1 = D 1 ( D 2 ) .
Each D k contains 1 × 1 channel reduction, spatial-and-channel squeeze-and-excitation (SCSE) recalibration, transposed-convolution upsampling, and a 1 × 1 output projection. For a decoder feature V, the SCSE response is
SCSE ( V ) = V s c ( V ) + V s s ( V ) ,
and the implementation adds this response residually to V before upsampling. The final head upsamples D 1 , applies a 3 × 3 refinement convolution, and produces two-class logits P:
P = H ( D 1 ) , Y ^ ( p ) = arg max c { 0 , 1 } P c ( p ) .
The positive class is the task-specific effective reclaimed-cropland transition. The reported model has one prediction head and does not use auxiliary multi-scale outputs or deep supervision.

2.7. Missing-Modality Robust Learning

2.7.1. Mixed Single-Phase Missing Simulation

The representation-learning phase and the controlled ablation experiments apply missing simulation per sample. Let b Bernoulli ( ρ ) indicate whether a sample is made incomplete. If b = 0 , both time points remain complete. If b = 1 , the removed modality q { o , s } is drawn with Pr ( q = o ) = π o , and one time point is selected. For optical missing, the second time point is selected with probability π 2 ; for SAR missing, the two phases are sampled uniformly. The resulting masks satisfy
( m 1 q , m 2 q ) { ( 0 , 1 ) , ( 1 , 0 ) } , m 1 q ¯ = m 2 q ¯ = 1 .
In the base phase, ρ = 0.5 , π o = 0.6 , and π 2 = 0.6 . Thus, in expectation, 50% of samples are complete, 30% are optical-missing, and 20% are SAR-missing. The optical emphasis reflects the target deployment scenario, whereas the SAR-missing samples regularize the network against over-reliance on one sensor. Figure 5 summarizes the training and inference protocol.

2.7.2. Condition-Balanced Multi-State Optimization

The proposed optimization strategy first establishes mixed-missing representations and then presents every mini-batch under Full, Missing-O, and Missing-S conditions during condition-balanced refinement. A single missing time point is sampled for each item and reused for the two incomplete conditions, so condition differences are not confounded by different batch content. With condition weights ( λ F , λ O , λ S ) = ( 1.15 , 1.60 , 1.25 ) , the final objective is
L rob = λ F L ce ( f θ ( T F ( X ) ) , Y ) + λ O L ce ( f θ ( T O ( X ) ) , Y ) + λ S L ce ( f θ ( T S ( X ) ) , Y ) λ F + λ O + λ S ,
where T F leaves the pair unchanged and T O , T S apply Equation (12) to one optical or SAR phase, respectively. The larger λ O directly targets the operationally central and empirically harder Missing-O condition while retaining substantial Full and Missing-S supervision.

2.7.3. Inference Conditions

The trained model is evaluated without condition-specific parameter updates. Full inference uses ( m 1 o , m 2 o , m 1 s , m 2 s ) = ( 1 , 1 , 1 , 1 ) . Missing- O 1 and Missing- O 2 remove one optical observation while retaining both SAR inputs; Missing-S removes one SAR observation. The removed tensor is never loaded by the condition transform. Instead, Equation (12) uses only the available counterpart from the other date, and the same selected checkpoint is used for all three conditions.

2.8. Optimization Objective and Implementation Details

2.8.1. Optimization Objective

MARC-Net has one prediction head and uses pixel-wise cross-entropy. For sample i with logits P i and binary label Y i , the condition-specific loss is
L ce ( i ) = 1 | Ω | p Ω c = 0 1 1 [ Y i ( p ) = c ] log Softmax ( P i , c ( p ) ) .
Equation (28) is averaged over the batch and then combined across conditions by Equation (27). No reconstruction, adversarial, teacher–student, distillation, or auxiliary-scale loss is used in the proposed model.
Because Scene 003 has a sparse positive class, the sensitivity study additionally evaluates an equally weighted CE–Dice objective. Let q i ( p ) be the predicted probability of the positive class. The implemented smoothed foreground Dice loss is
L dice ( i ) = 1 2 p Ω q i ( p ) Y i ( p ) + 1 p Ω q i ( p ) + p Ω Y i ( p ) + 1 ,
and the alternative objective is
L ce + dice ( i ) = L ce ( i ) + L dice ( i ) .
The additive constant 1 stabilizes samples with few positive pixels. Dice directly optimizes foreground overlap, whereas CE retains pixel-wise probabilistic supervision. The quantitative comparison is reported in Section 3.

2.8.2. Training and Implementation Details

The model is implemented in PyTorch 2.0.1 and trained only on the reclaimed-cropland training split, without transferred task weights. The shared temporal encoder and multi-level difference decoder follow established change-detection design principles [17,19], whereas the availability-conditioned representation, residual adapter, temporal proxy, and condition-balanced controller form the proposed MARC-Net pathway. The complete two-phase optimization schedule comprises representation learning from random initialization with Adam, a learning rate of 1 × 10 4 , and mixed single-phase missing simulation, followed by condition-balanced optimization with full-precision AdamW, cosine learning-rate decay from 1 × 10 5 to 1 × 10 6 , weight decay 1 × 10 4 , gradient clipping at 2, and exponential moving averaging with decay 0.995. The phase-wise allocation is reported once in Table 3. Batch-normalization running statistics are frozen during condition-balanced optimization while affine parameters remain trainable. Both phases use a training batch size of 2, 512 × 512 patches, random horizontal and vertical flips, and seed 2022. The reported checkpoint is selected exclusively by the highest three-condition validation mean IoU; the held-out test set is evaluated once with argmax prediction and no threshold tuning.
The sensitivity study in Section 3 varies the base-phase settings ρ { 0.25 , 0.50 , 0.75 } , the missing phase in { t 1 , t 2 , either } , and the loss in { CE , CE + Dice } . These experiments diagnose the representation-learning protocol, while the complete proposed configuration is reported separately so that the controlled comparisons remain traceable.

2.9. Evaluation Metrics

Evaluation is based on pixel-level true positives (TP), false positives (FP), false negatives (FN), and true negatives (TN) for the effective reclaimed-cropland change class. For the mixed-scene test set, the four counts are accumulated over all 184 held-out patches before the metrics are calculated. Precision is defined as
Precision = TP TP + FP .
Precision measures the reliability of predicted change pixels. A decrease in Precision indicates that a larger fraction of predicted reclaimed-cropland changes are false alarms.
Recall is defined as
Recall = TP TP + FN .
Recall measures the fraction of reference change pixels recovered by the model. It is therefore particularly informative for determining whether missing optical information increases omission errors.
The F1 score and intersection over union (IoU) are
F 1 = 2 Precision Recall Precision + Recall ,
IoU = TP TP + FP + FN .
F1 is the harmonic mean of Precision and Recall, whereas IoU directly measures the overlap between the predicted and reference change sets. IoU is used as the primary ranking metric because it penalizes both false alarms and omissions without being dominated by the much larger background class. LOSO values are arithmetic means of the four held-out-scene IoUs. Full-scene averages are arithmetic means over the four complete scenes.

3. Results

The experiments first separate cross-scene generalization from matched-distribution robustness, then analyze the training protocol, model components, sensitivity factors, external baselines, controlled optical degradation, and full-scene deployment. Unless otherwise stated, all controlled comparisons use the mixed-scene split and the same Full, Missing-O, and Missing-S inference conditions.

3.1. Data-Protocol Effect

The base-phase protocol study in Table 4 reports the four LOSO folds together with the matched mixed-scene test result. Averaged across the held-out scenes, LOSO obtains Full/Missing-O/Missing-S IoUs of 0.5898/0.3917/0.5655, whereas the mixed-scene protocol obtains 0.7760/0.6049/0.7659. The gap confirms that mixed-scene testing is useful for controlled robustness diagnosis but cannot replace evaluation on geographically unseen scenes. Scene 003 is the most challenging LOSO case, with a Full IoU of 0.2819 and a Missing-O IoU of 0.1150, consistent with its fragmented parcels and 3.9% positive-label ratio.

3.2. Missing-Modality Training Protocol

The base-phase training-protocol analysis in Table 5 shows a strong directional effect. Training only with missing optical input gives a Missing-O IoU of 0.5985 but reduces Full and Missing-S IoUs to 0.6704 and 0.6057. Training only with missing SAR input reaches 0.6600 under Missing-S inference but decreases to 0.2647 under Missing-O inference. Mixed optical+SAR missing simulation provides the best balance, reaching 0.7760/0.6049/0.7659 and matching or exceeding the single-target protocols in every reported IoU column. For this model, optical absence primarily increases omissions: Recall decreases from 0.8984 to 0.7281, whereas Precision decreases from 0.8506 to 0.7814.

3.3. Component Ablation and Sensitivity

The base-phase component ablations in Table 6 reveal complementary rather than uniformly monotonic effects. Removing the input adapter causes the largest Missing-O decrease, from 0.6049 to 0.5106, demonstrating the importance of stabilizing the distribution shift created by an unavailable optical phase. Removing the dilated context block reduces all three IoUs. Removing the difference skips or fixing the availability map can improve one condition slightly, but both reduce the three-condition mean. The complete MARC-Net therefore provides the most balanced component configuration, with the highest mean IoU of 0.7156.
The base-phase sensitivity study in Table 7 shows that robustness depends strongly on the exposure schedule. Increasing the missing probability from 0.50 to 0.75 raises Missing-O IoU from 0.6049 to 0.6325 and the mean from 0.7156 to 0.7248, with a small Full-IoU decrease to 0.7732. Fixed-phase masking is less effective than randomly selecting either phase, and the CE+ Dice objective reduces the mean to 0.6900. We therefore retain the pre-specified ρ = 0.50 /either/CE setting for the unified baseline comparison and report ρ = 0.75 as a robustness-oriented sensitivity point.

3.4. External Baselines

All methods in Table 8 use the same 1104/184/184 split, random seed 2022, deterministic test missing schedule, and argmax decision rule without test-set threshold tuning. The external baselines use an audited 25-epoch from-scratch schedule with ρ = 0.50 mixed missing simulation, while MARC-Net uses the complete two-phase optimization strategy specified in Section 2.5. Checkpoints are selected exclusively from validation performance. SwinSUNet uses its released encoder, temporal fusion, and decoder with a shared learned 1 × 1 projection from six channels to three. ChangeMamba uses the official CMBackbone and CMDecoder with the same shared 6 3 projection.
MARC-Net attains the highest IoU in all three inference conditions, reaching Full/Missing-O/Missing-S IoUs of 0.7826/0.7039/0.7780 and a three-condition mean of 0.7548. The strongest external mean is 0.7381 from SwinSUNet, giving MARC-Net a 0.0168 absolute mean-IoU advantage. The official-source ChangeMamba run obtains 0.7434/0.3038/0.4468 (mean 0.4980) under the same split and decision rule. The condition-balanced objective raises Missing-O IoU by 0.0990 relative to the variant without this component while also improving Full and Missing-S IoU. MARC-Net uses 41.08 million trainable parameters, fewer than SwinSUNet (43.58 million) and the adapted official ChangeMamba model (48.56 million).
Two unsupervised heterogeneous detectors are also evaluated on all 184 mixed-scene test samples. For Missing- O 1 , their input is ( S 1 , O 2 ) ; for Missing- O 2 , it is ( O 1 , S 2 ) . Table 9 reports these methods separately because they use one heterogeneous pair and no target labels, whereas MARC-Net is supervised and retains three temporal–modal observations. Their substantially lower IoUs show that directly applying heterogeneous pairwise detectors does not solve the target missing-optical setting, but the protocol difference prevents a capacity-equivalent ranking.
The released RIEM pipeline was executed as provided. For INLPG, the released DB4 wavelet-fusion binary was incompatible with the tested Windows/MATLAB R2016a environment. We therefore retained its official forward and backward distance-map computation and replaced only the final fusion step with symmetric averaging followed by Otsu thresholding. The INLPG result is consequently reported as a documented compatibility-adapted evaluation rather than a bitwise reproduction of the released environment.

3.5. Qualitative Robustness and Optical Degradation

Figure 6 presents six spatially diverse cases generated from the reported checkpoint, with consistently high minimum and mean IoU across Full, Missing-O, and Missing-S inference. The complete eight-column panel includes both optical dates, both SAR dates, the reference mask, and all three predictions, making the qualitative evidence directly traceable to the quantitative result.
To complement complete channel removal, deterministic synthetic degradation is applied to exactly one optical phase while keeping its availability flag at 1 and leaving both SAR observations unchanged. For the reported checkpoint, clean/Cloud-20/Cloud-40/Cloud-60 IoUs are 0.7826/0.5980/0.4482/0.3016; haze, shadow, and radiometric degradation yield 0.5260, 0.7064, and 0.4338. This controlled stress test quantifies the response to progressive observation-quality degradation and complements future evaluation on real cloud-labeled imagery. Table 10 summarizes these controlled-degradation results.

3.6. Full-Scene Deployment Evaluation

For deployment-scale diagnosis, the mixed-scene-trained base-phase MARC-Net is applied to each complete GeoTIFF using 512 × 512 tiles, a 256-pixel stride, and Hann-window overlap blending. Table 11 reports only the deployment results; the LOSO values remain available in Table 4. Average full-scene IoU is 0.7204/0.4435/0.7139 under Full/Missing-O/Missing-S inference. The small Full-to-Missing-S decrease indicates a complementary SAR contribution, whereas the much larger Missing-O decrease confirms that optical observations remain the principal semantic source.
Scene 003 is examined separately in Figure 7 as the most challenging deployment case. Its Missing-O IoU is 0.1410 at full-scene scale. In the selected 1024 × 1024 hotspot, the error map is dominated by missed fragmented parcels rather than a uniform increase in false alarms, indicating that sparse labels, scene shift, and parcel geometry jointly limit optical-missing inference.
Figure 8 shows Scene 002, the most robust deployment case. Full and Missing-S predictions remain spatially coherent, whereas the Missing-O result contains localized omissions along narrow parcel boundaries and spectrally ambiguous regions.

4. Discussion

The experiments show that missing-modality robustness is jointly determined by the data protocol, training exposure, and architectural mechanism. Mixed-scene testing provides a controlled environment for comparing missing-input behavior, whereas LOSO evaluates the complementary challenge of unseen-scene generalization. The 0.1862 Full-IoU difference reflects the distinct questions answered by matched-distribution and cross-scene evaluation and supports reporting both protocols together.
The ablations identify residual input stabilization as the most important component for Missing-O inference, while dilated context and multi-level difference skips improve the cross-condition balance. Training exposure is equally important. Randomly masking either temporal phase is markedly stronger than fixed-phase masking, and increasing ρ to 0.75 raises the mean IoU to 0.7248. These findings support the design choice of combining an explicit availability signal with distribution stabilization and mixed single-phase missing simulation rather than relying on zero filling alone. Building on the base-phase evidence, temporal proxy stabilization reduces the sentinel-to-observation gap using only the available other-date signal, and the condition-balanced objective converts this mechanism into a substantial Missing-O improvement without sacrificing the other conditions.
The proposed condition-balanced strategy establishes a strong and balanced operating point. MARC-Net reaches a three-condition mean IoU of 0.7548, compared with 0.7381 for the strongest external mean (SwinSUNet) and 0.4980 for the official-source ChangeMamba run. The gain is largest under Missing-O inference, while Full and Missing-S performance are preserved. This improvement is attributable to condition-aware temporal stabilization and joint multi-state optimization rather than a larger backbone: MARC-Net retains 41.08 million parameters. These results establish the effectiveness of the proposed design on the evaluated reclaimed-cropland scenes and motivate broader geographic validation.
The controlled degradation experiment further separates complete absence from corrupted-but-present optical input. The monotonic cloud-severity trend and the retained performance under shadow perturbation show that MARC-Net responds systematically to observation quality rather than only to the binary availability flag. Continuous optical-quality estimation and evaluation on real cloud- and haze-labeled imagery are promising extensions of the present framework.
Scene 003 is the most challenging case because its positive class is sparse and its hilly parcels are fragmented. The CE + Dice experiment indicates that overlap supervision alone is insufficient for this setting, motivating scene-aware sampling, class-balanced objectives, temporal-mismatch evaluation, and broader geographic transfer. Across the four complete scenes, overlap-blended sliding-window inference produces spatially coherent maps without obvious tiling artifacts.

5. Conclusions

This study formulates missing-optical bi-temporal optical–SAR change detection for effective reclaimed-cropland monitoring and proposes MARC-Net, which integrates availability-conditioned inputs, condition-aware temporal proxy stabilization, residual adaptation, shared temporal encoding, signed multi-level differences, dilated context refinement, and hierarchical SCSE decoding. The proposed model reaches Full/Missing-O/Missing-S IoUs of 0.7826/0.7039/0.7780 and a three-condition mean of 0.7548. It provides the strongest balanced result in the completed comparison while retaining one checkpoint and fewer parameters than the evaluated SwinSUNet and ChangeMamba implementations. LOSO and deployment analyses further characterize the remaining challenge of unseen-scene transfer.
The results demonstrate that robust incomplete-input monitoring is feasible without reconstructing the unavailable optical image. MARC-Net provides a practical foundation for reclaimed-cropland monitoring under variable sensor availability; future work will extend optical-quality conditioning, real-atmospheric evaluation, and transfer to sparse and fragmented unseen scenes.

Author Contributions

Conceptualization, Y.Z. and X.D.; methodology, C.F. and X.D.; software, Z.W.; validation, J.M.; formal analysis, Z.W. and Y.Z.; investigation, Y.Z. and H.Y.; data curation, H.Y. and X.W.; writing—original draft preparation, Y.Z. and J.M.; writing—review and editing, X.D. and F.H.; visualization, Z.W.; supervision, C.F. and F.H.; project administration, F.H. All authors have read and agreed to the published version of the manuscript.

Funding

This project was supported by the Open Fund of the Key Laboratory of National Geographical Census and Monitoring, Ministry of Natural Resources (Grant No. 2024NGCM02); the “Pioneer, Leading Goose + X” R&D Program of Zhejiang under Grant No. 2025C02052; and the Ministry–Province Cooperation Project under the Ministry of Natural Resources (Grant No. 2024ZRBSHZ001).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The datasets and derived experimental records used in this study are available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

LOSOLeave-one-scene-out
SARSynthetic aperture radar
IoUIntersection over union

References

  1. Zhang, L.; Zhang, L. Artificial intelligence for remote sensing data analysis: A review of challenges and opportunities. IEEE Geosci. Remote Sens. Mag. 2022, 10, 270–294. [Google Scholar] [CrossRef] [Scilit]
  2. Liu, G.; Li, Y.; Chen, Y.; Lu, Y.; Jiang, D.; Xu, A.; Zhong, Y.; Yin, H. Mapping Abandoned Cropland in Tropical/Subtropical Monsoon Areas with Multiple Crop Maturity Patterns. Int. J. Appl. Earth Obs. Geoinf. 2024, 127, 103674. [Google Scholar] [CrossRef] [Scilit]
  3. Yang, Y.; Wu, Z.; Xiao, W.; Zhou, Y.; Huang, Q.; Wu, T.; Luo, J.; Wang, H. Abandoned Land Mapping Based on Spatiotemporal Features from PolSAR Data via Deep Learning Methods. Remote Sens. 2023, 15, 3942. [Google Scholar] [CrossRef] [Scilit]
  4. Gui, S.; Li, J.; Chen, G.; Zhao, J.; Tang, B.; Li, L. Identification of Abandoned Cropland and Global–Local Driving Mechanism Analysis via Multi-Source Remote Sensing Data and Multi-Objective Optimization. Remote Sens. 2025, 17, 3086. [Google Scholar] [CrossRef] [Scilit]
  5. Liu, C.; Sun, Y.; Xu, Y.; Sun, Z.; Zhang, X.; Lei, L.; Kuang, G. A Review of Optical and SAR Image Deep Feature Fusion in Semantic Segmentation. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 12910–12930. [Google Scholar] [CrossRef] [Scilit]
  6. Xu, C.; Geng, Z.; Wu, L.; Zhu, D. Enhanced Semantic Segmentation in Remote Sensing Images with SAR-Optical Image Fusion and Image Translation. Sci. Rep. 2025, 15, 35433. [Google Scholar] [CrossRef] [Scilit]
  7. Kang, W.; Xiang, Y.; Wang, F.; You, H. CFNet: A Cross Fusion Network for Joint Land Cover Classification Using Optical and SAR Images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 1562–1574. [Google Scholar] [CrossRef] [Scilit]
  8. Li, X.; Lei, L.; Sun, Y.; Li, M.; Kuang, G. Multimodal Bilinear Fusion Network with Second-Order Attention-Based Channel Selection for Land Cover Classification. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2020, 13, 1011–1026. [Google Scholar] [CrossRef] [Scilit]
  9. Wan, L.; Xiang, Y.; You, H. A Post-Classification Comparison Method for SAR and Optical Images Change Detection. IEEE Geosci. Remote Sens. Lett. 2019, 16, 1026–1030. [Google Scholar] [CrossRef] [Scilit]
  10. Wan, L.; Xiang, Y.; You, H. An Object-Based Hierarchical Compound Classification Method for Change Detection in Heterogeneous Optical and SAR Images. IEEE Trans. Geosci. Remote Sens. 2019, 57, 9941–9959. [Google Scholar] [CrossRef]
  11. Long, J.; Shelhamer, E.; Darrell, T. Fully Convolutional Networks for Semantic Segmentation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Boston, MA, USA, 7–12 June 2015; IEEE: New York, MA, USA, 2015; pp. 3431–3440. [Google Scholar]
  12. Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Proceedings of the 18th International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI 2015), Munich, Germany, 5–9 October 2015; Springer International Publishing: Berlin/Heidelberg, Germany, 2015; pp. 234–241. [Google Scholar]
  13. Chen, L.-C.; Papandreou, G.; Kokkinos, I.; Murphy, K.; Yuille, A.L. DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 40, 834–848. [Google Scholar] [CrossRef] [Scilit]
  14. He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar]
  15. Woo, S.; Park, J.; Lee, J.-Y.; Kweon, I.S. CBAM: Convolutional Block Attention Module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; Springer International Publishing: Berlin/Heidelberg, Germany, 2018; pp. 3–19. [Google Scholar]
  16. Daudt, R.C.; Le Saux, B.; Boulch, A. Fully Convolutional Siamese Networks for Change Detection. In Proceedings of the 25th IEEE International Conference on Image Processing (ICIP); IEEE: New York, NY, USA, 2018; pp. 4063–4067. [Google Scholar] [CrossRef] [Scilit]
  17. Liu, Y.; Pang, C.; Zhan, Z.; Zhang, X.; Yang, X. Building Change Detection for Remote Sensing Images Using a Dual-Task Constrained Deep Siamese Convolutional Network Model. IEEE Geosci. Remote Sens. Lett. 2021, 18, 811–815. [Google Scholar] [CrossRef] [Scilit]
  18. Chen, H.; Qi, Z.; Shi, Z. Remote Sensing Image Change Detection with Transformers. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5607514. [Google Scholar] [CrossRef] [Scilit]
  19. Bandara, W.G.C.; Patel, V.M. A Transformer-Based Siamese Network for Change Detection. In Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (IGARSS); IEEE: New York, NY, USA, 2022; pp. 207–210. [Google Scholar] [CrossRef] [Scilit]
  20. Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV); IEEE: New York, NY, USA, 2021; pp. 10012–10022. [Google Scholar]
  21. Zhang, C.; Wang, L.; Cheng, S.; Li, Y. SwinSUNet: Pure Transformer Network for Remote Sensing Image Change Detection. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5224713. [Google Scholar] [CrossRef] [Scilit]
  22. Chen, H.; Song, J.; Han, C.; Xia, J.; Yokoya, N. ChangeMamba: Remote Sensing Change Detection with Spatiotemporal State Space Model. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4409720. [Google Scholar] [CrossRef] [Scilit]
  23. Zhang, H.; Chen, K.; Liu, C.; Chen, H.; Zou, Z.; Shi, Z. CDMamba: Incorporating Local Clues Into Mamba for Remote Sensing Image Binary Change Detection. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4405016. [Google Scholar] [CrossRef] [Scilit]
  24. Ding, L.; Guo, H.; Liu, S.; Mou, L.; Zhang, J.; Bruzzone, L. Bi-Temporal Semantic Reasoning for the Semantic Change Detection in HR Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2022, 60, 5620014. [Google Scholar] [CrossRef] [Scilit]
  25. Jiang, L.; Li, F.; Huang, L.; Peng, F.; Hu, L. TTNet: A Temporal-Transform Network for Semantic Change Detection Based on Bi-Temporal Remote Sensing Images. Remote Sens. 2023, 15, 4555. [Google Scholar] [CrossRef] [Scilit]
  26. Wang, L.; Sun, Q.; Pei, J.; Khan, M.A.; Al Dabel, M.M.; Al-Otaibi, Y.D.; Bashir, A.K. Bitemporal Remote Sensing Change Detection with State-Space Models. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2025, 18, 14942–14958. [Google Scholar] [CrossRef] [Scilit]
  27. Ni, Y.; Liu, S.; Guo, T.; Xia, M. TiBT-Net: A High-Resolution Remote Sensing Image Change Detection Network Integrating Bi-Temporal Space Enhancement and Token Interaction. Remote Sens. 2026, 18, 805. [Google Scholar] [CrossRef] [Scilit]
  28. Liu, Y.; Chen, K.; Liu, C.; Qin, Z.; Luo, Z.; Wang, J. Structured Knowledge Distillation for Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, CA, USA, 16–20 June 2019; IEEE: New York, NY, USA, 2019. [Google Scholar]
  29. Yang, C.; Zhou, H.; An, Z.; Jiang, X.; Xu, Y.; Zhang, Q. Cross-Image Relational Knowledge Distillation for Semantic Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 21–24 June 2022; IEEE: New York, NY, USA, 2022; pp. 12319–12328. [Google Scholar]
  30. Yuan, J.; Phan, M.H.; Liu, L.; Liu, Y. FAKD: Feature Augmented Knowledge Distillation for Semantic Segmentation. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, HI, USA, 3–8 January 2024; IEEE: New York, NY, USA, 2024; pp. 595–605. [Google Scholar]
  31. Dong, Z.; Gao, G.; Liu, T.; Gu, Y.; Zhang, X. Distilling Segmenters From CNNs and Transformers for Remote Sensing Images’ Semantic Segmentation. IEEE Trans. Geosci. Remote Sens. 2023, 61, 5613814. [Google Scholar] [CrossRef] [Scilit]
  32. Jiang, X.; Li, G.; Liu, Y.; Zhang, X.-P.; He, Y. Change Detection in Heterogeneous Optical and SAR Remote Sensing Images Via Deep Homogeneous Feature Fusion. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2020, 13, 1551–1566. [Google Scholar] [CrossRef] [Scilit]
  33. Saha, S.; Shahzad, M.; Ebel, P.; Zhu, X.X. Supervised Change Detection Using Prechange Optical-SAR and Postchange SAR Data. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2022, 15, 8170–8179. [Google Scholar] [CrossRef] [Scilit]
  34. Sun, Z.; Zhi, S.; Li, R.; Xia, J.; Liu, Y.; Jiang, W. GDROS: A Geometry-Guided Dense Registration Framework for Optical–SAR Images Under Large Geometric Transformations. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5650315. [Google Scholar] [CrossRef] [Scilit]
  35. Hu, F.; Wang, J. Detecting Spatio-Temporal Urban Surface Changes Using Identified Temporary Coherent Scatterers. J. Syst. Eng. Electron. 2021, 32, 1304–1317. [Google Scholar] [CrossRef] [Scilit]
  36. Quan, Y.; Zhang, R.; Li, J.; Ji, S.; Guo, H.; Yu, A. Learning SAR-Optical Cross Modal Features for Land Cover Classification. Remote Sens. 2024, 16, 431. [Google Scholar] [CrossRef] [Scilit]
  37. Shao, Y.; Zhang, X.; Zhang, S.; Chen, J.; Huang, Y.; Zhu, Z. A Multilevel Optical-SAR Imagery Fusion Network for Land Cover Classification. In Proceedings of SPIE 13818, Second International Conference on Remote Sensing and Global Positioning Algorithm; SPIE: Bellingham, WA, USA, 2025; p. 1381812. [Google Scholar]
  38. Wang, Y.; Zhang, W.; Chen, W.; Chen, C.; Liang, Z. MFFnet: Multimodal Feature Fusion Network for Synthetic Aperture Radar and Optical Image Land Cover Classification. Remote Sens. 2024, 16, 2459. [Google Scholar] [CrossRef] [Scilit]
  39. Ren, B.; Liu, B.; Wang, Q.; Hou, B.; Yang, C.; Jiao, L. DCIFNet: Cross-Modal Fusion with Correction and Interaction for Optical–SAR Land Cover Classification. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5643218. [Google Scholar] [CrossRef] [Scilit]
  40. Wu, W.; Guo, S.; Shao, Z.; Li, D. CroFuseNet: A Semantic Segmentation Network for Urban Impervious Surface Extraction Based on Cross Fusion of Optical and SAR Images. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2023, 16, 3801–3814. [Google Scholar] [CrossRef] [Scilit]
  41. Chang, H.; Fu, X.; Guo, K.; Dong, J.; Guan, J.; Liu, C. SOLSTM: Multisource Information Fusion Semantic Segmentation Network Based on SAR-OPT Matching Attention and Long Short-Term Memory Network. IEEE Geosci. Remote Sens. Lett. 2025, 22, 4004705. [Google Scholar] [CrossRef] [Scilit]
  42. Tan, Y.; Li, M.; Xu, K.; Lai, G. HAFNet: A Heterogeneous Adaptive Fusion Network of Optical and SAR Imagery for Improved Land Use Classification. Photogramm. Rec. 2025, 40, e70028. [Google Scholar] [CrossRef] [Scilit]
  43. Han, W.; Jiang, W.; Geng, J.; Bao, Y. Semantic Segmentation of Remote Sensing Images with Inconsistent Resolutions via a Spectral-Geometric Iterative Fusion Network. IEEE Trans. Geosci. Remote Sens. 2025, 63, 4419314. [Google Scholar] [CrossRef] [Scilit]
  44. Gong, Y.; Chen, C.; Zheng, Y. Hybrid Deep Learning Model for Multi-Source Remote Sensing Data Fusion: Integrating DenseNet and Swin Transformer for Spatial Alignment and Feature Extraction. Informatica 2025, 49, 197–208. [Google Scholar] [CrossRef] [Scilit]
  45. Zhao, Z.; Yan, J.; Li, C.; Wang, X.; Jiang, P.; Tang, J. DehazeMamba: SAR-Guided Optical Remote Sensing Image Dehazing with Adaptive State Space Model. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2026, 19, 778–794. [Google Scholar] [CrossRef] [Scilit]
  46. Zhou, X.; Fang, Q.; Gong, X.; Yang, S.; Lu, T.; Wan, Y.; Ma, A.; Zhong, Y. AFR-CR: An Adaptive Frequency Domain Feature Reconstruction-Based Method for Cloud Removal via SAR-Assisted Remote Sensing Image Fusion. Remote Sens. 2026, 18, 201. [Google Scholar] [CrossRef] [Scilit]
  47. Guo, Z.; Wu, L.; Huang, Y.; Guo, Z.; Zhao, J.; Li, N. Water-Body Segmentation for SAR Images: Past, Current, and Future. Remote Sens. 2022, 14, 1752. [Google Scholar] [CrossRef] [Scilit]
  48. Kim, M.U.; Oh, H.; Lee, S.-J.; Choi, Y.; Han, S. Deep Learning Based Water Segmentation Using KOMPSAT-5 SAR Images. In Proceedings of he 2021 IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Brussels, Belgium, 11–16 July 2021; IEEE: New York, NY, USA, 2021; pp. 4055–4058. [Google Scholar]
  49. Huang, B.; Li, P.; Lu, H.; Yin, J.; Li, Z.; Wang, H. WaterDetectionNet: A New Deep Learning Method for Flood Mapping with SAR Image Convolutional Neural Network. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 14471–14485. [Google Scholar] [CrossRef] [Scilit]
  50. Chen, Y.; Zhao, M.; Bruzzone, L. A Novel Approach to Incomplete Multimodal Learning for Remote Sensing Data Fusion. IEEE Trans. Geosci. Remote Sens. 2024, 62, 5404914. [Google Scholar] [CrossRef] [Scilit]
  51. Liu, X.; Jin, F.; Wang, S.; Rui, J.; Zuo, X.; Yang, X.; Cheng, C. Multimodal Online Knowledge Distillation Framework for Land Use/Cover Classification Using Full or Missing Modalities. IEEE Trans. Geosci. Remote Sens. 2024, 62, 4406817. [Google Scholar] [CrossRef] [Scilit]
  52. Wei, S.; Luo, Y.; Ma, X.; Ren, P.; Luo, C. MSH-Net: Modality-Shared Hallucination with Joint Adaptation Distillation for Remote Sensing Image Classification Using Missing Modalities. IEEE Trans. Geosci. Remote Sens. 2023, 61, 4402615. [Google Scholar] [CrossRef] [Scilit]
  53. Kampffmeyer, M.; Salberg, A.-B.; Jenssen, R. Urban Land Cover Classification with Missing Data Modalities Using Deep Convolutional Neural Networks. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2018, 11, 1758–1768. [Google Scholar] [CrossRef] [Scilit]
  54. Wang, Z.; Ma, J.; Wang, Z.; Hu, F. MICA-Net: A Missing-Modality Invariant Cross-Attention Network for Land Cover Mapping. In Proceedings of the IEEE International Geoscience and Remote Sensing Symposium (IGARSS), Washington, DC, USA, 9–14 August 2026. [Google Scholar]
  55. Qu, J.; Yang, Y.; Dong, W.; Yang, Y. LDS2AE: Local Diffusion Shared-Specific Autoencoder for Multimodal Remote Sensing Image Classification with Arbitrary Missing Modalities. In Proceedings of the the 38th AAAI Conference on Artificial Intelligence, Vancouver, BC, Canada, 20–27 February 2024. [Google Scholar]
  56. Zhou, Y.; Ma, A.; Wang, J.; Chen, Z.; Zhong, Y. Remote Sensing Meta Modal Representation for Missing Modality Land Cover Mapping: From EarthMiss Dataset to MetaRS Method. Remote Sens. Environ. 2026, 333, 115132. [Google Scholar] [CrossRef] [Scilit]
  57. Zhu, S.; Ablameyko, S.; Li, J. Dual-Strategy Improvement of YOLOv11n for Multi-Scale Object Detection in Remote Sensing Images. Comput. Mater. Contin. 2026, 88, 56. [Google Scholar] [CrossRef] [Scilit]
  58. Zhao, J.; Sun, W.; Chen, Z.; Hu, Q.; Li, Y.; Shi, H.; Niu, Y.; Lu, Z. ArgusSAR: A Multidimensional Perception and Robust Detection Method for Multiscale Objects in SAR Images. IEEE Trans. Geosci. Remote Sens. 2026, 64, 5204317. [Google Scholar] [CrossRef] [Scilit]
  59. Huang, Y.; Wang, Z.; Tang, T.; Ohtsuki, T.; Gui, G. Dual-Stream Multimodal Fusion with Local–Global Attention for Remote-Sensing Object Detection. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2026, 19, 1691–1702. [Google Scholar] [CrossRef] [Scilit]
  60. Sun, Y.; Lei, L.; Li, X.; Tan, X.; Kuang, G. Structure Consistency-Based Graph for Unsupervised Change Detection with Homogeneous and Heterogeneous Remote Sensing Images. IEEE Trans. Geosci. Remote Sens. 2022, 60, 4700221. [Google Scholar] [CrossRef] [Scilit]
  61. Sun, Y.; Lei, L.; Li, Z.; Kuang, G.; Yu, Q. Detecting Changes without Comparing Images: Rules Induced Change Detection in Heterogeneous Remote Sensing Images. ISPRS J. Photogramm. Remote Sens. 2025, 230, 241–257. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Dataset construction pipeline for effective reclaimed cropland change detection. The figure illustrates the organization of O 1 , S 1 , O 2 , and S 2 , the construction of binary reclaimed-cropland change labels, patch generation, and the training/validation/test split.
Figure 1. Dataset construction pipeline for effective reclaimed cropland change detection. The figure illustrates the organization of O 1 , S 1 , O 2 , and S 2 , the construction of binary reclaimed-cropland change labels, patch generation, and the training/validation/test split.
Remotesensing 18 02960 g001
Figure 2. Geographic overview of the study area and the four dataset scenes. A Sentinel-2 true-color image acquired on 17 May 2024 serves as the background. The red, blue, green, and orange rectangles denote Scenes 001, 002, 003, and 004, respectively. The four labeled rectangles indicate the spatial extent of each scene, covering a total area of approximately 13 km (east–west) by 30 km (north–south) in the Hangzhou region, Zhejiang Province, China.
Figure 2. Geographic overview of the study area and the four dataset scenes. A Sentinel-2 true-color image acquired on 17 May 2024 serves as the background. The red, blue, green, and orange rectangles denote Scenes 001, 002, 003, and 004, respectively. The four labeled rectangles indicate the spatial extent of each scene, covering a total area of approximately 13 km (east–west) by 30 km (north–south) in the Hangzhou region, Zhejiang Province, China.
Remotesensing 18 02960 g002
Figure 3. Task definition for effective reclaimed cropland detection. The figure shows how bi-temporal optical–SAR observations support the identification of effective reclaimed cropland change, and how the binary target map Y is derived from the task definition. Blue arrows indicate the optical–SAR observations entering the change-analysis module, green arrows and regions indicate effective reclaimed-cropland transitions and the positive class, and gray arrows and regions indicate excluded changes and the negative class.
Figure 3. Task definition for effective reclaimed cropland detection. The figure shows how bi-temporal optical–SAR observations support the identification of effective reclaimed cropland change, and how the binary target map Y is derived from the task definition. Blue arrows indicate the optical–SAR observations entering the change-analysis module, green arrows and regions indicate effective reclaimed-cropland transitions and the positive class, and gray arrows and regions indicate excluded changes and the negative class.
Remotesensing 18 02960 g003
Figure 4. Architecture of MARC-Net. Each time point is represented by four optical channels, one SAR channel, and one availability channel. When a single phase is unavailable, condition-aware temporal proxy stabilization uses only the corresponding available observation from the other date before shared residual adaptation. The shared SE-ResNet34 encoder produces four feature levels; signed bi-temporal differences are refined by sequential dilated convolution and decoded through hierarchical SCSE blocks with multi-level difference skips. A single head predicts the full-resolution change map, while mixed missing simulation and the condition-balanced multi-state objective are used only during training.
Figure 4. Architecture of MARC-Net. Each time point is represented by four optical channels, one SAR channel, and one availability channel. When a single phase is unavailable, condition-aware temporal proxy stabilization uses only the corresponding available observation from the other date before shared residual adaptation. The shared SE-ResNet34 encoder produces four feature levels; signed bi-temporal differences are refined by sequential dilated convolution and decoded through hierarchical SCSE blocks with multi-level difference skips. A single head predicts the full-resolution change map, while mixed missing simulation and the condition-balanced multi-state objective are used only during training.
Remotesensing 18 02960 g004
Figure 5. Training and inference protocol. Complete samples are mixed with single-phase optical-missing and SAR-missing samples during training. The primary evaluation removes O 1 or O 2 while retaining both SAR observations; missing-SAR inference is reported as a complementary stress test. No missing image is reconstructed.
Figure 5. Training and inference protocol. Complete samples are mixed with single-phase optical-missing and SAR-missing samples during training. The primary evaluation removes O 1 or O 2 while retaining both SAR observations; missing-SAR inference is reported as a complementary stress test. No missing image is reconstructed.
Remotesensing 18 02960 g005
Figure 6. Spatially diverse robust examples from the reported MARC-Net checkpoint under Full, Missing-O, and Missing-S inference. Each row shows O 1 , S 1 , O 2 , S 2 , the ground truth, and the three predictions.
Figure 6. Spatially diverse robust examples from the reported MARC-Net checkpoint under Full, Missing-O, and Missing-S inference. Each row shows O 1 , S 1 , O 2 , S 2 , the ground truth, and the three predictions.
Remotesensing 18 02960 g006
Figure 7. Focused challenging-case analysis for Scene 003. The panels show O 1 , O 2 , S 1 , S 2 , the ground truth, the Full prediction, the Missing-O prediction, and the Missing-O error map. Green, red, and blue denote true positives, false negatives, and false positives. Full-scene IoUs are 0.6900, 0.1410, and 0.6798 under Full, Missing-O, and Missing-S inference.
Figure 7. Focused challenging-case analysis for Scene 003. The panels show O 1 , O 2 , S 1 , S 2 , the ground truth, the Full prediction, the Missing-O prediction, and the Missing-O error map. Green, red, and blue denote true positives, false negatives, and false positives. Full-scene IoUs are 0.6900, 0.1410, and 0.6798 under Full, Missing-O, and Missing-S inference.
Remotesensing 18 02960 g007
Figure 8. Full-scene prediction maps for Scene 002: (a) ground truth; (b) Full prediction; (c) Missing-O prediction; and (d) Missing-S prediction. White denotes detected change, and black denotes background.
Figure 8. Full-scene prediction maps for Scene 002: (a) ground truth; (b) Full prediction; (c) Missing-O prediction; and (d) Missing-S prediction. White denotes detected change, and black denotes background.
Remotesensing 18 02960 g008
Table 1. Summary of the four dataset scenes. Scene IDs are used consistently throughout the experiments.
Table 1. Summary of the four dataset scenes. Scene IDs are used consistently throughout the experiments.
SceneSize (Pixels)Label Ratio (%)
00111,500 × 420027.0
00211,300 × 210041.7
003 3200 × 7300 3.9
004 2100 × 2700 25.9
Table 2. Dataset organization used for bi-temporal optical–SAR reclaimed cropland change detection.
Table 2. Dataset organization used for bi-temporal optical–SAR reclaimed cropland change detection.
SubsetPairsPatch SizeModalitiesSplit Protocol/Description
Training1104 512 × 512 O 1 , S 1 , O 2 , S 2 Mixed-scene training with complete and simulated single-phase optical- or SAR-missing samples.
Validation184 512 × 512 O 1 , S 1 , O 2 , S 2 Complete-input model selection on the fixed mixed-scene validation split.
Testing184 512 × 512 O 1 , S 1 , O 2 , S 2 with controlled missing settingsThe same fixed test pairs are evaluated under Full, Missing-O, and Missing-S conditions.
Table 3. Training and implementation settings of the reported MARC-Net.
Table 3. Training and implementation settings of the reported MARC-Net.
ItemSetting
Time-point input4 optical + 1 SAR + 1 joint-availability channel
Shared encoderSE-ResNet34, layer configuration [ 3 , 4 , 6 , 3 ]
Deep contextSequential dilated convolutions, rates 1 , 2 , 4 , 8
DecoderFour SCSE upsampling blocks with three difference skips
Prediction headsOne two-class full-resolution head
Patch size 512 × 512
Representation-learning phaseAdam, 1 × 10 4 , 25 epochs, mixed single-phase missing simulation
Condition-balanced phaseFP32 AdamW, 1 × 10 5 1 × 10 6 , 5 epochs
Condition-balanced stabilizationWeight decay 10 4 ; gradient clip 2; EMA 0.995; frozen BN statistics
Batch sizesTrain 2/validation 8
Random seed2022
Missing target/modeMixed optical+SAR/single-phase either
ρ / π o / π 2 0.5/0.6/0.6
Condition weightsFull 1.15/Missing-O 1.60/Missing-S 1.25
Temporal-proxy coefficients α o = 0.25 / α s = 1.00
Checkpoint ruleHighest three-condition validation mean IoU
Data augmentationRandom horizontal and vertical flips
Table 4. Base-phase LOSO scene-level IoU and matched mixed-scene test result for MARC-Net. LOSO averages are arithmetic means over the four held-out scenes. Bold values indicate the highest IoU in each inference-condition column.
Table 4. Base-phase LOSO scene-level IoU and matched mixed-scene test result for MARC-Net. LOSO averages are arithmetic means over the four held-out scenes. Bold values indicate the highest IoU in each inference-condition column.
ProtocolEvaluation SetFull IoUMissing-O IoUMissing-S IoU
LOSOScene 0010.69700.37500.7215
LOSOScene 0020.77760.58640.7734
LOSOScene 0030.28190.11500.1812
LOSOScene 0040.60270.49040.5860
LOSOAverage0.58980.39170.5655
Mixed sceneOfficial test split0.77600.60490.7659
Table 5. Base-phase training-protocol comparison of MARC-Net on the mixed-scene test set. Each inference condition reports IoU, Precision (P), and Recall (R). Bold values indicate the highest IoU in each inference-condition column.
Table 5. Base-phase training-protocol comparison of MARC-Net on the mixed-scene test set. Each inference condition reports IoU, Precision (P), and Recall (R). Bold values indicate the highest IoU in each inference-condition column.
Training ProtocolFullMissing-OMissing-S
IoUPRIoUPRIoUPR
Train Missing Optical0.67040.84060.76810.59850.78490.71600.60570.76840.7410
Train Missing SAR0.65530.82720.75920.26470.48220.36980.66000.83580.7583
Train Missing Optical + SAR0.77600.85060.89840.60490.78140.72810.76590.85210.8833
Table 6. Base-phase component-level ablation of MARC-Net under the mixed-scene protocol. Each condition reports IoU, and Mean is the unweighted average of the three IoUs. Bold values indicate the highest IoU in each column.
Table 6. Base-phase component-level ablation of MARC-Net under the mixed-scene protocol. Each condition reports IoU, and Mean is the unweighted average of the three IoUs. Bold values indicate the highest IoU in each column.
VariantControlled ChangeFullMissing-OMissing-SMean
w/o availability signalAvailability map fixed to 10.76670.61830.74970.7116
w/o input adapterDirect six-channel encoder input0.77630.51060.76800.6850
w/o dilated contextNo deepest dilated block0.77190.58130.75690.7034
w/o difference skipsDeepest difference only0.77380.60920.73910.7074
MARC-NetComplete architecture0.77600.60490.76590.7156
Table 7. Base-phase sensitivity of MARC-Net to missing probability ρ , simulated missing phase, and optimization loss. Each entry reports IoU. Bold values indicate the highest IoU in each column.
Table 7. Base-phase sensitivity of MARC-Net to missing probability ρ , simulated missing phase, and optimization loss. Each entry reports IoU. Bold values indicate the highest IoU in each column.
Training SettingFullMissing-OMissing-SMean
Reference: ρ = 0.50 , either, CE0.77600.60490.76590.7156
ρ = 0.25 , either, CE0.74780.55090.68610.6616
ρ = 0.75 , either, CE0.77320.63250.76870.7248
ρ = 0.50 , fixed t 1 , CE0.68330.35160.66160.5655
ρ = 0.50 , fixed t 2 , CE0.75380.47950.75300.6621
ρ = 0.50 , either, CE + Dice0.74960.57160.74890.6900
Table 8. Comparison of MARC-Net design variants and external supervised baselines under the mixed-scene protocol. Each inference condition reports IoU, Precision (P), and Recall (R). Bold and underlined IoU values indicate the best and second-best results, respectively, in each inference condition.
Table 8. Comparison of MARC-Net design variants and external supervised baselines under the mixed-scene protocol. Each inference condition reports IoU, Precision (P), and Recall (R). Bold and underlined IoU values indicate the best and second-best results, respectively, in each inference condition.
MethodFullMissing-OMissing-S
IoUPRIoUPRIoUPR
Initial formulation0.77010.84650.89500.29880.48630.43670.75010.83920.8760
FC-EF [16]0.73320.88010.81460.50920.86740.55220.76510.87950.8547
FC-Siam-diff [16]0.75040.82890.88780.28690.90780.29550.73480.80520.8936
DTCDSCN [17]0.78120.86870.88580.45780.91110.47920.77350.85950.8854
ChangeFormerV6 [19]0.77440.85130.89560.46880.80240.53000.77720.86200.8876
BIT [18]0.75910.90570.82430.28900.92140.29630.37920.91490.3931
SwinSUNet [21]0.77610.87990.86810.67220.85210.76100.76590.87460.8604
ChangeMamba [22] 0.7434 0.8687 0.8375 0.3038 0.5071 0.4310 0.4468 0.7822 0.5103
MARC-Net w/o balanced objective0.77600.85060.89840.60490.78140.72810.76590.85210.8833
MARC-Net + consistency regularization0.77280.85980.88430.53970.72690.67700.77200.86220.8807
MARC-Net (proposed) 0.7826 0.8944 0.8622 0.7039 0.8789 0.7795 0.7780 0.8929 0.8581
Table 9. Additional heterogeneous change-detection comparison under Missing-O inference on the 184-sample mixed-scene test set. Bold formatting identifies the proposed method.
Table 9. Additional heterogeneous change-detection comparison under Missing-O inference on the 184-sample mixed-scene test set. Bold formatting identifies the proposed method.
MethodSupervisionAvailable Inference InputPRF1IoU
INLPG [60]Unsupervised S 1 O 2 or O 1 S 2 0.19440.09380.12650.0675
RIEM [61]Unsupervised S 1 O 2 or O 1 S 2 0.21400.12040.15410.0835
MARC-Net (proposed)SupervisedRemaining optical image + S 1 , S 2 0.87890.77950.82620.7039
Table 10. Proposed MARC-Net under controlled single-phase optical degradation on the 184-sample mixed-scene test set. Affected is the mean fraction of modified pixels; the optical availability flag remains 1.
Table 10. Proposed MARC-Net under controlled single-phase optical degradation on the 184-sample mixed-scene test set. Affected is the mean fraction of modified pixels; the optical availability flag remains 1.
ConditionAffected (%)PRF1IoU
Clean0.00.89440.86220.87800.7826
Cloud 20%27.90.89330.64390.74840.5980
Cloud 40%51.00.90000.47170.61900.4482
Cloud 60%71.10.90950.31090.46340.3016
Haze100.00.88740.56360.68940.5260
Shadow40.00.91050.75910.82790.7064
Radiometric100.00.94140.44580.60510.4338
Table 11. Full-scene sliding-window IoU of the mixed-scene-trained base-phase MARC-Net. Bold formatting indicates the average across the four scenes.
Table 11. Full-scene sliding-window IoU of the mixed-scene-trained base-phase MARC-Net. Bold formatting indicates the average across the four scenes.
SceneFull IoUMissing-O IoUMissing-S IoU
0010.75400.57970.7520
0020.80510.69800.7948
0030.69000.14100.6798
0040.63230.35530.6289
Average0.72040.44350.7139
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zhan, Y.; Feng, C.; Deng, X.; Wang, Z.; Yu, H.; Ma, J.; Wang, X.; Hu, F. MARC-Net: A Modality-Availability-Aware Robust Change Network for Missing-Optical Bi-Temporal Optical–SAR Change Detection of Reclaimed Cropland. Remote Sens. 2026, 18, 2960. https://doi.org/10.3390/rs18172960

AMA Style

Zhan Y, Feng C, Deng X, Wang Z, Yu H, Ma J, Wang X, Hu F. MARC-Net: A Modality-Availability-Aware Robust Change Network for Missing-Optical Bi-Temporal Optical–SAR Change Detection of Reclaimed Cropland. Remote Sensing. 2026; 18(17):2960. https://doi.org/10.3390/rs18172960

Chicago/Turabian Style

Zhan, Yuanzeng, Cunjun Feng, Xiaoyuan Deng, Zhiyi Wang, Hui Yu, Junjie Ma, Xingkun Wang, and Fengming Hu. 2026. "MARC-Net: A Modality-Availability-Aware Robust Change Network for Missing-Optical Bi-Temporal Optical–SAR Change Detection of Reclaimed Cropland" Remote Sensing 18, no. 17: 2960. https://doi.org/10.3390/rs18172960

APA Style

Zhan, Y., Feng, C., Deng, X., Wang, Z., Yu, H., Ma, J., Wang, X., & Hu, F. (2026). MARC-Net: A Modality-Availability-Aware Robust Change Network for Missing-Optical Bi-Temporal Optical–SAR Change Detection of Reclaimed Cropland. Remote Sensing, 18(17), 2960. https://doi.org/10.3390/rs18172960

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Article metric data becomes available approximately 24 hours after publication online.
Back to TopTop