Next Article in Journal
Evolution of Land Use Suitability and Adaptation Strategies of the Agro-Pastoral Transitional Zone in Northern China Under Multiple Climate Change Scenarios
Previous Article in Journal
Assessing Public Experience of Waterfronts and Its Coupling with Urban Vitality: Evidence from Shanghai, China
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Mapping Landslide-Affected Land Surfaces in Complex Mountainous Landscapes Using a Twin-Path Multi-Scale Deep Learning Network

School of Geological Engineering, Qinghai University, Xining 810016, China
*
Author to whom correspondence should be addressed.
Land 2026, 15(7), 1298; https://doi.org/10.3390/land15071298
Submission received: 26 May 2026 / Revised: 16 July 2026 / Accepted: 17 July 2026 / Published: 20 July 2026

Abstract

Accurate mapping of landslide-affected land surfaces from very-high-resolution optical imagery is essential for mountainous land monitoring and hazard-related land management, yet it remains difficult in complex terrain because landslide bodies are fragmented, elongated, shadowed, and easily confused with bare soil, terraces, roads, and erosion features. This study presents TM-Net as a task-oriented decoder-centric architecture for landslide-affected land surface mapping. The Inception Token Mixer is adopted from InceptionNeXt as an encoder-side component, whereas TM-Block is newly designed as a decoder-side twin-path refinement module for boundary recovery, slender target preservation, and complex-background suppression. On the fused public benchmark, TM-Net achieved the highest IoU (73.56%) and F1 (84.76%) under the identical evaluation protocol. On XBLD, TM-Net achieved the highest IoU (44.94%) and F1 (62.01%). These results indicate that TM-Net provides a favorable balance between missed detections and false-positive predictions for mapping visually identifiable landslide-affected land surfaces in heterogeneous mountainous landscapes.

1. Introduction

Landslides are not only geological hazards but also important forms of rapid land surface disturbance in mountainous regions. They alter slope morphology, damage vegetation and farmland, disrupt transportation corridors, and reshape land-use safety patterns. Therefore, accurate mapping of landslide-affected land surfaces is essential for land degradation monitoring, post-disaster land assessment, and sustainable land management in mountain–basin transition zones.
Landslides are geological phenomena in which rock and soil on slopes lose stability, fail, and move downslope. They are common in mountainous regions worldwide and represent persistent threats to human life and critical infrastructure [1,2,3,4,5,6,7]. Located in northeastern Qinghai Province, the Xining Basin is prone to landslides because its fault-basin setting, loess-covered hillslopes, unconsolidated near-surface deposits, concentrated summer rainfall, and intensive engineering disturbance jointly promote slope instability [8,9]. As the main urban center of the region, Xining City is especially affected: under the combined influence of intense rainstorms and intensive engineering activities, geological hazards often occur abruptly and impact large areas. In recent years, landslide-related events have highlighted the need for reliable hazard interpretation in the Xining region. On 15 September 2022, a destructive mudstone landslide in Xining damaged a bridge and affected the Lanzhou–Xinjiang high-speed railway [10]. This event underscores the need for timely and accurate landslide identification and reliable pixel-level delineation for geological hazard interpretation and risk screening in the Xining region, where digital technologies are increasingly used to support landslide disaster risk research and management [11,12,13,14,15,16,17].
Driven by the rapid expansion of very-high-resolution optical remote sensing imagery, deep learning (DL) has accelerated the transition of landslide identification from manual visual interpretation to automated semantic segmentation [18,19]. DL methods extract hierarchical features automatically through models such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and autoencoders. These approaches are well suited for large-scale landslide identification, removing the need for manual feature construction and enabling efficient analysis of complex environments. Among them, CNNs have emerged as one of the most widely adopted techniques for landslide detection [20,21,22,23,24]. CNN-based architectures (e.g., U-Net, DeepLabV3+, ResNet variants, and Mask R-CNN) exhibit a strong capability to capture spatial and textural information across diverse mountainous settings and achieve competitive performance on public benchmarks, including recent landslide segmentation datasets [25,26,27,28,29]. However, optical remote sensing-based methods remain sensitive to topographic shadows, vegetation cover, and spectral confusion with look-alike backgrounds, which can lead to false detections and missed landslides, particularly on steep and shaded slopes [30,31,32,33]. Beyond data-related limitations, current DL models also face architectural trade-offs. CNNs are effective at reconstructing boundaries and fine details but are limited in modeling long-range context, whereas Transformer-based encoders capture global dependencies but often require substantial memory and computational resources for very-high-resolution imagery [26,34]. As a result, designing efficient hybrid architectures that integrate CNN-style local refinement with attention-enabled global context modeling, while maintaining a favorable balance between accuracy and efficiency, remains a key challenge in landslide segmentation [35]. Recent multi-scale and attention-enhanced frameworks have also been explored for landslide recognition in complex terrain [36]. In landslide segmentation, the dominant errors are often concentrated at boundaries and slender runout zones. Landslide surfaces are frequently characterized by weak, heterogeneous textures and irregular gully-aligned shapes, and they are easily confused with look-alike backgrounds. As a result, Transformer-dominated encoders may produce over-smoothed or fragmented masks for small and elongated failures, whereas CNN-based models, although better at local boundary reconstruction, may still fail to capture long-range geomorphic cues required to suppress false positives. Therefore, a decoder-centric design guided by global context is considered necessary for practical mapping in very-high-resolution scenes.
To address these issues, this study proposes TM-Net (Twin-path Multi-scale Network) for very-high-resolution landslide segmentation in complex mountainous basins. TM-Net integrates an adopted encoder-side Inception Token Mixer with a newly designed decoder-side TM-Block. The contribution of this study lies in the task-oriented twin-path refinement mechanism for landslide boundary recovery and target-background discrimination. Specifically, TM-Block couples parallel channel calibration and multi-scale spatial refinement through residual feature modulation to reduce fragmented boundaries, missed narrow runout zones, and visually confusing background responses.
For experimental validation, the fused public benchmark is constructed with spatial isolation and is used together with XBLD as an independent regional benchmark for loess-area landslides. The main contributions of this work are summarized as follows: (1) TM-Net is proposed with TM-Block as a newly designed decoder-side twin-path multi-scale refinement module for landslide boundary recovery and background suppression; (2) a spatially isolated public fused benchmark is constructed, and XBLD is used as an independent Xining loess-region validation dataset; and (3) eight-model comparisons, ablation experiments, qualitative visualizations, and efficiency comparisons show that TM-Net provides an advantageous overall segmentation balance reflected by IoU and F1.

2. Materials and Methods

2.1. Public Landslide Datasets

All datasets are provided at their original ground sampling distance (GSD), that is, their native spatial resolution. To evaluate robustness and in-domain performance, three public landslide datasets (Figure 1) are fused into a unified benchmark for comparative experiments. The fused public benchmark retains the original annotations of the constituent datasets. Because landslide types, image sources, regional backgrounds, and labeling philosophies differ across Bijie, GVLM, and SCLM, this benchmark is intended to evaluate robustness under heterogeneous cross-region conditions rather than to represent a fully harmonized geological inventory standard. XBLD, by contrast, is an independent regional dataset constructed under the unified interpretation rules used in this study. Accordingly, the XBLD target definition described below is not retroactively imposed on the public datasets; their original label semantics are retained for cross-region robustness evaluation.
The Bijie subset (released by Wuhan University) covers Bijie City, Guizhou Province, a mountainous and hilly transition zone with steep relief and extensive mosaics of cropland and bare soil [37]. The region is characterized by complex mountain river systems and a fragile ecological environment, making it highly prone to rainfall-induced landslides [17]. This subset is built from 0.8 m TripleSat RGB imagery acquired between May and August 2018; landslides are typically small and elongated, with boundaries blurred by vegetation and cultivated land, thus stressing thin object recovery and boundary fidelity. The GVLM subset is a global very-high-resolution landslide mapping dataset interpreted from 0.59 m Google Maps imagery, covering 17 subregions worldwide across varied climates and landforms [38]. It contains mixed rainfall-induced and earthquake-induced landslides as well as abundant landslide-like backgrounds such as cast shadows, talus, scree, and terraced fields, which pose a strong challenge to the cross-region robustness of segmentation models. The SCLM subset was constructed from high-resolution digital orthophotos and manually interpreted labels of landslide and debris-flow disasters in Sichuan Province and adjacent areas acquired between 2008 and 2020. The source imagery has a spatial resolution of approximately 0.2–0.9 m and covers diverse geomorphological and disaster settings. Because the dataset includes both landslide and debris-flow samples, it adds morphological variability and complex background conditions to the fused benchmark [39,40].

2.2. Study Area and Xining Basin Landslide Dataset

The Xining Basin is situated on the northeastern margin of the Qinghai–Tibet Plateau, at the transition to the Loess Plateau, and represents a complex mountain–basin environment [8,9]. It covers about 490 km2 between 101°33′45″–101°56′15″ E and 36°25′05″–36°47′30″ N, and Xining, the capital of Qinghai Province, lies within this high-altitude basin at a mean elevation of about 2261 m. The regional topography exhibits strong relief, with elevations generally decreasing from southwest to northeast. The basin lies within a faulted Mesozoic–Cenozoic depression of the Qilian Mountain fold system, and its upper deposits include sand, gravel, and thick loess. The loose, porous loess and steep hillslopes are sensitive to rainfall infiltration and human disturbance, creating conditions favorable for shallow landslides and debris flows [8,9,41]. The local climate is continental semi-arid, with long winters and short summers; a large proportion of annual rainfall occurs between June and September [8,41]. Xining has a spatially concentrated population distribution [42], while cut slopes, fills, terraces, and other disturbed surfaces form complex look-alike backgrounds in the optical imagery used for XBLD. The study area and illustrative examples, including representative image–label pairs and field photographs acquired by an unmanned aerial vehicle, are shown in Figure 2.
Within the above study area, we constructed XBLD from Gaofen-2 (GF-2) satellite imagery with a ground sampling distance of 0.8 m. GF-2 scenes acquired in 2022 and 2023 with low cloud cover were selected over landslide-prone zones, and landslide boundaries were manually delineated through visual interpretation to produce pixel-level labels.
For XBLD, the term landslide-affected land surface refers to the optically discernible ground surface altered by a discrete landslide at the time of image acquisition. It includes identifiable source, displaced-body, and runout or accumulation components when they can be interpreted as a spatially continuous landslide feature. The label does not represent the subsurface slip plane and is not equivalent to all bare land, disturbed ground, or engineering-exposed surfaces. Ordinary bare soil, farmland patches, road cuts, engineering slopes, terraces, and erosion gullies were treated as background unless morphology, slope position, texture, and spatial continuity jointly supported their interpretation as landslide foreground.
The inclusion criteria required that a landslide source area, main body, or transport/accumulation zone could be identified in the imagery; that the candidate had a reasonable geomorphic relation to slope position, gullies, road cuts, or disturbance boundaries; that continuous or semi-continuous landslide morphology, texture, or boundary evidence was present; and that the feature could be reliably separated from neighboring backgrounds. Excluded cases included highly ambiguous boundaries, erosion gullies or bare-soil patches lacking landslide morphology and spatial continuity, large composite failures extending beyond the patch context, and surface disturbances of uncertain origin. The labels were produced by manual visual interpretation based on GF-2 imagery, UAV photographs, and Google Earth 3D imagery to assess landslide morphology, slope position, boundary continuity, and geomorphic context. The main characteristics of the four datasets used in this study are summarized in Table 1. The statistical properties of the 898 cropped landslide instances, including area, aspect ratio, and foreground pixel ratio, are shown in Figure 3.

2.3. Preprocessing and Spatially Isolated Splits

For all datasets used in this study, a unified preprocessing pipeline is applied prior to model training. Images and labels are represented as 224 × 224 pixel three-channel RGB patches and binary labels, respectively, with landslide pixels encoded as 255 and background pixels as 0 (threshold value 127). Channel-wise z-score normalization statistics (mean and standard deviation) are estimated on the training split and applied to all images. During training, lightweight, label-consistent data augmentations are adopted, including horizontal and vertical flips, small rotations within the image plane (up to ±10°), and scale jittering in the range 0.7–1.4, whereas validation and test inference are performed without random augmentation, using a deterministic preprocessing pipeline [43]. For the three public datasets (Bijie, GVLM, and SCLM), the original scenes and labels are converted into this common patch format by padding and center cropping or by sliding window tiling and resampling. For each patch, the proportion of landslide pixels is computed, and only patches whose landslide ratio exceeds a small threshold are retained as positive samples, while background-dominated patches are kept as negatives to alleviate class imbalance. After preprocessing, geographic blocking is first performed at the scene level to ensure spatial isolation; the blocked units are then assigned to training and validation subsets to approximate an 8:2 ratio, and this fused public split is used for training and evaluating all baseline models and the proposed TM-Net.
For the self-constructed XBLD derived from GF-2 imagery, spatial isolation is performed at the original GF-2 scene level rather than by randomly splitting cropped patches. Five cloud-free GF-2 scenes covering landslide-prone areas of Xining are cropped into 224 × 224 patches based on the manually delineated landslide labels. Four scenes (Groups A, B, C, and E) are used to generate the training set, whereas the remaining scene (Group D) is reserved for validation, so the training and validation subsets do not share original scenes or spatially overlapping source imagery. Within each split, landslide-free patches are sampled from non-failure areas according to the original experimental setting, and all comparison models use the same XBLD split and evaluation protocol.

2.4. TM-Net Architecture

The proposed TM-Net is a decoder-centric hybrid architecture designed for landslide segmentation in very-high-resolution optical imagery. As shown in Figure 4, the network follows an encoder–decoder paradigm, where a hierarchical Transformer-based encoder with an adopted Inception Token Mixer [44] extracts multi-scale semantic features and a lightweight bottleneck bridges the encoder and decoder. In the decoding stage, multi-level features are progressively upsampled and fused with skip connections from the encoder, and each stage employs the newly designed TM-Block to support channel-wise recalibration and multi-scale spatial refinement. This design is intended to reduce landslide-like background interference and support recovery of thin, elongated landslide bodies.
This decoder-centric design is motivated by the observation that major segmentation errors in landslide mapping are often not caused by complete target absence but by boundary fragmentation, omission of narrow runout zones, and leakage into visually similar backgrounds.
On the encoder side, the Inception Token Mixer is adopted from InceptionNeXt at the input stage to replace the conventional single convolution operation. Its role is early low-cost anisotropic and multi-scale feature extraction, which helps represent elongated and gully-aligned landslide components before deep semantic abstraction. It is not claimed as an original module of this paper.
On the decoder side, TM-Block is the newly designed core module. TM-Block adopts a decoder-specific twin-path structure: a channel-calibration path performs adaptive global response recalibration, whereas a multi-scale spatial-refinement path restores local boundaries and texture details. The two paths are coupled through residual feature refinement to facilitate boundary continuity and reduce look-alike background responses.

2.5. Inception Token Mixer

The input feature map X is split channel-wise into four branches: X   =   X id ,   X hw ,   X w ,   X h In this study, the Inception Token Mixer at the encoder input is implemented using the IncepDWConv module proposed in InceptionNeXt, also termed IncStem. We use this module as an adopted lightweight stem to inject directional and mixed-scale priors at the earliest feature extraction stage. Its contribution in TM-Net is functional and task-oriented.
Y = Concat(Xid,DWConv3×3(Xhw)DWConv1×11(Xw),DWConv11×1(Xh)
where Y denotes the output feature map, Concat(·) denotes channel-wise concatenation, and DWConv(·) represents depthwise convolution with the corresponding kernel size.
By factorizing a large receptive field into several lightweight depthwise branches, a rich set of multi-scale, anisotropic local priors is provided at modest cost, which is particularly beneficial for capturing slender landslide bodies and narrow runout traces in the complex terrain of Xining City and for laying a solid foundation for subsequent hierarchical feature extraction. The Token Mixer block is shown in Figure 5.

2.6. TM-Block

In the decoding stage, TM-Block is employed to support recovery of fine-grained boundaries and texture details that are typically degraded in deep networks. TM-Block is a newly designed decoder-side Twin-path Multi-scale Block. Its channel-calibration path emphasizes globally relevant feature channels, while its multi-scale spatial-refinement path uses grouped multi-scale perception and channel interaction to sharpen local landslide boundaries and suppress confusing backgrounds. The overall architecture and processing workflow of TM-Block are illustrated in Figure 6.
In the channel attention branch, channel responses are recalibrated under global contextual guidance. Local responses are first extracted by a lightweight convolutional transformation, while global statistics are simultaneously captured by global average pooling (GAP) followed by a 1 × 1 convolution. The two streams are fused and passed through a bottleneck multilayer perceptron (MLP) to produce the channel attention map AC, which is defined as
Ac = σM(Conv1×1(X) + Conv1×1(GAP(X))
where σ(·) denotes the Sigmoid activation function and M(·) represents the bottleneck MLP structure, and AC denotes the channel attention map. This branch emphasizes landslide-related feature channels while suppressing less informative responses from complex backgrounds.
In the multi-scale spatial attention branch, which forms the spatial-refinement component of the proposed TM-Block, a “Compress–Split–Perceive” strategy is adopted. The input feature map X is first compressed and smoothed by a 7 × 7 convolution, generating an intermediate representation Z. Then, Z is split along the channel dimension into four subgroups, denoted as Z 1 ,   Z 2 ,   Z 3   and   Z 4 . These subgroups are processed in parallel to capture structures at different spatial scales:
Ẑ = Concat(Conv1×1(Z1),Conv3×3(Z2),Conv5×5(Z3),Z4)
where Ẑ denotes the concatenated multi-scale spatial representation. The 1 × 1, 3 × 3, and 5 × 5 convolution branches capture spatial responses with different receptive fields, while the identity branch Z4 preserves part of the original compressed representation.
To facilitate information exchange across the four branches, a Channel Shuffle operation S(·) is applied to Ẑ. The shuffled representation is then projected by a 1 × 1 convolution and normalized by the Sigmoid function to generate the spatial attention map AS:
As = σ(Conv1×1(S(Ẑ)))
Finally, the refined output Y is obtained by modulating the input X with both AC and AS, followed by a residual addition:
Y = X + X ⨀ Ac ⨀ As
where ⨀ denotes element-wise multiplication. Through this twin-path multi-scale design, TM-Block is intended to enhance landslide-relevant channel responses, refine local boundaries, and reduce background noise. This supports decoder-stage feature reconstruction for small, slender, and morphologically complex landslides in very-high-resolution scenes.

2.7. Hierarchical Feature Refinement Pipeline

To effectively integrate the proposed modules into the network, a structured feature refinement pipeline is employed within each decoder stage. At the beginning of each stage, the upsampled deep features are fused with the corresponding skip-connected encoder features through a lateral fusion block. The fused feature map is then processed by a TM-Block, which recalibrates channel responses and restores fine-grained spatial details at an early stage of decoding. Subsequently, a cross-attention module interacts with the global class token to inject long-range semantic guidance and suppress background noise. Finally, a Swin Transformer block further abstracts the refined representation and prepares it for propagation to the next stage. With this serial design, where local refinement is followed by global guidance, the features passed to subsequent stages are endowed with both sharp boundary delineation and consistent semantic structure, which is crucial for accurate segmentation of small and morphologically complex landslides.

2.8. Experimental Settings and Evaluation Metrics

TM-Net was implemented using the PyTorch 2.1 framework and trained on a single NVIDIA RTX 4090 GPU. For TM-Net, the AdamW optimizer was used with an initial learning rate of 1.0 × 10−3 and a weight decay of 0.05, together with cosine annealing after a 20-epoch warm-up. TM-Net was trained for 100 epochs with a batch size of 32. The compared models were trained using their respective implementation settings, including model-specific batch sizes. All models were evaluated using the same spatially isolated data splits and the same cumulative pixel-level metric calculation. The checkpoint with the best validation IoU was used for final evaluation. The training configuration of TM-Net is summarized in Table 2.
In this study, Intersection over Union (IoU), Precision (Pre), Recall (Rec), and F1 score were used to evaluate landslide segmentation performance. All metrics were calculated from the cumulative numbers of true positive (TP), false positive (FP), and false negative (FN) pixels over the complete validation set. This unified evaluation protocol ensures that F1 is mathematically consistent with both IoU and Precision/Recall and avoids inconsistencies caused by averaging rounded image-level values.
IoU = TP TP + FP + FN
F 1 = 2   ×   Pre   ×   Rec Pre   +   Rec
Pre = TP TP   +   FP
Rec = TP TP   +   FN
where TP denotes landslide pixels correctly predicted as landslides, FP denotes background pixels incorrectly predicted as landslides, and FN denotes landslide pixels incorrectly predicted as background. IoU measures the overlap between the predicted and reference landslide regions. Precision reflects the ability of the model to suppress false-positive predictions, Recall indicates the ability to recover true landslide pixels, and F1 provides a balanced assessment of Precision and Recall.

3. Results

3.1. Comparison with Representative Advanced Models

On the fused public benchmark, TM-Net achieved the highest IoU (73.56%) and F1 (84.76%) under the identical evaluation protocol (Table 3). The improvement over the strongest baseline was moderate but consistent: compared with SegFormer, TM-Net improved IoU by 0.89 percentage points and F1 by 0.59 percentage points. TM-Net also achieved the highest Pre (82.47%), whereas SegFormer produced the highest Rec (88.29%), indicating that TM-Net did not rank first in every individual metric.
These results support a restrained interpretation. TM-Net provided a more favorable overall balance between missed landslide pixels and false-positive predictions rather than comprehensively outperforming every model on every metric. The adopted Inception Token Mixer supports recovery of elongated targets, while the newly designed TM-Block strengthens decoder-side boundary refinement and confusing-background suppression.
All metrics were calculated from cumulative pixel statistics over the complete validation set. The relatively small numerical margin between TM-Net and the strongest baseline is also informative. On a fused benchmark containing different landslide types, image sources, spatial resolutions, and background conditions, large performance gaps are not necessarily expected. Instead, the results show that several models can achieve competitive performance while exhibiting different prediction tendencies. For example, a higher Recall may reflect more complete recovery of candidate landslide pixels, but it can also be accompanied by increased false-positive responses in visually similar terrain. Conversely, higher Precision may indicate more conservative suppression of background interference, while potentially excluding weak or fragmented target pixels. The joint consideration of IoU, F1, Precision, and Recall is therefore necessary for interpreting segmentation behavior in complex mountainous imagery.

3.1.1. Comparative Analysis of Segmentation Results

To qualitatively compare TM-Net with other state-of-the-art models, five representative validation patches are examined in Figure 7, where red, green, and blue denote TP, FP, and FN pixels, respectively. In the selected examples, TM-Net generally produced more continuous landslide masks with fewer fragmented boundaries, but these examples should be interpreted as qualitative support rather than proof of universal superiority.
The main visual differences among models can be summarized by three recurring error patterns: boundary shrinkage, fragmentation of narrow targets, and leakage into visually similar backgrounds. TM-Net was associated with fewer fragmented responses in several slender or low-contrast cases, while some baselines showed local omissions or false positives near forest edges, roads, bare soil, and terrace-like surfaces.
These qualitative observations are consistent with the quantitative comparison in Table 3. They suggest that TM-Net improves the balance between target recovery and false-positive suppression in the selected validation examples, while remaining subject to boundary uncertainty and background confusion in challenging scenes. In particular, the selected visual examples illustrate why a single metric cannot fully describe the mapping difficulty. For elongated or gully-aligned targets, a model may correctly identify the central part of a landslide while omitting the narrow terminal or lateral portions. Such errors can reduce the spatial continuity of the mapped surface even when the main body is recognized. In other cases, background elements with exposed soil, linear road margins, terrace boundaries, or shadow transitions may generate isolated false-positive patches. The qualitative comparison therefore complements Table 3 by showing how omission and commission errors are distributed spatially. It also highlights that the remaining uncertainty is concentrated mainly along weak boundaries and in locations where the optical appearance of landslide-affected surfaces overlaps with surrounding disturbed terrain.

3.1.2. Interpretability Analysis via Grad-CAM Visualization

Grad-CAM visualizations (Figure 8) are used as qualitative supporting evidence for how different models concentrate responses over representative validation scenes. TM-Net generally produced more compact responses over mapped landslide bodies in the selected examples, whereas several baselines showed more diffuse, fragmented, or displaced activation under complex textures.
These Grad-CAM patterns are consistent with the segmentation outputs and the numerical comparison, but they should not be interpreted as strict proof of the internal mechanism or causal behavior of the network. The activation maps should also be interpreted in relation to the heterogeneous appearance of the mapped targets. Landslide-affected surfaces may include exposed source areas, displaced materials, and runout components with different textures and illumination conditions within the same image patch. Consequently, a useful response pattern is not necessarily one that concentrates only on the visually brightest or most homogeneous part of a target. In the selected examples, TM-Net tended to maintain attention over a broader portion of the mapped landslide body while reducing diffuse responses in adjacent terrain. This observation is consistent with the more continuous masks shown in Figure 7, although the activation maps are not used as a quantitative measure of model reliability.

3.2. Ablation Study

To validate the effectiveness of the adopted Inception Token Mixer and the proposed TM-Block, a progressive ablation study was conducted on the fused public benchmark using the Baseline network. Three configurations were evaluated: adding TM-Block to the decoder, replacing the original stem with the Inception Token Mixer, and integrating both modules to form TM-Net. The results are summarized in Table 4.
Compared with the Baseline (IoU = 70.72%), adding TM-Block alone increased IoU to 73.06%, which is the largest single-module IoU improvement among the ablation variants. The TM-Block variant also achieved the highest Pre (84.15%), supporting its role in boundary discrimination and false-positive suppression.
Replacing the encoder stem with the Inception Token Mixer increased IoU to 72.16% and produced the highest single-module Rec (87.21%). This suggests that the adopted early anisotropic mixed-scale stem contributes more to target completeness and recovery of slender or low-contrast landslide regions.
When both modules were combined, TM-Net achieved the highest IoU (73.56%) and F1 (84.76%). These results support the interpretation that TM-Block primarily improves boundary discrimination and false-positive suppression, whereas the Inception Token Mixer contributes to target completeness. Their integration yields the best balance between omission and commission errors.

3.2.1. Qualitative Analysis of Ablation Study Results

TP-FP-FN visualizations on representative validation patches are provided in Figure 9 to qualitatively compare TM-Net with its ablation variants. The Baseline decoder can roughly localize landslide bodies, but the predictions tend to show boundary instability, fragmented narrow targets, and scattered false positives on visually similar backgrounds.
The Inception Token Mixer variant was associated with more complete recovery of some slender or low-contrast targets, consistent with its higher Rec. The TM-Block variant showed stronger suppression of confusing backgrounds in several examples, consistent with its higher Pre. These qualitative patterns should be viewed as supporting evidence for the numerical ablation results rather than as exhaustive proof across all scenes.
With both modules integrated, TM-Net showed the most favorable qualitative balance among the ablation variants in the selected examples. The masks were generally more continuous, boundary errors were more localized, and false positives were better constrained than in the Baseline or single-module variants. The qualitative ablation comparison also helps explain why the full configuration does not simply inherit the strongest individual metric from either single-module variant. The Inception Token Mixer variant produced higher Recall, suggesting a tendency to retain more weak or elongated target components, whereas the TM-Block variant produced higher Precision, indicating stronger control of background responses. When both components were used together, the resulting masks generally preserved the main landslide extent while limiting part of the over-expansion observed in difficult backgrounds. This pattern is consistent with the higher IoU and F1 of the complete model. However, residual errors remained in areas characterized by low contrast, incomplete optical exposure, or ambiguous boundaries, indicating that the combination reduces but does not eliminate the inherent uncertainty of optical landslide delineation.

3.2.2. Interpretability Analysis of the Ablation Study via Grad-CAM

To further interpret the contributions of the key architectural components, Grad-CAM visualizations are generated for four representative validation scenarios (Figure 10A–D). From top to bottom, the four rows correspond to Figure 10A, a large irregular landslide extending to the image boundary; Figure 10B, a long, narrow landslide body in forested terrain; Figure 10C, multiple closely spaced elongated landslide bodies in complex terrain; and Figure 10D, a small isolated landslide adjacent to a road. Collectively, these cases cover diverse landslide geometries and typical background interferences.
The Grad-CAM visualizations in Figure 10 provide qualitative support for the complementary roles suggested by Table 4. The Inception Token Mixer variant tended to retain broader responses over landslide targets, whereas the TM-Block variant was associated with more localized background suppression. The complete TM-Net combined these tendencies, but the visualizations are interpreted as supporting evidence rather than strict mechanistic proof.

3.3. Independent Regional Robustness Evaluation on XBLD

Under the same XBLD split and evaluation protocol, TM-Net achieved the highest IoU (44.94%) and F1 (62.01%), indicating the best overall segmentation balance on this dataset. TM-Net was not the best in every individual XBLD metric. DeepLabV3+ produced a slightly higher Pre (57.75%), while several models produced higher Rec, including Swin-UMamba (72.46%). These results show that TM-Net’s advantage is the overall balance between landslide recovery and false-positive control rather than dominance in every single metric.
The lower absolute metric values on XBLD relative to the fused public benchmark indicate that the independent regional task is substantially more difficult. This difference likely reflects the combined effects of regional domain shift, scene-level spatial isolation, and visually similar loess-slope backgrounds. XBLD contains small and fragmented landslide-affected surfaces embedded within loess slopes, cultivated land, terraces, road cuts, erosion features, and other disturbed surfaces. These elements often have partially overlapping spectral and textural characteristics, particularly where vegetation cover is sparse or where engineering disturbance has modified the slope surface. The results therefore should not be interpreted only as a ranking of models but also as evidence of the difficulty of applying optical segmentation methods to a realistic urban–rural fringe environment. Although TM-Net achieved the highest IoU and F1, the moderate values across all methods show that substantial uncertainty remains in distinguishing landslides from look-alike backgrounds.

3.3.1. Qualitative Comparison on XBLD

To qualitatively assess independent regional robustness on XBLD under complex real-world conditions, three representative samples are compared in Figure 11A–C. Figure 11A shows a small landslide near the upper image boundary, Figure 11B shows a partially visible landslide near the lower-left image boundary, and Figure 11C contains multiple spatially separated landslide bodies of different sizes.
Across the representative XBLD cases, TM-Net generally produced compact and spatially coherent responses while avoiding some extended false-positive patterns observed in several baselines. These observations are consistent with the XBLD results, where TM-Net achieved the highest IoU (44.94%) and F1 (62.01%), supporting a better balance between target recovery and background suppression under complex loess terrain. The three representative scenarios also reflect different sources of difficulty in the regional dataset. Isolated small landslides are vulnerable to omission because only a limited number of pixels represent the target. In scenes affected by strong texture interference, patches of exposed loess, cultivated surfaces, and erosion-related features can resemble disturbed landslide materials. Compound scenes introduce an additional challenge because adjacent failures and background disturbances may be spatially close while not belonging to the same landslide unit. In these situations, the visual objective is not only to recover as many target pixels as possible but also to preserve plausible object extent and avoid merging unrelated disturbed areas. The selected examples suggest that TM-Net can provide relatively coherent responses in such cases, although false positives and local omissions remain visible in difficult parts of the scenes.

3.3.2. Qualitative Attention Visualization on XBLD

Figure 12 shows Grad-CAM visualizations for three representative samples from XBLD, comparing TM-Net with mainstream baselines. These maps are used only as qualitative supporting evidence. Across the examples, TM-Net tended to produce compact responses over mapped landslide-affected surfaces, whereas several baselines showed more diffuse, fragmented, or displaced attention under look-alike loess backgrounds.
Overall, the Grad-CAM comparisons on XBLD are consistent with the segmentation outputs and Table 5 results. They suggest more coherent responses over landslide-affected surfaces in selected examples, but they do not constitute strict proof of the network’s internal mechanism.

3.4. Computational Efficiency Analysis

Table 6 summarizes the computational complexity and inference throughput of the compared models. TM-Net is not the smallest or highest-throughput model; instead, it provides a competitive accuracy–efficiency trade-off by combining the highest IoU and F1 in Table 3 with moderate parameter count, FLOPs, and FPS. The complexity results should be interpreted together with the intended mapping setting. TM-Net contains 35.85 M parameters and requires 4.18 G FLOPs, placing it between very lightweight architectures and larger convolutional models in terms of computational demand. Its inference speed of 71.97 FPS is lower than that of the fastest compared models but remains suitable for patch-based batch inference on a modern GPU. The table also shows that low parameter count does not necessarily correspond to low FLOPs or high throughput, because runtime behavior depends on feature resolution, decoder operations, attention mechanisms, and implementation details. For practical use, the relevant consideration is therefore not whether a model is universally fastest but whether its computational cost is acceptable relative to the quality and spatial coherence of the segmentation output. Under this perspective, TM-Net provides a usable compromise for high-resolution landslide mapping workflows that require both detailed boundary delineation and repeated processing of image patches.

4. Discussion

4.1. Mechanism: Synergy of Anisotropy and Refinement

A persistent difficulty in optical landslide segmentation is the simultaneous presence of strong background interference and fine boundary ambiguity. TM-Net addresses this problem through a task-oriented decoder-centric architecture. The adopted Inception Token Mixer introduces early anisotropic mixed-scale representation for elongated and gully-controlled targets, while the newly designed TM-Block performs decoder-stage residual refinement by coupling global channel recalibration with local multi-scale spatial modulation. This design is intended to improve landslide boundary recovery and target–background discrimination, especially where bare soil, terraces, roads, erosion gullies, shadows, and engineering slopes resemble landslide-affected surfaces. The ablation results support different but complementary roles: TM-Block contributes more to boundary discrimination and false-positive suppression, whereas the Inception Token Mixer contributes to target completeness. Grad-CAM is used only as qualitative supporting evidence for these observations. The observed behavior can also be understood from the spatial organization of landslide-affected surfaces in optical imagery. A landslide rarely appears as a uniform object with a single stable texture. Source areas may be brighter because of exposed soil, displaced material may show rougher texture, and runout zones may be partially covered by vegetation or affected by cast shadows. At the same time, adjacent non-landslide surfaces can contain similar exposed materials, linear boundaries, and local roughness. This spatial heterogeneity explains why coarse semantic representations alone may be insufficient for accurate delineation. The results suggest that maintaining local detail during decoding is particularly important where the mapping decision depends on weak boundary cues rather than on a strong spectral contrast. Nevertheless, optical evidence remains incomplete in some settings, especially where the original slope morphology is obscured, the image is shadowed, or the affected surface has already undergone partial recovery.

4.2. Implications for Mountainous Land Monitoring and Hazard-Related Land Management

The comparative results in Table 3 indicate that TM-Net achieved the highest IoU (73.56%) and F1 (84.76%) on the fused public benchmark, while maintaining competitive but not highest Recall. This pattern suggests that TM-Net provides a useful balance between preserving subtle landslide pixels and suppressing false positives. For mountainous land monitoring, such behavior is valuable because missed landslide-affected pixels can lead to incomplete post-disaster land damage inventories, fragmented candidate zones, or truncated runout boundaries.
In practical land-management workflows, TM-Net can support candidate-area extraction, preliminary mapping, boundary drafting, and decision support before expert review or field verification. The XBLD experiment reflects a realistic loess-region setting where landslides are embedded within complex backgrounds such as terraces, roads, erosion features, bare soil, and shadows. TM-Net should therefore be regarded as a deep-learning-assisted tool for improving the mapping efficiency of visually interpretable landslide-affected surfaces, not as a replacement for full geological interpretation, field investigation, or expert judgment. A practical application workflow would involve first applying the model to produce a preliminary probability map or binary candidate map over areas of interest, followed by visual review of locations with uncertain boundaries or strong background interference. The resulting products could help prioritize expert interpretation, identify areas requiring field verification, and provide an initial spatial framework for post-event land-surface inventorying. This is particularly relevant in mountain–basin transition zones, where the number of potentially unstable slopes may be large and manual interpretation of very-high-resolution imagery is time-consuming. The model output may also be useful for comparing mapped disturbance patterns between image dates when consistent data are available, although such comparisons should be interpreted cautiously because apparent changes can be influenced by illumination, vegetation, image quality, and seasonal surface conditions. In all cases, the segmentation result should be combined with terrain context, geological knowledge, and independent observations before it is used for engineering or hazard-management decisions.

4.3. Limitations and Future Work

Although TM-Net showed the best overall IoU and F1 balance across the public and XBLD evaluations, several limitations remain. First, the performance gap between the fused public benchmark and XBLD indicates that domain shift and look-alike backgrounds remain substantial. In loess-dominated geomorphology, erosion gullies, terrace edges, engineering cut slopes, and naturally exposed bare soil can present spectral and textural signatures similar to landslide scars. Because TM-Net primarily relies on mono-temporal RGB appearance cues, these ambiguous surfaces can still generate false positives when explicit geomorphic or temporal constraints are unavailable. Figure 13 illustrates a representative XBLD case of spectral confusion between landslide-affected surfaces and erosion features.
The fused public benchmark and XBLD also play different experimental roles. The public benchmark retains the original annotations of Bijie, GVLM, and SCLM, so its label semantics are heterogeneous across regions, image sources, landslide types, and publication-specific annotation practices. It is therefore used to evaluate robustness under heterogeneous cross-region conditions, not as a fully harmonized geological inventory standard. XBLD, by contrast, represents landslide-affected land surfaces that can be reliably interpreted from high-resolution imagery under the unified rules used in this study.
XBLD intentionally retains small, fragmented, shadow-affected, and urban–rural transition cases, so it is not a simple dataset. However, extremely ambiguous objects for which landslides could not be reliably distinguished from engineering disturbance or bare soil were excluded. The labels were generated through manual visual interpretation based on GF-2 imagery, UAV photographs, and Google Earth 3D imagery; however, a formal inter-interpreter consistency assessment was not conducted. This choice improves supervised label reliability but may underestimate the true recognition difficulty under the most ambiguous real-world conditions. Future work should incorporate inter-interpreter consistency assessment, more ambiguous boundary samples, multi-temporal optical imagery, DEM-derived terrain attributes, InSAR deformation products, and systematic field verification.
Second, the current framework is optimized for single-date optical segmentation. As a result, it cannot directly characterize landslide evolution, and it may fail to capture subtle, slow-moving, or dormant instabilities that lack clear optical expression. Future work will integrate multimodal information, such as DEM-derived terrain attributes and InSAR deformation products, to provide geometric and kinematic constraints for separating spectrally similar classes. The model will also be extended to time-series settings to enable change-aware mapping and life-cycle monitoring of landslide activity.
Third, the current model was evaluated mainly on small- and medium-sized landslides at a fixed spatial resolution. Its transferability to substantially different resolutions, large composite landslides, and extremely small-sample settings remains to be tested. Future studies should examine multi-resolution transfer learning, introduce object- or scene-level constraints for large landslide complexes, and test few-shot or semi-supervised learning strategies for regions with limited labeled data.

5. Conclusions

In this study, we proposed TM-Net for landslide-affected land surface mapping and segmentation from very-high-resolution optical imagery in complex mountainous landscapes. TM-Net combines an adopted encoder-side Inception Token Mixer with a newly designed decoder-side TM-Block. The two components are complementary rather than redundant: the Inception Token Mixer supports early anisotropic mixed-scale representation, while TM-Block improves decoder-side boundary refinement and confusing-background suppression. The results show that TM-Net achieved the highest IoU and F1 on both the fused public benchmark (73.56% and 84.76%) and XBLD (44.94% and 62.01%), indicating the best overall balance reflected by these metrics. TM-Net should be regarded as a deep learning-assisted mapping tool for visually identifiable landslide-affected surfaces, not as a complete replacement for expert geological interpretation or field verification.

Author Contributions

Conceptualization, H.Y. and W.L.; methodology, H.Y.; software, H.Y.; validation, H.Y.; formal analysis, H.Y.; investigation, H.Y.; resources, Y.L.; data curation, H.Y.; writing—original draft preparation, H.Y.; writing—review and editing, W.L.; visualization, H.Y.; supervision, W.L.; project administration, Y.L.; funding acquisition, Y.L. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the Science and Technology Program of Qinghai Province of China, Grant Number 2024-SF-129.

Data Availability Statement

Due to licensing restrictions associated with commercial satellite imagery, the data are not publicly available. Researchers interested in accessing the data may contact the corresponding author upon reasonable request. The code used in this study is available from the corresponding author upon reasonable request.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Brabb, E.E.; Harrod, B.L. (Eds.) Landslides: Extent and Economic Significance. In Proceedings of the 28th International Geological Congress, Symposium on Landslides, Washington, DC, USA, 17 July 1989; A.A. Balkema: Rotterdam, The Netherlands, 1989. [Google Scholar]
  2. Salvati, P.; Petrucci, O.; Rossi, M.; Bianchi, C.; Pasqua, A.A.; Guzzetti, F. Gender, age and circumstances analysis of flood and landslide fatalities in Italy. Sci. Total Environ. 2018, 610–611, 867–879. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Bahadır, Ü.; Hacıefendioğlu, K.; Kartal, M.E.; Toğan, V.; Varol, N. Automatic landslide detection and visualization by using deep ensemble learning method. Neural Comput. Appl. 2024, 36, 10761–10776. [Google Scholar] [CrossRef] [Scilit]
  4. Guzzetti, F.; Mondini, A.C.; Cardinali, M.; Fiorucci, F.; Santangelo, M.; Chang, K.-T. Landslide inventory maps: New tools for an old problem. Earth-Sci. Rev. 2012, 112, 42–66. [Google Scholar] [CrossRef] [Scilit]
  5. Koks, E.E.; Rozenberg, J.; Zorn, C.; Tariverdi, M.; Vousdoukas, M.; Fraser, S.A.; Hall, J.W.; Hallegatte, S. A global multi-hazard risk analysis of road and railway infrastructure assets. Nat. Commun. 2019, 10, 2677. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Kirschbaum, D.B.; Adler, R.; Hong, Y.; Hill, S.; Lerner-Lam, A. A global landslide catalog for hazard applications: Method, results, and limitations. Nat. Hazards 2010, 52, 561–575. [Google Scholar]
  7. Dikau, R.; Cavallin, A.; Jäger, S. Databases and GIS for landslide research in Europe. Geomorphology 1996, 15, 227–237. [Google Scholar] [CrossRef] [Scilit]
  8. Hu, X.S.; Brierley, G.; Zhu, H.L.; Li, G.R.; Fu, J.T.; Mao, X.Q.; Yu, Q.Q.; Qiao, N. An exploratory analysis of vegetation strategies to reduce shallow landslide activity on loess hillslopes, Northeast Qinghai–Tibet Plateau, China. J. Mt. Sci. 2013, 10, 668–686. [Google Scholar] [CrossRef] [Scilit]
  9. He, L.; Wu, X.; He, Z.; Xue, D.; Bai, W.; Kang, G.; Chen, X.; Zhang, Y. Landslide identification and deformation monitoring analysis in Xining City based on the time series InSAR of Sentinel-1A with ascending and descending orbits. Bull. Eng. Geol. Environ. 2024, 83, 255. [Google Scholar] [CrossRef] [Scilit]
  10. Wang, F.W.; Chen, Y.; Yan, K.M. A destructive mudstone landslide hit a high-speed railway on 15 September 2022 in Xining City, Qinghai Province, China. Landslides 2023, 20, 871–874. [Google Scholar] [CrossRef] [Scilit]
  11. Bao, H.; Zeng, C.; Peng, Y.; Wu, S. The use of digital technologies for landslide disaster risk research and disaster risk management: Progress and prospects. Environ. Earth Sci. 2022, 81, 446. [Google Scholar] [CrossRef] [Scilit]
  12. Peng, J.; Tong, X.; Wang, S.; Ma, P. Three-dimensional geological structures and sliding factors and modes of loess landslides. Environ. Earth Sci. 2018, 77, 675. [Google Scholar] [CrossRef] [Scilit]
  13. Li, Y.; Ji, P.; Liu, S.; Zhao, J.; Yang, Y. Susceptibility evaluation of highway landslide disasters based on SBAS-InSAR: A case study of S211 Highway in Lanping County. Nat. Hazards 2025, 121, 2587–2612. [Google Scholar]
  14. Keefer, D.K. Landslides caused by earthquakes. Geol. Soc. Am. Bull. 1984, 95, 406–421. [Google Scholar] [CrossRef] [Scilit]
  15. Collins, B.D.; Znidarcic, D. Stability analyses of rainfall induced landslides. J. Geotech. Geoenviron Eng. 2004, 130, 362–372. [Google Scholar] [CrossRef] [Scilit]
  16. McColl, S.T. Landslide causes and triggers. In Landslide Hazards, Risks, and Disasters; Elsevier: Amsterdam, The Netherlands, 2022; pp. 13–41. [Google Scholar]
  17. Ma, W.; Dong, J.; Wei, Z.; Peng, L.; Wu, Q.; Wang, X.; Dong, Y.; Wu, Y. Landslide susceptibility assessment using the certainty factor and deep neural network. Front. Earth Sci. 2023, 10, 1091560. [Google Scholar] [CrossRef] [Scilit]
  18. Ma, Z.; Mei, G.; Piccialli, F. Machine learning for landslides prevention: A survey. Neural Comput. Appl. 2021, 33, 10881–10907. [Google Scholar]
  19. Chen, W.; Chen, Z.; Song, D.; He, H.; Li, H.; Zhu, Y. Landslide detection using the unsupervised domain-adaptive image segmentation method. Land 2024, 13, 928. [Google Scholar] [CrossRef] [Scilit]
  20. Su, Z.; Chow, J.K.; Tan, P.S.; Wu, J.; Ho, Y.K.; Wang, Y.-H. Deep convolutional neural network-based pixel-wise landslide inventory mapping. Landslides 2021, 18, 1421–1443. [Google Scholar]
  21. Xu, Q.; Ouyang, C.; Jiang, T.; Yuan, X.; Fan, X.; Cheng, D. MFFENet and ADANet: A robust deep transfer learning method and its application in high precision and fast cross-scene recognition of earthquake-induced landslides. Landslides 2022, 19, 1617–1647. [Google Scholar] [CrossRef] [Scilit]
  22. Liu, P.; Wei, Y.; Wang, Q.; Chen, Y.; Xie, J. Research on post-earthquake landslide extraction algorithm based on improved U-Net model. Remote Sens. 2020, 12, 894. [Google Scholar] [CrossRef] [Scilit]
  23. Zhang, X.; Feng, X.; Agrawal, V. Deep learning for processing and analysis of remote sensing big data: A technical review. Big Earth Data 2022, 6, 378–417. [Google Scholar]
  24. Tang, X.; Tu, Z.; Wang, Y.; Liu, M.; Li, D.; Fan, X. Automatic detection of coseismic landslides using a new transformer method. Remote Sens. 2022, 14, 2884. [Google Scholar] [CrossRef] [Scilit]
  25. Song, Y.; Zou, Y.; Li, Y.; He, Y.; Wu, W.; Niu, R.; Xu, S. Enhancing landslide detection with SBConv-optimized U-Net architecture based on multisource remote sensing data. Land 2024, 13, 835. [Google Scholar] [CrossRef] [Scilit]
  26. Chen, C.; He, Y.; Zhang, L.; Yao, S.; Yang, W.; Fang, Y.; Liu, Y.; Gao, B. A landslide extraction method of channel attention mechanism U-Net network based on Sentinel-2A remote sensing images. Int. J. Digit. Earth 2023, 16, 552–577. [Google Scholar] [CrossRef] [Scilit]
  27. Chen, L.C.; Zhu, Y.; Papandreou, G.; Schroff, F.; Adam, H. Encoder-decoder with atrous separable convolution for semantic image segmentation. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 833–851. [Google Scholar]
  28. Ghorbanzadeh, O.; Xu, Y.; Ghamisi, P.; Kopp, M.; Kreil, D. Landslide4Sense: Reference benchmark data and deep learning models for landslide detection. IEEE Trans. Geosci. Remote Sens. 2022, 60, 1–17. [Google Scholar] [CrossRef] [Scilit]
  29. Qin, D.; Li, Q.; Fang, L. Fanet: Landslide recognition in remote sensing images based on multi-source data. Environ. Earth Sci. 2026, 85, 205. [Google Scholar] [CrossRef] [Scilit]
  30. Meena, S.R.; Soares, L.P.; Grohmann, C.H.; van Westen, C.; Bhuyan, K.; Singh, R.P.; Floris, M.; Catani, F. Landslide detection in the Himalayas using machine learning algorithms and U-Net. Landslides 2022, 19, 1209–1229. [Google Scholar] [CrossRef] [Scilit]
  31. Ullo, S.L.; Mohan, A.; Sebastianelli, A.; Ahamed, S.E.; Kumar, B.; Dwivedi, R.; Sinha, G. A new Mask R-CNN-based method for improved landslide detection. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2021, 14, 3799–3810. [Google Scholar] [CrossRef] [Scilit]
  32. Wang, L.; Fan, X.; Xu, Q. The post-failure spatiotemporal deformation of certain translational landslides may follow the pre-failure pattern. Remote Sens. 2022, 14, 2333. [Google Scholar] [CrossRef] [Scilit]
  33. Chrysafi, A.-A.; Tsangaratos, P.; Ilia, I.; Chen, W. Rapid landslide detection following an extreme rainfall event using remote sensing indices, synthetic aperture radar imagery, and probabilistic methods. Land 2025, 14, 21. [Google Scholar]
  34. Gao, Y.; Zhang, C.; He, Q.; Wang, Z. A deep learning semantic segmentation method for landslide scene based on transformer architecture. Sustainability 2022, 14, 16311. [Google Scholar] [CrossRef] [Scilit]
  35. Tang, X.; Lu, Z.; Fan, X.; Yan, X.; Yuan, X.; Li, D.; Li, H.; Li, H.; Meena, S.R.; Novellino, A.; et al. Mamba for landslide detection: A lightweight model for mapping landslides with very high-resolution images. IEEE Trans. Geosci. Remote Sens. 2025, 63, 5637117. [Google Scholar] [CrossRef] [Scilit]
  36. Wang, B.; Su, J.; Xi, J.; Chen, Y.; Cheng, H.; Li, H.; Chen, C.; Shang, H.; Yang, Y. Landslide detection with MSTA-YOLO in remote sensing images. Remote Sens. 2025, 17, 2795. [Google Scholar] [CrossRef] [Scilit]
  37. Ji, S.; Yu, D.; Shen, C.; Li, W.; Xu, Q. Landslide detection from an open satellite imagery and digital elevation model dataset using attention boosted convolutional neural networks. Landslides 2020, 17, 1337–1352. [Google Scholar] [CrossRef] [Scilit]
  38. Zhang, X.; Yu, W.; Pun, M.-O.; Shi, W. Cross-domain landslide mapping from large-scale remote sensing images using prototype-guided domain-aware progressive representation learning. ISPRS J. Photogramm. Remote Sens. 2023, 197, 1–17. [Google Scholar] [CrossRef] [Scilit]
  39. Zeng, C.; Cao, Z.; Su, F.; Zeng, Z.; Yu, C. A dataset of high-precision aerial imagery and interpretation of landslide and debris flow disaster in Sichuan and surrounding areas between 2008 and 2020. China Sci. Data 2022, 7, 195–205. [Google Scholar] [CrossRef] [Scilit]
  40. Zeng, C.; Cao, Z.; Su, F.; Zeng, Z.; Yu, C. High-Precision Aerial Imagery and Interpretation Dataset of Landslide and Debris Flow Disaster in Sichuan and Surrounding Areas; Science Data Bank: Beijing, China, 2021. [Google Scholar] [CrossRef] [Scilit]
  41. Wei, G.; Yan, J.; Xia, Z.; Li, B.; Qi, H. Research on the instability mechanism of loess landslides based on preferential infiltration of rainfall. Front. Earth Sci. 2025, 13, 1586275. [Google Scholar] [CrossRef] [Scilit]
  42. Wei, B.Y.; Su, G.W.; Liu, F.G. Dynamic assessment of spatiotemporal population distribution based on mobile phone data: A case study in Xining City, China. Int. J. Disaster Risk Sci. 2023, 14, 649–665. [Google Scholar] [CrossRef] [Scilit]
  43. Yi, Y.; Zhang, W. A new deep-learning-based approach for earthquake-triggered landslide detection from single-temporal RapidEye satellite imagery. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2020, 13, 6166–6176. [Google Scholar] [CrossRef] [Scilit]
  44. Yu, W.; Zhou, P.; Yan, S.; Wang, X. InceptionNeXt: When Inception Meets ConvNeXt. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 17–21 June 2024; pp. 5672–5683. [Google Scholar]
Figure 1. Representative samples from the public landslide datasets used in this study. The top row shows optical image patches and the bottom row shows the corresponding binary landslide labels. From left to right, the three groups correspond to the Bijie, GVLM, and SCLM datasets.
Figure 1. Representative samples from the public landslide datasets used in this study. The top row shows optical image patches and the bottom row shows the corresponding binary landslide labels. From left to right, the three groups correspond to the Bijie, GVLM, and SCLM datasets.
Land 15 01298 g001
Figure 2. Study area and illustrative examples. (a) Geographic location of the Xining Basin (study area) within Qinghai Province, China. (b) Overview map of the study area overlaid on optical imagery, showing the expressway, main rivers, and mapped landslides. (c) Representative image–label pairs from the dataset (top: image patches; bottom: corresponding landslide labels). (d) Field photographs acquired by an unmanned aerial vehicle showing typical landslides in the study area.
Figure 2. Study area and illustrative examples. (a) Geographic location of the Xining Basin (study area) within Qinghai Province, China. (b) Overview map of the study area overlaid on optical imagery, showing the expressway, main rivers, and mapped landslides. (c) Representative image–label pairs from the dataset (top: image patches; bottom: corresponding landslide labels). (d) Field photographs acquired by an unmanned aerial vehicle showing typical landslides in the study area.
Land 15 01298 g002
Figure 3. Statistical properties of the cropped landslide instances (n = 898). (a) Distribution of landslide area in logarithmic scale (median = 375.4 m2; P95 = 3130 m2). (b) Distribution of aspect ratio (length/width), where 23.8% of instances are slender (ratio > 3.0). (c) Distribution of foreground (landslide) pixel ratio per patch, showing severe foreground sparsity (mean ratio = 3.44%).
Figure 3. Statistical properties of the cropped landslide instances (n = 898). (a) Distribution of landslide area in logarithmic scale (median = 375.4 m2; P95 = 3130 m2). (b) Distribution of aspect ratio (length/width), where 23.8% of instances are slender (ratio > 3.0). (c) Distribution of foreground (landslide) pixel ratio per patch, showing severe foreground sparsity (mean ratio = 3.44%).
Land 15 01298 g003
Figure 4. The architecture of the TM-Net.
Figure 4. The architecture of the TM-Net.
Land 15 01298 g004
Figure 5. Structure of the Inception Token Mixer block.
Figure 5. Structure of the Inception Token Mixer block.
Land 15 01298 g005
Figure 6. Workflow of the TM-Block.
Figure 6. Workflow of the TM-Block.
Land 15 01298 g006
Figure 7. Qualitative comparison of segmentation results for representative validation samples in the comparative experiments. (A) A partially visible landslide near the image boundary; (B) multiple slender landslide bodies; (C) a long, narrow landslide body; (D) a small isolated landslide; and (E) a large irregular landslide in a heterogeneous background. For each sample, the columns from left to right show the original image, ground-truth label, and segmentation results generated by the evaluated models. True-positive (TP), false-positive (FP), and false-negative (FN) pixels are shown in red, green, and blue, respectively.
Figure 7. Qualitative comparison of segmentation results for representative validation samples in the comparative experiments. (A) A partially visible landslide near the image boundary; (B) multiple slender landslide bodies; (C) a long, narrow landslide body; (D) a small isolated landslide; and (E) a large irregular landslide in a heterogeneous background. For each sample, the columns from left to right show the original image, ground-truth label, and segmentation results generated by the evaluated models. True-positive (TP), false-positive (FP), and false-negative (FN) pixels are shown in red, green, and blue, respectively.
Land 15 01298 g007
Figure 8. Grad-CAM visualizations for the representative validation samples shown in Figure 7. (A) A partially visible landslide near the image boundary; (B) multiple slender landslide bodies; (C) a long, narrow landslide body; (D) a small isolated landslide; and (E) a large irregular landslide in a heterogeneous background. For each sample, the columns from left to right show the original image, ground-truth label, and Grad-CAM activation maps generated by the evaluated models. Warmer colors indicate stronger activation, whereas cooler colors indicate weaker activation.
Figure 8. Grad-CAM visualizations for the representative validation samples shown in Figure 7. (A) A partially visible landslide near the image boundary; (B) multiple slender landslide bodies; (C) a long, narrow landslide body; (D) a small isolated landslide; and (E) a large irregular landslide in a heterogeneous background. For each sample, the columns from left to right show the original image, ground-truth label, and Grad-CAM activation maps generated by the evaluated models. Warmer colors indicate stronger activation, whereas cooler colors indicate weaker activation.
Land 15 01298 g008
Figure 9. Qualitative comparison of segmentation results for representative validation samples in the ablation study. (A) A large irregular landslide extending to the image boundary; (B) a long, narrow landslide body in forested terrain; (C) multiple closely spaced elongated landslide bodies in complex terrain; and (D) a small isolated landslide adjacent to a road. The columns compare the Baseline decoder, Baseline + Inception Token Mixer, Baseline + TM-Block, and TM-Net. True-positive (TP), false-positive (FP), and false-negative (FN) pixels are shown in red, green, and blue, respectively.
Figure 9. Qualitative comparison of segmentation results for representative validation samples in the ablation study. (A) A large irregular landslide extending to the image boundary; (B) a long, narrow landslide body in forested terrain; (C) multiple closely spaced elongated landslide bodies in complex terrain; and (D) a small isolated landslide adjacent to a road. The columns compare the Baseline decoder, Baseline + Inception Token Mixer, Baseline + TM-Block, and TM-Net. True-positive (TP), false-positive (FP), and false-negative (FN) pixels are shown in red, green, and blue, respectively.
Land 15 01298 g009
Figure 10. Grad-CAM visualizations for the representative validation samples shown in Figure 9. (A) A large irregular landslide extending to the image boundary; (B) a long, narrow landslide body in forested terrain; (C) multiple closely spaced elongated landslide bodies in complex terrain; and (D) a small isolated landslide adjacent to a road. The columns show the Grad-CAM activation maps generated by the Baseline decoder, Baseline + Inception Token Mixer, Baseline + TM-Block, and TM-Net. Warmer colors indicate stronger activation, whereas cooler colors indicate weaker activation.
Figure 10. Grad-CAM visualizations for the representative validation samples shown in Figure 9. (A) A large irregular landslide extending to the image boundary; (B) a long, narrow landslide body in forested terrain; (C) multiple closely spaced elongated landslide bodies in complex terrain; and (D) a small isolated landslide adjacent to a road. The columns show the Grad-CAM activation maps generated by the Baseline decoder, Baseline + Inception Token Mixer, Baseline + TM-Block, and TM-Net. Warmer colors indicate stronger activation, whereas cooler colors indicate weaker activation.
Land 15 01298 g010
Figure 11. Qualitative comparison of segmentation results for representative samples from XBLD. (A) A small landslide near the upper image boundary; (B) a partially visible landslide near the lower-left image boundary; and (C) multiple spatially separated landslide bodies of different sizes. For each sample, the columns from left to right show the original image, ground-truth label, and segmentation results generated by the evaluated models. The yellow circles in the original images indicate the locations of the target landslides for visual reference only and do not represent an additional class or annotation. True-positive (TP), false-positive (FP), and false-negative (FN) pixels are shown in red, green, and blue, respectively.
Figure 11. Qualitative comparison of segmentation results for representative samples from XBLD. (A) A small landslide near the upper image boundary; (B) a partially visible landslide near the lower-left image boundary; and (C) multiple spatially separated landslide bodies of different sizes. For each sample, the columns from left to right show the original image, ground-truth label, and segmentation results generated by the evaluated models. The yellow circles in the original images indicate the locations of the target landslides for visual reference only and do not represent an additional class or annotation. True-positive (TP), false-positive (FP), and false-negative (FN) pixels are shown in red, green, and blue, respectively.
Land 15 01298 g011
Figure 12. Grad-CAM visualizations for the representative XBLD samples shown in Figure 11. (A) A small landslide near the upper image boundary; (B) a partially visible landslide near the lower-left image boundary; and (C) multiple spatially separated landslide bodies of different sizes. For each sample, the columns from left to right show the original image, ground-truth label, and Grad-CAM activation maps generated by the evaluated models. The yellow circles in the original images indicate the locations of the target landslides for visual reference only. Warmer colors indicate stronger activation, whereas cooler colors indicate weaker activation.
Figure 12. Grad-CAM visualizations for the representative XBLD samples shown in Figure 11. (A) A small landslide near the upper image boundary; (B) a partially visible landslide near the lower-left image boundary; and (C) multiple spatially separated landslide bodies of different sizes. For each sample, the columns from left to right show the original image, ground-truth label, and Grad-CAM activation maps generated by the evaluated models. The yellow circles in the original images indicate the locations of the target landslides for visual reference only. Warmer colors indicate stronger activation, whereas cooler colors indicate weaker activation.
Land 15 01298 g012
Figure 13. Visualizing the spectral confusion between landslides and erosion features identified by TM-Net in the Xining Basin.
Figure 13. Visualizing the spectral confusion between landslides and erosion features identified by TM-Net in the Xining Basin.
Land 15 01298 g013
Table 1. Summary of the landslide datasets used in this study.
Table 1. Summary of the landslide datasets used in this study.
DatasetRegionPeriodSourceNative GSD (m)Channels
BijieBijie City, Guizhou, ChinaMay–Aug 2018TripleSat optical0.8RGB
GVLM17 subregions worldwideRelease 2023Google Maps0.59RGB
SCLMSichuan and surrounding areas, China2008–2020Digital orthophotos0.2–0.9RGB
XBLDXining, Qinghai, China2022–2023Gaofen-20.8RGB
Table 2. Hyperparameter setting of TM-Net.
Table 2. Hyperparameter setting of TM-Net.
HyperparameterSetting
Image size224 × 224
Batch size32
Max Epoch100
OptimizerAdamW
Learning Rate0.001
LossComposite Loss (Dice & BCE with OHEM)
Table 3. Comparison of landslide segmentation performance on the fused public benchmark.
Table 3. Comparison of landslide segmentation performance on the fused public benchmark.
ModelIoU (%)F1 (%)Pre (%)Rec (%)
DeepLabV3+71.6883.5082.1084.95
SegFormer72.6784.1780.4388.29
VMUNet71.9883.7180.7986.85
WiT-UNet68.6181.3877.2585.98
PAM-UNet71.8083.5982.2185.01
Swin-UMamba68.4581.2780.7481.82
CSWin-UNet67.6780.7277.7483.94
TM-Net (Ours)73.5684.7682.4787.20
Table 4. Ablation study of the adopted Inception Token Mixer and the proposed TM-Block.
Table 4. Ablation study of the adopted Inception Token Mixer and the proposed TM-Block.
ModelIoU (%)F1 (%)Pre (%)Rec (%)
Baseline70.7282.8580.1785.71
Baseline + Inception Token Mixer72.1683.8380.7087.21
Baseline + TM-Block73.0684.4384.1584.72
TM-Net (Ours)73.5684.7682.4787.20
Table 5. Cross-architecture benchmarking on XBLD using the same data split and evaluation protocol.
Table 5. Cross-architecture benchmarking on XBLD using the same data split and evaluation protocol.
ModelIoU (%)F1 (%)Pre (%)Rec (%)
DeepLabV3+43.8160.9357.7564.49
SegFormer42.6459.7953.8367.23
VMUNet36.9753.9947.7062.19
WiT-UNet8.1515.0611.0123.83
PAM-UNet37.7854.8445.8768.17
Swin-UMamba9.9818.1510.3772.46
CSWin-UNet39.5556.6849.5766.17
TM-Net (Ours)44.9462.0157.6767.07
Table 6. Computational complexity and inference throughput of the compared models.
Table 6. Computational complexity and inference throughput of the compared models.
ModelParameters (M)FLOPs (G)FPS
DeepLabV3+54.6115.82153.14
SegFormer7.722.51165.52
VMUNet60.368.4040.44
WiT-UNet4.9963.3175.02
PAM-UNet5.8911.61189.89
Swin-UMamba27.494.7355.03
CSWin-UNet16.953.329.44
TM-Net (Ours)35.854.1871.97
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Yang, H.; Liu, W.; Liu, Y. Mapping Landslide-Affected Land Surfaces in Complex Mountainous Landscapes Using a Twin-Path Multi-Scale Deep Learning Network. Land 2026, 15, 1298. https://doi.org/10.3390/land15071298

AMA Style

Yang H, Liu W, Liu Y. Mapping Landslide-Affected Land Surfaces in Complex Mountainous Landscapes Using a Twin-Path Multi-Scale Deep Learning Network. Land. 2026; 15(7):1298. https://doi.org/10.3390/land15071298

Chicago/Turabian Style

Yang, Heming, Wenhui Liu, and Yabin Liu. 2026. "Mapping Landslide-Affected Land Surfaces in Complex Mountainous Landscapes Using a Twin-Path Multi-Scale Deep Learning Network" Land 15, no. 7: 1298. https://doi.org/10.3390/land15071298

APA Style

Yang, H., Liu, W., & Liu, Y. (2026). Mapping Landslide-Affected Land Surfaces in Complex Mountainous Landscapes Using a Twin-Path Multi-Scale Deep Learning Network. Land, 15(7), 1298. https://doi.org/10.3390/land15071298

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop