Next Article in Journal
Integrating Gross Error Identification with Deep Learning for InSAR Topography-Dependent Delay Correction: A Case Study of the Baihetan Hydropower Station Area
Previous Article in Journal
From Detection to Functional Analysis: Evaluating Vehicle Detection Models in High-Resolution Earth Observation Imagery
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Prior-Informed Local Landslide Susceptibility Modelling Using National-Scale Landslide Information: A Case Study of the Parlung Tsangpo Alpine Gorge Basin

1
College of Geography and Planning, Chengdu University of Technology, Chengdu 610059, China
2
State Key Laboratory of Geohazard Prevention and Geoenvironment Protection, Chengdu University of Technology, Chengdu 610059, China
3
College of Computer Science and Cybersecurity, Chengdu University of Technology, Chengdu 610059, China
*
Author to whom correspondence should be addressed.
Remote Sens. 2026, 18(13), 2167; https://doi.org/10.3390/rs18132167
Submission received: 9 June 2026 / Revised: 24 June 2026 / Accepted: 29 June 2026 / Published: 3 July 2026
(This article belongs to the Section AI Remote Sensing)

Highlights

What are the main findings?
  • National-scale landslide information can support local susceptibility modelling in data-scarce alpine gorge basins once it is transformed into a probability prior.
  • Direct transfer of a national model, or weighted source–target joint training, may introduce probability bias and lead to spatial overprediction of susceptibility.
What are the implications of the main findings?
  • Converting national-model outputs into locally recalibrated probability priors provides better regional adaptability than either direct transfer or joint sample training.
  • The prior-informed model improves discrimination, probability calibration, and mapping reliability, while yielding more spatially concentrated high-susceptibility patches.

Abstract

Landslides pose widespread threats to mountainous communities and infrastructure worldwide, yet susceptibility mapping in alpine gorge basins is often constrained by sparse and incomplete local inventories. Whether national-scale landslide information can be transformed into reliable local knowledge remains unclear, particularly where strong topographic and environmental heterogeneity limits the direct transfer of broad-scale models. Here, we use the Parlung Tsangpo Basin on the southeastern Tibetan Plateau as a test case and compare four strategies for introducing national-scale information: local baseline modelling, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling. The experiments use the same conditioning factors, data-processing workflow, spatially independent test set, and evaluation metrics, allowing the transfer strategies to be assessed under controlled conditions. The prior-informed strategy treats the susceptibility probability produced by the national model as a geoscientifically interpretable external prior, which is then relearned and recalibrated by local samples, and achieves the best overall performance. On the independent test set, it reaches an area under the receiver operating characteristic curve (AUC) of 0.901 and reduces the expected calibration error (ECE) to 0.055, outperforming the local baseline, direct transfer, and joint training strategies. Its susceptibility map shows an extensive low-susceptibility background with spatially concentrated high-susceptibility patches, thereby reducing broad-scale overprediction while preserving local landslide-prone zones. These results indicate that national-scale landslide information is more effective when converted into a locally recalibrated probability prior than when transferred directly, providing a practical pathway for susceptibility assessment in data-scarce mountainous basins.

1. Introduction

Landslides are a global geohazard that cause casualties, infrastructure disruption, river blockage, and long-term geomorphic disturbance in mountain regions [1]. In alpine gorge environments, slope failure may further evolve into cascading hazards, including long-runout movement, landslide-dammed lakes, outburst floods, and debris flow–river blockage chains [2,3]. Such hazards are particularly critical in High Asia and around the Tibetan Plateau, where climate warming, cryospheric degradation, and extreme precipitation are expected to increase landslide-related risks [1,4,5]. Reliable identification of landslide-susceptible terrain in steep, poorly accessible, and hazard-prone mountain basins is therefore essential for disaster prevention, infrastructure planning, and engineering safety assessment.
Landslide susceptibility assessment estimates the spatial likelihood of landslide occurrence by relating landslide inventories to conditioning factors such as topography, lithology, hydrology, vegetation, climate, and human activity [6]. Machine-learning methods have been increasingly adopted for this task because they can capture nonlinear relationships between landslide occurrence and multi-source environmental variables. For example, in susceptibility mapping over the Qinghai–Tibetan Plateau, random forest outperformed deep neural networks, logistic regression, naïve Bayes, and support vector machine models, demonstrating strong predictive capability under complex environmental conditions [6]. However, the reliability of data-driven models depends strongly on the completeness, positional accuracy, and spatial representativeness of landslide inventories [7,8]. In high-mountain basins, local inventories are often limited by poor accessibility, topographic shadow, vegetation cover, and incomplete historical records, which may lead to overfitting, unstable probability estimation, and uncertain susceptibility maps.
National- or large-regional-scale landslide inventories provide a potential way to alleviate local sample scarcity. Compared with a small-basin inventory, broad-scale records cover a wider range of geomorphological, lithological, climatic, and anthropogenic conditions, so the resulting national model can encode cross-regional statistical associations between landslide occurrence and environmental factors. These associations are useful as source-domain background information, but they should not be interpreted as universally transferable landslide-development rules. They may partly reflect the environments that are better represented in the national inventory and may not be directly applicable to the alpine gorge setting at the southeastern margin of the Tibetan Plateau. Differences in tectonic setting, geomorphic evolution, rainfall regime, snow and ice processes, and engineering disturbance may therefore cause direct national-model transfer or simple source–target sample merging to produce weak regional adaptability, probability bias, and excessive expansion of high-susceptibility zones [9,10]. Therefore, when sample scarcity and regional heterogeneity coexist, the key issue is how to transform national-scale landslide information into locally reliable susceptibility knowledge.
Previous studies have explored cross-regional landslide knowledge transfer through transfer learning, domain adaptation, case-based reasoning, sample selection, and joint training [9,11], highlighting the potential value of source-domain information for data-scarce regions. However, two issues remain. First, the general applicability of any transfer strategy still needs to be examined through independent validation in multiple basins. Second, when different studies use different study areas, factor systems, sample-construction rules, and evaluation schemes, it is difficult to isolate the effect of the source information utilisation strategy itself. The present study, therefore, focuses on a controlled comparison within one representative alpine gorge basin, using the same conditioning factors, data-processing workflow, spatially independent target-domain test set, and evaluation metrics. Less attention has been paid to converting the output of a national-scale model into an external probability prior that can be further learned and recalibrated by local samples, despite the importance of calibrated probability estimates for susceptibility mapping [12]. In this study, the two-stage computation is regarded as an implementation pathway rather than the main methodological emphasis. The key idea is to treat the national-model probability as a geoscientifically interpretable prior that expresses broad-scale landslide-background conditions in the target basin, while allowing target-domain samples to determine its effective weight and correction direction. This differs from generic stacking or model fusion [13], where first-level outputs mainly function as algorithmic ensemble features, and from feature- or sample-level transfer [9,11], where source-domain information is transferred more directly.
The Parlung Tsangpo Basin, located on the southeastern margin of the Tibetan Plateau, is a typical alpine gorge basin characterised by strong topographic relief, deeply incised river valleys, active fault systems, and pronounced river erosion. These conditions provide a favourable environment for frequent landslide occurrence and cascading hazard development. Deeply incised valleys, active tectonics, complex lithology, monsoonal precipitation, snow and ice meltwater, and ice–rock avalanche processes jointly control slope instability in this region. Repeated cascading events, including landslide-induced river blockage, landslide-dammed lake outburst flooding, and ice–rock avalanche–debris flow–river blockage chains, have occurred in the basin and adjacent areas [2,3,14]. Meanwhile, strong topographic obstruction, limited field verification, and incomplete historical records make it difficult for local landslide samples to cover the full range of conditioning environments. The basin, therefore, provides a suitable target domain for evaluating how national-scale landslide information can be used in local susceptibility modelling under sample-limited conditions.
In this study, the Parlung Tsangpo Basin is used as the target domain, and national-scale landslide records are introduced as external source-domain information. Under a unified conditioning-factor system, data-processing workflow, target-domain test set, and evaluation metrics, four strategies are compared: a local baseline model, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling. The objectives are to: (1) test whether national-scale landslide information can improve susceptibility prediction in a data-scarce alpine gorge basin; (2) compare how different source information utilisation strategies perform in terms of discrimination ability, classification accuracy, probability calibration, and spatial mapping; and (3) evaluate whether national-model output can serve as an effective probability prior for producing locally adapted susceptibility maps. This study provides a controlled within-basin comparison framework for converting broad-scale landslide information into basin-scale susceptibility assessment, while leaving the broader cross-basin transferability of the proposed strategy to be tested in future studies.

2. Study Area and Data

2.1. Study Area and Application Context

The Parlung Tsangpo Basin lies on the southeastern margin of the Tibetan Plateau and is an alpine gorge basin within the Yarlung Tsangpo River system (Figure 1a). As Figure 1a shows, the study area has strong topographic relief, a deeply incised main valley, and pronounced slope cutting by the trunk river and its tributaries; active faults, transportation corridors, and hydropower-related infrastructure are distributed across the basin. Under the combined effects of strong tectonic activity, river incision, high and steep slopes, and lithological contrasts, the region presents the geomorphological and geological conditions typical of landslide development [15,16].
As shown in Figure 1c, the stratigraphic and lithological assemblages of the basin are complex, and slope structure, weathering and unloading conditions, and shear-strength characteristics differ markedly among geological units. Historically, large landslides, river blockages, and landslide-dammed lake events have occurred in the study area and in adjacent basins, reflecting the setting in which landslides and cascading hazards such as landslide–river blockage–outburst flood chains develop in this region [17,18].
Because of the poor terrain accessibility of alpine gorge regions and the effects of topographic shadow and vegetation cover on remote-sensing interpretation, local landslide samples are typically limited in number and unevenly distributed in space. This study treats the Parlung Tsangpo Basin as the target domain and introduces national-scale landslide records as source-domain information. As shown in Figure 2, the national-scale records provide broad but spatially uneven coverage across China: they are relatively dense in eastern and central regions but much sparser in western high-mountain areas, including the Tibetan Plateau. Therefore, these records provide useful broad-scale source-domain information, while their uneven spatial representation also needs to be considered when transferring national-scale knowledge to the target alpine gorge basin.

2.2. Geospatial Data Sources

To support landslide susceptibility modelling in the Parlung Tsangpo Basin and the introduction of national-scale source-domain information, this study integrates multi-source geospatial data covering basic geography, geology, geomorphology, topography, climate, vegetation, land cover, soil, and landslide labels. The environmental data are used to construct the landslide conditioning-factor system, and the label data are used to generate the target-domain and source-domain samples. From these data, 19 candidate conditioning factors were derived, spanning environmental characteristics relevant to landslide development, including topographic relief, valley incision, tectonic and lithological conditions, geomorphological and soil backgrounds, climatic conditions, vegetation cover, land use, and human engineering disturbance. These factors were selected to represent the main static controls that can be mapped consistently for both the national source domain and the Parlung Tsangpo target basin. In particular, the factor set covers valley incision, fault and lithological contrasts, vegetation and land-cover conditions, soil and geomorphic backgrounds, and transportation-related engineering disturbance, which are closely related to the alpine gorge setting of the study area. The principal use, source, spatial specification, and temporal range of each dataset are summarised in Table 1.
The landslide-label data comprise the target-domain inventory of the Parlung Tsangpo Basin and the national-scale landslide records. The target-domain inventory consists of 249 historical landslide points from the National Geological Hazard Database and 33 points additionally interpreted in this study, giving 282 records in total. The national-scale records were derived from the National Geological Hazard Point Inventory compiled by the Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences. The original inventory contained 109,607 records, and 105,821 valid records were retained after data cleaning and used in this study. These records cover the period from 1995 to 2024, with a spatial accuracy generally within 20 m. The inventory includes different landslide types, such as rainfall-induced landslides, earthquake-induced landslides, and landslides related to human engineering activities, and some records also include information on landslide scale and economic losses. Records outside the target domain were used for source-domain information extraction. The target-domain inventory characterises local landslide development within the study area, whereas the national-scale records provide the sample basis for introducing cross-regional source-domain information.
The national inventory covers landslides developed under a wide range of geomorphic, climatic, tectonic, and anthropogenic conditions. This broad coverage provides the data basis for extracting national-scale landslide-background information, while also introducing source-domain heterogeneity relative to the Parlung Tsangpo Basin. The target basin is an alpine gorge region where landslide occurrence is closely associated with deeply incised valleys, active tectonics, complex lithology, monsoonal precipitation, snow and ice meltwater, freeze–thaw processes, and local engineering corridors. By comparison, the national records also include landslides from hilly terrains and cases more directly related to rainfall, earthquakes, or human engineering activities. These source–target differences provide the background for comparing direct national-model transfer, weighted source–target joint training, and locally recalibrated prior-informed modelling in this study.

3. Materials and Methods

The methodological framework of this study is shown in Figure 3. For clarity, the four experimental strategies are denoted as: E1 (local baseline), E2 (direct national-model transfer), E3 (weighted source–target joint training), and E4 (prior-informed local modelling). The Parlung Tsangpo Basin is treated as the target domain and the national-scale landslide samples outside it as the source domain. Around the central question of how national-scale landslide information can be transformed into local susceptibility modelling, four controlled experiments are constructed. The overall workflow comprises multi-source data preparation, conditioning-factor derivation and preprocessing, the design of cross-scale information introduction strategies, model validation and robustness assessment, and landslide susceptibility mapping and interpretation.
To ensure that the results of the different experiments are comparable, E1–E4 share the same original conditioning-factor system, data-preprocessing workflow, sample-construction method, spatial partitioning scheme, and evaluation metrics; only the way in which national-scale information is introduced is varied. Let the target-domain samples be
D T = { ( x i , y i ) } i = 1 n T ,
and the source-domain samples be
D S = { ( x i , y i ) } i = 1 n S ,
where y i { 0,1 } denotes a non-landslide or landslide sample, x i R p denotes the model input-feature vector after preprocessing and one-hot encoding of categorical variables, and p is the actual feature dimension after encoding. Before entering a model, categorical variables such as lithology, land-use type, soil type, and geomorphological type are expanded into higher-dimensional features through one-hot encoding.
On the basis of this unified setting, the study compares four strategies—the local baseline model, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling—in order to identify effective ways of using national-scale information when local samples are limited.

3.1. Experimental Design and Overall Rationale

To examine the local applicability of national-scale landslide information in data-scarce basins, four controlled modelling strategies are defined: the local baseline model, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling. The four strategies share the same independent target-domain test set and evaluation framework, ensuring that performance differences arise mainly from the way cross-scale information is introduced rather than from differences in sample partitioning, factor systems, or evaluation methods.
E1 is the local baseline. Only samples from the Parlung Tsangpo Basin and the 19 original conditioning factors are used to train the model, characterising conventional local modelling performance in the absence of external information [19,20]. Letting A ( ) denote a given learner together with its training procedure, E1 can be written as
f ^ E 1 = A D T , X ,
where X denotes the model input features formed from the 19 original conditioning factors after preprocessing.
E2 is the direct national-model transfer strategy: national samples outside the target domain are used to train a model that is then applied directly to the Parlung Tsangpo Basin, in order to examine the generalisation ability of the national-scale model in the local geomorphological environment [9,21,22]. This process can be expressed as
f ^ E 2 = A D S , X ,
p ^ E 2 x = f ^ E 2 x ,   x D T .
Because no target-domain samples are used in training, this strategy serves to evaluate the direct-transfer ability of the source-domain model in the target domain.
E3 is the weighted source–target joint training strategy: national and target-domain samples are used jointly for training, with higher weights assigned to the target-domain samples to reduce the risk of negative transfer caused by source–target distribution mismatch [9,23,24]. Its training objective can be written as
f ^ E 3 = a r g   m i n f i D S l ( y i , f ( x i ) ) + α i D T l ( y i , f ( x i ) )
where l is the classification loss and α is the target-domain sample weight. In this study α is searched over 1 ,   2 ,   4 ,   6 ,   8 to balance the external information provided by the national samples against the local environmental characteristics reflected by the target-domain samples.
E4 is the prior-informed local modelling strategy. A source-domain model is first trained on the national samples and used to generate national-scale susceptibility probabilities over the target domain. These probabilities are then used as an additional probability-prior variable and fed, together with the local conditioning factors, into the target-domain model, allowing the national-scale information to be relearned and recalibrated under the constraint of local samples [13,25,26]. This process can be expressed as
p N x = f ^ S x ,
x i * = x i , p N x i ,
f ^ E 4 = A D T , X * ,
where p N ( x ) is the susceptibility probability output by the national model at target-domain sample or pixel locations, and x i * is the target-domain input feature after the national probability prior has been appended. Unlike E2, E4 does not adopt the national-model output as the final result; instead, it treats that output as a learnable probability-prior variable that is locally recalibrated by the target-domain samples. Unlike generic stacking, this variable is not introduced merely as an algorithmic ensemble feature, but represents the national-scale model’s probabilistic expression of broad-scale landslide-background conditions at each target-domain location. Its effective contribution and correction direction are therefore estimated from the local samples together with the original conditioning factors.
In summary, E1 provides the local baseline without external information, E2 examines the direct-transfer ability of the national model, E3 evaluates the gain from weighted source–target joint training, and E4 tests the effectiveness of the national probability prior after relearning with local samples. Together, the four strategies form a controlled comparison spanning no transfer, direct transfer, sample-level fusion, and probability-prior fusion.
The distinctions among the three information-introduction strategies are summarised as follows. Direct transfer (E2) applies the national-scale model predictions directly to the target basin without further adaptation. Joint training (E3) introduces source-domain information by combining national and local samples into a unified training set and fitting a new model jointly. In contrast, prior-informed local modelling (E4) does not directly use the national prediction as the final result; instead, it treats the national model output as a probabilistic prior that is further learned, reweighted, and recalibrated using target-domain samples together with the original conditioning factors.

3.2. Remote-Sensing and Spatial Data Preprocessing

Before sample extraction, the multi-source geospatial data were uniformly preprocessed to ensure consistency of the input variables in spatial scale and raster registration [27]. All data were brought to a common coordinate reference system, spatial resolution, and raster-alignment framework; continuous raster variables were resampled by cubic convolution to better preserve smooth spatial variation in continuous surfaces, whereas categorical variables were resampled by nearest-neighbour resampling to preserve original class labels and avoid the creation of artificial category values [28,29].
For unordered categorical variables such as lithology, land-use type, soil type, and geomorphological type, one-hot encoding was used to convert them into binary indicator variables, thereby preventing the model from misinterpreting category codes as numerical magnitudes or ordinal relationships [30]. Continuous variables were scaled according to learner characteristics: Z-score standardisation was applied to the logistic regression and SVM-RBF models, whereas the tree ensemble models retained the original variable scales [31,32]. For a continuous variable x j , its standardised form is
x i j = x i j μ j σ j ,
where μ j and σ j are the mean and standard deviation of the j -th variable in the training data and are subsequently applied to the corresponding validation set, test set, and mapping raster.

3.3. Feature Engineering

3.3.1. Selection of Conditioning Factors

Combining commonly used landslide susceptibility factors with the alpine-gorge characteristics of the Parlung Tsangpo Basin, 19 conditioning factors were selected from topography, hydrology, geology, geomorphology, climate, vegetation, land cover, and human disturbance [33,34,35,36]. To make the data-to-factor relationship explicit, all variables were derived on the unified raster grid described in Section 3.2. Let s denote a raster cell, z ( s )   the DEM elevation, W s a local moving window, and F a vector layer such as rivers, faults, roads, or railways. DEM-based factors were derived from elevation, local window statistics, or surface derivatives; for example,
R ( s ) = m a x   u W s z ( u ) m i n   u W s z ( u ) ,
T P I ( s ) = z ( s ) 1 W s u W s z ( u ) .
The topographic wetness index was calculated as
T W I ( s ) = l n a ( s ) t a n   β ( s ) ,
where a ( s ) is the upslope contributing area and β s   is slope [37]. Distance factors were calculated as Euclidean distances,
D F ( s ) = m i n u F s u .
Aspect was represented by s i n A ( s )   and c o s A ( s )   to avoid angular discontinuity [38]. Categorical factors were rasterised and one-hot encoded as I k s   =   1   if C s   =   k , and I k s   =   0   otherwise. The selected factors and their derivation basis are summarised in Table 2.

3.3.2. Feature Screening

Because this study first compared multiple candidate learners with different sensitivities to correlated predictors, feature screening was conducted on the common conditioning-factor set before final learner selection. Tolerance (TOL) and the variance inflation factor (VIF) were used to diagnose the 19 candidate factors [42,43]. For the j -th candidate factor, an auxiliary regression was fitted using the remaining factors as explanatory variables, yielding the coefficient of determination R j 2 . The tolerance and variance inflation factor are defined as
T O L j = 1 R j 2 ,
V I F j = 1 T O L j = 1 1 R j 2 .
Following common criteria, VIF > 10 or TOL < 0.1 was adopted as the threshold for severe collinearity [44]. Factors below these thresholds were retained for model training; where severe collinearity was present, factors were removed or adjusted with reference to both their geoscientific meaning and model stability.

3.4. Construction of Experimental Samples

The study-area and national-scale landslide samples were built with a unified construction workflow. To reduce information overlap among spatially adjacent samples and the influence of spatial autocorrelation, the landslide points were spatially thinned before modelling. The 250 m minimum retained separation was selected as an a priori operational threshold based on the 30 m DEM-derived mapping resolution, the local clustering characteristics of landslides in the alpine gorge setting, and the need to retain sufficient local positive samples for model training. This distance corresponds to approximately eight DEM pixels, which helps reduce repeated sampling of neighbouring points with highly similar terrain conditions. In the Parlung Tsangpo Basin, landslide records are strongly clustered along valley sides and transportation corridors; therefore, a much smaller threshold would retain many samples from the same local terrain context, whereas a much larger threshold would unnecessarily reduce the already limited local landslide records. The same 250 m distance was also used as the buffer width during spatial cross-validation, so that sample thinning and validation partitioning treated local spatial dependence consistently. Neighbouring points closer than this threshold were removed [45,46,47,48,49].
Negative-sample construction balanced geoscientific plausibility against spatial independence. First, water bodies, built-up land, glaciers, and areas immediately adjacent to landslide points were excluded, reducing the chance that potential unlabelled landslides or unsuitable background areas would be sampled as negatives [46,50,51]. Geomorphological type was then used as the stratification unit, and negative samples were drawn at random within each sampleable area, with their spatial distribution controlled to weaken the influence of sample clustering on the model decision boundary [52,53,54]. Finally, positive and negative samples were constructed at a 1:1 ratio to limit the effect of class imbalance on model training and evaluation [55,56]. The above procedure was used as the main negative-sampling strategy and is denoted as A0. To evaluate whether the model comparison depended on this particular negative-sampling rule, two additional negative-sampling schemes were designed as sensitivity tests and reported in the Supplementary Materials. In A1, negative samples were resampled within slope strata so that the comparison was less affected by differences in the slope-frequency structure. In A2, hard negative samples were drawn from the slope range overlapping the positive samples, thereby forcing the models to distinguish stable and unstable terrain under similar gradient conditions. The same positive samples, conditioning factors, learner settings, and evaluation metrics were kept for E1–E4, and only the negative-sampling rule was changed.

3.5. Dataset Partitioning

To reduce performance overestimation caused by spatial autocorrelation, a spatial-blocking partitioning strategy was adopted. A 2 km × 2 km grid was used at the study-area scale and a 20 km × 20 km grid at the national scale, strengthening the spatial independence among the training, validation, and independent test sets [45,47,48].
Sample partitioning followed the principles of spatial independence and class balance. An independent test set was first constructed by holding out entire spatial blocks, so that the test samples were spatially isolated from the training samples while the numbers of positive and negative samples were kept equal. The remaining samples were used for model training and cross-validation. During cross-validation, four-fold spatial outer cross-validation was built using the spatial grids as basic units, and a 250 m spatial buffer was placed between the training and validation sets; samples falling within the buffer were excluded from the corresponding training or validation fold, reducing the risk of information leakage from neighbouring samples [45,48,49].

3.6. Model Training and Hyperparameter Tuning

To reduce the dependence of model selection on the assumptions of any single learner, four candidate models were compared under a unified sample, factor, and spatial-validation framework: L2-regularised logistic regression, random forest, XGBoost, and stacking [19,57]. Logistic regression provides a linear probabilistic baseline [58]; random forest and XGBoost represent bagging and boosting tree ensembles, respectively [59,60]; and stacking fuses the complementary information of several nonlinear learners [13,61,62]. In the stacking model, the first-level learners are Extra Trees, HistGradientBoosting, and SVM-RBF, and the second-level learner is L2-regularised logistic regression. Letting the model output be the landslide-occurrence probability p ^ i = P ( y i = 1 x i ) , the core expression of each learner is given in Table 3.
Model training followed the spatial-partitioning and preprocessing rules described above. Logistic regression and SVM-RBF were standardised using Z-score parameters estimated from the training fold, whereas the tree ensembles retained the original scales of the continuous variables; categorical variables were transformed using one-hot encoding rules determined from the training fold, so as to avoid data leakage [66,67,68]. Hyperparameters were tuned by Bayesian optimisation with a maximum of 50 iterations [69,70,71]. Letting θ denote a hyperparameter combination, Θ the search space, and S C V ( θ ) the spatial-cross-validation performance, the optimal hyperparameters are
θ * = a r g   m a x   θ Θ S C V ( θ ) .
The main search spaces of each model are listed in Table 4. After model selection, the learner type, preprocessing workflow, and tuning budget were fixed across E1–E4, and only the way of introducing national-scale information was varied.

3.7. Model Evaluation and Susceptibility Mapping

Model performance was evaluated from three angles: discrimination ability, threshold-based classification performance, and probability-output quality. Discrimination ability was measured by the area under the ROC curve (AUC) and average precision (AP), which describe how well a model ranks landslide samples above non-landslide samples [72,73]:
A U C = P p ^ + > p ^ ,
A P = k ( R k R k 1 ) P k ,
where p ^ + and p ^ are the predicted probabilities of landslide and non-landslide samples, respectively, and P k and R k are the precision and recall at the k -th threshold of the precision–recall curve.
Under a unified classification threshold τ , threshold-based performance was described by accuracy, precision, recall, and F1-score [74]:
A c c u r a c y = T P + T N T P + T N + F P + F N ,
P r e c i s i o n = T P T P + F P ,   R e c a l l = T P T P + F N ,
F 1 = 2 P r e c i s i o n R e c a l l P r e c i s i o n + R e c a l l .
Probability output quality was assessed using the Brier score, log loss, and expected calibration error (ECE) [75,76,77]:
B r i e r = 1 n i = 1 n ( p ^ i y i ) 2 ,
L o g L o s s = 1 n i = 1 n y i log p ^ i + 1 y i log 1 p ^ i ,
E C E = m = 1 M B m n a c c ( B m ) c o n f ( B m ) .
Here, TP, TN, FP, and FN denote true positives, true negatives, false positives, and false negatives, respectively. For the ECE calculation, the probability interval [0, 1] was divided into M = 10 equal-width bins with edges of 0.0, 0.1, …, and 1.0. The first nine bins followed a lower-closed and upper-open rule, i.e., [lower, upper), whereas the last bin included the upper endpoint, i.e., [0.9, 1.0]. B m   denotes the m -th probability bin, B m is the number of test samples in that bin, a c c ( B m )   denotes the observed landslide frequency, and c o n f ( B m ) denotes the mean predicted susceptibility probability. ECE was calculated as a sample-weighted average using B m / n , and the same bin edges were used for E1–E4. Empty bins, if present, were not plotted in the reliability diagrams and contributed zero to ECE. These metrics correspond to ranking discrimination, classification discrimination, and probabilistic reliability, respectively, and are used to compare the experimental strategies [75,76,77].
During susceptibility mapping, the input factors, scaling, and model-training workflow were kept consistent. The processed raster factors were fed into the model pixel by pixel to obtain continuous landslide susceptibility probabilities for the study area:
p ^ s = f ^ x s ,   s Ω ,
where s is a spatial pixel, Ω is the study-area extent, and x ( s ) is the conditioning-factor vector for that pixel. The continuous probabilities were then divided at intervals of 0.2 into five classes—very low, low, moderate, high, and very high—to represent the spatial differentiation of susceptibility [33]; blue, grey-blue, pale yellow, light orange, and coral red were used to denote these classes in order of increasing susceptibility.
To further examine the spatial credibility of the final susceptibility map, a spatial-level validation was conducted using the independent test landslide points that were not involved in model training or hyperparameter tuning. These landslide points were overlaid on the five susceptibility classes of the final E4 map. For each susceptibility class, the area proportion, number and proportion of independent landslide points, landslide density, and frequency ratio were calculated. The frequency ratio was defined as the ratio between the landslide-point proportion and the area proportion of a given susceptibility class. In addition, representative local susceptibility insets and a UAV photograph for Site E were used to compare the mapped high-susceptibility patches with recorded landslide locations, inferred landslide morphology, and local alpine-gorge geomorphic settings.

4. Results

4.1. Sample, Factor, and Data Partitioning Results

After spatial verification, removal of duplicate points, and 250 m spatial thinning, 273 valid positive samples were retained from the 282 original landslide points in the study area, with a minimum pairwise distance of 256.85 m, satisfying the spatial-independence constraint. Negative samples were drawn from background pixels at which all factors had valid values, after excluding water bodies, built-up land, glaciers, and the 1000 m buffers around landslide points, so as to reduce the chance that potential unlabelled landslides or unsuitable background areas would be sampled as negatives. The final sample set was constructed at a 1:1 positive-to-negative ratio, providing a class-balanced basis for the subsequent binary classification and strategy comparison.
To test the spatial generalisation ability of the models, an independent test set was constructed by holding out entire spatial blocks. The test set comprised 15 blocks of 2 km grid size and 112 samples in total (56 positive and 56 negative); the minimum distance between the test set and the remaining positive samples was 335.8 m, and no sample IDs were duplicated across subsets. After the test set was removed, the remaining samples were used for model training and four-fold spatial outer cross-validation. A 250 m buffer was placed between the training and validation sets in each fold, and samples within the buffer were excluded from the corresponding fold; the numbers of positive and negative samples in each fold were kept equal. The detailed partitioning is shown in Table 5.
Nineteen candidate conditioning factors were then constructed to characterise the topographic relief, valley incision, tectonic and lithological conditions, geomorphological and soil backgrounds, climatic conditions, vegetation cover, land use, and human engineering disturbance of the Parlung Tsangpo Basin. Their spatial distributions are shown in Figure 4: the continuous factors include topographic, climatic, hydrological, vegetation, and distance-related variables, while the categorical factors include lithology, land-use type, soil type, and geomorphology. Overall, the factors exhibit clear spatial heterogeneity within the study area and adequately reflect the complex landslide-conditioning environment of the alpine gorge basin. To further examine whether the negative samples represented only an easily separable low-susceptibility background, we compared the distributions of positive samples, A0 main negative samples, and sampleable background pixels for key conditioning factors (Supplementary Figure S1 and Supplementary Tables S2 and S3). The sampleable background pixels and A0 negative samples were not dominated by very gentle terrain: their median slopes were 21.55° and 20.84°, respectively, whereas the median slope of the positive samples was 14.05°. The A0 negative samples also covered broad ranges of elevation, relief, TPI, NDVI, and distances to rivers, roads, and active faults. These results indicate that the main negative set retained substantial environmental variability rather than representing only valley-floor or uniformly low-susceptibility pixels.
The multicollinearity diagnostics in Table 6 show that the VIF values of the 19 factors range from 1.000 to 7.720 and the tolerance values from 0.130 to 0.997, none reaching the severe-collinearity thresholds of VIF > 10 or TOL < 0.1. This result indicates that the common conditioning-factor set did not contain severe multicollinearity. Therefore, no factor was removed at this screening step, allowing the same 19 conditioning factors to be retained for both candidate-learner comparison and the subsequent RF-based E1–E4 strategy comparison.

4.2. Learner Selection and Overall Performance of the Strategies

To determine the base learner for prior-informed local modelling, the independent-test performance of random forest, XGBoost, L2-regularised logistic regression, and stacking was first compared under the E4 setting (Figure 5c). Random forest achieved the best overall performance in discrimination, classification, and calibration, with AUC, AP, and F1-scores of 0.901, 0.895, and 0.828, respectively, and the lowest log loss and ECE, at 0.401 and 0.055. Although XGBoost, logistic regression, and stacking showed advantages in individual metrics—for example, higher recall—their precision, probability error, or calibration was comparatively weaker. Taking discrimination, misclassification control, and probabilistic reliability together, random forest was selected as the base learner for the subsequent strategy comparison, robustness analysis, and susceptibility mapping. The selected random-forest configuration used for the E4 prior-informed local model was n_estimators = 970, max_depth = None, min_samples_split = 2, min_samples_leaf = 2, and max_features = sqrt. The remaining fixed settings were criterion = gini, bootstrap = True, class_weight = None, and random_state = 42.
On the basis of the unified learner, target-domain test set, and evaluation metrics, the overall performance of the four information-introduction strategies E1–E4 was then compared. The ROC curves show that E4 attained an AUC of 0.901, higher than the 0.848, 0.845, and 0.864 of E1, E2, and E3, indicating stronger sample discrimination for the prior-informed strategy. The remaining metrics confirm that E4 achieved the best AP, F1-score, precision, accuracy, log loss, and ECE, so that it improved not only classification performance but also the quality of the probabilistic predictions.
In terms of the predicted-probability distributions, E4 assigned overall lower probabilities to negative samples while keeping the probabilities of positive samples high, giving clearer separation between the two classes. Compared with E2 (direct national-model transfer), E4 attained higher precision and lower log loss at the same recall, indicating that it reduced misclassification risk and improved the stability of the probability outputs while preserving landslide-detection ability. Overall, the prior-informed local modelling strategy performed best on the independent test set and was therefore adopted as the final strategy for the subsequent robustness analysis, calibration evaluation, and susceptibility mapping.

4.3. Robustness and Sample-Size Sensitivity

To test the robustness of E4 relative to the local baseline E1, repeated spatial validation and sensitivity analysis under different proportions of target-domain training samples were conducted (Figure 6). The repeated-validation results show that E4 exceeded E1 in AUC, AP, and F1-score while reducing the Brier score and log loss, indicating that the prior-informed strategy improved sample discrimination and classification performance as well as the accuracy of the probabilistic predictions.
As the training-sample fraction increased from 25% to 100%, the advantage of E4 over E1 remained broadly stable. Across sample sizes, E4 consistently showed higher discrimination and lower probability error, indicating that the national-scale probability prior provides stable supplementary information when local samples are limited. This further indicates that the improvement of E4 does not depend on a single data partition but reflects good spatial generalisation and sample adaptability.
The alternative negative-sampling tests further support this interpretation (Supplementary Figure S2 and Supplementary Tables S5 and S6). Under the A1 slope-stratified strategy, E4 achieved the highest AUC, AP, and F1-score among E1–E4 and also showed the lowest Brier score and log loss. Under the more demanding A2 slope-overlap hard-negative strategy, the discrimination difference between E4 and E1 became very small, with nearly identical AUC values. However, E4 still achieved a slightly higher F1-score and recall and lower Brier score, log loss, and ECE than E1. Therefore, the advantage of the prior-informed strategy does not rely on separating landslides from an artificially easy, gentle-slope background. Under harder negative-sampling conditions, its main benefit lies in maintaining comparable discrimination while improving the reliability of the predicted probabilities.

4.4. Probability Calibration Performance

To evaluate the reliability of the susceptibility probabilities produced by the different strategies, the calibration performance of E1–E4 was compared on the same spatially independent target-domain test set ((N = 112), including 56 landslide and 56 non-landslide samples; Figure 7). The reliability diagrams were constructed using the same 10 equal-width probability bins as the ECE calculation, and the sample size of each bin is shown in the diagram to aid interpretation. The predicted probabilities of E4 lie closest overall to the perfect-calibration line, indicating better agreement between the predicted susceptibility probabilities and the observed landslide frequency.

4.5. Landslide Susceptibility Mapping and Spatial-Level Validation Results

The susceptibility maps and study-area-wide predicted-probability distributions of the four strategies are shown in Figure 8. All four can identify high-susceptibility areas along gullies, slopes, and linear geomorphological units, but they differ clearly in spatial expression. E2 shows overall higher predicted probabilities, with moderate- and high-susceptibility areas distributed more diffusely, indicating that direct national-model transfer may overestimate risk in the target domain. By contrast, E4 expresses the low-susceptibility background more fully and concentrates its high-susceptibility patches more tightly, suppressing large-area overestimation while retaining local high-risk areas.
The probability-distribution statistics are consistent with the spatial patterns. The mean and median predicted probabilities of E4 are lower than those of the other strategies, and its share of low-probability area is the highest, while a small number of tightly concentrated very-high-susceptibility areas are still retained. Overall, E4 exhibits a mapping signature of an extensive low-susceptibility background with concentrated high-susceptibility patches, in line with the independent-test performance and calibration results above, and indicating that the prior-informed model generates more stable and more spatially discriminating susceptibility maps.
To further verify the spatial reliability of the final susceptibility map, the independent test landslide points were overlaid on the E4 susceptibility classes, and the spatial concentration of landslides in each class was quantified (Table 7 and Figure 9). The high and very high susceptibility classes occupied only 3.49% of the study area, but contained 40 of the 56 independent test landslide points, accounting for 71.43% of the independent landslide occurrences. In contrast, the very low class occupied 87.95% of the study area but contained only two independent landslide points. Landslide density increased markedly from 0.008 points/100 km2 in the very low class to 7.378 points/100 km2 in the very high class. The frequency ratio also increased monotonically from 0.041 in the very low class to 39.602 in the very high class. The landslide density of the high and very high classes was 107.46 times that of the very low and low classes combined. These results indicate that the E4 susceptibility map has clear map-level discriminatory ability and that the mapped high-susceptibility zones effectively concentrate independent landslide occurrences.
Representative local sites provide further site-scale evidence for the susceptibility zoning (Figure 9). Sites A–E are located within mapped high-probability susceptibility patches, with E4-predicted probabilities of 0.959, 0.958, 0.933, 0.889, and 0.818, respectively. These sites are mainly distributed along the deeply incised valley system and local slope-transition zones, where steep valley-side slopes, river erosion, and engineering corridors may jointly contribute to slope instability. The enlarged local susceptibility maps show that the recorded landslide locations are embedded within or adjacent to continuous high- and very-high-susceptibility belts. For Site E, the local susceptibility inset is further compared with the UAV photograph: the inferred landslide perimeter and movement direction correspond to a steep valley-side slope section that is mapped as high to very high susceptibility. These representative examples indicate that the final E4 susceptibility map is not only statistically reliable but also spatially consistent with recorded landslide occurrence, visible slope-failure morphology, and plausible alpine-gorge geomorphic controls.

5. Discussion

5.1. Local Adaptation of National-Scale Information: From Rigid Transfer to a Probability Prior

This study shows that national-scale landslide information can improve local susceptibility modelling in data-scarce basins, but that its effectiveness depends on how it is introduced. For alpine gorge regions such as the Parlung Tsangpo Basin, national-scale information is best not used directly as a fixed model, nor forcibly transferred as source-domain samples; rather, it is better transformed into a probability prior that can be relearned and recalibrated by local samples. In other words, the national model does not directly determine the final susceptibility pattern of the target basin but provides the local model with a background risk judgement that can be selected, corrected, and reweighted.
Traditional susceptibility studies rely on the statistical relationships between a target-area landslide inventory and conditioning factors, and their reliability is strongly affected by inventory completeness, sample size, spatial representativeness, and validation method [33,78]. Incomplete or systematically biassed inventories have been shown to affect statistical susceptibility models and may distort spatial predictions [79]. With only 273 valid positive samples retained, the Parlung Tsangpo Basin cannot fully cover the diverse conditioning environments shaped by deeply incised valleys, complex lithology, tectonic fracturing, fluvial erosion, and engineering disturbance. Training a model on local samples alone can therefore leave the results dominated by a limited sample structure, reducing the stability of the susceptibility maps.
In recent years, global-, national-, or large-regional-scale susceptibility studies have provided important supplementary information for local assessment [53,80]. However, national-scale landslide–environment relationships are learned from mixed geomorphic settings and triggering contexts, and they cannot be assumed to be directly equivalent to the controlling mechanisms of a specific alpine gorge basin. This issue is particularly relevant here because the source domain includes landslides developed under diverse terrain, climatic, tectonic, and anthropogenic conditions, whereas the Parlung Tsangpo Basin is dominated by deeply incised valleys, active tectonics, lithological contrasts, glacier- and snowmelt-related water supply, freeze–thaw processes, monsoonal rainfall, and local engineering disturbance.
Such source–target differences may introduce negative transfer when national information is introduced in a rigid form. In E2, the national model is applied directly to the target basin, so the discriminative rules learned from heterogeneous source-domain samples directly determine the target-domain predictions. If part of these rules reflects hilly terrain landslides or engineering-related failures rather than alpine-gorge instability, the transferred model may produce biassed probabilities or diffuse high-susceptibility areas. In E3, assigning higher weights to target-domain samples can reduce this problem [23,24], but sample-level joint training still allows source-domain patterns with different triggering or developmental mechanisms to enter the decision boundary. Therefore, the weaker performance of E2 and E3 does not simply indicate that national information is useless; rather, it shows that the form in which source-domain information is transferred is critical.
The uneven spatial distribution of the national inventory further helps explain this result. Although the source-domain records are national in extent, Figure 2 shows that they are much denser in eastern and central China than in western high-mountain regions, including the Tibetan Plateau. Therefore, the national model may be trained mainly on environmental settings that are better represented in the inventory, whereas the target alpine gorge conditions receive fewer source-domain samples. This imbalance can limit the direct applicability of E2 and may also dilute local information when national and target samples are simply merged in E3.
The contrast between E2 and E4 further clarifies both the value and the limitation of the probability-prior strategy. On the one hand, the improvement of E4 relative to the local baseline indicates that the national model contains useful broad-scale information on landslide-prone environmental backgrounds. On the other hand, the poorer calibration and more diffuse susceptibility pattern of E2 show that the national model may also carry source-domain bias and should not be used as a fixed national predictor. In E4, this probability is implemented as an additional input variable rather than as the final prediction. Therefore, local correction is achieved only through the relearning process: the target-domain model can downweight or reshape the prior when local samples and the original conditioning factors provide contradictory evidence. If the national model assigns high probabilities to locally stable terrain, or low probabilities to locally unstable terrain, and such contradictory cases are insufficiently represented in the target-domain samples, E4 may not fully remove the bias and may even propagate part of it. Thus, the proposed strategy should be interpreted as a locally constrained recalibration mechanism rather than as a guarantee that national-model bias is completely eliminated. The better calibration, lower probability error, and less diffuse susceptibility pattern of E4 suggest that, in the present basin, the available local samples were sufficient to correct part of the national-model bias while retaining useful broad-scale information.
Selecting source-domain samples that are environmentally similar to the target basin is also a reasonable way to reduce domain mismatch, and it is closely related to domain adaptation and area-of-applicability analysis [9,10,21]. Nevertheless, using similarity-based source selection as the main strategy would introduce additional methodological choices, including the definition of similarity metrics, thresholds, and factor weights. It could also reduce the diversity of national-scale landslide-background information and exclude some less frequent but potentially relevant alpine or tectonic cases. For this reason, the present study retained the full national source domain and used E2 and E3 as diagnostic comparisons for direct and sample-level transfer, while E4 was designed to transform national information into a locally correctable probability prior. The results are consistent with the interpretation that probability-prior transfer can weaken the misleading influence of source-domain samples whose controlling mechanisms differ from those of the target basin, without discarding the useful broad-scale information contained in the national inventory.

5.2. Discrimination, Probability Calibration, and Mapping Reliability

Landslide susceptibility assessment requires not only that a model distinguish landslide from non-landslide samples, but also that its output probabilities support continuous mapping and class division. AUC, AP, and F1-score reflect ranking and classification ability but cannot reveal whether the predicted probabilities match empirical occurrence frequencies [81]. For studies whose final products are susceptibility-probability maps, probabilistic reliability should therefore be assessed with the Brier score, log loss, and ECE in addition to the discrimination metrics.
On the spatially independent test set, the AUC, AP, F1-score, and accuracy of E4 were 0.901, 0.895, 0.828, and 0.821, respectively, generally better than those of E1, E2, and E3. Recent susceptibility studies in mountainous and gorge regions have shown that machine-learning models can usually attain high discrimination, but that performance is affected by study scale, inventory quality, negative-sample construction, and validation method [34,81,82,83,84]; AUC values from different regional studies should therefore not be ranked in an absolute sense. The significance of the present results is that, under a unified factor system, a common test set, and spatially independent validation, E4 still outperformed the other three strategies consistently, indicating that the improvement stems mainly from expressing national-scale information as a probability prior rather than from differences in the evaluation settings.
In terms of probabilistic reliability, the Brier score, log loss, and ECE of E4 fell to 0.127, 0.401, and 0.055, respectively—the best of the four strategies. A recent study of geological-hazard susceptibility along a railway corridor also used the Brier score and ECE to assess probability reliability, with its random-forest model achieving AUC = 0.85, Brier score = 0.1589, and ECE = 0.0572 on the validation set [85]. Compared with such studies, E4 maintained high discrimination while achieving lower probability error, indicating that its continuous susceptibility probabilities are better suited to spatial zoning and the identification of key patches.
The contrast between E2 and E4 further shows that a higher landslide-detection rate does not guarantee reliable mapping. At equal recall, the precision, log loss, and ECE of E2 were all inferior to those of E4, indicating that direct national-model transfer is prone to probability bias and misclassification. By relearning and recalibrating the national probability prior with local samples, E4 improved probability credibility while preserving landslide-detection ability. Its advantage is therefore not merely that it classifies more accurately, but that it outputs more reliable continuous susceptibility probabilities, providing a more robust basis for susceptibility class division, hazard screening, and risk interpretation.

5.3. Geomorphological Interpretation of the Spatial Susceptibility Pattern

According to Figure 8, the susceptibility map produced by E4 displays an extensive low-susceptibility background with concentrated high-susceptibility patches. The high-susceptibility areas occur mainly on both sides of the main valley, at tributary confluences, on strongly incised valley slopes, and near transportation corridors, indicating that the prior-informed model does not simply amplify the national-scale risk background but confines the high-probability areas to slope units with clear conditioning factors.
The spatial-level validation further supports this geomorphological interpretation. Although the high and very high susceptibility classes account for only 3.49% of the study area, they contain 71.43% of the independent test landslide points, and their landslide density is 107.46 times that of the very low and low classes combined. This concentration pattern confirms that the high-probability patches identified by E4 are not randomly distributed, but are closely associated with observed landslide occurrence and key alpine-gorge conditioning environments.
Figure 9 provides site-scale visual evidence for the spatial validation results. In the enlarged maps, the recorded landslide locations of Sites A–E are located within or close to high-probability patches along valley-side slopes. At Site E, the UAV image shows a slope-failure area, and its inferred perimeter and movement direction are consistent with the mapped high- and very-high-susceptibility belt. These observations support the geomorphic plausibility of the E4 susceptibility map.
This pattern is consistent with current understanding of the southeastern Tibetan Plateau and adjacent alpine gorge regions. Studies in the Yarlung Tsangpo River Basin and the Three Parallel Rivers region have shown that landslide occurrence is controlled not by a single factor but by the long-term coupling of tectonic uplift, fluvial incision, glacial processes, and climatic weathering [16,86]. In the Parlung Tsangpo Basin, active faults continuously fracture and weaken rock masses, creating structural discontinuities that reduce slope strength. At the same time, rapid river incision and lateral erosion by the Parlung Tsangpo River and its tributaries remove support at slope toes, promote stress redistribution, and accelerate gravitational deformation of steep valley-side slopes. These tectonic and fluvial processes jointly generate highly unstable topographic conditions that favour landslide initiation.
In addition, the basin is located in a transition zone strongly influenced by modern glaciers and seasonal cryospheric processes. Repeated freeze–thaw cycles can enlarge existing fractures, increase rock-mass disintegration, and facilitate the detachment of unstable material from steep rock slopes. Glacier retreat and meltwater infiltration may further alter slope hydrology and reduce the mechanical stability of weathered materials. Consequently, fault-controlled rock weakening, river erosion, glacial activity, and freeze–thaw weathering act together to control the spatial concentration of landslides in the basin. The high-susceptibility patches identified by E4 are mainly distributed along deeply incised valleys, tributary junctions, and fault-influenced slope sections, which is consistent with this coupled geomorphic and geological control mechanism.
The high susceptibility near transportation corridors is also regionally consistent. A study of the Yunnan–Tibet transportation corridor found that high- and very-high-susceptibility areas were distributed mainly along major river valleys and highway belts, reflecting the superimposed effects of gorge terrain and road-engineering disturbance [87]. The present results likewise show that some high-susceptibility patches concentrate in valley–road-adjacent areas, indicating that the mapped pattern is consistent with spatial associations between natural geomorphic processes and human disturbance. Together with existing cascading-hazard studies in the Parlung Tsangpo and Yarlung Tsangpo Grand Canyon regions, the coupling of glaciers, freeze–thaw processes, avalanches, channel transport, and river blockage further explains the high susceptibility in the main-valley and tributary areas [88,89].
By comparison, the high-probability areas of E2 are more diffuse, indicating that direct national-model transfer tends to project broad-scale average patterns onto the target basin and raise the background susceptibility. By recalibrating the national probability prior with local samples, E4 concentrates the high-susceptibility areas where landslides actually cluster and where the geomorphic setting is plausible. Its maps therefore offer not only better statistical performance but also clearer geomorphological interpretability, providing a more stable spatial basis for identifying key slope sections, screening transportation-corridor risks, and supporting engineering route selection in alpine gorge basins.

5.4. Limitations and Future Work

Although E4 performed well in discrimination, calibration, and spatial mapping, this study has several limitations. First, the target-domain inventory is small and the samples participate in modelling mainly as points. The completeness, spatial accuracy, and representation of an inventory affect susceptibility-model training and zoning [90,91], and in alpine gorge regions such as the Parlung Tsangpo Basin, remote-sensing interpretation is further hampered by topographic shadow, vegetation cover, bare rock, and valley sediments. Future work should combine high-resolution remote sensing, unmanned-aerial-vehicle surveys, InSAR monitoring, and field verification to improve the inventory and should gradually extend point samples to polygon-based inventories so as to represent landslide boundaries, scales, and types more accurately.
Second, this study mainly evaluates spatial landslide susceptibility. The conditioning factors used in the model are dominated by relatively stable spatial variables, including topography, geology, geomorphology, vegetation, land cover, and human activity. Annual mean temperature was retained as a climatic background variable to describe the regional thermal setting and elevation-related climatic gradients. Rainfall and freeze–thaw processes were not included as independent dynamic predictors in the final model. The main reason is that the landslide inventory used in this study consists mainly of multi-year historical records and interpreted landslide points, which are difficult to match reliably with specific rainfall events or freeze–thaw periods. Annual precipitation was examined during the initial factor preparation stage, but it was not retained because it showed spatial redundancy with the retained climatic and topographic background variables. In addition, annual precipitation totals provide limited information on short-duration or extreme rainfall events that commonly trigger landslides. Future work should incorporate event-dated landslide inventories, time-series rainfall and extreme-precipitation indices, freeze–thaw frequency or intensity indicators, glacier/snow-cover change or distance-to-glacier/moraine variables, ground-motion parameters, fluvial-erosion rates, and records of engineering activity, thereby extending static susceptibility assessment towards dynamic susceptibility or hazard assessment [88,89,92,93].
Third, this study evaluates the proposed strategy only in the Parlung Tsangpo Basin. Although this basin is suitable for testing how national-scale information can be transformed and locally recalibrated under sample-limited alpine-gorge conditions, the results should be interpreted as evidence from a representative target basin rather than as proof of cross-basin general applicability. Model generalisation can easily be overestimated if spatial autocorrelation and independent validation are ignored [47], whereas unified datasets and reproducible benchmark workflows are important for improving the comparability of cross-regional results [84]. Future work should therefore conduct external tests in multiple representative basins, combining spatial cross-validation, independent-basin validation, and unified benchmark workflows to delimit the applicability boundaries of the national-probability-prior and local-sample-correction strategy more precisely.
Finally, this study has not yet fully decomposed the sources of uncertainty in the susceptibility results. Subsequent research could evaluate the effects of inventory error, negative-sample construction, factor resolution, model structure, calibration, and regional-transfer differences, and could express prediction confidence through multi-model ensembles, bootstrap resampling, spatial-blocking sensitivity analysis, and uncertainty mapping, thereby providing a more robust basis for disaster prevention and mitigation, engineering route selection, and the screening of priority hazards.

6. Conclusions

To address the cross-scale problem of incorporating national landslide information into local susceptibility assessment under limited target-domain samples, this study proposed a prior-informed modelling strategy in which national-scale susceptibility probabilities are transformed into probabilistic knowledge and then relearned and corrected using local samples. This design avoids treating the national model as directly transferable and instead allows the target-domain model to adapt external knowledge to the geomorphic and environmental conditions of the Parlung Tsangpo alpine gorge basin.
The experimental results demonstrate that the prior-informed strategy achieved the most robust performance among the tested modelling pathways. On the independent test set, E4 obtained the highest AUC, AP, F1-score, and accuracy, with values of 0.901, 0.895, 0.828, and 0.821, respectively. Compared with the local baseline, its AUC and AP increased by 0.053 and 0.052, indicating that the national-scale probability prior improved both discrimination ability and prediction reliability under sample-limited conditions. The low log loss, Brier score, and ECE further suggest that the proposed strategy produced better-calibrated susceptibility probabilities rather than only improving classification accuracy.
The susceptibility maps further show that E4 maintained a low-probability background while concentrating high-susceptibility zones along the main valley, tributary confluences, strongly incised slopes, and transportation corridors. This spatial pattern is consistent with the combined effects of deep valley incision, tectonic and lithological contrasts, fluvial erosion, and local engineering disturbance in the Parlung Tsangpo Basin. Therefore, the prior-informed strategy not only improved quantitative prediction performance but also enhanced the geomorphic interpretability and practical mapping value of the susceptibility results.
Future work should further test the transferability of this national-probability-prior and local-sample-correction framework in other alpine gorge basins with different geomorphic settings and landslide inventories. In addition, integrating multi-temporal remote sensing observations, dynamic triggering factors such as rainfall and seismicity, and uncertainty-aware modelling may help extend the present static susceptibility framework towards more transferable and operational landslide hazard assessment.
In addition, future studies could explicitly incorporate source-domain similarity screening or area-of-applicability analysis before transfer. Stratifying the national inventory by geomorphic setting, triggering mechanism, or human-disturbance intensity may help determine which types of source-domain samples are most transferable to alpine gorge basins, and would further clarify the applicability boundary of the proposed prior-informed framework.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/rs18132167/s1, Figure S1: Distributions of selected conditioning factors for positive samples, A0 main negative samples, and sampleable background pixels; Figure S2: Sensitivity of E1–E4 to alternative negative-sampling strategies; Table S1: Geological symbols and corresponding lithostratigraphic or lithological descriptions of the geological units shown in Figure 1c and Figure 4p; Table S2: Distributions of continuous conditioning factors for positive samples, A0 main negative samples, and sampleable background pixels; Table S3: Distributions of categorical conditioning factors for positive samples, A0 main negative samples, and sampleable background pixels; Table S4: Soil-type codes and corresponding soil classes; Table S5: Model performance under the main and alternative negative-sampling strategies; Table S6: Performance differences between E4 and E1 under different negative-sampling strategies.

Author Contributions

Conceptualization, Y.G. and X.D.; methodology, Y.G. and X.D.; software, Y.G.; validation, Y.G., Y.L., Z.J. and X.D.; formal analysis, Y.G.; investigation, Y.G. and Z.J.; resources, X.D., W.L. and G.Q.; data curation, Y.G.; writing—original draft preparation, Y.G.; writing—review and editing, Y.G., Y.L., Z.J., G.Q., W.L. and X.D.; visualisation, Y.G. and Z.J.; supervision, X.D.; project administration, Y.G. and X.D.; funding acquisition, Y.G. and X.D. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the National College Students Innovation and Entrepreneurship Training Program (grant numbers 202510616011 and D202604040948415079), the Open Project of the Middle Yarlung Zangbo River Natural Resources Observation and Research Station (grant number 2024YJZKF005), the Spatial Information Acquisition and Application Joint Laboratory of Anhui Province (grant number 2024tlxykjxx02), and the Project of China Power Engineering Consulting Group Northwest Engineering Co., Ltd. (grant number XBY-ZDKJ-2023-9).

Data Availability Statement

The processed data and derived products supporting the findings of this study are openly available in Zenodo at https://doi.org/10.5281/zenodo.20595741. The deposited dataset includes the data and supporting materials used for landslide susceptibility modelling, model evaluation, and susceptibility mapping in the Parlung Tsangpo Basin. Some original third-party geospatial datasets used to derive the conditioning factors are not redistributed in raw form and should be obtained from the original data providers cited in the manuscript.

Conflicts of Interest

The authors declare that they have no financial or personal relationships with other people or organisations that could inappropriately influence this work. The authors also declare that there is no professional or other personal interest of any nature or kind regarding any product, service, or company that could be construed as influencing the position presented in, or the interpretation of, the manuscript. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results.

Abbreviations

The following abbreviations are used in this manuscript:
APAverage precision
AUCArea under the receiver operating characteristic curve
DEMDigital elevation model
E1Local baseline model
E2Direct national-model transfer model
E3Weighted source–target joint training model
E4Prior-informed local model
ECEExpected calibration error
ESAEuropean Space Agency
F1F1-score
FNFalse negative
FPFalse positive
GLO-30Copernicus global 30 m digital elevation model product
L2L2 regularisation
LRLogistic regression
NDVINormalised difference vegetation index
RFRandom forest
ROCReceiver operating characteristic
SVM-RBFSupport vector machine with a radial basis function kernel
TNTrue negative
TOLTolerance
TPTrue positive
TPITopographic position index
TWITopographic wetness index
VIFVariance inflation factor
XGBoostExtreme gradient boosting
YTHPYarlung Tsangpo Hydropower Project

References

  1. Alcántara-Ayala, I. Landslides in a changing world. Landslides 2025, 22, 2851–2865. [Google Scholar] [CrossRef]
  2. Yang, W.; Wang, Z.; An, B.; Chen, Y.; Zhao, C.; Li, C.; Wang, Y.; Wang, W.; Li, J.; Wu, G.; et al. Early warning system for ice collapses and river blockages in the Sedongpu Valley, southeastern Tibetan Plateau. Nat. Hazards Earth Syst. Sci. 2023, 23, 3015–3029. [Google Scholar] [CrossRef]
  3. Deng, Y.; Gao, Q.; Wang, X.; Fan, X. A large-scale rock avalanche-debris flow cascading hazard in the Sedongpu catchment, southeastern Tibetan Plateau. Landslides 2025, 22, 109–120. [Google Scholar] [CrossRef]
  4. Stanley, T.A.; Soobitsky, R.B.; Amatya, P.M.; Kirschbaum, D.B. Landslide hazard is projected to increase across High Mountain Asia. Earths Future 2024, 12, e2023EF004325. [Google Scholar] [CrossRef]
  5. Wang, X.; Wang, Y.; Lin, Q.; Yang, X. Assessing global landslide casualty risk under moderate climate change based on multiple GCM projections. Int. J. Disaster Risk Sci. 2023, 14, 751–767. [Google Scholar] [CrossRef]
  6. Sajadi, P.; Sang, Y.-F.; Gholamnia, M.; Bonafoni, S.; Mukherjee, S. Evaluation of the landslide susceptibility and its spatial difference in the whole Qinghai-Tibetan Plateau region by five learning algorithms. Geosci. Lett. 2022, 9, 9. [Google Scholar] [CrossRef]
  7. Merghadi, A.; Yunus, A.P.; Dou, J.; Whiteley, J.; ThaiPham, B.; Bui, D.T.; Avtar, R.; Abderrahmane, B. Machine learning methods for landslide susceptibility studies: A comparative overview of algorithm performance. Earth Sci. Rev. 2020, 207, 103225. [Google Scholar] [CrossRef]
  8. Tan, J.; Yang, C.; Wang, Y.; Xiong, H.; Ma, C. A hybrid model to overcome landslide inventory incompleteness issue for landslide susceptibility prediction. Geocarto Int. 2024, 39, 2322066. [Google Scholar] [CrossRef]
  9. Hong, H.; Wang, D.; Zhu, A.-X.; Wang, Y. Landslide susceptibility mapping based on the reliability of landslide and non-landslide sample. Expert Syst. Appl. 2024, 243, 122933. [Google Scholar] [CrossRef]
  10. Wang, Z.; Goetz, J.; Brenning, A. Transfer learning for landslide susceptibility modeling using domain adaptation and case-based reasoning. Geosci. Model Dev. 2022, 15, 8765–8784. [Google Scholar] [CrossRef]
  11. Meyer, H.; Pebesma, E. Predicting into unknown space? Estimating the area of applicability of spatial prediction models. Methods Ecol. Evol. 2021, 12, 1620–1633. [Google Scholar] [CrossRef]
  12. Singh, A.; Dhiman, N.; Niraj, K.C.; Shukla, D.P. Ensembled transfer learning approach for error reduction in landslide susceptibility mapping of the data scare region. Sci. Rep. 2024, 14, 29060. [Google Scholar] [CrossRef] [PubMed]
  13. Wolpert, D.H. Stacked generalization. Neural Netw. 1992, 5, 241–259. [Google Scholar] [CrossRef]
  14. Guo, C.; Pleiss, G.; Sun, Y.; Weinberger, K.Q. On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning, Sydney, NSW, Australia, 6–11 August 2017; Proceedings of Machine Learning Research—PMLR: Sydney, Australia, 2017; Volume 70, pp. 1321–1330. [Google Scholar]
  15. Hu, K.; Wu, C.; Wei, L.; Zhang, X.; Zhang, Q.; Liu, W.; Yanites, B.J. Geomorphic effects of recurrent outburst superfloods in the Yigong River on the southeastern margin of Tibet. Sci. Rep. 2021, 11, 15577. [Google Scholar] [CrossRef] [PubMed]
  16. Zhao, B.; Su, L.; Wang, Y.; Li, W.; Wang, L. Insights into some large-scale landslides in the southeastern margin of the Qinghai-Tibet Plateau. J. Rock Mech. Geotech. Eng. 2023, 15, 1960–1985. [Google Scholar] [CrossRef]
  17. Zhao, B.; Su, L. Complex spatial and size distributions of landslides in the Yarlung Tsangpo River (YTR) Basin. J. Rock Mech. Geotech. Eng. 2025, 17, 897–914. [Google Scholar] [CrossRef]
  18. Guo, C.; Montgomery, D.R.; Zhang, Y.; Zhong, N.; Fan, C.; Wu, R.; Yang, Z.; Ding, Y.; Jin, J.; Yan, Y. Evidence for repeated failure of the giant Yigong landslide on the edge of the Tibetan Plateau. Sci. Rep. 2020, 10, 14371. [Google Scholar] [CrossRef] [PubMed]
  19. Delaney, K.B.; Evans, S.G. The 2000 Yigong landslide (Tibetan Plateau), rockslide-dammed lake and outburst flood: Review, remote sensing analysis, and process modelling. Geomorphology 2015, 246, 377–393. [Google Scholar] [CrossRef]
  20. Abdelkader, M.M.; Csámer, Á. Comparative assessment of machine learning models for landslide susceptibility mapping: A focus on validation and accuracy. Nat. Hazards 2025, 121, 10299–10321. [Google Scholar] [CrossRef]
  21. Wang, H.; Wang, L.; Zhang, L. Transfer learning improves landslide susceptibility assessment. Gondwana Res. 2023, 123, 238–254. [Google Scholar] [CrossRef]
  22. Zhang, W.; Liu, S.; Wang, L.; Sun, W.; Zhang, Y.; Nie, W. Improvement of large-scale-region landslide susceptibility mapping accuracy by transfer learning. J. Cent. South Univ. 2024, 31, 3823–3837. [Google Scholar] [CrossRef]
  23. Zhiyong, F.; Changdong, L.; Wenmin, Y. Landslide susceptibility assessment through TrAdaBoost transfer learning models using two landslide inventories. CATENA 2023, 222, 106799. [Google Scholar] [CrossRef]
  24. Dai, W.; Yang, Q.; Xue, G.-R.; Yu, Y. Boosting for transfer learning. In Proceedings of the 24th International Conference on Machine Learning, Corvallis, OR, USA, 20–24 June 2007; ACM: New York, NY, USA, 2007; pp. 193–200. [Google Scholar] [CrossRef]
  25. Phillips, R.V.; Van Der Laan, M.J.; Lee, H.; Gruber, S. Practical considerations for specifying a super learner. Int. J. Epidemiol. 2023, 52, 1276–1285. [Google Scholar] [CrossRef] [PubMed]
  26. Tian, Z.; Wang, Y.; Fang, Z.; Wang, Y.; Du, B. Spatiotemporal landslide susceptibility modeling based on integrated transfer learning. Nat. Hazards 2025, 121, 22403–22427. [Google Scholar] [CrossRef]
  27. Claverie, M.; Ju, J.; Masek, J.G.; Dungan, J.L.; Vermote, E.F.; Roger, J.-C.; Skakun, S.V.; Justice, C. The Harmonized Landsat and Sentinel-2 surface reflectance data set. Remote Sens. Environ. 2018, 219, 145–161. [Google Scholar] [CrossRef]
  28. He, X.; Zhou, C.; Gao, M.; Sun, S.; Lyu, C.; Han, X. A method for preserving three spatial features in the upscaling of categorical raster data. Comput. Geosci. 2025, 201, 105933. [Google Scholar] [CrossRef]
  29. Keys, R. Cubic convolution interpolation for digital image processing. IEEE Trans. Acoust. Speech Signal Process. 1981, 29, 1153–1160. [Google Scholar] [CrossRef]
  30. Cerda, P.; Varoquaux, G.; Kégl, B. Similarity encoding for learning with dirty categorical variables. Mach. Learn. 2018, 107, 1477–1494. [Google Scholar] [CrossRef]
  31. Nkikabahizi, C.; Cheruiyot, W.; Kibe, A. Chaining zscore and feature scaling methods to improve neural networks for classification. Appl. Soft Comput. 2022, 123, 108908. [Google Scholar] [CrossRef]
  32. Niño-Adan, I.; Landa-Torres, I.; Portillo, E.; Manjarres, D. Influence of statistical feature normalisation methods on k-nearest neighbours and k-means in the context of Industry 4.0. Eng. Appl. Artif. Intell. 2022, 111, 104807. [Google Scholar] [CrossRef]
  33. Reichenbach, P.; Rossi, M.; Malamud, B.D.; Mihir, M.; Guzzetti, F. A review of statistically-based landslide susceptibility models. Earth Sci. Rev. 2018, 180, 60–91. [Google Scholar] [CrossRef]
  34. Chicas, S.D.; Li, H.; Mizoue, N.; Ota, T.; Du, Y.; Somogyvári, M. Landslide susceptibility mapping core-base factors and models’ performance variability: A systematic review. Nat. Hazards 2024, 120, 12573–12593. [Google Scholar] [CrossRef]
  35. Du, G.; Zhang, Y.; Gu, L.; Yang, Z.; Ren, S.; Yuan, S. A systematic approach to landslide susceptibility assessment in regions of rapid geomorphic evolution: A case study of the Yarlung Zangbo Grand Bend. Bull. Eng. Geol. Environ. 2025, 84, 581. [Google Scholar] [CrossRef]
  36. Yu, G.; Lu, J.; Li, Z.; Hou, W. Geomorphic effects of debris flows in high mountain areas of the Parlung Zangbo Basin, southeast Tibet under the influence of climate change. Acta Geogr. Sin. 2022, 77, 619–634. [Google Scholar] [CrossRef]
  37. Amatulli, G.; McInerney, D.; Sethi, T.; Strobl, P.; Domisch, S. Geomorpho90m, empirical evaluation and accuracy assessment of global high-resolution geomorphometric layers. Sci. Data 2020, 7, 162. [Google Scholar] [CrossRef] [PubMed]
  38. Qin, C.-Z.; Zhu, A.-X.; Pei, T.; Li, B.-L.; Scholten, T.; Behrens, T.; Zhou, C.-H. An approach to computing topographic wetness index based on maximum downslope gradient. Precis. Agric. 2011, 12, 32–43. [Google Scholar] [CrossRef]
  39. Zhang, Z.; Jiang, Y.; Lu, X.; Ren, T.; Lu, H.; Zhao, W.; Liu, Y. Landslide susceptibility assessment and analysis of main controlling factors in Ranwu–Tongmai section of Parlung Zangbo, Xizang. J. Glaciol. Geocryol. 2025, 47, 1401–1415. [Google Scholar] [CrossRef]
  40. Wang, D.; Hao, M.; Chen, S.; Meng, Z.; Jiang, D.; Ding, F. Assessment of landslide susceptibility and risk factors in China. Nat. Hazards 2021, 108, 3045–3059. [Google Scholar] [CrossRef]
  41. He, Q.; Wu, S.; Zhao, X.; Hui, Z.; Wang, Z.; Tsangaratos, P.; Ilia, I.; Chen, W.; Chen, Y.; Hao, Y. Evaluation of landslide susceptibility of mountain highway based on RF and SVM models. Sci. Rep. 2025, 15, 24991. [Google Scholar] [CrossRef] [PubMed]
  42. Dormann, C.F.; Elith, J.; Bacher, S.; Buchmann, C.; Carl, G.; Carré, G.; Marquéz, J.R.G.; Gruber, B.; Lafourcade, B.; Leitão, P.J.; et al. Collinearity: A review of methods to deal with it and a simulation study evaluating their performance. Ecography 2013, 36, 27–46. [Google Scholar] [CrossRef]
  43. Kim, J.H. Multicollinearity and misleading statistical results. Korean J. Anesthesiol. 2019, 72, 558–569. [Google Scholar] [CrossRef] [PubMed]
  44. O’Brien, R.M. A caution regarding rules of thumb for variance inflation factors. Qual. Quant. 2007, 41, 673–690. [Google Scholar] [CrossRef]
  45. Wang, Y.; Khodadadzadeh, M.; Zurita-Milla, R. Spatial+: A new cross-validation method to evaluate geospatial machine learning models. Int. J. Appl. Earth Obs. Geoinf. 2023, 121, 103364. [Google Scholar] [CrossRef]
  46. Agboola, G.; Beni, L.H.; Elbayoumi, T.; Thompson, G. Optimizing landslide susceptibility mapping using machine learning and geospatial techniques. Ecol. Inform. 2024, 81, 102583. [Google Scholar] [CrossRef]
  47. Kumar, C.; Walton, G.; Santi, P.; Luza, C. Random cross-validation produces biased assessment of machine learning performance in regional landslide susceptibility prediction. Remote Sens. 2025, 17, 213. [Google Scholar] [CrossRef]
  48. Gu, T.; Duan, P.; Wang, M.; Li, J.; Zhang, Y. Effects of non-landslide sampling strategies on machine learning models in landslide susceptibility mapping. Sci. Rep. 2024, 14, 7201. [Google Scholar] [CrossRef] [PubMed]
  49. Liu, L.-L.; Xiao, H.; Zhang, Y.-L.; Yang, C. An improved buffer-controlled sampling strategy for landslide susceptibility assessment considering the spatial heterogeneity of conditioning factors. Bull. Eng. Geol. Environ. 2024, 83, 512. [Google Scholar] [CrossRef]
  50. Sun, D.; Chen, D.; Zhang, J.; Mi, C.; Gu, Q.; Wen, H. Landslide susceptibility mapping based on interpretable machine learning from the perspective of geomorphological differentiation. Land 2023, 12, 1018. [Google Scholar] [CrossRef]
  51. Lin, Q.; Lima, P.; Steger, S.; Glade, T.; Jiang, T.; Zhang, J.; Liu, T.; Wang, Y. National-scale data-driven rainfall induced landslide susceptibility mapping for China by accounting for incomplete landslide data. Geosci. Front. 2021, 12, 101248. [Google Scholar] [CrossRef]
  52. Zhang, G.; Liu, Y.; Chen, Z.; Xu, Z.; Yuan, Y.; Wang, S.; Lian, W.; Xu, H.; Ding, Z.; Wang, R. Production and analysis of a landslide susceptibility map covering entire China. Remote Sens. 2025, 17, 1615. [Google Scholar] [CrossRef]
  53. Fu, Y.; Fan, Z.; Li, X.; Wang, P.; Sun, X.; Ren, Y.; Cao, W. The influence of non-landslide sample selection methods on landslide susceptibility prediction. Land 2025, 14, 722. [Google Scholar] [CrossRef]
  54. Wu, B.; Shi, Z.; Zheng, H.; Peng, M.; Meng, S. Impact of sampling for landslide susceptibility assessment using interpretable machine learning models. Bull. Eng. Geol. Environ. 2024, 83, 461. [Google Scholar] [CrossRef]
  55. Barella, C.F.; Zêzere, J.L.; Fernandes, N.F. Validation and prediction challenges in data-driven landslide susceptibility mapping: Insights from random sampling of training and test data in machine learning models. Bull. Eng. Geol. Environ. 2026, 85, 253. [Google Scholar] [CrossRef]
  56. Mahoney, M.J.; Johnson, L.K.; Silge, J.; Frick, H.; Kuhn, M.; Beier, C.M. Assessing the performance of spatial cross-validation approaches for models of spatially structured data. arXiv 2023, arXiv:2303.07334. [Google Scholar] [CrossRef]
  57. Zeng, T.; Wu, L.; Peduto, D.; Glade, T.; Hayakawa, Y.S.; Yin, K. Ensemble learning framework for landslide susceptibility mapping: Different basic classifier and ensemble strategy. Geosci. Front. 2023, 14, 101645. [Google Scholar] [CrossRef]
  58. Hosmer, D.W.; Lemeshow, S.; Sturdivant, R.X. Applied Logistic Regression, 3rd ed.; Wiley: Hoboken, NJ, USA, 2013. [Google Scholar] [CrossRef]
  59. Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; ACM: New York, NY, USA, 2016; pp. 785–794. [Google Scholar] [CrossRef]
  60. Breiman, L. Random forests. Mach. Learn. 2001, 45, 5–32. [Google Scholar] [CrossRef]
  61. Huan, Y.; Song, L.; Khan, U.; Zhang, B. Stacking ensemble of machine learning methods for landslide susceptibility mapping in Zhangjiajie City, Hunan Province, China. Environ. Earth Sci. 2023, 82, 35. [Google Scholar] [CrossRef]
  62. Van Der Laan, M.J.; Polley, E.C.; Hubbard, A.E. Super learner. Stat. Appl. Genet. Mol. Biol. 2007, 6, 25. [Google Scholar] [CrossRef] [PubMed]
  63. Geurts, P.; Ernst, D.; Wehenkel, L. Extremely randomized trees. Mach. Learn. 2006, 63, 3–42. [Google Scholar] [CrossRef]
  64. Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef]
  65. Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar] [CrossRef]
  66. Koukaras, P.; Tjortjis, C. Data preprocessing and feature engineering for data mining: Techniques, tools, and best practices. AI 2025, 6, 257. [Google Scholar] [CrossRef]
  67. Bouke, M.A.; Abdullah, A. An empirical study of pattern leakage impact during data preprocessing on machine learning-based intrusion detection models reliability. Expert Syst. Appl. 2023, 230, 120715. [Google Scholar] [CrossRef]
  68. Apicella, A.; Isgrò, F.; Prevete, R. Don’t push the button! Exploring data leakage risks in machine learning and transfer learning. Artif. Intell. Rev. 2025, 58, 339. [Google Scholar] [CrossRef]
  69. Shahriari, B.; Swersky, K.; Wang, Z.; Adams, R.P.; De Freitas, N. Taking the human out of the loop: A review of Bayesian optimization. Proc. IEEE 2016, 104, 148–175. [Google Scholar] [CrossRef]
  70. Wu, J.; Chen, X.-Y.; Zhang, H.; Xiong, L.-D.; Lei, H.; Deng, S.-H. Hyperparameter optimization for machine learning models based on Bayesian optimization. J. Electron. Sci. Technol. 2019, 17, 26–40. [Google Scholar] [CrossRef]
  71. Wang, S.; Zhuang, J.; Zheng, J.; Fan, H.; Kong, J.; Zhan, J. Application of Bayesian hyperparameter optimized random forest and XGBoost model for landslide susceptibility mapping. Front. Earth Sci. 2021, 9, 712240. [Google Scholar] [CrossRef]
  72. Hanley, J.A.; McNeil, B.J. The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 1982, 143, 29–36. [Google Scholar] [CrossRef] [PubMed]
  73. Davis, J.; Goadrich, M. The relationship between precision-recall and ROC curves. In Proceedings of the 23rd International Conference on Machine Learning, Pittsburgh, PA, USA, 25–29 June 2006; ACM: New York, NY, USA, 2006; pp. 233–240. [Google Scholar] [CrossRef]
  74. Sokolova, M.; Lapalme, G. A systematic analysis of performance measures for classification tasks. Inf. Process. Manag. 2009, 45, 427–437. [Google Scholar] [CrossRef]
  75. Pakdaman Naeini, M.; Cooper, G.; Hauskrecht, M. Obtaining well calibrated probabilities using Bayesian binning. Proc. AAAI Conf. Artif. Intell. 2015, 29, 2901–2907. [Google Scholar] [CrossRef]
  76. Good, I.J. Rational decisions. J. R. Stat. Soc. Ser. B Stat. Methodol. 1952, 14, 107–114. [Google Scholar] [CrossRef]
  77. Brier, G.W. Verification of forecasts expressed in terms of probability. Mon. Weather Rev. 1950, 78, 1–3. [Google Scholar] [CrossRef]
  78. Van Westen, C.J.; Van Asch, T.W.J.; Soeters, R. Landslide hazard and risk zonation—Why is it still so difficult? Bull. Eng. Geol. Environ. 2006, 65, 167–184. [Google Scholar] [CrossRef]
  79. Steger, S.; Brenning, A.; Bell, R.; Glade, T. The influence of systematically incomplete shallow landslide inventories on statistical susceptibility models and suggestions for improvements. Landslides 2017, 14, 1767–1781. [Google Scholar] [CrossRef]
  80. Stanley, T.; Kirschbaum, D.B. A heuristic approach to global landslide susceptibility mapping. Nat. Hazards 2017, 87, 145–164. [Google Scholar] [CrossRef] [PubMed]
  81. Ado, M.; Amitab, K.; Maji, A.K.; Jasińska, E.; Gono, R.; Leonowicz, Z.; Jasiński, M. Landslide susceptibility mapping using machine learning: A literature survey. Remote Sens. 2022, 14, 3029. [Google Scholar] [CrossRef]
  82. Deng, H.; Wu, X.; Zhang, W.; Liu, Y.; Li, W.; Li, X.; Zhou, P.; Zhuo, W. Slope-unit scale landslide susceptibility mapping based on the random forest model in deep valley areas. Remote Sens. 2022, 14, 4245. [Google Scholar] [CrossRef]
  83. Dou, H.; Huang, S.; Jian, W.; Wang, H. Landslide susceptibility mapping of mountain roads based on machine learning combined model. J. Mt. Sci. 2023, 20, 1232–1248. [Google Scholar] [CrossRef]
  84. Alvioli, M.; Loche, M.; Jacobs, L.; Grohmann, C.H.; Abraham, M.T.; Gupta, K.; Satyam, N.; Scaringi, G.; Bornaetxea, T.; Rossi, M.; et al. A benchmark dataset and workflow for landslide susceptibility zonation. Earth Sci. Rev. 2024, 258, 104927. [Google Scholar] [CrossRef]
  85. Liang, J.; Qi, W.; Xu, C.; Wang, P.; Sun, J.; Zhang, X.; Xue, Z.; Chen, J.; Cui, Y.; Pan, J.; et al. Evaluation of geological hazards susceptibility along a key railway based on machine learning. Sci. Rep. 2025, 15, 42497. [Google Scholar] [CrossRef] [PubMed]
  86. Yang, Z.; Pang, B.; Dong, W.; Li, D.; Huang, Z. Interaction of landslide spatial patterns and river canyon landforms: Insights into the Three Parallel Rivers area, southeastern Tibetan Plateau. Sci. Total Environ. 2024, 914, 169935. [Google Scholar] [CrossRef] [PubMed]
  87. Wang, S.; Ling, S.; Wu, X.; Wen, H.; Huang, J.; Wang, F.; Sun, C. Key predisposing factors and susceptibility assessment of landslides along the Yunnan–Tibet traffic corridor, Tibetan Plateau: Comparison with the LR, RF, NB, and MLP techniques. Front. Earth Sci. 2023, 10, 1100363. [Google Scholar] [CrossRef]
  88. Zhang, T.; Gao, Y.; Li, B.; Yin, Y.; Liu, X.; Gao, H.; Yang, W. Characteristics of rock-ice avalanches and geohazard-chains in the Parlung Zangbo Basin, Tibet, China. Geomorphology 2023, 422, 108549. [Google Scholar] [CrossRef]
  89. Bai, L.; Jiang, Y.; Mori, J. Source processes associated with the 2021 glacier collapse in the Yarlung Tsangpo Grand Canyon, southeastern Tibetan Plateau. Landslides 2023, 20, 421–426. [Google Scholar] [CrossRef]
  90. Huang, F.; Mao, D.; Jiang, S.-H.; Zhou, C.; Fan, X.; Zeng, Z.; Catani, F.; Yu, C.; Chang, Z.; Huang, J.; et al. Uncertainties in landslide susceptibility prediction modeling: A review on the incompleteness of landslide inventory and its influence rules. Geosci. Front. 2024, 15, 101886. [Google Scholar] [CrossRef]
  91. Segoni, S.; Ajin, R.S.; Nocentini, N.; Fanti, R. Insights gained from the review of landslide susceptibility assessment studies in Italy. Remote Sens. 2024, 16, 4491. [Google Scholar] [CrossRef]
  92. Li, B.; Liu, K.; Wang, M.; He, Q.; Jiang, Z.; Zhu, W.; Qiao, N. Global dynamic rainfall-induced landslide susceptibility mapping using machine learning. Remote Sens. 2022, 14, 5795. [Google Scholar] [CrossRef]
  93. Han, Y.; Semnani, S.J. Important considerations in machine learning-based landslide susceptibility assessment under future climate conditions. Acta Geotech. 2025, 20, 475–500. [Google Scholar] [CrossRef]
Figure 1. Study area setting of the Parlung Tsangpo Basin. (a) Topography, landslides, faults, drainage, highways, and the location of the Yarlung Tsangpo Hydropower Project (YTHP); (b) regional location within the Tibetan Plateau river-basin system; and (c) geological units. The geological symbols in panel (c) represent different stratigraphic and lithological units, and their detailed meanings are provided in Supplementary Table S1.
Figure 1. Study area setting of the Parlung Tsangpo Basin. (a) Topography, landslides, faults, drainage, highways, and the location of the Yarlung Tsangpo Hydropower Project (YTHP); (b) regional location within the Tibetan Plateau river-basin system; and (c) geological units. The geological symbols in panel (c) represent different stratigraphic and lithological units, and their detailed meanings are provided in Supplementary Table S1.
Remotesensing 18 02167 g001
Figure 2. National-scale source-domain landslide inventory and topographic context. Grey points denote the national landslide records used for source-domain information extraction; the shaded background represents elevation across China. Provincial boundaries and the South China Sea inset are shown for reference.
Figure 2. National-scale source-domain landslide inventory and topographic context. Grey points denote the national landslide records used for source-domain information extraction; the shaded background represents elevation across China. Provincial boundaries and the South China Sea inset are shown for reference.
Remotesensing 18 02167 g002
Figure 3. Research framework and experimental workflow. The workflow comprises multi-source data preparation, conditioning-factor derivation, four controlled modelling strategies, model validation and robustness assessment, and susceptibility mapping. E1–E4 denote the local baseline, national transfer, weighted source–target joint training, and prior-informed local models, respectively.
Figure 3. Research framework and experimental workflow. The workflow comprises multi-source data preparation, conditioning-factor derivation, four controlled modelling strategies, model validation and robustness assessment, and susceptibility mapping. E1–E4 denote the local baseline, national transfer, weighted source–target joint training, and prior-informed local models, respectively.
Remotesensing 18 02167 g003
Figure 4. Spatial distribution of landslide conditioning factors in the Parlung Tsangpo Basin. Panels (ao) show continuous factors: (a) topographic position index, (b) annual mean temperature, (c) slope, (d) profile curvature, (e) topographic wetness index, (f) elevation, (g) relief, (h) plan curvature, (i) normalised difference vegetation index, (j) distance to railways, (k) distance to rivers, (l) distance to active faults, (m) distance to roads, (n) sine of aspect, and (o) cosine of aspect. Panels (ps) show categorical factors: (p) lithology, (q) land-use type, (r) soil type, and (s) geomorphology. The lithological symbols in panel (p) and soil-type codes in panel (r) are defined in Tables S1 and S4, respectively. Categorical variables were encoded before model training.
Figure 4. Spatial distribution of landslide conditioning factors in the Parlung Tsangpo Basin. Panels (ao) show continuous factors: (a) topographic position index, (b) annual mean temperature, (c) slope, (d) profile curvature, (e) topographic wetness index, (f) elevation, (g) relief, (h) plan curvature, (i) normalised difference vegetation index, (j) distance to railways, (k) distance to rivers, (l) distance to active faults, (m) distance to roads, (n) sine of aspect, and (o) cosine of aspect. Panels (ps) show categorical factors: (p) lithology, (q) land-use type, (r) soil type, and (s) geomorphology. The lithological symbols in panel (p) and soil-type codes in panel (r) are defined in Tables S1 and S4, respectively. Categorical variables were encoded before model training.
Remotesensing 18 02167 g004
Figure 5. Learner selection and independent-test performance of the national information introduction strategies. (a) ROC curves of E1–E4 under the random-forest learner; (b) predicted-probability distributions of negative and positive samples under E1–E4; and (c) independent-test metrics for strategy comparison and learner selection, including AUC, AP, F1-score, precision, recall, accuracy, log loss, and ECE. The threshold-based metrics in panel (c), including F1-score, precision, recall, and accuracy, were calculated using a common probability threshold of τ = 0.50, with samples assigned to the landslide class when the predicted probability was ≥0.50. E1–E4 denote the local baseline, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling, respectively.
Figure 5. Learner selection and independent-test performance of the national information introduction strategies. (a) ROC curves of E1–E4 under the random-forest learner; (b) predicted-probability distributions of negative and positive samples under E1–E4; and (c) independent-test metrics for strategy comparison and learner selection, including AUC, AP, F1-score, precision, recall, accuracy, log loss, and ECE. The threshold-based metrics in panel (c), including F1-score, precision, recall, and accuracy, were calculated using a common probability threshold of τ = 0.50, with samples assigned to the landslide class when the predicted probability was ≥0.50. E1–E4 denote the local baseline, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling, respectively.
Remotesensing 18 02167 g005
Figure 6. Robustness and sample-size sensitivity of the prior-informed local model relative to the local baseline. (a) Performance gains or error reductions in E4 over E1 across repeated spatial-validation runs; (b) changes in the E4 improvements under different target-domain training-sample fractions; (c) mean ± standard-deviation performance of E1 and E4; and (d) E4 gains or error reductions under different training fractions. In the boxplots, orange horizontal lines indicate medians, green triangles indicate means, and black points indicate outliers; in the line plots, colors correspond to the performance metrics shown in the legend. Positive values for the Brier score and log loss indicate lower probabilistic-prediction error.
Figure 6. Robustness and sample-size sensitivity of the prior-informed local model relative to the local baseline. (a) Performance gains or error reductions in E4 over E1 across repeated spatial-validation runs; (b) changes in the E4 improvements under different target-domain training-sample fractions; (c) mean ± standard-deviation performance of E1 and E4; and (d) E4 gains or error reductions under different training fractions. In the boxplots, orange horizontal lines indicate medians, green triangles indicate means, and black points indicate outliers; in the line plots, colors correspond to the performance metrics shown in the legend. Positive values for the Brier score and log loss indicate lower probabilistic-prediction error.
Remotesensing 18 02167 g006
Figure 7. Probability calibration of the national information introduction strategies on the target-domain test set. (a) Reliability diagrams for E1–E4 constructed using 10 equal-width probability bins over [0, 1], with the dashed line denoting perfect calibration. The ECE values of the four strategies are reported in the legend. (b) Summary of calibration metrics, including Brier score, log loss, and ECE. (c) Sample counts n m in each predicted susceptibility probability bin for E1–E4. The binning scheme in (a,c) is identical to that used for the ECE calculation, with n m / N used as the ECE weighting factor. Lower Brier score, log loss, and ECE values indicate better probabilistic prediction and calibration.
Figure 7. Probability calibration of the national information introduction strategies on the target-domain test set. (a) Reliability diagrams for E1–E4 constructed using 10 equal-width probability bins over [0, 1], with the dashed line denoting perfect calibration. The ECE values of the four strategies are reported in the legend. (b) Summary of calibration metrics, including Brier score, log loss, and ECE. (c) Sample counts n m in each predicted susceptibility probability bin for E1–E4. The binning scheme in (a,c) is identical to that used for the ECE calculation, with n m / N used as the ECE weighting factor. Lower Brier score, log loss, and ECE values indicate better probabilistic prediction and calibration.
Remotesensing 18 02167 g007
Figure 8. Landslide susceptibility maps and mapped probability distributions under the national information introduction strategies. (a) Susceptibility maps of E1–E4 classified into five levels; (b) mapped susceptibility-probability distributions on a square-root scale; and (c) summary statistics of the mapped probabilities, including central tendency, upper quantiles, and low- and high-probability area fractions.
Figure 8. Landslide susceptibility maps and mapped probability distributions under the national information introduction strategies. (a) Susceptibility maps of E1–E4 classified into five levels; (b) mapped susceptibility-probability distributions on a square-root scale; and (c) summary statistics of the mapped probabilities, including central tendency, upper quantiles, and low- and high-probability area fractions.
Remotesensing 18 02167 g008
Figure 9. Spatial-level validation and representative site-scale evidence for the E4 susceptibility map. The left upper panel shows the basin-scale E4 susceptibility classes and the locations of representative Sites A–E. The local enlarged panels show the susceptibility patterns around Sites A–D and Site E, with the labelled p values indicating the E4-predicted susceptibility probabilities at the recorded landslide locations. The UAV photograph of Site E shows the observed slope-failure area, where the red outline denotes the inferred landslide perimeter and the red arrow indicates the inferred movement direction. Together, these examples illustrate the spatial correspondence between recorded landslide locations, mapped high- and very-high-susceptibility patches, and local alpine-gorge geomorphic conditions.
Figure 9. Spatial-level validation and representative site-scale evidence for the E4 susceptibility map. The left upper panel shows the basin-scale E4 susceptibility classes and the locations of representative Sites A–E. The local enlarged panels show the susceptibility patterns around Sites A–D and Site E, with the labelled p values indicating the E4-predicted susceptibility probabilities at the recorded landslide locations. The UAV photograph of Site E shows the observed slope-failure area, where the red outline denotes the inferred landslide perimeter and the red arrow indicates the inferred movement direction. Together, these examples illustrate the spatial correspondence between recorded landslide locations, mapped high- and very-high-susceptibility patches, and local alpine-gorge geomorphic conditions.
Remotesensing 18 02167 g009
Table 1. Geospatial datasets and landslide-label records used for conditioning-factor derivation and source–target modelling.
Table 1. Geospatial datasets and landslide-label records used for conditioning-factor derivation and source–target modelling.
Data GroupDerived Variables/UseSource Dataset or ConstructionSpatial SpecificationTime Period/Sample Size
Basic geographyDistance to roads, railways, and riversNational 1:250,000 Public Basic Geographic DatasetVector, 1:250,000
GeologyLithology; distance to active faultsChina 1:2,500,000 Geological MapVector, 1:2,500,000
GeomorphologyGeomorphological type1:4,000,000 Geomorphological Map of ChinaVector, 1:4,000,000
TopographyElevation, slope, relief, Topographic position index (TPI), plan/profile curvature, Topographic wetness index (TWI), aspect componentsESA Copernicus Digital Elevation Model Global 30 m (DEM GLO-30)Raster, 30 m
ClimateAnnual mean temperature1 km Monthly Mean Temperature Dataset for ChinaRaster, 1 km1901–2024
VegetationNormalised Difference Vegetation Index (NDVI)Annual 30 m Maximum-Value Composite NDVI Dataset of ChinaRaster, 30 m1985–2024
Land coverLand-use type30 m Annual Land Cover Dataset of China and its DynamicsRaster, 30 m1985–2025
SoilSoil typeHarmonised World Soil DatabaseRaster, 30 arc-second
Target-domain landslide inventoryLandslide records in the Parlung Tsangpo BasinHistorical records and visual interpretationVector points249 historical; 33 interpreted
National landslide recordsSource-domain landslide information extractionNational Geological Hazard Point Inventory compiled by the Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of SciencesVector point records; spatial accuracy generally within 20 m1995–2024; 109,607 original records, 105,821 valid records after cleaning
Table 2. Derivation and selection basis of landslide conditioning factors in the Parlung Tsangpo Basin.
Table 2. Derivation and selection basis of landslide conditioning factors in the Parlung Tsangpo Basin.
Factor GroupSelected FactorsDerivation and Selection Basis
Topographic and slope conditionsElevation, slope, relief, TPI, plan curvature, profile curvature, sine of aspect, cosine of aspectElevation was extracted from the DEM; slope, curvature, relief, TPI, and aspect components were derived from DEM surface derivatives. These factors describe terrain height, slope steepness, local topographic position, incision, and slope morphology [33,34,35,38].
Hydrological and valley conditionsTWI, distance to riversTWI was derived from the upslope contributing area and slope; distance to rivers was calculated as the Euclidean distance to the nearest river. These factors represent runoff convergence, wetness, river incision, and toe erosion [35,36,37].
Geological and tectonic conditionsLithology, distance to active faultsLithology was rasterised from the geological map and one-hot encoded; distance to active faults was calculated as the Euclidean distance to the nearest fault. These factors reflect rock-mass properties and tectonic fracturing [33,34,35].
Geomorphological and soil backgroundsGeomorphology, soil typeGeomorphological and soil maps were rasterised and one-hot encoded to represent geomorphic units, near-surface materials, soil texture, and infiltration conditions [33,34].
Climatic backgroundAnnual mean temperatureMonthly temperature data were aggregated to annual mean temperature, representing the regional thermal background and elevation-related climatic gradients [39].
Vegetation and land-cover conditionsNDVI, land-use typeNDVI was extracted from the annual maximum-value composite product; land-use type was rasterised and one-hot encoded. These factors describe vegetation cover and surface-cover differences [33,34,40].
Human disturbance and accessibilityDistance to roads, distance to railwaysDistances to roads and railways were calculated as Euclidean distances, representing engineering disturbance, slope cutting, drainage modification, and record accessibility [34,40,41].
Table 3. Candidate learners and their core expressions.
Table 3. Candidate learners and their core expressions.
ModelCore ExpressionMethodological Role
L2-regularised logistic regression [58] p ^ i = σ β 0 + x i β , σ ( z ) = 1 / ( 1 + e z ) Linear probabilistic baseline; second-level learner in stacking
Random forest [60] p ^ i = 1 B b = 1 B T b ( x i ) Bagging tree ensemble; candidate primary model
Extra Trees [63] p ^ i = 1 B b = 1 B T b E T ( x i ) Strongly randomised tree ensemble; first-level learner in stacking
HistGradientBoosting [64] F m x = F m 1 x + η h m x Gradient-boosting model; first-level learner in stacking
SVM-RBF [65] K ( x i , x j ) = e x p   ( γ x i x j 2 ) Kernel-based nonlinear classifier; first-level learner in stacking
XGBoost [59] L t = i l y i , y ^ i t 1 + f t x i + Ω f t Regularised gradient-boosting trees; candidate primary model
Stacking [13] z i = p ^ i 1 , , p ^ i M ,   p ^ i = σ ( θ 0 + z i θ ) Two-level ensemble learning; candidate primary model
Table 4. Main hyperparameter search spaces of the candidate models.
Table 4. Main hyperparameter search spaces of the candidate models.
ModelMain Hyperparameters and Search Ranges
L2-regularised logistic regressionC: 10−4–103
Random forestn_estimators: 300–1200; max_depth: 5–50 or None; min_samples_split: 2–50; min_samples_leaf: 1–20; max_features: sqrt, log2, or 0.3–1.0
Extra Treesn_estimators: 300–1200; max_depth: 5–50 or None; min_samples_split: 2–50; min_samples_leaf: 1–20; max_features: sqrt, log2, or 0.3–1.0
HistGradientBoostinglearning_rate: 0.01–0.3; max_iter: 100–800; max_leaf_nodes: 15–127; max_depth: 3–20 or None; l2_regularization: 10−6–10
SVM-RBFC: 10−3–103; γ: 10−4–10
XGBoostn_estimators: 200–1500; max_depth: 2–10; learning_rate: 0.005–0.3; subsample: 0.5–1.0; colsample_bytree: 0.5–1.0; min_child_weight: 1–20; reg_alpha: 10−8–10; reg_lambda: 10−3–100
StackingFirst-level learners used their respective optimised hyperparameters; the second-level learner was L2-regularised logistic regression with C: 10−4–103.
Table 5. Sample partitioning of the four-fold spatial outer cross-validation.
Table 5. Sample partitioning of the four-fold spatial outer cross-validation.
Outer FoldTraining SetValidation Set
Fold 1306 = 153 positive/153 negative110 = 55 positive/55 negative
Fold 2312 = 156 positive/156 negative108 = 54 positive/54 negative
Fold 3302 = 151 positive/151 negative108 = 54 positive/54 negative
Fold 4310 = 155 positive/155 negative108 = 54 positive/54 negative
Table 6. Multicollinearity diagnostics of the candidate conditioning factors.
Table 6. Multicollinearity diagnostics of the candidate conditioning factors.
VariableVIFTolerance
Topographic position index (TPI)7.7200.130
Annual mean temperature6.0400.166
Slope5.7000.176
Profile curvature4.6900.213
Topographic wetness index (TWI)4.6800.214
Elevation4.3700.229
Lithology3.2100.311
Relief3.1100.322
Plan curvature2.5200.396
Normalised difference vegetation index (NDVI)1.5700.639
Distance to railways1.4200.706
Land-use type1.2300.813
Distance to rivers1.2300.814
Soil type1.2200.823
Geomorphology1.1000.907
Distance to active faults1.0600.945
Distance to roads1.0500.948
Sine of aspect1.0100.995
Cosine of aspect1.0000.997
Table 7. Spatial-level validation of the E4 susceptibility map using independent test landslide points.
Table 7. Spatial-level validation of the E4 susceptibility map using independent test landslide points.
Susceptibility ClassProbability IntervalArea (km2)Area Proportion (%)Independent Landslide PointsLandslide-Point Proportion (%)Landslide Density (Points/100 km2)Frequency Ratio
Very low0–0.226,434.31887.94623.5710.0080.041
Low0.2–0.41718.6865.718814.2860.4652.498
Moderate0.4–0.6856.5952.850610.7140.7003.760
High0.6–0.8641.3602.1341017.8571.5598.369
Very high0.8–1.0406.6051.3533053.5717.37839.602
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Guo, Y.; Jiang, Z.; Lin, Y.; Qu, G.; Li, W.; Dai, X. Prior-Informed Local Landslide Susceptibility Modelling Using National-Scale Landslide Information: A Case Study of the Parlung Tsangpo Alpine Gorge Basin. Remote Sens. 2026, 18, 2167. https://doi.org/10.3390/rs18132167

AMA Style

Guo Y, Jiang Z, Lin Y, Qu G, Li W, Dai X. Prior-Informed Local Landslide Susceptibility Modelling Using National-Scale Landslide Information: A Case Study of the Parlung Tsangpo Alpine Gorge Basin. Remote Sensing. 2026; 18(13):2167. https://doi.org/10.3390/rs18132167

Chicago/Turabian Style

Guo, Yantong, Zhongxiang Jiang, Yuxuan Lin, Ge Qu, Weile Li, and Xiaoai Dai. 2026. "Prior-Informed Local Landslide Susceptibility Modelling Using National-Scale Landslide Information: A Case Study of the Parlung Tsangpo Alpine Gorge Basin" Remote Sensing 18, no. 13: 2167. https://doi.org/10.3390/rs18132167

APA Style

Guo, Y., Jiang, Z., Lin, Y., Qu, G., Li, W., & Dai, X. (2026). Prior-Informed Local Landslide Susceptibility Modelling Using National-Scale Landslide Information: A Case Study of the Parlung Tsangpo Alpine Gorge Basin. Remote Sensing, 18(13), 2167. https://doi.org/10.3390/rs18132167

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop