1. Introduction
Landslides are a global geohazard that cause casualties, infrastructure disruption, river blockage, and long-term geomorphic disturbance in mountain regions [
1]. In alpine gorge environments, slope failure may further evolve into cascading hazards, including long-runout movement, landslide-dammed lakes, outburst floods, and debris flow–river blockage chains [
2,
3]. Such hazards are particularly critical in High Asia and around the Tibetan Plateau, where climate warming, cryospheric degradation, and extreme precipitation are expected to increase landslide-related risks [
1,
4,
5]. Reliable identification of landslide-susceptible terrain in steep, poorly accessible, and hazard-prone mountain basins is therefore essential for disaster prevention, infrastructure planning, and engineering safety assessment.
Landslide susceptibility assessment estimates the spatial likelihood of landslide occurrence by relating landslide inventories to conditioning factors such as topography, lithology, hydrology, vegetation, climate, and human activity [
6]. Machine-learning methods have been increasingly adopted for this task because they can capture nonlinear relationships between landslide occurrence and multi-source environmental variables. For example, in susceptibility mapping over the Qinghai–Tibetan Plateau, random forest outperformed deep neural networks, logistic regression, naïve Bayes, and support vector machine models, demonstrating strong predictive capability under complex environmental conditions [
6]. However, the reliability of data-driven models depends strongly on the completeness, positional accuracy, and spatial representativeness of landslide inventories [
7,
8]. In high-mountain basins, local inventories are often limited by poor accessibility, topographic shadow, vegetation cover, and incomplete historical records, which may lead to overfitting, unstable probability estimation, and uncertain susceptibility maps.
National- or large-regional-scale landslide inventories provide a potential way to alleviate local sample scarcity. Compared with a small-basin inventory, broad-scale records cover a wider range of geomorphological, lithological, climatic, and anthropogenic conditions, so the resulting national model can encode cross-regional statistical associations between landslide occurrence and environmental factors. These associations are useful as source-domain background information, but they should not be interpreted as universally transferable landslide-development rules. They may partly reflect the environments that are better represented in the national inventory and may not be directly applicable to the alpine gorge setting at the southeastern margin of the Tibetan Plateau. Differences in tectonic setting, geomorphic evolution, rainfall regime, snow and ice processes, and engineering disturbance may therefore cause direct national-model transfer or simple source–target sample merging to produce weak regional adaptability, probability bias, and excessive expansion of high-susceptibility zones [
9,
10]. Therefore, when sample scarcity and regional heterogeneity coexist, the key issue is how to transform national-scale landslide information into locally reliable susceptibility knowledge.
Previous studies have explored cross-regional landslide knowledge transfer through transfer learning, domain adaptation, case-based reasoning, sample selection, and joint training [
9,
11], highlighting the potential value of source-domain information for data-scarce regions. However, two issues remain. First, the general applicability of any transfer strategy still needs to be examined through independent validation in multiple basins. Second, when different studies use different study areas, factor systems, sample-construction rules, and evaluation schemes, it is difficult to isolate the effect of the source information utilisation strategy itself. The present study, therefore, focuses on a controlled comparison within one representative alpine gorge basin, using the same conditioning factors, data-processing workflow, spatially independent target-domain test set, and evaluation metrics. Less attention has been paid to converting the output of a national-scale model into an external probability prior that can be further learned and recalibrated by local samples, despite the importance of calibrated probability estimates for susceptibility mapping [
12]. In this study, the two-stage computation is regarded as an implementation pathway rather than the main methodological emphasis. The key idea is to treat the national-model probability as a geoscientifically interpretable prior that expresses broad-scale landslide-background conditions in the target basin, while allowing target-domain samples to determine its effective weight and correction direction. This differs from generic stacking or model fusion [
13], where first-level outputs mainly function as algorithmic ensemble features, and from feature- or sample-level transfer [
9,
11], where source-domain information is transferred more directly.
The Parlung Tsangpo Basin, located on the southeastern margin of the Tibetan Plateau, is a typical alpine gorge basin characterised by strong topographic relief, deeply incised river valleys, active fault systems, and pronounced river erosion. These conditions provide a favourable environment for frequent landslide occurrence and cascading hazard development. Deeply incised valleys, active tectonics, complex lithology, monsoonal precipitation, snow and ice meltwater, and ice–rock avalanche processes jointly control slope instability in this region. Repeated cascading events, including landslide-induced river blockage, landslide-dammed lake outburst flooding, and ice–rock avalanche–debris flow–river blockage chains, have occurred in the basin and adjacent areas [
2,
3,
14]. Meanwhile, strong topographic obstruction, limited field verification, and incomplete historical records make it difficult for local landslide samples to cover the full range of conditioning environments. The basin, therefore, provides a suitable target domain for evaluating how national-scale landslide information can be used in local susceptibility modelling under sample-limited conditions.
In this study, the Parlung Tsangpo Basin is used as the target domain, and national-scale landslide records are introduced as external source-domain information. Under a unified conditioning-factor system, data-processing workflow, target-domain test set, and evaluation metrics, four strategies are compared: a local baseline model, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling. The objectives are to: (1) test whether national-scale landslide information can improve susceptibility prediction in a data-scarce alpine gorge basin; (2) compare how different source information utilisation strategies perform in terms of discrimination ability, classification accuracy, probability calibration, and spatial mapping; and (3) evaluate whether national-model output can serve as an effective probability prior for producing locally adapted susceptibility maps. This study provides a controlled within-basin comparison framework for converting broad-scale landslide information into basin-scale susceptibility assessment, while leaving the broader cross-basin transferability of the proposed strategy to be tested in future studies.
3. Materials and Methods
The methodological framework of this study is shown in
Figure 3. For clarity, the four experimental strategies are denoted as: E1 (local baseline), E2 (direct national-model transfer), E3 (weighted source–target joint training), and E4 (prior-informed local modelling). The Parlung Tsangpo Basin is treated as the target domain and the national-scale landslide samples outside it as the source domain. Around the central question of how national-scale landslide information can be transformed into local susceptibility modelling, four controlled experiments are constructed. The overall workflow comprises multi-source data preparation, conditioning-factor derivation and preprocessing, the design of cross-scale information introduction strategies, model validation and robustness assessment, and landslide susceptibility mapping and interpretation.
To ensure that the results of the different experiments are comparable, E1–E4 share the same original conditioning-factor system, data-preprocessing workflow, sample-construction method, spatial partitioning scheme, and evaluation metrics; only the way in which national-scale information is introduced is varied. Let the target-domain samples be
and the source-domain samples be
where
denotes a non-landslide or landslide sample,
denotes the model input-feature vector after preprocessing and one-hot encoding of categorical variables, and
is the actual feature dimension after encoding. Before entering a model, categorical variables such as lithology, land-use type, soil type, and geomorphological type are expanded into higher-dimensional features through one-hot encoding.
On the basis of this unified setting, the study compares four strategies—the local baseline model, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling—in order to identify effective ways of using national-scale information when local samples are limited.
3.1. Experimental Design and Overall Rationale
To examine the local applicability of national-scale landslide information in data-scarce basins, four controlled modelling strategies are defined: the local baseline model, direct national-model transfer, weighted source–target joint training, and prior-informed local modelling. The four strategies share the same independent target-domain test set and evaluation framework, ensuring that performance differences arise mainly from the way cross-scale information is introduced rather than from differences in sample partitioning, factor systems, or evaluation methods.
E1 is the local baseline. Only samples from the Parlung Tsangpo Basin and the 19 original conditioning factors are used to train the model, characterising conventional local modelling performance in the absence of external information [
19,
20]. Letting
denote a given learner together with its training procedure, E1 can be written as
where
denotes the model input features formed from the 19 original conditioning factors after preprocessing.
E2 is the direct national-model transfer strategy: national samples outside the target domain are used to train a model that is then applied directly to the Parlung Tsangpo Basin, in order to examine the generalisation ability of the national-scale model in the local geomorphological environment [
9,
21,
22]. This process can be expressed as
Because no target-domain samples are used in training, this strategy serves to evaluate the direct-transfer ability of the source-domain model in the target domain.
E3 is the weighted source–target joint training strategy: national and target-domain samples are used jointly for training, with higher weights assigned to the target-domain samples to reduce the risk of negative transfer caused by source–target distribution mismatch [
9,
23,
24]. Its training objective can be written as
where
is the classification loss and
is the target-domain sample weight. In this study
is searched over
to balance the external information provided by the national samples against the local environmental characteristics reflected by the target-domain samples.
E4 is the prior-informed local modelling strategy. A source-domain model is first trained on the national samples and used to generate national-scale susceptibility probabilities over the target domain. These probabilities are then used as an additional probability-prior variable and fed, together with the local conditioning factors, into the target-domain model, allowing the national-scale information to be relearned and recalibrated under the constraint of local samples [
13,
25,
26]. This process can be expressed as
where
is the susceptibility probability output by the national model at target-domain sample or pixel locations, and
is the target-domain input feature after the national probability prior has been appended. Unlike E2, E4 does not adopt the national-model output as the final result; instead, it treats that output as a learnable probability-prior variable that is locally recalibrated by the target-domain samples. Unlike generic stacking, this variable is not introduced merely as an algorithmic ensemble feature, but represents the national-scale model’s probabilistic expression of broad-scale landslide-background conditions at each target-domain location. Its effective contribution and correction direction are therefore estimated from the local samples together with the original conditioning factors.
In summary, E1 provides the local baseline without external information, E2 examines the direct-transfer ability of the national model, E3 evaluates the gain from weighted source–target joint training, and E4 tests the effectiveness of the national probability prior after relearning with local samples. Together, the four strategies form a controlled comparison spanning no transfer, direct transfer, sample-level fusion, and probability-prior fusion.
The distinctions among the three information-introduction strategies are summarised as follows. Direct transfer (E2) applies the national-scale model predictions directly to the target basin without further adaptation. Joint training (E3) introduces source-domain information by combining national and local samples into a unified training set and fitting a new model jointly. In contrast, prior-informed local modelling (E4) does not directly use the national prediction as the final result; instead, it treats the national model output as a probabilistic prior that is further learned, reweighted, and recalibrated using target-domain samples together with the original conditioning factors.
3.2. Remote-Sensing and Spatial Data Preprocessing
Before sample extraction, the multi-source geospatial data were uniformly preprocessed to ensure consistency of the input variables in spatial scale and raster registration [
27]. All data were brought to a common coordinate reference system, spatial resolution, and raster-alignment framework; continuous raster variables were resampled by cubic convolution to better preserve smooth spatial variation in continuous surfaces, whereas categorical variables were resampled by nearest-neighbour resampling to preserve original class labels and avoid the creation of artificial category values [
28,
29].
For unordered categorical variables such as lithology, land-use type, soil type, and geomorphological type, one-hot encoding was used to convert them into binary indicator variables, thereby preventing the model from misinterpreting category codes as numerical magnitudes or ordinal relationships [
30]. Continuous variables were scaled according to learner characteristics: Z-score standardisation was applied to the logistic regression and SVM-RBF models, whereas the tree ensemble models retained the original variable scales [
31,
32]. For a continuous variable
, its standardised form is
where
and
are the mean and standard deviation of the
-th variable in the training data and are subsequently applied to the corresponding validation set, test set, and mapping raster.
3.3. Feature Engineering
3.3.1. Selection of Conditioning Factors
Combining commonly used landslide susceptibility factors with the alpine-gorge characteristics of the Parlung Tsangpo Basin, 19 conditioning factors were selected from topography, hydrology, geology, geomorphology, climate, vegetation, land cover, and human disturbance [
33,
34,
35,
36]. To make the data-to-factor relationship explicit, all variables were derived on the unified raster grid described in
Section 3.2. Let
denote a raster cell,
the DEM elevation,
a local moving window, and
a vector layer such as rivers, faults, roads, or railways. DEM-based factors were derived from elevation, local window statistics, or surface derivatives; for example,
The topographic wetness index was calculated as
where
is the upslope contributing area and
is slope [
37]. Distance factors were calculated as Euclidean distances,
Aspect was represented by
and
to avoid angular discontinuity [
38]. Categorical factors were rasterised and one-hot encoded as
if
, and
otherwise. The selected factors and their derivation basis are summarised in
Table 2.
3.3.2. Feature Screening
Because this study first compared multiple candidate learners with different sensitivities to correlated predictors, feature screening was conducted on the common conditioning-factor set before final learner selection. Tolerance (TOL) and the variance inflation factor (VIF) were used to diagnose the 19 candidate factors [
42,
43]. For the
-th candidate factor, an auxiliary regression was fitted using the remaining factors as explanatory variables, yielding the coefficient of determination
. The tolerance and variance inflation factor are defined as
Following common criteria, VIF > 10 or TOL < 0.1 was adopted as the threshold for severe collinearity [
44]. Factors below these thresholds were retained for model training; where severe collinearity was present, factors were removed or adjusted with reference to both their geoscientific meaning and model stability.
3.4. Construction of Experimental Samples
The study-area and national-scale landslide samples were built with a unified construction workflow. To reduce information overlap among spatially adjacent samples and the influence of spatial autocorrelation, the landslide points were spatially thinned before modelling. The 250 m minimum retained separation was selected as an a priori operational threshold based on the 30 m DEM-derived mapping resolution, the local clustering characteristics of landslides in the alpine gorge setting, and the need to retain sufficient local positive samples for model training. This distance corresponds to approximately eight DEM pixels, which helps reduce repeated sampling of neighbouring points with highly similar terrain conditions. In the Parlung Tsangpo Basin, landslide records are strongly clustered along valley sides and transportation corridors; therefore, a much smaller threshold would retain many samples from the same local terrain context, whereas a much larger threshold would unnecessarily reduce the already limited local landslide records. The same 250 m distance was also used as the buffer width during spatial cross-validation, so that sample thinning and validation partitioning treated local spatial dependence consistently. Neighbouring points closer than this threshold were removed [
45,
46,
47,
48,
49].
Negative-sample construction balanced geoscientific plausibility against spatial independence. First, water bodies, built-up land, glaciers, and areas immediately adjacent to landslide points were excluded, reducing the chance that potential unlabelled landslides or unsuitable background areas would be sampled as negatives [
46,
50,
51]. Geomorphological type was then used as the stratification unit, and negative samples were drawn at random within each sampleable area, with their spatial distribution controlled to weaken the influence of sample clustering on the model decision boundary [
52,
53,
54]. Finally, positive and negative samples were constructed at a 1:1 ratio to limit the effect of class imbalance on model training and evaluation [
55,
56]. The above procedure was used as the main negative-sampling strategy and is denoted as A0. To evaluate whether the model comparison depended on this particular negative-sampling rule, two additional negative-sampling schemes were designed as sensitivity tests and reported in the
Supplementary Materials. In A1, negative samples were resampled within slope strata so that the comparison was less affected by differences in the slope-frequency structure. In A2, hard negative samples were drawn from the slope range overlapping the positive samples, thereby forcing the models to distinguish stable and unstable terrain under similar gradient conditions. The same positive samples, conditioning factors, learner settings, and evaluation metrics were kept for E1–E4, and only the negative-sampling rule was changed.
3.5. Dataset Partitioning
To reduce performance overestimation caused by spatial autocorrelation, a spatial-blocking partitioning strategy was adopted. A 2 km × 2 km grid was used at the study-area scale and a 20 km × 20 km grid at the national scale, strengthening the spatial independence among the training, validation, and independent test sets [
45,
47,
48].
Sample partitioning followed the principles of spatial independence and class balance. An independent test set was first constructed by holding out entire spatial blocks, so that the test samples were spatially isolated from the training samples while the numbers of positive and negative samples were kept equal. The remaining samples were used for model training and cross-validation. During cross-validation, four-fold spatial outer cross-validation was built using the spatial grids as basic units, and a 250 m spatial buffer was placed between the training and validation sets; samples falling within the buffer were excluded from the corresponding training or validation fold, reducing the risk of information leakage from neighbouring samples [
45,
48,
49].
3.6. Model Training and Hyperparameter Tuning
To reduce the dependence of model selection on the assumptions of any single learner, four candidate models were compared under a unified sample, factor, and spatial-validation framework: L2-regularised logistic regression, random forest, XGBoost, and stacking [
19,
57]. Logistic regression provides a linear probabilistic baseline [
58]; random forest and XGBoost represent bagging and boosting tree ensembles, respectively [
59,
60]; and stacking fuses the complementary information of several nonlinear learners [
13,
61,
62]. In the stacking model, the first-level learners are Extra Trees, HistGradientBoosting, and SVM-RBF, and the second-level learner is L2-regularised logistic regression. Letting the model output be the landslide-occurrence probability
, the core expression of each learner is given in
Table 3.
Model training followed the spatial-partitioning and preprocessing rules described above. Logistic regression and SVM-RBF were standardised using Z-score parameters estimated from the training fold, whereas the tree ensembles retained the original scales of the continuous variables; categorical variables were transformed using one-hot encoding rules determined from the training fold, so as to avoid data leakage [
66,
67,
68]. Hyperparameters were tuned by Bayesian optimisation with a maximum of 50 iterations [
69,
70,
71]. Letting
denote a hyperparameter combination,
the search space, and
the spatial-cross-validation performance, the optimal hyperparameters are
The main search spaces of each model are listed in
Table 4. After model selection, the learner type, preprocessing workflow, and tuning budget were fixed across E1–E4, and only the way of introducing national-scale information was varied.
3.7. Model Evaluation and Susceptibility Mapping
Model performance was evaluated from three angles: discrimination ability, threshold-based classification performance, and probability-output quality. Discrimination ability was measured by the area under the ROC curve (AUC) and average precision (AP), which describe how well a model ranks landslide samples above non-landslide samples [
72,
73]:
where
and
are the predicted probabilities of landslide and non-landslide samples, respectively, and
and
are the precision and recall at the
-th threshold of the precision–recall curve.
Under a unified classification threshold
, threshold-based performance was described by accuracy, precision, recall, and F1-score [
74]:
Probability output quality was assessed using the Brier score, log loss, and expected calibration error (ECE) [
75,
76,
77]:
Here, TP, TN, FP, and FN denote true positives, true negatives, false positives, and false negatives, respectively. For the ECE calculation, the probability interval [0, 1] was divided into M = 10 equal-width bins with edges of 0.0, 0.1, …, and 1.0. The first nine bins followed a lower-closed and upper-open rule, i.e., [lower, upper), whereas the last bin included the upper endpoint, i.e., [0.9, 1.0].
denotes the
-th probability bin,
is the number of test samples in that bin,
denotes the observed landslide frequency, and
denotes the mean predicted susceptibility probability. ECE was calculated as a sample-weighted average using
, and the same bin edges were used for E1–E4. Empty bins, if present, were not plotted in the reliability diagrams and contributed zero to ECE. These metrics correspond to ranking discrimination, classification discrimination, and probabilistic reliability, respectively, and are used to compare the experimental strategies [
75,
76,
77].
During susceptibility mapping, the input factors, scaling, and model-training workflow were kept consistent. The processed raster factors were fed into the model pixel by pixel to obtain continuous landslide susceptibility probabilities for the study area:
where
is a spatial pixel,
is the study-area extent, and
is the conditioning-factor vector for that pixel. The continuous probabilities were then divided at intervals of 0.2 into five classes—very low, low, moderate, high, and very high—to represent the spatial differentiation of susceptibility [
33]; blue, grey-blue, pale yellow, light orange, and coral red were used to denote these classes in order of increasing susceptibility.
To further examine the spatial credibility of the final susceptibility map, a spatial-level validation was conducted using the independent test landslide points that were not involved in model training or hyperparameter tuning. These landslide points were overlaid on the five susceptibility classes of the final E4 map. For each susceptibility class, the area proportion, number and proportion of independent landslide points, landslide density, and frequency ratio were calculated. The frequency ratio was defined as the ratio between the landslide-point proportion and the area proportion of a given susceptibility class. In addition, representative local susceptibility insets and a UAV photograph for Site E were used to compare the mapped high-susceptibility patches with recorded landslide locations, inferred landslide morphology, and local alpine-gorge geomorphic settings.
4. Results
4.1. Sample, Factor, and Data Partitioning Results
After spatial verification, removal of duplicate points, and 250 m spatial thinning, 273 valid positive samples were retained from the 282 original landslide points in the study area, with a minimum pairwise distance of 256.85 m, satisfying the spatial-independence constraint. Negative samples were drawn from background pixels at which all factors had valid values, after excluding water bodies, built-up land, glaciers, and the 1000 m buffers around landslide points, so as to reduce the chance that potential unlabelled landslides or unsuitable background areas would be sampled as negatives. The final sample set was constructed at a 1:1 positive-to-negative ratio, providing a class-balanced basis for the subsequent binary classification and strategy comparison.
To test the spatial generalisation ability of the models, an independent test set was constructed by holding out entire spatial blocks. The test set comprised 15 blocks of 2 km grid size and 112 samples in total (56 positive and 56 negative); the minimum distance between the test set and the remaining positive samples was 335.8 m, and no sample IDs were duplicated across subsets. After the test set was removed, the remaining samples were used for model training and four-fold spatial outer cross-validation. A 250 m buffer was placed between the training and validation sets in each fold, and samples within the buffer were excluded from the corresponding fold; the numbers of positive and negative samples in each fold were kept equal. The detailed partitioning is shown in
Table 5.
Nineteen candidate conditioning factors were then constructed to characterise the topographic relief, valley incision, tectonic and lithological conditions, geomorphological and soil backgrounds, climatic conditions, vegetation cover, land use, and human engineering disturbance of the Parlung Tsangpo Basin. Their spatial distributions are shown in
Figure 4: the continuous factors include topographic, climatic, hydrological, vegetation, and distance-related variables, while the categorical factors include lithology, land-use type, soil type, and geomorphology. Overall, the factors exhibit clear spatial heterogeneity within the study area and adequately reflect the complex landslide-conditioning environment of the alpine gorge basin. To further examine whether the negative samples represented only an easily separable low-susceptibility background, we compared the distributions of positive samples, A0 main negative samples, and sampleable background pixels for key conditioning factors (
Supplementary Figure S1 and Supplementary Tables S2 and S3). The sampleable background pixels and A0 negative samples were not dominated by very gentle terrain: their median slopes were 21.55° and 20.84°, respectively, whereas the median slope of the positive samples was 14.05°. The A0 negative samples also covered broad ranges of elevation, relief, TPI, NDVI, and distances to rivers, roads, and active faults. These results indicate that the main negative set retained substantial environmental variability rather than representing only valley-floor or uniformly low-susceptibility pixels.
The multicollinearity diagnostics in
Table 6 show that the VIF values of the 19 factors range from 1.000 to 7.720 and the tolerance values from 0.130 to 0.997, none reaching the severe-collinearity thresholds of VIF > 10 or TOL < 0.1. This result indicates that the common conditioning-factor set did not contain severe multicollinearity. Therefore, no factor was removed at this screening step, allowing the same 19 conditioning factors to be retained for both candidate-learner comparison and the subsequent RF-based E1–E4 strategy comparison.
4.2. Learner Selection and Overall Performance of the Strategies
To determine the base learner for prior-informed local modelling, the independent-test performance of random forest, XGBoost, L2-regularised logistic regression, and stacking was first compared under the E4 setting (
Figure 5c). Random forest achieved the best overall performance in discrimination, classification, and calibration, with AUC, AP, and F1-scores of 0.901, 0.895, and 0.828, respectively, and the lowest log loss and ECE, at 0.401 and 0.055. Although XGBoost, logistic regression, and stacking showed advantages in individual metrics—for example, higher recall—their precision, probability error, or calibration was comparatively weaker. Taking discrimination, misclassification control, and probabilistic reliability together, random forest was selected as the base learner for the subsequent strategy comparison, robustness analysis, and susceptibility mapping. The selected random-forest configuration used for the E4 prior-informed local model was n_estimators = 970, max_depth = None, min_samples_split = 2, min_samples_leaf = 2, and max_features = sqrt. The remaining fixed settings were criterion = gini, bootstrap = True, class_weight = None, and random_state = 42.
On the basis of the unified learner, target-domain test set, and evaluation metrics, the overall performance of the four information-introduction strategies E1–E4 was then compared. The ROC curves show that E4 attained an AUC of 0.901, higher than the 0.848, 0.845, and 0.864 of E1, E2, and E3, indicating stronger sample discrimination for the prior-informed strategy. The remaining metrics confirm that E4 achieved the best AP, F1-score, precision, accuracy, log loss, and ECE, so that it improved not only classification performance but also the quality of the probabilistic predictions.
In terms of the predicted-probability distributions, E4 assigned overall lower probabilities to negative samples while keeping the probabilities of positive samples high, giving clearer separation between the two classes. Compared with E2 (direct national-model transfer), E4 attained higher precision and lower log loss at the same recall, indicating that it reduced misclassification risk and improved the stability of the probability outputs while preserving landslide-detection ability. Overall, the prior-informed local modelling strategy performed best on the independent test set and was therefore adopted as the final strategy for the subsequent robustness analysis, calibration evaluation, and susceptibility mapping.
4.3. Robustness and Sample-Size Sensitivity
To test the robustness of E4 relative to the local baseline E1, repeated spatial validation and sensitivity analysis under different proportions of target-domain training samples were conducted (
Figure 6). The repeated-validation results show that E4 exceeded E1 in AUC, AP, and F1-score while reducing the Brier score and log loss, indicating that the prior-informed strategy improved sample discrimination and classification performance as well as the accuracy of the probabilistic predictions.
As the training-sample fraction increased from 25% to 100%, the advantage of E4 over E1 remained broadly stable. Across sample sizes, E4 consistently showed higher discrimination and lower probability error, indicating that the national-scale probability prior provides stable supplementary information when local samples are limited. This further indicates that the improvement of E4 does not depend on a single data partition but reflects good spatial generalisation and sample adaptability.
The alternative negative-sampling tests further support this interpretation (
Supplementary Figure S2 and Supplementary Tables S5 and S6). Under the A1 slope-stratified strategy, E4 achieved the highest AUC, AP, and F1-score among E1–E4 and also showed the lowest Brier score and log loss. Under the more demanding A2 slope-overlap hard-negative strategy, the discrimination difference between E4 and E1 became very small, with nearly identical AUC values. However, E4 still achieved a slightly higher F1-score and recall and lower Brier score, log loss, and ECE than E1. Therefore, the advantage of the prior-informed strategy does not rely on separating landslides from an artificially easy, gentle-slope background. Under harder negative-sampling conditions, its main benefit lies in maintaining comparable discrimination while improving the reliability of the predicted probabilities.
4.4. Probability Calibration Performance
To evaluate the reliability of the susceptibility probabilities produced by the different strategies, the calibration performance of E1–E4 was compared on the same spatially independent target-domain test set ((N = 112), including 56 landslide and 56 non-landslide samples;
Figure 7). The reliability diagrams were constructed using the same 10 equal-width probability bins as the ECE calculation, and the sample size of each bin is shown in the diagram to aid interpretation. The predicted probabilities of E4 lie closest overall to the perfect-calibration line, indicating better agreement between the predicted susceptibility probabilities and the observed landslide frequency.
4.5. Landslide Susceptibility Mapping and Spatial-Level Validation Results
The susceptibility maps and study-area-wide predicted-probability distributions of the four strategies are shown in
Figure 8. All four can identify high-susceptibility areas along gullies, slopes, and linear geomorphological units, but they differ clearly in spatial expression. E2 shows overall higher predicted probabilities, with moderate- and high-susceptibility areas distributed more diffusely, indicating that direct national-model transfer may overestimate risk in the target domain. By contrast, E4 expresses the low-susceptibility background more fully and concentrates its high-susceptibility patches more tightly, suppressing large-area overestimation while retaining local high-risk areas.
The probability-distribution statistics are consistent with the spatial patterns. The mean and median predicted probabilities of E4 are lower than those of the other strategies, and its share of low-probability area is the highest, while a small number of tightly concentrated very-high-susceptibility areas are still retained. Overall, E4 exhibits a mapping signature of an extensive low-susceptibility background with concentrated high-susceptibility patches, in line with the independent-test performance and calibration results above, and indicating that the prior-informed model generates more stable and more spatially discriminating susceptibility maps.
To further verify the spatial reliability of the final susceptibility map, the independent test landslide points were overlaid on the E4 susceptibility classes, and the spatial concentration of landslides in each class was quantified (
Table 7 and
Figure 9). The high and very high susceptibility classes occupied only 3.49% of the study area, but contained 40 of the 56 independent test landslide points, accounting for 71.43% of the independent landslide occurrences. In contrast, the very low class occupied 87.95% of the study area but contained only two independent landslide points. Landslide density increased markedly from 0.008 points/100 km
2 in the very low class to 7.378 points/100 km
2 in the very high class. The frequency ratio also increased monotonically from 0.041 in the very low class to 39.602 in the very high class. The landslide density of the high and very high classes was 107.46 times that of the very low and low classes combined. These results indicate that the E4 susceptibility map has clear map-level discriminatory ability and that the mapped high-susceptibility zones effectively concentrate independent landslide occurrences.
Representative local sites provide further site-scale evidence for the susceptibility zoning (
Figure 9). Sites A–E are located within mapped high-probability susceptibility patches, with E4-predicted probabilities of 0.959, 0.958, 0.933, 0.889, and 0.818, respectively. These sites are mainly distributed along the deeply incised valley system and local slope-transition zones, where steep valley-side slopes, river erosion, and engineering corridors may jointly contribute to slope instability. The enlarged local susceptibility maps show that the recorded landslide locations are embedded within or adjacent to continuous high- and very-high-susceptibility belts. For Site E, the local susceptibility inset is further compared with the UAV photograph: the inferred landslide perimeter and movement direction correspond to a steep valley-side slope section that is mapped as high to very high susceptibility. These representative examples indicate that the final E4 susceptibility map is not only statistically reliable but also spatially consistent with recorded landslide occurrence, visible slope-failure morphology, and plausible alpine-gorge geomorphic controls.
5. Discussion
5.1. Local Adaptation of National-Scale Information: From Rigid Transfer to a Probability Prior
This study shows that national-scale landslide information can improve local susceptibility modelling in data-scarce basins, but that its effectiveness depends on how it is introduced. For alpine gorge regions such as the Parlung Tsangpo Basin, national-scale information is best not used directly as a fixed model, nor forcibly transferred as source-domain samples; rather, it is better transformed into a probability prior that can be relearned and recalibrated by local samples. In other words, the national model does not directly determine the final susceptibility pattern of the target basin but provides the local model with a background risk judgement that can be selected, corrected, and reweighted.
Traditional susceptibility studies rely on the statistical relationships between a target-area landslide inventory and conditioning factors, and their reliability is strongly affected by inventory completeness, sample size, spatial representativeness, and validation method [
33,
78]. Incomplete or systematically biassed inventories have been shown to affect statistical susceptibility models and may distort spatial predictions [
79]. With only 273 valid positive samples retained, the Parlung Tsangpo Basin cannot fully cover the diverse conditioning environments shaped by deeply incised valleys, complex lithology, tectonic fracturing, fluvial erosion, and engineering disturbance. Training a model on local samples alone can therefore leave the results dominated by a limited sample structure, reducing the stability of the susceptibility maps.
In recent years, global-, national-, or large-regional-scale susceptibility studies have provided important supplementary information for local assessment [
53,
80]. However, national-scale landslide–environment relationships are learned from mixed geomorphic settings and triggering contexts, and they cannot be assumed to be directly equivalent to the controlling mechanisms of a specific alpine gorge basin. This issue is particularly relevant here because the source domain includes landslides developed under diverse terrain, climatic, tectonic, and anthropogenic conditions, whereas the Parlung Tsangpo Basin is dominated by deeply incised valleys, active tectonics, lithological contrasts, glacier- and snowmelt-related water supply, freeze–thaw processes, monsoonal rainfall, and local engineering disturbance.
Such source–target differences may introduce negative transfer when national information is introduced in a rigid form. In E2, the national model is applied directly to the target basin, so the discriminative rules learned from heterogeneous source-domain samples directly determine the target-domain predictions. If part of these rules reflects hilly terrain landslides or engineering-related failures rather than alpine-gorge instability, the transferred model may produce biassed probabilities or diffuse high-susceptibility areas. In E3, assigning higher weights to target-domain samples can reduce this problem [
23,
24], but sample-level joint training still allows source-domain patterns with different triggering or developmental mechanisms to enter the decision boundary. Therefore, the weaker performance of E2 and E3 does not simply indicate that national information is useless; rather, it shows that the form in which source-domain information is transferred is critical.
The uneven spatial distribution of the national inventory further helps explain this result. Although the source-domain records are national in extent,
Figure 2 shows that they are much denser in eastern and central China than in western high-mountain regions, including the Tibetan Plateau. Therefore, the national model may be trained mainly on environmental settings that are better represented in the inventory, whereas the target alpine gorge conditions receive fewer source-domain samples. This imbalance can limit the direct applicability of E2 and may also dilute local information when national and target samples are simply merged in E3.
The contrast between E2 and E4 further clarifies both the value and the limitation of the probability-prior strategy. On the one hand, the improvement of E4 relative to the local baseline indicates that the national model contains useful broad-scale information on landslide-prone environmental backgrounds. On the other hand, the poorer calibration and more diffuse susceptibility pattern of E2 show that the national model may also carry source-domain bias and should not be used as a fixed national predictor. In E4, this probability is implemented as an additional input variable rather than as the final prediction. Therefore, local correction is achieved only through the relearning process: the target-domain model can downweight or reshape the prior when local samples and the original conditioning factors provide contradictory evidence. If the national model assigns high probabilities to locally stable terrain, or low probabilities to locally unstable terrain, and such contradictory cases are insufficiently represented in the target-domain samples, E4 may not fully remove the bias and may even propagate part of it. Thus, the proposed strategy should be interpreted as a locally constrained recalibration mechanism rather than as a guarantee that national-model bias is completely eliminated. The better calibration, lower probability error, and less diffuse susceptibility pattern of E4 suggest that, in the present basin, the available local samples were sufficient to correct part of the national-model bias while retaining useful broad-scale information.
Selecting source-domain samples that are environmentally similar to the target basin is also a reasonable way to reduce domain mismatch, and it is closely related to domain adaptation and area-of-applicability analysis [
9,
10,
21]. Nevertheless, using similarity-based source selection as the main strategy would introduce additional methodological choices, including the definition of similarity metrics, thresholds, and factor weights. It could also reduce the diversity of national-scale landslide-background information and exclude some less frequent but potentially relevant alpine or tectonic cases. For this reason, the present study retained the full national source domain and used E2 and E3 as diagnostic comparisons for direct and sample-level transfer, while E4 was designed to transform national information into a locally correctable probability prior. The results are consistent with the interpretation that probability-prior transfer can weaken the misleading influence of source-domain samples whose controlling mechanisms differ from those of the target basin, without discarding the useful broad-scale information contained in the national inventory.
5.2. Discrimination, Probability Calibration, and Mapping Reliability
Landslide susceptibility assessment requires not only that a model distinguish landslide from non-landslide samples, but also that its output probabilities support continuous mapping and class division. AUC, AP, and F1-score reflect ranking and classification ability but cannot reveal whether the predicted probabilities match empirical occurrence frequencies [
81]. For studies whose final products are susceptibility-probability maps, probabilistic reliability should therefore be assessed with the Brier score, log loss, and ECE in addition to the discrimination metrics.
On the spatially independent test set, the AUC, AP, F1-score, and accuracy of E4 were 0.901, 0.895, 0.828, and 0.821, respectively, generally better than those of E1, E2, and E3. Recent susceptibility studies in mountainous and gorge regions have shown that machine-learning models can usually attain high discrimination, but that performance is affected by study scale, inventory quality, negative-sample construction, and validation method [
34,
81,
82,
83,
84]; AUC values from different regional studies should therefore not be ranked in an absolute sense. The significance of the present results is that, under a unified factor system, a common test set, and spatially independent validation, E4 still outperformed the other three strategies consistently, indicating that the improvement stems mainly from expressing national-scale information as a probability prior rather than from differences in the evaluation settings.
In terms of probabilistic reliability, the Brier score, log loss, and ECE of E4 fell to 0.127, 0.401, and 0.055, respectively—the best of the four strategies. A recent study of geological-hazard susceptibility along a railway corridor also used the Brier score and ECE to assess probability reliability, with its random-forest model achieving AUC = 0.85, Brier score = 0.1589, and ECE = 0.0572 on the validation set [
85]. Compared with such studies, E4 maintained high discrimination while achieving lower probability error, indicating that its continuous susceptibility probabilities are better suited to spatial zoning and the identification of key patches.
The contrast between E2 and E4 further shows that a higher landslide-detection rate does not guarantee reliable mapping. At equal recall, the precision, log loss, and ECE of E2 were all inferior to those of E4, indicating that direct national-model transfer is prone to probability bias and misclassification. By relearning and recalibrating the national probability prior with local samples, E4 improved probability credibility while preserving landslide-detection ability. Its advantage is therefore not merely that it classifies more accurately, but that it outputs more reliable continuous susceptibility probabilities, providing a more robust basis for susceptibility class division, hazard screening, and risk interpretation.
5.3. Geomorphological Interpretation of the Spatial Susceptibility Pattern
According to
Figure 8, the susceptibility map produced by E4 displays an extensive low-susceptibility background with concentrated high-susceptibility patches. The high-susceptibility areas occur mainly on both sides of the main valley, at tributary confluences, on strongly incised valley slopes, and near transportation corridors, indicating that the prior-informed model does not simply amplify the national-scale risk background but confines the high-probability areas to slope units with clear conditioning factors.
The spatial-level validation further supports this geomorphological interpretation. Although the high and very high susceptibility classes account for only 3.49% of the study area, they contain 71.43% of the independent test landslide points, and their landslide density is 107.46 times that of the very low and low classes combined. This concentration pattern confirms that the high-probability patches identified by E4 are not randomly distributed, but are closely associated with observed landslide occurrence and key alpine-gorge conditioning environments.
Figure 9 provides site-scale visual evidence for the spatial validation results. In the enlarged maps, the recorded landslide locations of Sites A–E are located within or close to high-probability patches along valley-side slopes. At Site E, the UAV image shows a slope-failure area, and its inferred perimeter and movement direction are consistent with the mapped high- and very-high-susceptibility belt. These observations support the geomorphic plausibility of the E4 susceptibility map.
This pattern is consistent with current understanding of the southeastern Tibetan Plateau and adjacent alpine gorge regions. Studies in the Yarlung Tsangpo River Basin and the Three Parallel Rivers region have shown that landslide occurrence is controlled not by a single factor but by the long-term coupling of tectonic uplift, fluvial incision, glacial processes, and climatic weathering [
16,
86]. In the Parlung Tsangpo Basin, active faults continuously fracture and weaken rock masses, creating structural discontinuities that reduce slope strength. At the same time, rapid river incision and lateral erosion by the Parlung Tsangpo River and its tributaries remove support at slope toes, promote stress redistribution, and accelerate gravitational deformation of steep valley-side slopes. These tectonic and fluvial processes jointly generate highly unstable topographic conditions that favour landslide initiation.
In addition, the basin is located in a transition zone strongly influenced by modern glaciers and seasonal cryospheric processes. Repeated freeze–thaw cycles can enlarge existing fractures, increase rock-mass disintegration, and facilitate the detachment of unstable material from steep rock slopes. Glacier retreat and meltwater infiltration may further alter slope hydrology and reduce the mechanical stability of weathered materials. Consequently, fault-controlled rock weakening, river erosion, glacial activity, and freeze–thaw weathering act together to control the spatial concentration of landslides in the basin. The high-susceptibility patches identified by E4 are mainly distributed along deeply incised valleys, tributary junctions, and fault-influenced slope sections, which is consistent with this coupled geomorphic and geological control mechanism.
The high susceptibility near transportation corridors is also regionally consistent. A study of the Yunnan–Tibet transportation corridor found that high- and very-high-susceptibility areas were distributed mainly along major river valleys and highway belts, reflecting the superimposed effects of gorge terrain and road-engineering disturbance [
87]. The present results likewise show that some high-susceptibility patches concentrate in valley–road-adjacent areas, indicating that the mapped pattern is consistent with spatial associations between natural geomorphic processes and human disturbance. Together with existing cascading-hazard studies in the Parlung Tsangpo and Yarlung Tsangpo Grand Canyon regions, the coupling of glaciers, freeze–thaw processes, avalanches, channel transport, and river blockage further explains the high susceptibility in the main-valley and tributary areas [
88,
89].
By comparison, the high-probability areas of E2 are more diffuse, indicating that direct national-model transfer tends to project broad-scale average patterns onto the target basin and raise the background susceptibility. By recalibrating the national probability prior with local samples, E4 concentrates the high-susceptibility areas where landslides actually cluster and where the geomorphic setting is plausible. Its maps therefore offer not only better statistical performance but also clearer geomorphological interpretability, providing a more stable spatial basis for identifying key slope sections, screening transportation-corridor risks, and supporting engineering route selection in alpine gorge basins.
5.4. Limitations and Future Work
Although E4 performed well in discrimination, calibration, and spatial mapping, this study has several limitations. First, the target-domain inventory is small and the samples participate in modelling mainly as points. The completeness, spatial accuracy, and representation of an inventory affect susceptibility-model training and zoning [
90,
91], and in alpine gorge regions such as the Parlung Tsangpo Basin, remote-sensing interpretation is further hampered by topographic shadow, vegetation cover, bare rock, and valley sediments. Future work should combine high-resolution remote sensing, unmanned-aerial-vehicle surveys, InSAR monitoring, and field verification to improve the inventory and should gradually extend point samples to polygon-based inventories so as to represent landslide boundaries, scales, and types more accurately.
Second, this study mainly evaluates spatial landslide susceptibility. The conditioning factors used in the model are dominated by relatively stable spatial variables, including topography, geology, geomorphology, vegetation, land cover, and human activity. Annual mean temperature was retained as a climatic background variable to describe the regional thermal setting and elevation-related climatic gradients. Rainfall and freeze–thaw processes were not included as independent dynamic predictors in the final model. The main reason is that the landslide inventory used in this study consists mainly of multi-year historical records and interpreted landslide points, which are difficult to match reliably with specific rainfall events or freeze–thaw periods. Annual precipitation was examined during the initial factor preparation stage, but it was not retained because it showed spatial redundancy with the retained climatic and topographic background variables. In addition, annual precipitation totals provide limited information on short-duration or extreme rainfall events that commonly trigger landslides. Future work should incorporate event-dated landslide inventories, time-series rainfall and extreme-precipitation indices, freeze–thaw frequency or intensity indicators, glacier/snow-cover change or distance-to-glacier/moraine variables, ground-motion parameters, fluvial-erosion rates, and records of engineering activity, thereby extending static susceptibility assessment towards dynamic susceptibility or hazard assessment [
88,
89,
92,
93].
Third, this study evaluates the proposed strategy only in the Parlung Tsangpo Basin. Although this basin is suitable for testing how national-scale information can be transformed and locally recalibrated under sample-limited alpine-gorge conditions, the results should be interpreted as evidence from a representative target basin rather than as proof of cross-basin general applicability. Model generalisation can easily be overestimated if spatial autocorrelation and independent validation are ignored [
47], whereas unified datasets and reproducible benchmark workflows are important for improving the comparability of cross-regional results [
84]. Future work should therefore conduct external tests in multiple representative basins, combining spatial cross-validation, independent-basin validation, and unified benchmark workflows to delimit the applicability boundaries of the national-probability-prior and local-sample-correction strategy more precisely.
Finally, this study has not yet fully decomposed the sources of uncertainty in the susceptibility results. Subsequent research could evaluate the effects of inventory error, negative-sample construction, factor resolution, model structure, calibration, and regional-transfer differences, and could express prediction confidence through multi-model ensembles, bootstrap resampling, spatial-blocking sensitivity analysis, and uncertainty mapping, thereby providing a more robust basis for disaster prevention and mitigation, engineering route selection, and the screening of priority hazards.
6. Conclusions
To address the cross-scale problem of incorporating national landslide information into local susceptibility assessment under limited target-domain samples, this study proposed a prior-informed modelling strategy in which national-scale susceptibility probabilities are transformed into probabilistic knowledge and then relearned and corrected using local samples. This design avoids treating the national model as directly transferable and instead allows the target-domain model to adapt external knowledge to the geomorphic and environmental conditions of the Parlung Tsangpo alpine gorge basin.
The experimental results demonstrate that the prior-informed strategy achieved the most robust performance among the tested modelling pathways. On the independent test set, E4 obtained the highest AUC, AP, F1-score, and accuracy, with values of 0.901, 0.895, 0.828, and 0.821, respectively. Compared with the local baseline, its AUC and AP increased by 0.053 and 0.052, indicating that the national-scale probability prior improved both discrimination ability and prediction reliability under sample-limited conditions. The low log loss, Brier score, and ECE further suggest that the proposed strategy produced better-calibrated susceptibility probabilities rather than only improving classification accuracy.
The susceptibility maps further show that E4 maintained a low-probability background while concentrating high-susceptibility zones along the main valley, tributary confluences, strongly incised slopes, and transportation corridors. This spatial pattern is consistent with the combined effects of deep valley incision, tectonic and lithological contrasts, fluvial erosion, and local engineering disturbance in the Parlung Tsangpo Basin. Therefore, the prior-informed strategy not only improved quantitative prediction performance but also enhanced the geomorphic interpretability and practical mapping value of the susceptibility results.
Future work should further test the transferability of this national-probability-prior and local-sample-correction framework in other alpine gorge basins with different geomorphic settings and landslide inventories. In addition, integrating multi-temporal remote sensing observations, dynamic triggering factors such as rainfall and seismicity, and uncertainty-aware modelling may help extend the present static susceptibility framework towards more transferable and operational landslide hazard assessment.
In addition, future studies could explicitly incorporate source-domain similarity screening or area-of-applicability analysis before transfer. Stratifying the national inventory by geomorphic setting, triggering mechanism, or human-disturbance intensity may help determine which types of source-domain samples are most transferable to alpine gorge basins, and would further clarify the applicability boundary of the proposed prior-informed framework.