Next Article in Journal
Deformation Laws of Coal Mining-Affected Slopes in Loess Gully Area
Next Article in Special Issue
Spatial Leakage in Classifying NASA FIRMS Thermal Anomalies as Wildfire Incidents: A Leakage-Controlled Evaluation of Radiometric, Temporal, and Spatiotemporal Features
Previous Article in Journal
A Systematic Comparison of Statistical and Machine-Learning Models for Mapping Landslide Susceptibility: Evidence from the 2018 Rainfall-Induced Landslides in Hiroshima
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Comparative Landslide Susceptibility Mapping in Longchuan, Guangdong Province, China, Using Explainable Machine Learning

Faculty of Humanities and Social Sciences, Macao Polytechnic University, Macao 999078, China
*
Author to whom correspondence should be addressed.
GeoHazards 2026, 7(3), 88; https://doi.org/10.3390/geohazards7030088
Submission received: 5 June 2026 / Revised: 24 June 2026 / Accepted: 29 June 2026 / Published: 19 July 2026
(This article belongs to the Special Issue Machine Learning and AI in Geohazard Detection and Prediction)

Abstract

Landslide susceptibility assessment is essential for disaster-risk reduction, land-use regulation, and territorial spatial planning in mountainous and hilly regions. However, the practical application of machine learning-based susceptibility models is often limited by the trade-off between predictive accuracy and model interpretability, as well as the instability of factor importance across different algorithms and study areas. Taking Longchuan County in northeastern Guangdong Province, China, as a case study, this research develops a comparative explainable machine learning framework to evaluate landslide susceptibility and examine the cross-model stability of SHAP-based factor attribution under local geo-environmental conditions. Fifteen conditioning factors were initially derived from multi-source geological, topographic, hydrological, environmental, and anthropogenic datasets. After multicollinearity screening using Pearson correlation analysis, twelve key factors were retained for model construction. A total of 363 historical landslide points and an equal number of non-landslide samples were divided into training and testing datasets using a stratified 70:30 sampling strategy. Eight machine learning models were optimized through grid-search parameter tuning and then comparatively evaluated. The results show that all models achieved strong predictive performance, with test-set AUC values exceeding 0.938. Among them, the Gradient Boosting Decision Tree model performed best, with an AUC of 0.9520 and the most stable control of overfitting, followed closely by CatBoost with an AUC of 0.9512. SHAP-based interpretation further revealed that the normalized difference water index, relief, and distance to rivers were the dominant factors controlling landslide susceptibility in the study area, with the normalized difference water index serving as a key explanatory factor across models. The proposed framework improves the transparency and reliability of landslide susceptibility assessment and provides a methodological reference for localized, explainable machine learning applications in geohazard risk management.

1. Introduction

Landslides are among the most widely distributed and frequently destructive geological hazards in the mountainous and hilly regions of China. Under intense or prolonged rainfall, they often form cascading hazards together with collapses and debris flows, posing serious threats to human life, property, and regional sustainable development. Scientific and reliable landslide susceptibility assessment, which delineates high-risk spatial units, provides a scientific basis for disaster-risk management, prevention and mitigation decisions and territorial spatial planning. Over the past two decades, the methodology of landslide susceptibility assessment has shifted from expert-weighted approaches, bivariate statistics, and multivariate statistical regression to data-driven machine learning and even deep learning, with steadily improving discrimination accuracy. Yet as model complexity increases, the tension between high-accuracy black boxes and transparent explanation has become increasingly acute. Disaster management requires traceable and accountable factor attribution, whereas it is often difficult to make the decision mechanisms of high-accuracy models transparent. Post hoc explainability methods such as SHAP, which have rapidly developed in the past three years, have partly bridged the gap between accuracy and explanation. Existing studies, however, usually validate explanatory effectiveness with a single algorithm and a single study area. Key issues, including whether factor importance drifts across algorithms and whether explanatory conclusions are physically consistent, remain insufficiently tested. As a result, explainability often becomes an accessory to accuracy rather than independent evidence of reliability.
At the same time, substantial local geological-mechanism knowledge and disaster inventories have accumulated in China’s loess regions, reservoir-bank regions, and earthquake-damaged regions. Yet these studies have mostly remained at the level of statistical modeling or disaster identification. Few systematic efforts have integrated local geological priors, horizontal multi-model comparison, and explainable attribution within a single framework. Longchuan County, located in northeastern Guangdong Province, in the upper reaches of the Dongjiang and Hanjiang rivers, contains interlaced mountains and hills, reservoir-bank slopes, and Danxia and granite-weathering landforms. Shallow landslides occur frequently under heavy rainfall, but the area has long lacked a targeted high-quality landslide inventory and a multi-model comparative assessment constrained by explainability.
Against this background, this study takes the Longchuan basin as the research object and constructs a multi-model comparative framework for explainable machine learning. On the one hand, a horizontal comparison of eight mainstream machine learning models is used to test the stability of SHAP-based factor attribution across models, responding to blind spots in the international literature concerning explanatory credibility. On the other hand, local terrain, geological, hydrological, and human-activity factors in Longchuan are incorporated into explainable modeling, bridging the gap between local empirical research and frontier methods. The goal is to provide an accurate, transparent, and accountable scientific basis for regional landslide prevention and mitigation.
The mainstream approach to landslide susceptibility modeling couples historical landslide inventories with topographic, geological, hydrological, vegetation, and human-activity factors and then trains classifiers to output spatial probabilities. In algorithm selection, international research has long compared random forest (RF), support vector machine (SVM), logistic regression (LR), and gradient-boosting models. Saha et al. systematically compared single and ensemble models in the Rudraprayag district of the Garhwal Himalayas and found that the ANN-RF-LR ensemble outperformed any single algorithm in both goodness of fit and predictive ability, establishing an empirical consensus that ensembles outperform individual models [1]. Dou et al. built bagging, boosting, and stacking frameworks using SVM as the base learner in northern Kyushu, Japan, verified the best performance of SVM-Boosting, and identified rainfall as the dominant triggering factor [2]. Such studies established the basic paradigm of multi-model comparison: using AUC/ROC as a common metric to judge algorithmic performance on the same dataset.
After adopting this paradigm, research quickly turned toward more complex heterogeneous ensembles and stacking strategies closely linked to practical data from typical hazard areas in China. Hu et al. used Luxi County in western Yunnan as an example and adopted a stacking ensemble with SVM, ANN, NB, and LR models as base learners; the maximum AUC reached 0.94, proving that stacking models are superior to single algorithms for accurately delineating susceptible areas [3]. Because reservoir-bank landslides induced by impoundment and drawdown are dense in the Three Gorges Reservoir area, this region has become an important testing ground for ensemble learning in China. Gong et al. proposed three ensemble schemes in Fengjie County by integrating statistical models and machine learning [4]. Zhou et al. combined non-landslide sample optimization with heterogeneous stacking in Fengjie County. Guo et al. systematically compared the generalization ability of homogeneous and heterogeneous ensembles in Wanzhou District [5]. Wu and Wang further introduced EasyEnsemble to handle class imbalance and verified the superiority of gcForest [6]. Liu et al., in Anhua County in Hunan; Liu et al., in Enshi City in Hubei; and Tian et al., in mountainous Chongqing, promoted the hierarchical and multi-source development of stacking ensembles through three-layer stacking and spectral-feature fusion [7,8,9]. Together, these studies show that Chinese scholars have not been satisfied with simply replicating international algorithm comparisons; instead, they have continuously deepened and broadened ensemble strategies to respond to China’s complex geological structures and diverse hazard types.
International research has also deepened its reflection on the internal structure of ensembles in recent years. Matougui et al. systematically compared homogeneous and heterogeneous ensembles in the Djebahia region of Algeria and showed that homogeneous ensembles performed better in AUC and threshold-based metrics but were less robust than heterogeneous schemes such as dynamic ensemble selection (DES) [10]. This finding offers a methodological reminder to the stacking frameworks commonly favored in China. Meanwhile, for small samples and class imbalance, Hussain et al. introduced generative adversarial networks (GANs) to synthesize balanced samples and used PS-InSAR for cross-validation [11]. Lucchese et al. proposed an RF-ANN hybrid bagging ensemble [12]. Zhang et al. compared optimization strategies such as PSO-SVM on the Qinghai–Tibet Plateau [13]. Kumar et al. constructed a stacking ensemble using multivariate index overlay in the Darjeeling Himalayas [14]. Li et al. combined information-value constrained sampling with stacking learning in Yiling District [15]. The core debate is whether ensemble gains come from algorithmic diversity itself or from hidden optimization in sample and feature engineering. International studies tend to treat ensembles as algorithmic contributions that can be independently tested, whereas Chinese studies more often evaluate ensembles together with localized sample and factor processing. The two traditions therefore differ in their understanding of comparability.
With the success of convolutional neural networks (CNNs) in remote-sensing image analysis, landslide susceptibility assessment has entered a deep learning stage. Wu et al. were among the first to combine SMOTE oversampling with CNNs in Wanzhou District in the Three Gorges Reservoir area, significantly improving assessment accuracy [16]. Youssef et al. compared SVM, CNN-1D, and CNN-2D models in the Asir region of Saudi Arabia and found that two-dimensional convolution achieved the highest accuracy because it could capture spatial neighborhood structure. This established the judgment that spatial structural information can be effectively used by deep models. Deep models have since evolved along two paths. The first is structural coupling: Zhao et al. proposed a shared-parameter multidimensional CNN to balance computational efficiency and overfitting control, and Wang Shouhua et al. coupled a certainty factor, CNN, and LSTM to achieve an AUC of 0.953 in Wuzhou, Guangxi [17]. The second is the introduction of attention and Transformer mechanisms: Zhao et al. built CTLGNet, integrating a CNN and Transformer to extract local and global features simultaneously [18]; Li et al. designed a lightweight CNN–Transformer multi-attention model that balances accuracy and deployability [19]; Qu et al. combined dynamic buffer-based sample expansion with a Transformer–CNN coupling in Baoji [20]; and Yuan and Chen proposed a three-stage ST–Transformer framework for national-scale assessment [21]. In China, Kong et al. coupled the information-value method with a CNN (IV-CNN) on the Loess Plateau, showing the localized grafting of statistical priors [22], while Liang et al. enhanced CNNs with bivariate methods in the Himalayas while considering interpretability [23].
The accuracy advantage of deep learning is accompanied by a collapse of interpretability, giving rise to a repeatedly discussed paradox: the deeper and more complex the model representation becomes, the less transparent its decision mechanism is, while disaster management requires traceable and accountable factor attribution. International research has responded by embedding XAI directly into deep learning frameworks. Alqadhi et al. combined CNN-LSTM with explainability methods in the Nainital region and identified terrain and human activities as key factors [24]. This path has also been adopted in China, but Chinese studies place greater emphasis on coupling deep models with InSAR deformation monitoring to improve dynamic performance and physical credibility. For example, Li et al. integrated SBAS-InSAR, YOLOv5n, and SHAP for detecting landslide hazards in open-pit mines [25]. The divergence lies in the source of trust. International studies regard deep learning interpretability mainly as a post hoc algorithmic attribution problem, whereas Chinese studies tend to provide external physical constraints for deep models through multi-source observations, especially deformation data, thereby using data credibility to compensate for model opacity.
The reliability of susceptibility assessment depends not only on models but also on the selection and representation of conditioning factors, the sampling strategy for positive and negative samples, and uncertainty in the results. At the factor level, international studies focus on statistical independence and information redundancy. Bravo-Lopez et al. compared multiple feature-selection algorithms in Cuenca, Ecuador, and achieved the best results with XGBoost [26]. Mind’je et al. analyzed the spatial correlations of ten factor categories in Rwanda using frequency ratios [27]. Kaynak systematically integrated multiple feature-selection methods and artificial neural networks, verifying the dual improvement that dimensionality reduction brings to accuracy and interpretability [28]. Chinese studies have questioned the representation of factors themselves more deeply. Huang et al. showed that replacing conventional linear factors with continuous spatial-density factors can significantly reduce prediction uncertainty [29]. Liu et al. compared five models and six feature-selection methods and confirmed that recursive feature elimination optimized RF performed best [30]. Liao et al. examined the interaction between raster resolution and dominant factors in Wushan-Wuxi [31]. Chang et al. used slope units as assessment units to handle spatial heterogeneity [32]. The central debate is whether factor contribution is an intrinsic property or a relative quantity that drifts with representation, spatial unit, and resolution. Extensive Chinese empirical research reveals the ubiquity of such drift and substantively challenges the implicit assumption, common in international studies, that factor importance can be transferred across regions.
Sample strategy, especially the selection of non-landslide negative samples, is another long-underestimated source of error. Chinese research has formed a relatively dense set of methods on this issue. Liu et al. revealed the influence of sampling strategy on model stability in Taojiang County and warned that AUC is not the only reliable criterion [33]. Ning and Tie proposed negative-sample reliability sampling based on fuzzy membership [34]. Yang et al. used frequency ratios to guide negative-sample selection and sample-ratio stratification [35]. Liu et al. innovatively used SHAP factor importance to guide Bayesian negative-sample optimization, improving the AUC by 8.2–9.0% in three counties [36]. Guo et al. verified the interaction between environmental-factor combinations and negative-sample strategies in granite collapse–gully areas [37]. This line of research shows that Chinese studies have elevated sample reliability from an empirical operation to a quantifiable and optimizable modeling component. By contrast, international studies treat negative samples more sporadically, usually embedding them in preprocessing rather than testing them independently.
At the levels of uncertainty and validation, international and Chinese studies also differ in orientation. International studies place more emphasis on statistical quantification of uncertainty. Le et al. compared uncertainty and interpretability across five model types using Bayesian optimization and evaluated stability through Monte Carlo simulation [38]. Chinese studies focus more on the geological sources and engineering implications of uncertainty. Xiao et al. characterized the spatial uncertainty of areal landslide susceptibility through weak-interlayer controlled sliding mechanisms [39]. Meng and Xing, and Wang et al., provided dynamic checks on susceptibility results using InSAR deformation [40,41]. Xiao et al. proposed uncertainty-aware ensembles and dynamic threshold optimization [42]. Huang et al. developed an ArcGIS V10.8-integrated SVM-LSM toolbox to improve mapping efficiency and reproducibility [43]. Comparative studies of rockfall susceptibility by Gull et al. and Pokharel et al. further suggest that statistical and physical models have different applicability boundaries across hazard types [44,45].
In the past three years, post hoc explainability methods represented by SHAP (Shapley Additive Explanations) have rapidly become a frontier in landslide susceptibility research. This directly challenges the naive assumption that factor importance is model-independent and highlights the necessity of explainability analysis in multi-model comparison. Since then, the combination of XAI and machine learning has spread rapidly worldwide. Khan et al. and Utthasini et al. used XGBoost with SHAP and variable selection in the Himalayas to reveal distance to roads, elevation, and rainfall as controlling factors [46]. Achu et al. combined multi-depth soil profiles with RF and SHAP and verified that slope units outperform grid units [47]. Shao et al. used SHAP in Italy to analyze the spatial heterogeneity of rainfall thresholds [48]. This represents another technical route, moving from post hoc explanation toward intrinsically interpretable models.
Chinese AI applications show two clear characteristics. First, SHAP is deeply embedded in ensemble and optimization frameworks. Huang et al. proposed a stacking–SHAP enhanced model that improves accuracy and interpretability simultaneously [49]. Yan et al. built a heterogeneous hybrid model to strengthen explanatory robustness [50]. Qiu et al. jointly used PDP, LIME, and SHAP to systematically improve model transparency [51].
Although AI has partly bridged the gap between accuracy and explanation, both Chinese and international studies still reveal a deeper tension. International studies tend to treat SHAP as a unified explanatory language transferable across models and regions, often validating the superiority of an algorithm–explanation combination in a single study area. Chinese studies, however, are strongly anchored in the local empirical realities of specific Chinese hazard regions, such as loess, reservoir-bank, and earthquake-damaged areas, and repeatedly find that factor importance, optimal spatial units, and even optimal models drift by study area. The warning raised by Abdelkader, Csamer, and others that a high ROC-AUC does not necessarily mean practical reliability has not been fully internalized by most AI studies. Many studies still use discrimination accuracy as an implicit proof of explanatory validity, while lacking systematic tests of the stability and physical consistency of explanatory conclusions themselves.
Another important line of Chinese landslide research is localized exploration oriented toward specific geological units, especially the Loess Plateau and the Three Gorges Reservoir area. In the loess region, Sun et al. systematically reviewed progress in geological-hazard investigation and prevention in western loess areas and proposed strengthening risk investigation and monitoring under new conditions of climate variability and human activity [52]. Zhang et al. revealed that hydrogeological conditions and saturated shear liquefaction of loess are key controls on persistent landslides in Heifangtai terraces and established liquefaction-susceptibility criteria [53]. Tang et al. used PCA coupled with statistical models on the Shanxi Loess Plateau to compare the causal factors of loess landslides and rockfalls [54]. These studies have accumulated substantial local geological knowledge, but their methods mostly remain at the level of statistical modeling or disaster identification and have not been fully integrated with the frontier framework of explainable machine learning multi-model comparison.
In summary, two nested gaps remain in current research. First, frontier international frameworks for explainable machine learning are not fully adapted to local conditions in terms of the stability and physical consistency of explanatory conclusions. Most studies validate explanatory effectiveness with a single algorithm and a single study area. As a result, explainability becomes an accessory to accuracy rather than independent evidence of reliability. Second, China’s rich local empirical research remains theoretically underdeveloped. Loess, reservoir-bank, and earthquake-damaged regions have accumulated abundant geological-mechanism knowledge and disaster inventories, but these are mostly tied to statistical modeling or disaster identification. Few systematic studies integrate local geological priors, multi-model comparison, and explainable attribution into one assessment framework. For specific basins such as Longchuan, there is still no targeted high-quality landslide inventory or factor system, nor has any multi-model comparative assessment under explainability constraints been carried out.
Therefore, this study constructs an explainable machine learning multi-model comparative framework for the Longchuan basin. It tests the cross-model stability of SHAP factor attribution through horizontal multi-model comparison, incorporates local geological conditions and factor systems into explainable modeling, and provides an accurate, transparent, and accountable scientific basis for regional landslide prevention and mitigation.

2. Materials and Methods

2.1. Study Area

As shown in Figure 1, Longchuan County is located in northeastern Guangdong Province, in the upper reaches of the Dongjiang and Hanjiang rivers. It borders Xunwu County and Dingnan County of Jiangxi Province to the north, Wuhua County and Xingning City to the east, Dongyuan County of Heyuan City to the south, and Heping County of Heyuan City to the west, with a total area of approximately 3081 km2. The distribution of landslide points and the terrain of the study area are shown in Figure 1. The region is dominated by mountains and hills and contains diverse geomorphic types. Continuous middle–low mountains with relatively high terrain are developed in the northeast and west. The terrain becomes gentler along the main Dongjiang river corridor in the central part, forming bead-like valley basins. Overall, the geomorphic pattern is characterized by mountains encircling the area, rivers crossing it, and basins scattered within it. The highest elevation in the county is approximately 1312 m and the lowest is approximately 43 m. Longchuan County has a mid-subtropical monsoon climate, with mild conditions, abundant rainfall, long summers and short winters, a long sunshine duration, a long frost-free period, and a distinct monsoon. The annual mean temperature is about 21 degrees C (21.5 degrees C in 2020), and annual precipitation is about 1378 mm. Precipitation is concentrated mainly from April to September, with heat and rainfall occurring in the same season and an uneven intra-annual distribution.
From the perspective of terrain and geomorphology, the county extends substantially from north to south, with low mountains and hills, valley basins, and reservoir-bank slopes interlaced. The Fengshuba Reservoir in the north and the Huoshan area in the central part form a distinctive combination of mountain water systems and Danxia and granite-weathering landforms. The upper Dongjiang and Hanjiang river systems, reservoir-bank slopes, and numerous mountain streams jointly control slope runoff and bank erosion. During heavy rainfall, collapses, landslides, debris flows, and other cascading hazards are easily triggered. This geomorphic combination gives landslide development clear slope-gradient zoning, proximity to water systems, and sensitivity to engineering disturbance. Especially under heavy or continuous rainfall, the water-content state of shallow residual slope deposits and weathered rock masses changes rapidly, readily inducing shallow landslides, collapse-slides, and small debris flows.

2.2. Data Sources

This study constructs a landslide susceptibility factor system from multi-source basic geo-environmental data, selecting fifteen feature factors in total. The terrain, geomorphology, and geology factors (eight factors) include elevation, aspect, relief, land-use/land-cover type (LULC), distance to faults (DTF), lithology, plan curvature (PlanCur), and profile curvature (ProfileCur). These factors identify slope geometry and geological-structure background and constitute the basic constraints on landslide occurrence. The hydrology and vegetation factors (five factors) include annual mean rainfall, distance to water systems (DTW), topographic wetness index (TWI), normalized difference water index (NDWI), and normalized difference vegetation index (NDVI). NDWI was computed as NDWI = (Green − NIR)/(Green + NIR), using the green and near-infrared bands of Landsat 5/7/8/9 surface-reflectance imagery. It enhances the spectral signal of surface water and near-surface wetness while suppressing vegetation and dry-soil reflectance; it is therefore used here as a proxy for surface wetness and water-body proximity rather than as a direct hydrological measurement. The human engineering activity factors (two factors) include distance to roads (DTR) and normalized difference built-up index (NDBI), representing the disturbance effects of engineering construction on slope stability. This factor system identifies landslide-forming environments from terrain, geology, hydrology, vegetation, and human activity and provides systematic support for susceptibility assessment. The data sources for each factor are listed in Table 1.

2.3. Machine Learning Models

Eight machine learning classifiers were comparatively evaluated: logistic regression (LR), support vector machine (SVM), random forest (RF), extremely randomized trees (ExtraTrees), gradient boosting decision tree (GBDT), extreme gradient boosting (XGBoost), light gradient boosting machine (LightGBM), and categorical boosting (CatBoost). LR and SVM were included as classical low-variance linear and kernel-based baselines. RF and ExtraTrees are bagging-type tree ensembles that aggregate randomized decision trees to reduce variance. GBDT, XGBoost, LightGBM, and CatBoost are boosting-type ensembles that build trees sequentially, with each new tree fitted to the residual error of the current ensemble [55]. Although these four boosting models share this sequential-correction principle, they differ in implementation: XGBoost adds explicit regularization, a second-order loss approximation, and sparsity-aware split finding [56]; LightGBM improves training efficiency through histogram-based splitting and leaf-wise growth; CatBoost uses ordered boosting and symmetric trees to reduce target leakage and improve stability for categorical or small-sample settings. Retaining these related but distinct models allowed model performance and overfitting behavior to be compared under the same Longchuan factor system.

2.4. Hyperparameter Tuning

Hyperparameters were optimized using GridSearchCV v0.19with three-fold cross-validation on the training set, using accuracy as the scoring criterion. The parameter combination with the highest mean cross-validation score was retained and then evaluated on the held-out test set. The main tuned or fixed hyperparameters available from the model record are summarized in Table 2. The same GridSearchCV workflow was applied to CatBoost, LightGBM, and ExtraTrees.

2.5. Sample Construction and Validation Strategy

A total of 363 historical landslide points from the Resource and Environmental Science Data Center, Chinese Academy of Sciences, accumulated to 2024, were used as positive samples. An equal number of non-landslide points were generated within the study area at a fixed step size, producing 726 samples with a positive-to-negative ratio close to 1:1. The samples were divided into a training set (508 samples, 70%) and an independent test set (218 samples, 30%) using stratified random sampling, with the landslide/non-landslide label as the stratification variable. This preserved class balance in both subsets. All model-comparison metrics reported below were computed on the held-out test set, which was not used during hyperparameter tuning. This strategy controls class balance but does not explicitly enforce spatial separation between training and test samples; the implications of spatial autocorrelation are discussed as a limitation in Section 5.

2.6. SHAP-Based Explainability Approach

SHAP (SHapley Additive exPlanations), implemented through the tree-specific TreeExplainer, was used to interpret the best-performing model, GBDT [57]. For each sample, SHAP decomposes the predicted landslide probability into an additive sum of per-feature contributions. Global importance was computed as the mean absolute SHAP value of each feature. In the dependence plots, point colors encode the value of each panel’s automatically selected dominant, i.e., most strongly interacting, factor. This interacting factor is identified separately for each focal feature as the variable with the highest approximate SHAP interaction strength; it therefore varies from panel to panel and is not, by definition, always the single most important factor overall. Pairwise SHAP interaction values and a network representation were then used to assess whether the main factors acted independently or through systematic hydrological, geomorphic, and engineering-disturbance interactions.

3. Results and Analysis

3.1. Analysis of Landslide Conditioning Factors

Terrain, geomorphology, and geological-structure factors show clear spatial selectivity in Longchuan. Landslides are concentrated mainly in low- to middle-elevation basins, mountain-front transition zones, moderate to steep slopes, areas of medium to high terrain roughness, larger relief areas, and belts close to faults. These conditions combine gravitational potential energy, valley incision, fractured rock masses, and stronger human disturbance, making hill–mountain transition zones and structurally fractured belts the key terrain–geological settings for landslide prevention.
Hydrological, vegetation, and human-activity factors further refine this pattern. Landslide points tend to occur against medium-to-high-rainfall backgrounds, near river valleys and runoff-convergence channels, in medium-to-high-TWI zones, in areas with medium to low vegetation cover, and close to roads or built-up fringe zones. These patterns indicate that slope water conditions, toe erosion, drainage modification, and construction disturbance interact with terrain conditions rather than acting as isolated controls.
The main descriptive ranges and spatial associations extracted from Figure 2 are summarized in Table 3. The table retains the factor-level evidence, while the text emphasizes the integrated physical interpretation: landslides in Longchuan are promoted by the superposition of relief, runoff concentration, fault-related fragmentation, and road- or settlement-related disturbance.

3.2. Factor Screening

To avoid the interference of multicollinearity with model performance, the Pearson correlation coefficient (PCC) was used to evaluate correlations among factors. As shown in Figure 3, factors with high correlations (PCCs) were removed according to this analysis.
Among the fifteen landslide susceptibility factors for Longchuan, twelve core feature factors were ultimately retained. The three factors that were removed were elevation, NDBI, and NDVI. The remaining twelve core feature factors were aspect, relief, LULC, DTF, lithology, PlanCur, ProfileCur, rainfall, DTW, TWI, NDWI, and DTR, as shown in Figure 3.

3.3. Model Performance Evaluation

This study constructs a classification evaluation system based on a confusion matrix. Let TP (true positives) denote the number of landslide samples correctly identified, TN (true negatives) the number of non-landslide samples correctly identified, FP (false positives) the number of non-landslide samples incorrectly identified as landslides, and FN (false negatives) the number of landslide samples missed by the model. The four classification metrics are defined as follows.
Accuracy measures the proportion of all samples that are correctly classified by the model:
A c c u r a c y = T P + T N T P + T N + F P + F N
Precision measures the proportion of samples predicted as landslides that are true landslides and reflects the model’s ability to control false alarms:
P r e c i s i o n = T P T P + F P
Recall, also called sensitivity, measures the proportion of actual landslide samples that are correctly identified and reflects the model’s ability to control omissions:
R e c a l l = T P T P + F N
F1-Score is the harmonic mean of precision and recall and comprehensively measures the balance between them:
F 1 = 2 × P r e c i s i o n × R e c a l l P r e c i s i o n + R e c a l l
In addition, the false-positive rate (FPR) and the true-positive rate (TPR) are the horizontal and vertical coordinates of the ROC curve, respectively, and are defined as follows:
F P R = F P F P + T N ; T P R = T P T P + F N
AUC (Area Under the Curve) is the area under the ROC curve. Its value ranges from 0.5 to 1, and a value closer to 1 indicates stronger overall ranking ability for positive and negative samples.
Figure 4 further presents the discrimination-performance differences among the eight models in the form of ROC curves.
Combining Table 4 and Figure 4, all models achieved a test-set AUC of at least 0.938, and the accuracy, precision, recall, and F1-Score all remained above 0.84. This indicates that all eight models have good landslide discrimination ability under the complex geological environment of Longchuan. Nevertheless, significant differences remain among models in terms of discrimination accuracy, overfitting control, and generalization stability. These differences are analyzed below.
As shown in Table 4, in terms of test-set AUC ranking, GBDT ranked first among the eight models with an AUC of 0.9520, followed closely by CatBoost (0.9512). The two models formed the first performance tier, and the AUC gap between them was only on the order of 0.001. LightGBM (0.9467), XGBoost (0.9434), RandomForest (0.9408), and ExtraTrees (0.9407) followed in order, forming the second tier. SVM (0.9391) and LR (0.9386), as classical non-ensemble discrimination models, had relatively lower AUC values and formed the third tier. From the overall shape of the ROC curves in Figure 4, all eight curves lie well above the diagonal random-discrimination baseline and overlap closely at a broad scale, confirming that all models have effective landslide discrimination ability. However, in the low false-positive-rate interval (FPR < 0.1) near the upper-left corner of the ROC plot, the GBDT and CatBoost curves rise more steeply, achieving a higher true-positive rate at a lower false-alarm cost and showing advantages in high-confidence prediction scenarios. The LR and SVM curves lag relatively in this interval, reflecting weaker discrimination ability under strict thresholds.
The discrimination performance of the eight models on the test set shows a clear gradient. GBDT leads with the highest AUC (0.9520). Its accuracy (0.8853), precision (0.8869), recall (0.8853), and F1-Score (0.8852) are all among the best and highly consistent, indicating excellent classification balance. In addition, it produced only one false positive and one false negative in the training set (training accuracy: 0.9961), making it the only one among the five gradient/tree ensemble models that did not completely memorize the training samples and showing the weakest overfitting tendency. CatBoost follows closely (AUC: 0.9512), with balanced metrics (accuracy: 0.8761, precision: 0.8777, F1: 0.8760). Through ordered boosting and a symmetric-tree structure, it kept training accuracy at 0.9528, making it one of the most restrained tree ensemble models in fitting. Together, GBDT and CatBoost form the first performance tier. LightGBM (AUC: 0.9467, metrics stable at 0.8807), XGBoost (AUC: 0.9434, metrics: 0.8532–0.8537), RandomForest (AUC: 0.9408, but with a relatively high number of test-set false positives, FP = 20), and ExtraTrees (AUC: 0.9407, accuracy: 0.8578) follow in performance. However, all four achieved completely correct classification on the training set (training accuracy: 1.0000), meaning that they memorized the training samples and showed a clear overfitting tendency. Their test performance therefore needs to be assessed together with cross-validation and spatial extrapolation ability. SVM (AUC: 0.9391, training accuracy: 0.8760) and LR (AUC: 0.9386, training accuracy: 0.8425), as classical non-ensemble discrimination models, had relatively lower AUC values, but their training and test performances were highly consistent and showed almost no overfitting. They therefore provide the most stable generalization and can be used as baseline models for measuring the performance gain of ensemble learning. Among them, LR is limited by its linear assumptions and insufficiently captures complex nonlinear relationships, giving it the lowest overall discrimination accuracy.
A comprehensive analysis of Table 4 and Figure 4 leads to the following conclusions. First, in terms of test-set AUC, GBDT and CatBoost form the first performance tier; the AUC difference between them is almost statistically indistinguishable, and both can be used as high-accuracy discrimination models. Second, in terms of overfitting control, GBDT and CatBoost are the two most restrained tree ensemble models in terms of training fit. GBDT had only two misclassified training samples, and CatBoost had a training accuracy of only 0.9528. Neither completely memorized the training samples, and both achieved the best balance between fitting and generalization, making them the most recommended models in overall performance. Third, RandomForest, ExtraTrees, LightGBM, and XGBoost all achieved completely correct classification on the training set (training accuracy: 1.0000), meaning that they had fully memorized the training samples during training. Their test-set performance should therefore be assessed together with cross-validation results and spatial extrapolation ability. Fourth, although LR and SVM have relatively lower AUC values, their training–test performance is highly consistent and their generalization is stable, so they can serve as baseline comparison models.
From the perspective of refined landslide-hazard assessment, this study recommends GBDT as the preferred model for landslide susceptibility modeling in Longchuan. Its advantages appear in three dimensions: the highest test-set AUC among the eight models (0.9520), balanced and excellent classification metrics, and the best overfitting-control ability among tree ensemble models. CatBoost and LightGBM can be used in parallel as validation models. The combination of these three boosting models ensures prediction accuracy and strengthens result reliability through mutual verification among models, providing robust decision support for precise landslide prevention and risk management in the mountainous and hilly regions of Longchuan.

4. Explainability Analysis Based on the SHAP Model

To deeply analyze the internal mechanism by which machine learning models identify landslide susceptibility and to reveal the specific patterns of the contribution of environmental factors to landslide occurrence, this study uses SHAP (SHapley Additive exPlanations), based on Shapley cooperative game theory, to interpret the GBDT model, which showed the best comprehensive performance. SHAP assigns a contribution value to each feature and decomposes the model output, namely, landslide occurrence probability, into the cumulative contribution of each feature. It can reveal the relative importance of factors at the global level and explain specific prediction results at the individual-sample level, effectively compensating for the black-box characteristics of traditional machine learning models. This section analyzes the decision mechanism of GBDT from four perspectives: global feature importance and impact direction, nonlinear response of key factors, two-factor interaction effects, and the systematic interaction network. The twelve environmental factors used are NDWI, relief, DTR, ProfileCur, LULC, DTW, DTF, PlanCur, TWI, rainfall, aspect, and lithology.
(1) Global feature importance and impact direction
Figure 5 and Table 5 integrates a SHAP-based global feature-importance bar chart and a beeswarm plot, jointly characterizing the global action patterns of factors from the perspectives of importance ranking and impact direction. The horizontal bar chart on the left side of Figure 5, corresponding to the upper mean |SHAP value| axis, shows that the twelve environmental factors are ranked by mean absolute SHAP value as follows: NDWI > relief > DTR > ProfileCur > LULC > DTW > DTF > PlanCur > TWI > rainfall > aspect > lithology. NDWI has the highest mean absolute SHAP value (about 0.20), followed by relief (about 0.14) and DTR (about 0.06). These three factors are clearly higher than the remaining factors and constitute the three dominant factors in landslide susceptibility prediction for Longchuan. The contributions of intermediate factors such as ProfileCur, LULC, DTW, DTF, and PlanCur decrease in sequence. Aspect and lithology have the lowest contributions, and the mean absolute SHAP value of lithology approaches zero. This distribution indicates that landslide occurrence in Longchuan is mainly controlled by the coupling of water-body characteristics, terrain relief, and water-system activity, while the direct effects of lithology and aspect are relatively limited.
The right side of Figure 5, including the beeswarm plot, further reveals the impact direction of each factor on model prediction. In this plot, the horizontal axis is the SHAP value, with positive values increasing landslide probability and negative values decreasing it, while point color indicates the feature value of each sample, with yellow representing high values and dark purple representing low values. The beeswarm distribution of NDWI shows a very wide two-direction spread. The positive SHAP region on the right is mainly occupied by dark-purple or brown low-value points, whereas high-value yellow-green points tend toward the negative SHAP region. This means that a lower NDWI, or weaker water-body characteristics, significantly increases landslide probability, whereas a high NDWI, representing water bodies themselves or stable low-lying floodplain areas, corresponds to lower risk. This is consistent with its physical meaning. Relief also shows strong two-direction influence. Low-value dark-purple points tend toward the positive SHAP region, while high values tend toward the negative region, suggesting that the effect of relief on landslides is clearly nonlinear and requires further analysis with dependence plots. In the beeswarm distribution for DTR, low-value points spread toward the positive SHAP region, indicating that proximity to roads increases landslide probability and reflecting the close link between road-cut disturbance and slope-drainage modification. DTW shows a pattern consistent with DTR, further confirming the key inducing role of water-system activity. Although rainfall has relatively low overall importance, its high-value yellow points tend toward the positive SHAP region, reflecting the positive triggering effect of high rainfall on landslides and matching the classical physical mechanism of rainfall-induced landslides.
(2) Nonlinear responses of key factors
Figure 6 presents SHAP dependence plots for the twelve environmental factors. Through scatter points and fitted curves, these plots show the nonlinear relationship between factor values on the horizontal axis and SHAP contributions on the vertical axis. The light-blue shading represents the 95% confidence interval. Point colors encode each panel’s automatically selected dominant, i.e., most strongly interacting, factor; this coloring variable is selected separately for each focal factor and therefore varies across panels rather than always representing NDWI.
As the most important factor, NDWI has a dependence curve that is already near zero with a slightly positive level in the lower NDWI interval. As NDWI increases, the curve remains generally flat and declines slightly at the high-value end. However, its confidence interval widens sharply in the low-NDWI zone, extending downward into clearly negative values. This indicates large prediction uncertainty for samples with weak water-body characteristics, whereas the high-NDWI zone, representing strong water-body characteristics, stably corresponds to low landslide risk and is a stable zone.
The dependence curve of relief shows a clear pattern of first high, then low, and then rising again. In the low-relief zone (standardized value < 0.1), SHAP values exceed 0.15 and strongly increase landslide probability. As relief increases, SHAP values continue to decline and reach a trough of about −0.2 to −0.25 in the 0.25–0.30 interval. When relief continues to increase (>0.35), SHAP values rise again. This pattern reflects differentiated landslide mechanisms in different geomorphic positions in Longchuan. Low-relief hilly piedmont areas are frequently affected by human engineering disturbances and are prone to shallow instability; medium-relief areas are relatively stable; and high-relief steep mountains again become landslide-prone under gravity and weathering.
The dependence curve of DTR shows a typical sharp decline followed by a rebound. In near-road areas with standardized DTR values below 0.1, SHAP values reach 0.2–0.4 and strongly increase landslide probability. As the distance from roads increases, SHAP values rapidly fall to near zero and remain low in the 0.2–0.4 interval, then rise slightly near 0.6–0.7 before declining again. This curve indicates that the immediate road-disturbance zone is a critical threshold for abrupt landslide-risk change. The dependence curve of DTW is highly similar and shows a monotonic abrupt decline: when DTW < 0.05, SHAP values reach 0.3–0.4, then drop rapidly and stabilize as the distance from water increases. This confirms that water-system activity is another core driver of landslides in Longchuan.
The dependence curve of ProfileCur shows an inverted-U single peak. In the middle interval (standardized values of 0.40–0.45), SHAP values reach a peak of about 0.03, whereas values at both ends turn negative and drop sharply. This indicates that both strongly concave and strongly convex profile forms are unfavorable for slope stability. PlanCur shows a similar single-peak pattern, with the peak near 0.65 and negative values at both ends. The dependence curve of aspect shows an obvious bell-shaped single peak: SHAP values are positive in the middle aspect interval (0.2–0.5) and become negative at both ends, reflecting the promoting effect of certain aspects on landslides. The dependence curve of TWI shows a monotonic upward trend, with high-TWI areas, where water is prone to accumulate, corresponding to higher positive SHAP contributions. Rainfall shows a weak peak pattern of first increasing and then decreasing, while the SHAP magnitudes of LULC and lithology are generally small and have relatively limited influence on final prediction.
It is worth noting that the point-color coding in the dependence plots further reveals second-order interaction effects. For example, colors in the DTR dependence plot vary with rainfall, colors in the DTW dependence plot vary with DTR, and colors in the TWI dependence plot vary with relief. These patterns suggest interactions among these factor pairs, which are quantified in the following subsection.
(3) Systematic interaction network
To more intuitively show the systematic relationships among factors in landslide susceptibility prediction, Figure 7 presents the twelve factor nodes and their interaction edges in a circular network diagram. Node size corresponds to feature importance, with a scale of 0.02–0.21, and node color corresponds to the signed overall impact direction, with yellow indicating positive effects and dark purple indicating negative effects. Edge thickness corresponds to interaction strength, and edge color corresponds to signed interaction direction.
Figure 8 clearly reveals a core–periphery structure in the network. The NDWI node on the left side of the network has the largest size and brightest yellow color, making it the hub node with the strongest importance and positive effect. The relief node in the lower right is relatively large and dark purple, indicating that it mainly contributes negatively, meaning that increasing relief generally lowers landslide probability. This is highly consistent with the preceding beeswarm and dependence-plot analyses. DTR, DTW, and LULC nodes are orange and occupy positions with moderate positive influence, whereas curvature factors such as ProfileCur and PlanCur are purple and mainly negative. In the edge topology, NDWI–relief forms one of the thickest connections in the figure, corresponding to the strongest interaction identified above, and its color indicates a signed strong interaction. Several relatively thick edges, including NDWI-DTR, NDWI-ProfileCur, NDWI-PlanCur, and NDWI-LULC, all originate from the NDWI node and connect the hydrology–terrain–curvature–cover system. DTR–relief forms another prominent connection across the network. The remaining edges are thin and pale, indicating weak interaction contributions for most factor pairs. This network topology further confirms that the landslide mechanism in Longchuan is a systematic process dominated by the three core nodes NDWI, relief, and DTR, with NDWI acting as the key hub node.
Based on the optimal GBDT machine learning model, landslide susceptibility was predicted, and the study area was divided into five susceptibility levels according to predicted susceptibility probability: very low, low, moderate, high, and very high.
To systematically evaluate the performance differences in GBDT-based landslide susceptibility zoning in Longchuan, in Figure 8. The predicted susceptibility probability was divided into five interpretable probability intervals (Table 6): very low (<0.2), low (0.2–0.4), moderate (0.4–0.6), high (0.6–0.8), and very high (>0.8). The frequency ratio (FR) was used as the core evaluation indicator. FR is defined as the ratio of the proportion of landslides in each class to the proportion of area in that class. It reflects the relative efficiency with which the model identifies historical landslide points. The closer the FR value is to 0, the closer the area is to a no-risk state; the more it exceeds 1, the higher the likelihood of landslide occurrence per unit area in that class.
Although the very-high-susceptibility class occupies 34.06% of the study area, the classification is not arbitrary: FR values increase monotonically from the very low to very high classes and reach 2.5642 in the very high class, which contains 87.33% of historical landslide points within only about one-third of the study area. The probability-threshold classification therefore remains strongly discriminative. Future work should compare these thresholds with quantile or natural-breaks alternatives using the full probability raster to evaluate the sensitivity of area proportions and FR values to the classification scheme.

5. Discussion

Physical credibility and potential representation effects of the NDWI-dominated pattern. The most notable finding of this study is that NDWI ranks first among the twelve factors with a mean absolute SHAP value of about 0.20 and becomes the hub factor in landslide discrimination through its synergy with relief, distance to rivers, and curvature factors. From the perspective of physical mechanisms, this result is reasonable. The SHAP beeswarm plot shows that a lower NDWI, or weaker water-body characteristics, significantly increases landslide probability, whereas high-NDWI areas, such as water bodies themselves and low-lying stable floodplains, correspond to lower risk. This is consistent with the physical understanding that low-lying near-water stable areas are less prone to instability, whereas slope sections with poor drainage conditions and strong fluctuations in water content are prone to instability. At the same time, NDWI, DTR, and DTW are highly synergistic and jointly point to the core driving chain of water-system activity, valley runoff convergence, and slope-toe erosion in Longchuan. However, NDWI is a remote-sensing spectral index, and its high importance for landslides should be treated cautiously. This may partly arise from its integrated representation of moisture background, valley terrain, and even land cover; in other words, it may statistically proxy several physical processes that are not explicitly included or have been removed by correlation screening. This phenomenon echoes Huang et al.’s argument in the introduction that factor representation itself affects prediction uncertainty, as well as the general drift of factor importance with representation, spatial unit, and resolution repeatedly revealed in Chinese research. In other words, the Longchuan case again suggests that SHAP importance ranking reflects the internal decision logic of a model under a specific factor system and study area, rather than a physical essence of factors that can be transferred across regions without verification.
The dominance of NDWI, relief, and DTR identified here is not universal across the explainable-ML landslide literature. For example, Himalayan XGBoost-SHAP studies cited in the Introduction identified distance to roads, elevation, and rainfall as leading factors [46], whereas frequency-ratio work in Rwanda emphasized different terrain and land-cover associations [27]. Such differences support the broader argument that SHAP-based factor importance reflects the joint effect of local geo-environmental setting, factor representation, spatial resolution, and model structure rather than a transferable physical ranking [29]. The strong role of water-related variables in Longchuan is plausibly linked to the county’s dense river-valley settlement pattern, reservoir-bank geomorphology, and rainfall-controlled shallow-slope processes.
A narrow absolute NDWI range does not imply limited discriminative power. SHAP importance reflects how strongly variation in a feature shifts the model’s predicted probability, not the size of the feature’s numerical range. A tree-based model such as GBDT can learn sharp nonlinear thresholds within a narrow value interval, as shown by the low-value transition in the NDWI dependence curve (Figure 6). In addition, NDWI integrates several correlated physical signals, including surface wetness, water-body proximity, and partly land-cover state; small index differences can therefore correspond to meaningfully different slope-hydrology conditions.
Regarding the significance of multi-model comparison for testing explanatory stability, the literature review in the Introduction points out that international studies tend to regard SHAP as a unified explanatory language transferable across models, while Al-Najjar et al. had already revealed the problem that explanatory direction drifts with algorithms. This study provides a partial response to this issue through a horizontal comparison of eight models. Although tree ensemble models such as GBDT, CatBoost, LightGBM, and XGBoost differ in AUC by only 0.001–0.01, their training–fitting behaviors differ significantly. GBDT had only two misclassified training samples, and CatBoost had a training accuracy of only 0.9528; these two models were the most restrained in balancing fit and generalization. By contrast, RandomForest, ExtraTrees, LightGBM, and XGBoost all achieved completely correct classification on the training set (training accuracy: 1.0000), showing an obvious tendency to memorize training samples. This comparison shows that simply using test-set AUC to judge model superiority is limited. Overfitting-control ability and training–test consistency should be included in a comprehensive evaluation. This is consistent with Liu et al.’s warning that AUC is not the only reliable criterion and Abdelkader et al.’s warning that a high ROC-AUC does not equal practical reliability. Accordingly, this study recommends GBDT, which has the best overfitting control, as the preferred model and uses CatBoost and LightGBM for parallel verification, thereby attempting to enhance the credibility of explanatory conclusions through mutual confirmation among models.
Regarding the nonlinear mechanism of relief and coupling with local geomorphology, the SHAP dependence curve for relief shows a complex pattern of first high, then low, and then rising again, revealing differentiated landslide mechanisms in Longchuan. Low-relief hilly piedmont areas are frequently disturbed by human engineering activities and are prone to shallow instability. Medium-relief areas are relatively stable. High-relief steep mountains again become landslide-prone under gravity and weathering. This nonlinearity is precisely where black-box models have advantages over linear models such as logistic regression, and it explains why LR had the lowest discrimination accuracy in this study. At the same time, the color coding of points in the dependence plots reveals second-order interactions such as DTR × rainfall, DTW × DTR, and TWI × relief, while the interaction-value matrix further quantifies NDWI × relief (0.036) as the strongest interaction. This indicates that landslides in Longchuan are not linearly driven by a single factor but are instead systematic processes involving hydrology, terrain, and land cover. Single-factor importance rankings therefore need to be interpreted together with the interaction network.
Several limitations of this study should be acknowledged. First, the sample size is relatively limited, with 363 landslide points and an equal number of non-landslide points. Non-landslide samples were generated using a fixed step size, and more refined strategies such as SHAP-guided or frequency-ratio-guided negative-sample optimization remain to be tested. Second, the train–test split was stratified by class label but was not explicitly spatially blocked. Because landslide-conditioning factors are spatially autocorrelated, conventional random or stratified-random splitting can place spatially proximate, environmentally similar samples on both sides of the split, which may inflate apparent test-set performance relative to fully spatially independent evaluation. The AUC values reported in Table 4 should therefore be interpreted as the held-out random-test performance rather than as a strict estimate of spatial-transfer performance. Third, the proportion of flat or low-slope terrain and model performance after excluding such terrain should be quantified using the slope raster and sample coordinates in future work because easy discrimination of flat areas may influence overall accuracy. Fourth, the factor system is dominated by static geo-environmental factors and lacks physical corroboration from dynamic observations such as InSAR deformation. Finally, the results are based on the single Longchuan basin, and the cross-regional transferability of SHAP attribution requires verification in additional study areas.

6. Conclusions

This study takes Longchuan County, Guangdong Province, as the study area and constructs a multi-model comparative framework for explainable machine learning. Based on twelve core factors retained from fifteen initially selected factors after correlation screening and 363 historical landslide samples, eight machine learning models were systematically compared. The optimal model was then used for SHAP explainability analysis and susceptibility zoning. The main conclusions are as follows:
In terms of model performance, all eight models showed good landslide discrimination ability under the complex geological environment of Longchuan. Test-set AUC values were all at least 0.938, and all classification metrics were stable above 0.84. GBDT ranked first with an AUC of 0.9520, showing balanced and excellent classification metrics and the best overfitting-control ability among tree ensemble models. CatBoost (0.9512) followed closely, and together they formed the first performance tier. RandomForest, ExtraTrees, LightGBM, and XGBoost had similar accuracy, but they completely memorized the training samples and showed obvious overfitting tendencies. SVM and LR showed stable generalization and can serve as baseline comparisons. Considering both accuracy and generalization ability, this study recommends GBDT as the preferred model for landslide susceptibility modeling in Longchuan, with CatBoost and LightGBM used for parallel verification.
In terms of factor mechanisms, SHAP analysis reveals that NDWI, relief, and DTR constitute the three dominant factors of landslide development in Longchuan, while the direct effects of lithology and aspect are relatively limited. NDWI not only has the highest importance but also becomes the hub factor in the systematic interaction network through strong synergy with relief, DTR, and curvature factors. The NDWI × relief interaction reaches 0.036, making it the strongest interaction. Key factors generally show significant nonlinear responses. Relief follows a pattern of first high, then low, and then rising again, while DTR and DTW show a pattern of sharp increase near water and rapid decline away from water. These results indicate that landslides in Longchuan are a systematic process driven by the synergy of water-system activity, relief, and human engineering disturbance.
In terms of susceptibility zoning, the GBDT-based assessment divides the study area into five levels: very low, low, moderate, high, and very high. The frequency ratio of the very high susceptibility zone reaches 2.5642 and includes approximately 87% of historical landslide points. The zoning results are highly consistent with the distribution of historical hazards, verifying the practical reliability of the model.
In terms of methodological significance, this study tests the cross-model stability of SHAP factor attribution through horizontal multi-model comparison and incorporates local terrain, geological, hydrological, and human-activity factors into explainable modeling. It responds to some blind spots in international research concerning explanatory credibility, bridges the gap between local empirical research and frontier explainable methods, and provides a precise, transparent, and accountable scientific basis for landslide prevention, risk management, and territorial spatial planning in the Longchuan basin and comparable mountainous and hilly regions.

Author Contributions

Conceptualization, R.C.; methodology, R.C. and X.W.; software, R.C.; formal analysis, S.Z.; investigation, R.C. and X.W.; data curation, X.W.; writing—original draft preparation, R.C.; writing—review and editing, R.C.; visualization, X.W.; supervision, X.W. and S.Z. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the research projects of Macao Polytechnic University, funding number: RP/FCHS-03/2022.

Data Availability Statement

The data that support the findings of this study will be available in Figshare at https://doi.org/10.6084/m9.figshare.32415018.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Saha, S.; Saha, A.; Hembram, T.K.; Pradhan, B.; Alamri, A.M. Evaluating the performance of individual and novel ensemble of machine learning and statistical models for landslide susceptibility assessment at rudraprayag district of garhwal himalaya. Appl. Sci. 2020, 10, 3772. [Google Scholar] [CrossRef] [Scilit]
  2. Dou, J.; Yunus, A.P.; Bui, D.T.; Merghadi, A.; Sahana, M.; Zhu, Z.; Chen, C.-W.; Han, Z.; Pham, B.T. Improved landslide assessment using support vector machine with bagging, boosting, and stacking ensemble machine learning framework in a mountainous watershed, japan. Landslides 2020, 17, 641–658. [Google Scholar] [CrossRef] [Scilit]
  3. Hu, X.; Zhang, H.; Mei, H.; Xiao, D.; Li, Y.; Li, M. Landslide susceptibility mapping using the stacking ensemble machine learning method in lushui, southwest china. Appl. Sci. 2020, 10, 4016. [Google Scholar] [CrossRef] [Scilit]
  4. Gong, W.; Hu, M.; Zhang, Y.; Tang, H.; Liu, D.; Song, Q. Gis-based landslide susceptibility mapping using ensemble methods for fengjie county in the three gorges reservoir region, China. Int. J. Environ. Sci. Technol. 2022, 19, 7803–7820. [Google Scholar] [CrossRef] [Scilit]
  5. Guo, W.; Liu, X.; Chen, B. Landslide susceptibility assessment based on ensemble learning strategies: A case study of wanzhou district in the three gorges reservoir area, china. Q. J. Eng. Geol. Hydrogeol. 2025, 58, qjegh2024-151. [Google Scholar] [CrossRef] [Scilit]
  6. Wu, X.; Wang, J. Application of bagging, boosting and stacking ensemble and easyensemble methods for landslide susceptibility mapping in the three gorges reservoir area of china. Int. J. Environ. Res. Public Health 2023, 20, 4977. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Liu, Y.; Zhang, G.; Chen, Z.; Xu, Z.; Yuan, Y.; Lian, W.; Wang, S. Landslide susceptibility assessment using a three-level stacking ensemble strategy. Geo-Spat. Inf. Sci. 2026, 1–22. [Google Scholar] [CrossRef] [Scilit]
  8. Tian, Y.; Zeng, T.; Wang, L.; Chen, G.; Yang, S.; Chen, H.; Wang, L. Spectral feature integration and ensemble learning optimization for regional-scale landslide susceptibility mapping in mountainous areas. Remote Sens. 2026, 18, 382. [Google Scholar] [CrossRef] [Scilit]
  9. Liu, L.-L.; Danish, A.; Wang, X.-M.; Zhu, W.-Q. Ensemble stacking: A powerful tool for landslide susceptibility assessment—A case study in anhua county, hunan province, china. Geocarto Int. 2024, 39, 2326005. [Google Scholar] [CrossRef] [Scilit]
  10. Matougui, Z.; Djerbal, L.; Bahar, R. A comparative study of heterogeneous and homogeneous ensemble approaches for landslide susceptibility assessment in the djebahia region, algeria. Environ. Sci. Pollut. Res. 2024, 31, 40554–40580. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Hussain, W.; Shu, H.; Abbas, H.; Hussain, S.; Kulsoom, I.; Hussain, S.; Mustafa, H.; Khan, A.A.; Ismail, M.; Iqbal, J. The generative adversarial neural network with multi-layers stack ensemble hybrid model for landslide prediction in case of training sample imbalance. Stoch. Environ. Res. Risk Assess. 2025, 39, 4507–4526. [Google Scholar] [CrossRef] [Scilit]
  12. Lucchese, L.V.; de Oliveira, G.G.; Pedrollo, O.C. A hybrid random forests and artificial neural networks bagging ensemble for landslide susceptibility modelling. Geocarto Int. 2022, 37, 16492–16511. [Google Scholar] [CrossRef] [Scilit]
  13. Zhang, Y.-B.; Xu, P.-Y.; Liu, J.; He, J.-X.; Yang, H.-T.; Zeng, Y.; He, Y.-Y.; Yang, C.-F. Comparison of lr, 5-cv svm, ga svm, and pso svm for landslide susceptibility assessment in tibetan plateau area, china. J. Mt. Sci. 2023, 20, 979–995. [Google Scholar] [CrossRef] [Scilit]
  14. Kumar, S.; Singh, G.; Karmakar, R.; Mishra, A.K. Stacking ensemble for improved landslide susceptibility mapping in darjeeling himalayas, india. J. Earth Syst. Sci. 2026, 135, 36. [Google Scholar] [CrossRef] [Scilit]
  15. Li, L.; Feng, Q.; Yu, B. Research on the application of stacking ensemble learning model with negative sample constraints in landslide susceptibility assessment. Nat. Hazards 2026, 122, 148. [Google Scholar] [CrossRef] [Scilit]
  16. Wu, X.; Yang, J.; Niu, R. A landslide susceptibility assessment method using smote and convolutional neural network. Wuhan Daxue Xuebao (Xinxi Kexue Ban)/Geomat. Inf. Sci. Wuhan Univ. 2020, 45, 1223–1232. [Google Scholar] [CrossRef]
  17. Ding, J.; Wang, X. Landslide and collapse susceptibility analysis in wenchuan earthquake-damaged area based on ensemble learning methods. Gongcheng Kexue Yu Jishu/Adv. Eng. Sci. 2025, 57, 52–61. [Google Scholar] [CrossRef]
  18. Zhao, Z.; Chen, T.; Dou, J.; Liu, G.; Plaza, A. Landslide susceptibility mapping considering landslide local-global features based on cnn and transformer. IEEE J. Sel. Top. Appl. Earth Obs. Remote Sens. 2024, 17, 7475–7489. [Google Scholar] [CrossRef] [Scilit]
  19. Li, M.; Ming, D.; Li, Z.; Ma, S.A.; Zhang, H. Lightweight cnn-transformer model with multi-attention mechanism to assess susceptibility to landslides. J. Chengdu Univ. Technol. (Sci. Technol. Ed.) 2025, 52, 1133–1150. [Google Scholar] [CrossRef]
  20. Qu, W.; Bian, Z.; Li, J.; Tang, X.; Chen, P. Landslide susceptibility assessment model based on positive sample buffering with transformer-cnn integration: A case study of baoji city. Wuhan Daxue Xuebao (xinxi Kexue Ban)/Geomat. Inf. Sci. Wuhan Univ. 2026, 51, 726–740. [Google Scholar] [CrossRef]
  21. Yuan, R.; Chen, J. A novel method based on deep learning model for national-scale landslide hazard assessment. Landslides 2023, 20, 2379–2403. [Google Scholar] [CrossRef] [Scilit]
  22. Mao, Z.; Yu, H.; Liang, W.; Ma, X.; Zhong, J.; Gao, G.; Shi, S.; Tian, Y. Identification and feature analysis of regional loess landslides based on uav tilt photogrammetry 3d modeling. Geol. China 2024, 51, 561–576. [Google Scholar] [CrossRef]
  23. Liang, Z.; Zhang, X.; Liu, Z.; Huang, H.; Liu, Y. Enhancing landslide susceptibility mapping through a hybrid model utilizing bivariate methods and convolutional neural networks. Front. Earth Sci. 2026, 14, 1792825. [Google Scholar] [CrossRef] [Scilit]
  24. Alqadhi, S.; Mallick, J.; Alkahtani, M.; Ahmad, I.; Alqahtani, D.; Hang, H.T. Developing a hybrid deep learning model with explainable artificial intelligence (xai) for enhanced landslide susceptibility modeling and management. Nat. Hazards 2024, 120, 3719–3747. [Google Scholar] [CrossRef] [Scilit]
  25. Li, M.; Li, R.; Sha, Z.; Su, Y.; Wang, Y. Augmented detection of potential landslides in open-pit mines by integrating interpretable deep learning and time-series insar. J. Geo-Inf. Sci. 2025, 27, 2927–2950. [Google Scholar] [CrossRef]
  26. Bravo-López, E.; Fernández Del Castillo, T.; Sellers, C.; Delgado-García, J. Analysis of conditioning factors in cuenca, ecuador, for landslide susceptibility maps generation employing machine learning methods. Land 2023, 12, 1135. [Google Scholar] [CrossRef] [Scilit]
  27. Mind’je, R.; Li, L.; Nsengiyumva, J.B.; Mupenzi, C.; Nyesheja, E.M.; Kayumba, P.M.; Gasirabo, A.; Hakorimana, E. Landslide susceptibility and influencing factors analysis in rwanda. Environ. Dev. Sustain. 2020, 22, 7985–8012. [Google Scholar] [CrossRef] [Scilit]
  28. Kaynak, T. A systematic framework for the integration of feature selection and artificial intelligence in landslide susceptibility assessment. Nat. Hazards 2026, 122, 91. [Google Scholar] [CrossRef] [Scilit]
  29. Huang, F.; Pan, L.; Fan, X.; Jiang, S.-H.; Huang, J.; Zhou, C. The uncertainty of landslide susceptibility prediction modeling: Suitability of linear conditioning factors. Bull. Eng. Geol. Environ. 2022, 81, 182. [Google Scholar] [CrossRef] [Scilit]
  30. Liu, L.L.; Yang, C.; Wang, X.M. Landslide susceptibility assessment using feature selection based machine learning models. Geomech. Eng. 2021, 25, 1–16. [Google Scholar] [CrossRef]
  31. Liao, M.; Wen, H.; Yang, L. Identifying the essential conditioning factors of landslide susceptibility models under different grid resolutions using hybrid machine learning: A case of wushan and wuxi counties, china. Catena 2022, 217, 106428. [Google Scholar] [CrossRef] [Scilit]
  32. Chang, Z.; Catani, F.; Huang, F.; Liu, G.; Meena, S.R.; Huang, J.; Zhou, C. Landslide susceptibility prediction using slope unit-based machine learning models considering the heterogeneity of conditioning factors. J. Rock Mech. Geotech. Eng. 2023, 15, 1127–1143. [Google Scholar] [CrossRef] [Scilit]
  33. Liu, X.-D.; Xiao, T.; Zhang, S.-H.; Sun, P.-H.; Liu, L.-L.; Peng, Z.-W. Comparative study of sampling strategies for machine learning-based landslide susceptibility assessment. Stoch. Environ. Res. Risk Assess. 2024, 38, 4935–4957. [Google Scholar] [CrossRef] [Scilit]
  34. Ning, Z.; Tie, Y. Sampling method based on fuzzy membership for computing negative sample credibility and its applications. Appl. Sci. 2025, 15, 7646. [Google Scholar] [CrossRef] [Scilit]
  35. Yang, Z.; Shi, M.; Mei, H.; Zheng, M.; Yuan, J.; Wang, L. Frequency ratio–guided optimization of negative sample selection and sample ratio for landslide susceptibility assessment: A case study of the heishui river basin, china. Appl. Sci. 2026, 16, 342. [Google Scholar] [CrossRef] [Scilit]
  36. Liu, L.-L.; Duan, C.; Gao, J.-H.; Xiao, H.; Zhu, W.-Q.; Yang, C. Landslide susceptibility assessment using machine learning with a novel shap-based sampling strategy. Geosci. Front. 2026, 17, 102188. [Google Scholar] [CrossRef] [Scilit]
  37. Guo, F.; Jiang, G.; Huang, X.; Wang, X.; Xia, D.; Chen, Y.; Li, X. Impact of environmental factor combinations and negative sample selection on benggang susceptibility assessment in granite areas. Nongye Gongcheng Xuebao/Trans. Chin. Soc. Agric. Eng. 2024, 40, 191–200. [Google Scholar] [CrossRef]
  38. Le, X.-H.; Choi, C.; Eu, S.; Yeon, M.; Lee, G. Quantitative evaluation of uncertainty and interpretability in machine learning-based landslide susceptibility mapping through feature selection and explainable ai. Front. Environ. Sci. 2024, 12, 1424988. [Google Scholar] [CrossRef] [Scilit]
  39. Feng, X.; Wang, Y.; Liu, Y.; Liu, Q.; Du, J.; Chai, B. Susceptibility assessment of a translational rockslide considering the control mechanism and spatial uncertainty of a weak interlayer: Application study in tiefeng township, wanzhou district. Bull. Geol. Sci. Technol. 2022, 41, 254–266. [Google Scholar] [CrossRef]
  40. Wang, H.; Deng, C.; Zhang, Z.; Jiang, Z.; Wei, Q.; Yi, W.; Chen, T.; Ma, J. Dynamic landslide susceptibility assessment integrating sbas-insar and interpretable machine learning: A case study of the baihetan reservoir area, southwest china. Remote Sens. 2026, 18, 578. [Google Scholar] [CrossRef] [Scilit]
  41. Meng, X.; Xing, Z. Landslide susceptibility assessment in the western changyang section of the qingjiang river basin based on insar technology and random forest algorithm method. Carsologica Sin. 2025, 44, 609–620. [Google Scholar] [CrossRef]
  42. Xiao, T.; Huang, W.; Wang, L.; Yang, B.; Qin, Z.; Liu, X.; Xiao, Y. Uncertainty-aware ensemble learning and dynamic threshold optimization for landslide susceptibility mapping. Comput. Geosci. 2026, 206, 106042. [Google Scholar] [CrossRef] [Scilit]
  43. Huang, W.; Ding, M.; Li, Z.; Zhuang, J.; Yang, J.; Li, X.; Meng, L.; Zhang, H.; Dong, Y. An efficient user-friendly integration tool for landslide susceptibility mapping based on support vector machines: Svm-lsm toolbox. Remote Sens. 2022, 14, 3408. [Google Scholar] [CrossRef] [Scilit]
  44. Gull, A.; Mahmood, S.; Ahamad, M.I.; Rehman, A.; Purohit, S. Comparative assessment of rockfall susceptibility in the potohar plateau using frequency ratio, analytical hierarchy process, and weights of evidence models. Earth Syst. Environ. 2025, 9, 1393–1412. [Google Scholar] [CrossRef] [Scilit]
  45. Pokharel, B.; Lim, S.; Bhattarai, T.N.; Alvioli, M. Rockfall susceptibility along pasang lhamu and galchhi-rasuwagadhi highways, rasuwa, central nepal. Bull. Eng. Geol. Environ. 2023, 82, 183. [Google Scholar] [CrossRef] [Scilit]
  46. Utthasini, M.; Ilampooranan, I.; Singh, S.K.; Kanga, S.; Kumar, P.; Halder, K.; Pradhan, B.; Srivastava, A.K.; Chatterjee, R.S.; Chakrabortty, R.; et al. Enhancing landslide susceptibility mapping in the himalayas: Geospatial and machine learning with explainable ai (xai). Gondwana Res. 2026, 149, 262–290. [Google Scholar] [CrossRef] [Scilit]
  47. Achu, A.L.; Aju, C.D.; Thomas, J.; Gopinath, G. Catena matters: Enhancing landslide prediction with soil profile characteristics and explainable ai. Eng. Geol. 2026, 364, 108599. [Google Scholar] [CrossRef] [Scilit]
  48. Shao, X.; Yan, W.; Yan, C.; Zhao, W.; Wang, Y.; Shi, X.; Dong, H.; Li, T.; Yu, J.; Zuo, P.; et al. Explainable machine learning for mapping rainfall-induced landslide thresholds in italy. Appl. Sci. 2025, 15, 7937. [Google Scholar] [CrossRef] [Scilit]
  49. Huang, X.; Ye, J.; Liu, C.; Zeng, Q.; Guo, W.; Guo, Z. A stacking-shap ensemble method for landslide susceptibility prediction with high accuracy and interpretability. Cehui Xuebao/Acta Geod. Cartogr. Sin. 2025, 54, 1826–1840. [Google Scholar] [CrossRef]
  50. Yan, X.; Zhang, D.; Han, Y.; Li, T.; Zhong, P.; Ning, Z.; Tan, S. Developing a hybrid model to enhance the robustness of interpretability for landslide susceptibility assessment. ISPRS Int. J. Geo-Inf. 2025, 14, 277. [Google Scholar] [CrossRef] [Scilit]
  51. Qiu, H.; Xu, Y.; Tang, B.; Su, L.; Li, Y.; Yang, D.; Ullah, M. Interpretable landslide susceptibility evaluation based on model optimization. Land 2024, 13, 639. [Google Scholar] [CrossRef] [Scilit]
  52. Sun, P.; Zhang, M.; Jia, J.; Cheng, X.; Zhu, L.; Xue, Q.; Wang, J. Geo-hazards research and investigation in the loess regions of western china. Northwest. Geol. 2022, 55, 96–107. [Google Scholar] [CrossRef]
  53. Zhang, F.; Wang, G.; Peng, J. Initiation and mobility of recurring loess flowslides on the heifangtai irrigated terrace in china: Insights from hydrogeological conditions and liquefaction criteria. Eng. Geol. 2022, 302, 106619. [Google Scholar] [CrossRef] [Scilit]
  54. Tang, Y.; Feng, F.; Guo, Z.; Feng, W.; Li, Z.; Wang, J.; Sun, Q.; Ma, H.; Li, Y. Integrating principal component analysis with statistically-based models for analysis of causal factors and landslide susceptibility mapping: A comparative study from the loess plateau area in shanxi (china). J. Clean. Prod. 2020, 277, 124159. [Google Scholar] [CrossRef] [Scilit]
  55. Friedman, J.H. Greedy function approximation: A gradient boosting machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef] [Scilit]
  56. Chen, T.; Guestrin, C. XGBoost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; Association for Computing Machinery: New York, NY, USA, 2016; pp. 785–794. [Google Scholar]
  57. Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. Adv. Neural Inf. Process. Syst. 2017, 30, 4765–4774. [Google Scholar]
Figure 1. Terrain and landslide location distribution in Longchuan.
Figure 1. Terrain and landslide location distribution in Longchuan.
Geohazards 07 00088 g001
Figure 2. Landslide susceptibility factor maps of Longchuan.
Figure 2. Landslide susceptibility factor maps of Longchuan.
Geohazards 07 00088 g002aGeohazards 07 00088 g002b
Figure 3. Pearson correlation coefficient heat map.
Figure 3. Pearson correlation coefficient heat map.
Geohazards 07 00088 g003
Figure 4. ROC curves of different landslide susceptibility models in Longchuan.
Figure 4. ROC curves of different landslide susceptibility models in Longchuan.
Geohazards 07 00088 g004
Figure 5. SHAP-based global feature importance analysis of the GBDT model.
Figure 5. SHAP-based global feature importance analysis of the GBDT model.
Geohazards 07 00088 g005
Figure 6. SHAP dependence plots of the 12 environmental factors.
Figure 6. SHAP dependence plots of the 12 environmental factors.
Geohazards 07 00088 g006
Figure 7. SHAP-based environmental factor interaction network.
Figure 7. SHAP-based environmental factor interaction network.
Geohazards 07 00088 g007
Figure 8. Landslide susceptibility assessment of Longchuan based on the GBDT machine learning model.
Figure 8. Landslide susceptibility assessment of Longchuan based on the GBDT machine learning model.
Geohazards 07 00088 g008
Table 1. Data types and data sources.
Table 1. Data types and data sources.
Data NameData TypeResolution/mData SourceTemporal Parameter
Historical landslide-point dataVector (shp) Resource and Environmental Science Data Center, Chinese Academy of SciencesAccumulated to 2024
ElevationRaster (tif)30 mCopernicusDEM (ESA)2016
AspectRaster (tif)30 mExtracted from DEM2016
ReliefRaster (tif)30 mExtracted from DEM2016
LULCRaster (tif)30 mWuhan University CLCD dataset2024
DTFVector (shp)Scale 1:2,500,000National Geological Archives of China-
LithologyVector (shp)Scale 1:2,500,000National Geological Archives of China-
PlanCurRaster (tif)30 mExtracted from DEM2016
ProfileCurRaster (tif)30 mExtracted from DEM2016
RainfallRaster (tif)1 kmNational Tibetan Plateau Data Center2001–2024
DTWVector (shp)No fixed scaleOpenStreetMap (OSM)2025
TWIRaster (tif)30 mExtracted from DEM2016
NDWIRaster (tif)30 mLandsat 5, 7, 8, and 9 (GEE platform)1985–2025
NDVIRaster (tif)30 mLandsat 5, 7, 8, and 9 (GEE platform)1985–2025
DTRVector (shp)No fixed scaleOpenStreetMap (OSM)2025
NDBIRaster (tif)30 mLandsat 5, 7, 8, and 9 (GEE platform)1985–2025
Table 2. Hyperparameter settings used for the machine learning models.
Table 2. Hyperparameter settings used for the machine learning models.
ModelKey Tuned or Fixed Hyperparameters
LRPenalty = L2; solver = liblinear; max iter = 1000; regularization strength tuned by grid search
RFn estimators = 200; max depth = none; min_samples_split = 10; class_weight = balanced
SVMKernel = RBF; C tuned by grid search; gamma = scale
GBDTLearning rate = 0.01; n estimators = 200; max depth = 5; subsample = 0.8
XGBoostLearning rate = 0.01; n estimators = 200; max depth = 5; subsample = 0.8
CatBoostGridSearchCV-tuned configuration
LightGBMGridSearchCV-tuned configuration
ExtraTreesGridSearchCV-tuned configuration
Table 3. Summary statistics and spatial association of conditioning factors in Longchuan.
Table 3. Summary statistics and spatial association of conditioning factors in Longchuan.
Factor GroupFactor or IndicatorObserved Range/ClassMain Spatial Association with Landslides
TerrainElevation43–1312 mLow-to-middle-elevation basins and mountain-front transition zones
TerrainSlope0–63.95 degreesModerate to relatively steep slope-transition zones
TerrainTerrain roughness1–2.28Medium to high roughness, fragmented slopes, and incised valleys
TerrainRelief0–163 mGreater relief and stronger terrain incision
GeologyDTF0–15,131.6 mBelts close to fault zones and fractured rock masses
Land coverLULCCategoricalForest/cultivated slopes and construction-fringe areas
HydrologyRainfall1568.71–1897.81 mmCentral rainfall-gradient transition zone
HydrologyDTW0–10,122.8 mNear to middle river-valley distances
HydrologyTWI2.70–26.50Runoff-convergence belts and adjacent slopes
HydrologyNDWI−0.33–0.74Medium to low NDWI intervals and changing slope-wetness background
VegetationNDVI−0.27–0.81Medium-to-low-vegetation-cover areas
Human activityDTR0–6397.88 mRoad-adjacent belts and disturbed slope toes
Human activityNDBI−0.53–0.18Low to medium built-up fringe zones
Table 4. Test-set performance comparison of eight models.
Table 4. Test-set performance comparison of eight models.
AccuracyPrecisionRecallF1-ScoreTest AUC
GBDT0.88530.88690.88530.88520.9520
CatBoost0.87610.87770.87610.87600.9512
LightGBM0.88070.88070.88070.88070.9467
XGBoost0.85320.85370.85320.85320.9434
RandomForest0.84860.85010.84860.84850.9408
ExtraTrees0.85780.85780.85780.85780.9407
SVM0.85780.85810.85780.85780.9391
LR0.84400.84590.84400.84380.9386
Table 5. Qualitative interpretation of SHAP-based factor importance for engineering and planning use.
Table 5. Qualitative interpretation of SHAP-based factor importance for engineering and planning use.
FactorMean|SHAP Value|/RankQualitative ImportancePractical Interpretation
NDWIAbout 0.20Very highPrimary hydrological/wetness proxy and interaction hub
ReliefAbout 0.14Very highMajor terrain-control variable with strong nonlinear response
DTRAbout 0.06HighRoad-disturbance and accessibility-related slope modification
ProfileCurRank 4HighCurvature control on convergence/divergence of slope flow
LULC, DTW, DTFRanks 5–7ModerateLand-cover, water-system proximity, and structural controls
PlanCur, TWI, RainfallRanks 8–10LowSecondary terrain-wetness and triggering-background effects
Aspect, LithologyRanks 11–12Very lowLimited direct contribution in this model and factor system
Table 6. Classification statistics of landslide susceptibility mapping results.
Table 6. Classification statistics of landslide susceptibility mapping results.
ModelLandslide Susceptibility ProbabilitySusceptibility ClassLandslide ProportionArea ProportionFrequency Ratio (FR)
GBDT model<0.2Very low0.0220.42460.0519
0.2~0.4Low0.00280.0940.0293
0.4~0.6Moderate0.01380.06680.2063
0.6~0.8High0.08820.07411.1895
>0.8Very high0.87330.34062.5642
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, X.; Cai, R.; Zhao, S. Comparative Landslide Susceptibility Mapping in Longchuan, Guangdong Province, China, Using Explainable Machine Learning. GeoHazards 2026, 7, 88. https://doi.org/10.3390/geohazards7030088

AMA Style

Wang X, Cai R, Zhao S. Comparative Landslide Susceptibility Mapping in Longchuan, Guangdong Province, China, Using Explainable Machine Learning. GeoHazards. 2026; 7(3):88. https://doi.org/10.3390/geohazards7030088

Chicago/Turabian Style

Wang, Xi, Rongjiang Cai, and Shufang Zhao. 2026. "Comparative Landslide Susceptibility Mapping in Longchuan, Guangdong Province, China, Using Explainable Machine Learning" GeoHazards 7, no. 3: 88. https://doi.org/10.3390/geohazards7030088

APA Style

Wang, X., Cai, R., & Zhao, S. (2026). Comparative Landslide Susceptibility Mapping in Longchuan, Guangdong Province, China, Using Explainable Machine Learning. GeoHazards, 7(3), 88. https://doi.org/10.3390/geohazards7030088

Article Metrics

Back to TopTop