Figure 1.
Geographic setting of the study area. The blue rectangle marks the extent of the 64 rows by 32 columns tile grid covering most of the Laonung (Laonong) Creek Watershed in southern Taiwan, and the red rectangle marks the pilot subregion (rows 1–4, columns 1–32) used in this study.
Figure 1.
Geographic setting of the study area. The blue rectangle marks the extent of the 64 rows by 32 columns tile grid covering most of the Laonung (Laonong) Creek Watershed in southern Taiwan, and the red rectangle marks the pilot subregion (rows 1–4, columns 1–32) used in this study.
Figure 2.
Spatial distribution of the 128 annotated tiles and the cross-validation fold assignments across the pilot subregion (rows 1–4, columns 1–32). The background shows the DEM-derived hillshade mosaic. Colored overlays indicate fold assignment: Fold 1 (blue, columns 1–8), Fold 2 (orange, columns 9–16), Fold 3 (green, columns 17–24), and Fold 4 (red, columns 25–32). Thick vertical lines mark fold boundaries.
Figure 2.
Spatial distribution of the 128 annotated tiles and the cross-validation fold assignments across the pilot subregion (rows 1–4, columns 1–32). The background shows the DEM-derived hillshade mosaic. Colored overlays indicate fold assignment: Fold 1 (blue, columns 1–8), Fold 2 (orange, columns 9–16), Fold 3 (green, columns 17–24), and Fold 4 (red, columns 25–32). Thick vertical lines mark fold boundaries.
Figure 3.
Representative annotation workflow for tile r01_c24. The panels show the DEM-derived hillshade, SPOT-6 natural color composite (RGB = bands 3, 2, 1), fused annotation composite, and expert-delineated landslide mask. The mask reflects the interpreted geomorphic-unit extent rather than the spectral bare-soil boundary alone.
Figure 3.
Representative annotation workflow for tile r01_c24. The panels show the DEM-derived hillshade, SPOT-6 natural color composite (RGB = bands 3, 2, 1), fused annotation composite, and expert-delineated landslide mask. The mask reflects the interpreted geomorphic-unit extent rather than the spectral bare-soil boundary alone.
Figure 4.
Precision–recall curves for the three classifiers under spatially blocked four-fold cross-validation. Faint lines show individual folds, bold lines show mean curves, and shaded bands indicate ±1 standard deviation. The dashed line marks the no-skill baseline for the resampled evaluation dataset, and filled circles indicate F1-maximizing thresholds estimated from the test-fold predictions. The inset magnifies the high-recall region (recall > 0.70), where the mean curves of the three models are most clearly separated and the F1-maximizing operating points are located.
Figure 4.
Precision–recall curves for the three classifiers under spatially blocked four-fold cross-validation. Faint lines show individual folds, bold lines show mean curves, and shaded bands indicate ±1 standard deviation. The dashed line marks the no-skill baseline for the resampled evaluation dataset, and filled circles indicate F1-maximizing thresholds estimated from the test-fold predictions. The inset magnifies the high-recall region (recall > 0.70), where the mean curves of the three models are most clearly separated and the F1-maximizing operating points are located.
Figure 5.
Receiver operating characteristic (ROC) curves for the three classifiers under spatially blocked four-fold cross-validation. Faint lines show individual folds, bold lines show mean curves, and shaded bands indicate ±1 standard deviation. The diagonal dashed line represents the no-skill reference.
Figure 5.
Receiver operating characteristic (ROC) curves for the three classifiers under spatially blocked four-fold cross-validation. Faint lines show individual folds, bold lines show mean curves, and shaded bands indicate ±1 standard deviation. The diagonal dashed line represents the no-skill reference.
Figure 6.
Radar chart comparing the three classifiers (logistic regression, random forest, and XGBoost) across the five performance metrics reported in
Table 3: precision, recall, F1, ROC-AUC, and average precision (AP). Values are the means across the four spatial cross-validation folds. The radial axis spans 0.6–1.0 to make differences among the closely performing models legible. Random Forest shows the highest precision and lowest recall, logistic regression the highest recall and lowest precision, and XGBoost an intermediate, balanced profile; all three models converge at high ROC-AUC and AP.
Figure 6.
Radar chart comparing the three classifiers (logistic regression, random forest, and XGBoost) across the five performance metrics reported in
Table 3: precision, recall, F1, ROC-AUC, and average precision (AP). Values are the means across the four spatial cross-validation folds. The radial axis spans 0.6–1.0 to make differences among the closely performing models legible. Random Forest shows the highest precision and lowest recall, logistic regression the highest recall and lowest precision, and XGBoost an intermediate, balanced profile; all three models converge at high ROC-AUC and AP.
Figure 7.
Normalized confusion matrices for the three classifiers across four spatial test folds. Each cell displays the row-normalized proportion in bold, the raw sampled pixel count in parentheses, and the metric name in italics (Specificity, FPR, FNR, or Recall). All subplots share a consistent color scale from 0 to 1. Rows represent models (Logistic Regression, Random Forest, XGBoost), and columns represent spatial test folds (Fold 1: columns 1–8; Fold 2: columns 9–16; Fold 3: columns 17–24; Fold 4: columns 25–32). The matrices show that Random Forest maintained high background specificity across folds while exhibiting lower landslide recall in Fold 1 than in the other folds.
Figure 7.
Normalized confusion matrices for the three classifiers across four spatial test folds. Each cell displays the row-normalized proportion in bold, the raw sampled pixel count in parentheses, and the metric name in italics (Specificity, FPR, FNR, or Recall). All subplots share a consistent color scale from 0 to 1. Rows represent models (Logistic Regression, Random Forest, XGBoost), and columns represent spatial test folds (Fold 1: columns 1–8; Fold 2: columns 9–16; Fold 3: columns 17–24; Fold 4: columns 25–32). The matrices show that Random Forest maintained high background specificity across folds while exhibiting lower landslide recall in Fold 1 than in the other folds.
Figure 8.
Mean feature importance for Random Forest (left, green) and XGBoost (right, red) across four spatial cross-validation folds. Error bars show the standard deviation across folds. Features are ordered according to Random Forest importance. SPOT-6 Band 3 (Red) is the most important feature in both models, accounting for approximately 21% of Random Forest importance and 57% of XGBoost importance. Random Forest distributes importance across several spectral and topographic variables, whereas XGBoost places much greater emphasis on Band 3 and, to a lesser extent, NDVI. Note that Random Forest importance is computed as the mean decrease in Gini impurity, whereas XGBoost importance is computed as normalized gain across all splits. These quantities are not directly numerically comparable across the two models, and the comparison is intended to reveal qualitative differences in feature use rather than differences in absolute importance magnitude.
Figure 8.
Mean feature importance for Random Forest (left, green) and XGBoost (right, red) across four spatial cross-validation folds. Error bars show the standard deviation across folds. Features are ordered according to Random Forest importance. SPOT-6 Band 3 (Red) is the most important feature in both models, accounting for approximately 21% of Random Forest importance and 57% of XGBoost importance. Random Forest distributes importance across several spectral and topographic variables, whereas XGBoost places much greater emphasis on Band 3 and, to a lesser extent, NDVI. Note that Random Forest importance is computed as the mean decrease in Gini impurity, whereas XGBoost importance is computed as normalized gain across all splits. These quantities are not directly numerically comparable across the two models, and the comparison is intended to reveal qualitative differences in feature use rather than differences in absolute importance magnitude.
Figure 9.
Mean absolute SHAP values for all 13 features, computed for the Random Forest model across the four spatial folds. For each fold, the model was retrained on the other three folds and SHAP values were computed on a stratified subsample of 2000 pixels (1000 positive and 1000 negative) from the held-out fold. Bars show the across-fold mean and error bars indicate ± one across-fold standard deviation (sample standard deviation, computed across the four folds; this reflects both sampling and model variation across spatial blocks and is not a bootstrap confidence interval). Features are sorted in descending order of across-fold mean |SHAP|. Bar colors indicate feature category: brown = topographic (DEM-derived), blue = raw SPOT-6 bands, and green = spectral indices.
Figure 9.
Mean absolute SHAP values for all 13 features, computed for the Random Forest model across the four spatial folds. For each fold, the model was retrained on the other three folds and SHAP values were computed on a stratified subsample of 2000 pixels (1000 positive and 1000 negative) from the held-out fold. Bars show the across-fold mean and error bars indicate ± one across-fold standard deviation (sample standard deviation, computed across the four folds; this reflects both sampling and model variation across spatial blocks and is not a bootstrap confidence interval). Features are sorted in descending order of across-fold mean |SHAP|. Bar colors indicate feature category: brown = topographic (DEM-derived), blue = raw SPOT-6 bands, and green = spectral indices.
Figure 10.
SHAP beeswarm summary plot for the Random Forest model, generated from a stratified sample of 1000 pixels in the fold 4 test set. Each point represents an individual pixel. The horizontal axis shows the SHAP value, with positive values contributing toward the landslide class and negative values contributing toward the background class. Point color represents the corresponding feature value, with red indicating high values and blue indicating low values. Features are ranked by mean absolute SHAP value.
Figure 10.
SHAP beeswarm summary plot for the Random Forest model, generated from a stratified sample of 1000 pixels in the fold 4 test set. Each point represents an individual pixel. The horizontal axis shows the SHAP value, with positive values contributing toward the landslide class and negative values contributing toward the background class. Point color represents the corresponding feature value, with red indicating high values and blue indicating low values. Features are ranked by mean absolute SHAP value.
Figure 11.
SHAP dependence plots for the three most influential features overall (SPOT-6 Band 3 (Red), NDVI, and SPOT-6 Band 1 (Blue)) and the two most influential topographic features (slope and DEM elevation). The x-axis shows the original pre-standardization feature values, and the y-axis represents the corresponding SHAP value for each pixel. Point color represents the value of the feature with the strongest detected interaction: NDVI for Band 3, SPOT-6 Band 3 for NDVI, NDWI for Band 1, NDVI for slope, and SPOT-6 Band 4 for DEM elevation. Dashed horizontal lines mark SHAP . The symbol “#” denotes the sequential number of each displayed feature and does not necessarily indicate its overall feature-importance rank.
Figure 11.
SHAP dependence plots for the three most influential features overall (SPOT-6 Band 3 (Red), NDVI, and SPOT-6 Band 1 (Blue)) and the two most influential topographic features (slope and DEM elevation). The x-axis shows the original pre-standardization feature values, and the y-axis represents the corresponding SHAP value for each pixel. Point color represents the value of the feature with the strongest detected interaction: NDVI for Band 3, SPOT-6 Band 3 for NDVI, NDWI for Band 1, NDVI for slope, and SPOT-6 Band 4 for DEM elevation. Dashed horizontal lines mark SHAP . The symbol “#” denotes the sequential number of each displayed feature and does not necessarily indicate its overall feature-importance rank.
Figure 12.
Spatial distribution of per-tile classification errors for the Random Forest model across the pilot tile grid. (Left panel): false negative rate (FNR), defined as the fraction of sampled landslide pixels missed in each tile. (Right panel): false positive rate (FPR), defined as the fraction of sampled background pixels incorrectly classified as landslide. The two panels use independent color scales. The FNR panel spans the full 0–1.0 range, whereas the FPR panel uses a data-adaptive range of 0–0.12 (maximum observed per-tile FPR ), restoring visual differentiation among tiles in the FPR map. Gray cells indicate annotated tiles without landslides and are not included in the per-tile landslide error summaries. Thick vertical lines mark fold boundaries (Fold 1: columns 1–8, Fold 2: columns 9–16, Fold 3: columns 17–24, and Fold 4: columns 25–32). In the (Left panel), blue borders indicate tiles with FNR > 0.5, meaning that most landslide pixels were missed. In the (Right panel), blue borders indicate tiles with FPR > 0.05.
Figure 12.
Spatial distribution of per-tile classification errors for the Random Forest model across the pilot tile grid. (Left panel): false negative rate (FNR), defined as the fraction of sampled landslide pixels missed in each tile. (Right panel): false positive rate (FPR), defined as the fraction of sampled background pixels incorrectly classified as landslide. The two panels use independent color scales. The FNR panel spans the full 0–1.0 range, whereas the FPR panel uses a data-adaptive range of 0–0.12 (maximum observed per-tile FPR ), restoring visual differentiation among tiles in the FPR map. Gray cells indicate annotated tiles without landslides and are not included in the per-tile landslide error summaries. Thick vertical lines mark fold boundaries (Fold 1: columns 1–8, Fold 2: columns 9–16, Fold 3: columns 17–24, and Fold 4: columns 25–32). In the (Left panel), blue borders indicate tiles with FNR > 0.5, meaning that most landslide pixels were missed. In the (Right panel), blue borders indicate tiles with FPR > 0.05.
![Sustainability 18 07779 g012 Sustainability 18 07779 g012]()
Figure 13.
Relationship between per-tile landslide fraction and false negative rate (FNR) for the Random Forest model. Each point represents one annotated landslide-containing tile, and colors indicate the spatial fold. The dashed line shows the fitted linear trend (Pearson , , Spearman , ), indicating that smaller landslide fractions were associated with higher miss rates in this dataset. Open circles mark three tiles flagged as potential outliers (|externally studentized residual| ). The dotted line shows the linear fit with these tiles excluded (Pearson , , Spearman , ), confirming that the negative relationship is robust to these leverage points.
Figure 13.
Relationship between per-tile landslide fraction and false negative rate (FNR) for the Random Forest model. Each point represents one annotated landslide-containing tile, and colors indicate the spatial fold. The dashed line shows the fitted linear trend (Pearson , , Spearman , ), indicating that smaller landslide fractions were associated with higher miss rates in this dataset. Open circles mark three tiles flagged as potential outliers (|externally studentized residual| ). The dotted line shows the linear fit with these tiles excluded (Pearson , , Spearman , ), confirming that the negative relationship is robust to these leverage points.
Figure 14.
Random Forest spatial predictions on two held-out test tiles from Fold 1. Each row shows, from left to right, the SPOT-6 natural-color composite (RGB = bands 3, 2, 1), the expert-delineated ground-truth landslide mask, the predicted landslide probability (shared 0–1 color scale), and the predicted binary class at the default threshold of 0.5. The top row (tile r02_c08) is a representative strong case (FNR ): predicted probability concentrates on the annotated landslide scars and most landslide pixels are recovered at the default threshold, with some scattered false positives on spectrally similar bare surfaces outside the mask (per-tile FPR ). The bottom row (tile r01_c07) is a representative difficult case from Row 1 (FNR ): the annotated units are small and partially revegetated, produce only weak predicted probabilities, and are largely missed at the default threshold. Predictions are genuinely out-of-sample, generated by the Random Forest model trained on Folds 2–4. Panels are shown at tile-relative extent without axis ticks or scale bars.
Figure 14.
Random Forest spatial predictions on two held-out test tiles from Fold 1. Each row shows, from left to right, the SPOT-6 natural-color composite (RGB = bands 3, 2, 1), the expert-delineated ground-truth landslide mask, the predicted landslide probability (shared 0–1 color scale), and the predicted binary class at the default threshold of 0.5. The top row (tile r02_c08) is a representative strong case (FNR ): predicted probability concentrates on the annotated landslide scars and most landslide pixels are recovered at the default threshold, with some scattered false positives on spectrally similar bare surfaces outside the mask (per-tile FPR ). The bottom row (tile r01_c07) is a representative difficult case from Row 1 (FNR ): the annotated units are small and partially revegetated, produce only weak predicted probabilities, and are largely missed at the default threshold. Predictions are genuinely out-of-sample, generated by the Random Forest model trained on Folds 2–4. Panels are shown at tile-relative extent without axis ticks or scale bars.
![Sustainability 18 07779 g014 Sustainability 18 07779 g014]()
Figure 15.
Reliability diagram for the Random Forest model, computed on the 3:1 resampled evaluation distribution (25% positive prevalence) with out-of-fold test predictions pooled across the four spatial folds. (Top panel): observed positive fraction versus mean predicted probability over ten equal-width bins, with marker size proportional to bin count and the dashed line indicating perfect calibration; the curve lying above the diagonal indicates under-confidence. (Bottom panel): histogram of predicted probabilities, with the 0.25 base rate marked. Brier score, base-rate Brier score, expected calibration error (ECE), and maximum calibration error (MCE) are annotated. Calibration is prevalence-dependent, so these values characterize the resampled evaluation setting and would differ under the natural class prevalence of the full watershed.
Figure 15.
Reliability diagram for the Random Forest model, computed on the 3:1 resampled evaluation distribution (25% positive prevalence) with out-of-fold test predictions pooled across the four spatial folds. (Top panel): observed positive fraction versus mean predicted probability over ten equal-width bins, with marker size proportional to bin count and the dashed line indicating perfect calibration; the curve lying above the diagonal indicates under-confidence. (Bottom panel): histogram of predicted probabilities, with the 0.25 base rate marked. Brier score, base-rate Brier score, expected calibration error (ECE), and maximum calibration error (MCE) are annotated. Calibration is prevalence-dependent, so these values characterize the resampled evaluation setting and would differ under the natural class prevalence of the full watershed.
Table 1.
Feature set used for pixel-level landslide classification. B1–B4 refer to the SPOT-6 Blue, Green, Red, and Near-Infrared (NIR) bands, respectively.
Table 1.
Feature set used for pixel-level landslide classification. B1–B4 refer to the SPOT-6 Blue, Green, Red, and Near-Infrared (NIR) bands, respectively.
| Feature | Category | Definition |
|---|
| DEM | Topographic | Digital Elevation Model; elevation (m) |
| Slope | Topographic | Local slope angle (degrees) |
| Curvature | Topographic | Profile curvature: the second derivative of the elevation surface in the direction of maximum slope, quantifying the rate of change in slope angle along the steepest descent direction |
| SPOT B1 | Spectral | Blue band (450–525 nm), UInt16 |
| SPOT B2 | Spectral | Green band (530–590 nm), UInt16 |
| SPOT B3 | Spectral | Red band (625–695 nm), UInt16 |
| SPOT B4 | Spectral | Near-Infrared band (760–890 nm), UInt16 |
| NDVI | Index | Normalized Difference Vegetation Index, |
| SAVI | Index | Soil-Adjusted Vegetation Index, |
| EVI | Index | Enhanced Vegetation Index, |
| BI | Index | Brightness Index, ; this red–NIR formulation was used as a simple spectral brightness measure and may differ from other BI definitions in the literature |
| BSI | Index | Bare Soil Index, ; this is a SPOT-6-adapted visible–near-infrared formulation because SPOT-6 does not provide a shortwave-infrared band |
| NDWI | Index | Normalized Difference Water Index, ; this corresponds to the green–NIR formulation of NDWI |
Table 2.
Spatial cross-validation fold composition. LS = landslide.
Table 2.
Spatial cross-validation fold composition. LS = landslide.
| Fold | Columns | Tiles | Mean LS Fraction | Pixels (Sampled) |
|---|
| 1 | 1–8 | 23 | 0.014 | 1,140,448 |
| 2 | 9–16 | 24 | 0.014 | 1,254,980 |
| 3 | 17–24 | 20 | 0.011 | 760,212 |
| 4 | 25–32 | 29 | 0.025 | 2,653,868 |
| Total | 1–32 | 96 | 0.017 | 5,809,508 |
Table 3.
Cross-validated classification performance (mean ± standard deviation across four spatial folds). AP = average precision. All metrics were computed on the resampled evaluation dataset.
Table 3.
Cross-validated classification performance (mean ± standard deviation across four spatial folds). AP = average precision. All metrics were computed on the resampled evaluation dataset.
| Model | Precision | Recall | F1 | ROC-AUC | AP |
|---|
| Logistic Regression | | | | | |
| Random Forest | | | | | |
| XGBoost | | | | | |
Table 4.
Variance inflation factors (VIFs) for the 13-feature set, computed on a random subsample of 100,000 pixels with an intercept term included. Values are sorted in descending order. High VIFs among the spectral features reflect their construction from a shared set of four SPOT-6 bands.
Table 4.
Variance inflation factors (VIFs) for the 13-feature set, computed on a random subsample of 100,000 pixels with an intercept term included. Values are sorted in descending order. High VIFs among the spectral features reflect their construction from a shared set of four SPOT-6 bands.
| Feature | VIF |
|---|
| SAVI | > |
| NDVI | > |
| BI | 5049.3 |
| SPOT B4 | 4917.0 |
| BSI | 1155.5 |
| NDWI | 359.3 |
| SPOT B3 | 345.4 |
| SPOT B2 | 331.7 |
| SPOT B1 | 122.4 |
| DEM | 1.8 |
| slope | 1.1 |
| EVI | 1.0 |
| curvature | 1.0 |
Table 5.
Classification performance at the default threshold (0.5) and at the fold-optimized F1 threshold for each model. Values are mean ± standard deviation across four spatial folds. Optimal thresholds were identified from test-fold predictions and should be interpreted as optimistic upper bounds on threshold-optimized performance.
Table 5.
Classification performance at the default threshold (0.5) and at the fold-optimized F1 threshold for each model. Values are mean ± standard deviation across four spatial folds. Optimal thresholds were identified from test-fold predictions and should be interpreted as optimistic upper bounds on threshold-optimized performance.
| Model | Default F1 | Optimal F1 | F1 | Optimal Threshold |
|---|
| Logistic Regression | | | +0.014 | |
| Random Forest | | | +0.042 | |
| XGBoost | | | +0.004 | |
Table 6.
Random Forest precision and recall at selected decision thresholds across two recommended operating point ranges. Values are mean ± standard deviation across four spatial folds. The default threshold of 0.5 is included for reference; the corresponding F1 score at this threshold is reported in
Table 5.
Table 6.
Random Forest precision and recall at selected decision thresholds across two recommended operating point ranges. Values are mean ± standard deviation across four spatial folds. The default threshold of 0.5 is included for reference; the corresponding F1 score at this threshold is reported in
Table 5.
| Range | Threshold | Precision | Recall |
|---|
| High recall | 0.20 | 0.738 ± 0.058 | 0.853 ± 0.060 |
| High recall | 0.25 | 0.767 ± 0.053 | 0.826 ± 0.068 |
| High recall | 0.30 | 0.790 ± 0.048 | 0.800 ± 0.074 |
| Reference | 0.50 | 0.858 ± 0.033 | 0.681 ± 0.095 |
| High precision | 0.60 | 0.883 ± 0.030 | 0.614 ± 0.100 |
| High precision | 0.65 | 0.895 ± 0.028 | 0.577 ± 0.102 |
| High precision | 0.70 | 0.907 ± 0.027 | 0.537 ± 0.101 |
Table 7.
Key results from the companion U-Net study [
49], reproduced here so that the comparison in
Section 5.1 can be followed without consulting the companion paper. Matched-pixel AP values are computed on the stratified pixel sample shared with the present study and are elevated relative to full-tile AP; they are therefore not directly comparable to the full-tile input-modality AP values in the lower panel. All values are as reported in the published companion study, including the matched-pixel AP of the present study’s Random Forest model, which was evaluated there on the shared test tiles.
Table 7.
Key results from the companion U-Net study [
49], reproduced here so that the comparison in
Section 5.1 can be followed without consulting the companion paper. Matched-pixel AP values are computed on the stratified pixel sample shared with the present study and are elevated relative to full-tile AP; they are therefore not directly comparable to the full-tile input-modality AP values in the lower panel. All values are as reported in the published companion study, including the matched-pixel AP of the present study’s Random Forest model, which was evaluated there on the shared test tiles.
| Quantity | Value |
|---|
| Matched-pixel comparison (RF vs. U-Net, 29 shared test tiles) |
| Random Forest matched-pixel AP (this study) | 0.824 |
| U-Net matched-pixel AP (Input D, companion) | 0.847 |
| Mean per-tile AP difference (DL − RF) | +0.019 |
| Bootstrap 95% CI of per-tile difference | [−0.034, +0.058] |
| U-Net full-tile test AP by input modality (companion) |
| Input A—fused DEM–optical composite | 0.511 |
| Input B—SPOT-6 natural color (spectral only) | 0.263 |
| Input C—DEM stack (terrain only) | 0.152 |
| Input D—multi-source six-channel stack | 0.556 |