The central message is methodological: whether an event-based LSM map looks geomorphically plausible against independent satellite evidence is decided by how the validation footprint is constructed. A raw co-event footprint, contaminated by the concurrent flooding and agriculture, makes a model with AUC ≈ 0.945 indistinguishable from a random mask; a process-targeted, landslide-relevant footprint shows the same model capturing observed disturbance at about twice the chance (
Section 4.2). Both readings carry the same warning: high inventory-split skill is not equivalent to geomorphic plausibility—even on the favorable footprint, the realized skill (lift ≈ 2×) is far below what AUC ≈ 0.945 implies. That a statistically strong model can be spatially over-rated is consistent with the geomorphic-plausibility literature [
10,
11]; our contribution is to operationalize the check with two independent sensors and to show that footprint construction is itself decisive and currently neglected. The well-known design effects—random-vs.-spatial CV optimism and negative-sampling sensitivity—are reported as applied good practice and as a further reason absolute skill must not be over-read; they are not claimed as novel [
6,
7,
8,
9]. SHAP shows the signal is governed by the triggering event, rainfall, and terrain, with the clay-shale lithotype as the leading geological control—coherent with rainfall-induced earth-flow behavior.
5.2. Limitations
ΔNDVI threshold calibration circularity. The −0.15 operating threshold is calibrated on the same inventory (RER2023) that is then used to construct the validation footprint—the distribution of inside-polygon ΔNDVI fixes the threshold, and the resulting footprint is compared against the model predictions. This is methodologically a form of self-consistency rather than full independence: a hold-out subset of the inventory used exclusively for threshold calibration would be a stricter design. We do not implement such a hold-out in the present analysis because the inventory is already partitioned for spatial CV (folds of pixels), and a further inventory-level hold-out would either weaken the within-event training signal or impose a temporal split, which we have already declared as out of scope. The sensitivity table (
Table 4) provides partial mitigation: the four thresholds in the operationally plausible range [−0.10, −0.25] all yield bootstrap 95% CIs that exclude 1 on the masked footprint, with point lifts between 1.71× and 2.24×. The qualitative conclusion—masked lift exceeds chance robustly—does not depend on the specific operating threshold. A stricter independent calibration on a regionally adjacent prior-event inventory is the natural follow-up. Null-threshold sanity check: a circular self-fulfilling design would also produce high lift in non-discriminating threshold regimes. We swept the threshold across {−0.30, …, +0.10}, including the positive-Δ (vegetation-gain) regime. Masked lift in the disturbance regime (−0.30 to −0.10) is monotonically 1.71–2.24×; in the vegetation-gain regime (+0.05, +0.10) it drops to 1.11–1.37×; at the boundary (any negative) it falls to 0.69×. The test does not fire in the null regime. The positive-Δ lift is above 1, not exactly 1; this residual reflects baseline spatial autocorrelation between High + Very-High predicted classes and mountain-area phenology. The headline disturbance lift (2.00× at −0.15) is therefore ≈1.46× above this null-regime baseline (2.00/1.37) rather than 2.00 above the random expectation.
Two senses of independence. The footprint is independent of the model—the satellite-derived disturbance never enters training or feature selection (enforced by a runtime leakage guard), so the model–footprint comparison is leakage-safe. It is not independent of the inventory: the disturbance threshold is anchored on the inside-polygon ΔNDVI distribution, and the enrichment, χ2, and recall evidence that the footprint tracks real landslides is itself measured against the RER2023 polygons. The footprint should therefore be read as an inventory-anchored but model-independent validation reference—it tests whether the model reproduces the optically detectable disturbance of the mapped failures, not whether that disturbance exists independently of the inventory. A fully inventory-independent reference would require a disturbance threshold calibrated and labeled without RER2023 (a regionally adjacent prior-event inventory, or an unsupervised change-detection threshold), which we identify as the natural next step.
Macro-geographic-segregation hypothesis (test). We further tested whether the AUC ≈ 0.945 could be explained primarily by macro-geographic segregation (mountain vs. plain) rather than by genuine geomorphic skill, given the 30 m grid and one-point-per-polygon design. We re-ran the spatial-CV XGBoost ablation with progressively reduced feature sets: elevation alone yields AUC = 0.761 (±0.043); elevation + dist_road yields 0.809 (±0.038); elevation + three distance features yields 0.831 (±0.033); the existing terrain-only re-model (slope, aspect, curvature, elevation, TRI, TWI;
Section 4.7 (iv)) yields 0.9039; the full 36-feature model yields 0.945. The macro-geographic features alone reach only AUC 0.76–0.83; the within-terrain micromorphology (slope, curvature, TRI, TWI) adds an extra ≈0.07 to reach 0.9039; rainfall, lithology, land-cover, and NDVI together add a further ≈0.041 to reach 0.945. The model’s skill is therefore primarily driven by micromorphology and conditioning factors, not by macro-geographic segregation; this macro-geographic-segregation explanation is not supported at the Δ AUC ≈ 0.18 level (full vs. elevation-only).
Operational implication of the lithology-lift pattern.
Section 4.7 (vi) and
Figure 9 show that argille-scagliose-dominated 10 km blocks have low per-block lift; the direct empirical decomposition we report there (r(clay-shale fraction, predicted H + VH base rate) = −0.33,
n = 38) finds that, contrary to a simple “model saturation” narrative, these blocks also have a low H + VH base rate, so the compressed lift is not a simple consequence of the model flagging every clay-shale pixel as High. The operational implication is nevertheless concrete: a regional aggregate lift can be lower than the local pixel-level skill of the model in geologically stratified terrain—because aggregating to a 10 km block averages over within-block factors (rainfall, slope) that the model uses to discriminate, while the lithology one-hot is constant across the block. Practitioners who apply satellite-derived footprints to an LSM map at coarse aggregation without geological stratification will therefore systematically understate model skill in clay-shale basins (Northern Apennines) and probably overstate it elsewhere. A direct decomposition of the 10 km blocks supports this reading: the below-chance majority is concentrated in blocks with a higher clay-shale fraction (mean 0.13 versus 0.03 in the above-chance minority) and a low predicted High + Very-High base rate (mean 0.16 versus 0.39)—terrain where the cohesive argille-scagliose units fail diffusely rather than in spatially concentrated patches. Footprint density and forest cover, by contrast, do not separate low- from high-lift blocks, and cultivated and built-up land is excluded from the masked domain by construction, so the collapse is a lithological and model-confidence effect rather than residual agricultural or infrastructural noise. The above-chance upper tail corresponds to moderately steeper flysch blocks where the model concentrates High + Very-High predictions and the observed disturbance coincides with them.
Modeling grid scale. The 5 m LiDAR-derived DTM is block-aggregated to 30 m to match the other predictors. Because the median RER2023 landslide polygon (479 m
2) is smaller than a 30 m pixel, each polygon is represented by a single point in this design. This one-point-per-polygon design eliminates pseudo-replication at the cost of also masking the sub-pixel geomorphic niches (very local curvature/TRI variation, micro-channel geometry) that a tree-based model such as XGBoost could in principle exploit if it were trained on the full native 5 m DTM. Slope-unit or object-based segmentation exploiting the full 5 m DTM is the natural next step for closing this scale gap. Beyond the grid-scale mismatch, the inventory itself carries delineation and positional uncertainty that propagates into statistical susceptibility models [
9]; the single-point-per-polygon sampling and the 200 m negative buffer absorb, rather than remove, this sub-polygon positional error.
Independent footprint coverage and recall bias. The phenology-matched optical view is coverage-limited (clear post-event view over ≈76% of the domain, ≈43% for the same-year April pre-event composite). Sentinel-1 12-day coherence carries a high vegetation-decorrelation baseline, and ΔNBR is cloud-shadow-sensitive, so both stay near chance even on the masked footprint—the optical ΔNDVI is the informative proxy. The optical footprint recovers only ≈17% of mapped RER2023 landslides overall, with recall both size- and type-dependent (full breakdown in
Section 4.7 (i) and
Table 9); small or rocky failures leave little ΔNDVI signature, so the test is a consistency check on the optically detectable subset, not an exhaustive detector.
Mixed-mechanism binary labeling. The susceptibility model treats all RER2023 failures as a single positive class, yet the inventory spans kinematically and physically distinct mechanisms—from low-angle earth flows and earth/debris slides to high-angle rock-slope failures—that respond to different combinations of lithology, slope, and hydrology and leave different surface signatures. A single binary classifier therefore blends predisposition signals that a mechanism-stratified model would separate, and this blending propagates into the independent validation: because the optical footprint detects disturbance through vegetation removal, it preferentially registers the vegetation-stripping earth-flow/clay-shale failures (per-type recall 28.1% for Earth Flow versus 5.5% for the Rock-Slide complex;
Section 4.7 (vi)) and barely registers rocky or deep-seated failures. The ≈2× masked lift is therefore a plausibility statement for the optically detectable, predominantly flow-type subset of the event, not for its full kinematic spectrum; a mechanism-stratified model validated against mechanism-specific footprints (for example, SAR or DEM-differencing for rocky failures) is the route to resolving this limitation.
Rigid topographic masking. The landslide-relevant footprint applies a fixed slope cutoff (≥10°, with ≥5° reported for comparison), which by construction removes the low-angle distal shear-deposition and debris-flow runout zones that debouch onto flatter ground. We quantified this blind spot against the inventory: only 4.0% of mapped RER2023 landslide pixels (3.6% of polygons in part, 0.8% entirely) fall below 10°, and 0.5% below 5°—the mapped failures are overwhelmingly steep-source features (median pixel slope 26.8°). The cutoff therefore excludes only a small distal fraction of the mapped landslide area, and because it removes real landslide pixels, it can only lower the measured capture, biasing the lift downward (conservatively); the consistency of the ≥5° and ≥10° variants (
Section 4.2) confirms the result is not an artifact of the specific cutoff. Event settings dominated by long-runout flows on valley floors would instead require a runout-aware mask rather than a single slope threshold.
Heterogeneity of the masked lift. The masked aggregate (95% CI 1.25–2.72×; canonical bootstrap) is genuinely spatially heterogeneous: most 10 km blocks score near chance, and the aggregate is driven by an upper tail (
Figure 8; per-lithotype panel
Figure 9). The lift should be read as a methodological demonstration that footprint construction controls the verdict, not as a transferable per-region plausibility magnitude.
AUC on the high side; not a negative-pool artifact. XGBoost AUC ≈ 0.945 is on the high side of the LSM literature because ≈25% of the convex hull domain is easy, non-susceptible terrain. The masked-footprint analysis already recomputes the chance baseline on the susceptible-only domain (11.9% → 16.5%), so the reported lift is not measured against an inflated baseline. Re-training the four models with the negative pool itself restricted to the same susceptible domain (strategy terr_susc) moves the spatial-CV XGBoost AUC by only −0.0005 and preserves the masked-footprint lift (2.45× vs. 2.00×;
Section 4.7); the gap is therefore not an artifact of the negative pool.
Multicollinearity/SHAP attribution. The iterative VIF prune designed in the pipeline has been applied empirically as a SHAP-attribution robustness check (
Appendix A,
Figure A1). Five features with VIF > 10 are removed (clc_bare, clc_tree, clc_wetland, rain_ep1, slope); held-out random-fold AUC moves by +0.0001 (XGBoost) and ≈0 (RF); SHAP top-8 overlap is 7/8 (XGBoost) and 6/8 (RF); Spearman rank correlation of mean(|SHAP|) is ρ = 0.967 (XGBoost)/ρ = 0.998 (RF). The substantive SHAP interpretation is therefore not an artifact of multicollinearity. A full VIF-pruned re-run on all four models with corresponding susceptibility mapping remains an interesting follow-up.
Sentinel-1 sensor context. The result is conditioned on a single-satellite, 12-day repeat configuration: in May 2023, Sentinel-1A was the only operational SAR platform (Sentinel-1B failed in December 2021; Sentinel-1C was not yet launched or integrated into the HyP3 archive), so the smallest accessible interferometric pair was a 12-day S1A–S1A pair. Over forested Apennines at peak spring phenology, 12-day vegetated-target decorrelation is the a priori expectation rather than an artifact of these particular acquisitions. A 6-day pair (S1A + S1B or S1A + S1C, when available), differential/ascending+descending fused coherence, or arid/sparsely vegetated terrain could push S1 from a coverage check toward a true validator; these configurations are not tested here.
Operational coregistration of the ascending track. The ascending track-117 InSAR pair failed HyP3 GAMMA coregistration because the AOI sits near the swath edge for that track, leaving insufficient common reference/event frame overlap to support sub-pixel coregistration (
Section 4.2). The HyP3 on-demand workflow exposes only the canonical GAMMA pipeline; custom DEM substitution, manual coregistration windows, or unconventional baseline/Doppler tolerances are not available through the cloud service, and processing the same acquisitions in a local GAMMA installation would step outside the open-pipeline scope we adopt here. We accept the swath-edge failure of track-117 as the standard SAR pre-condition rather than as a recoverable parameter problem.
Event-level temporal scope. The model is trained on the RER2023 inventory of the same event whose footprint serves as the independent validation. The leakage safeguard removes any direct contamination at the pixel/feature level—the satellite-derived disturbance never enters training and is enforced by a runtime guard on the feature matrix—but the inventory and the footprint share, by construction, the same event-level conditioning. A stricter temporal independence test would train on a prior event in the same region and validate on May 2023; this is out of scope here (the RER2023 inventory is the first systematic, polygon-resolved event-conditioned LSM benchmark for this basin) but is the natural next step for the protocol. The headline result of this study is therefore best read as a within-event plausibility test on a single regional event; cross-event temporal transferability of the protocol—and of the magnitude of the lift—has not been tested and is not claimed.
Alternative SAR configurations not tested. The 12-day S1A-only configuration used here is the cloud-independent baseline that was operationally accessible in May 2023. Higher-cadence or longer-wavelength alternatives that could plausibly turn S1 into a landslide discriminator in this setting—and which we explicitly do not test—include: (i) sub-daily commercial X-band SAR (ICEYE, Capella Space), whose short revisit times reduce vegetation decorrelation for direct interferometry; (ii) ALOS-2/future NISAR-class L-band SAR, where the longer wavelength penetrates vegetation canopies and tolerates the 12–14-day decorrelation budgets that C-band cannot support over forested Apennines at peak phenology; and (iii) Sentinel-1 6-day pairs (S1A + S1B or S1A + S1C, once the constellation is fully operational again). Recommending one particular configuration would be premature without acquisition-cost and access-rights comparisons that are outside the scope of this study; we flag the three families above as the natural candidates for the next iteration of the protocol.