1. Introduction
Surface soil moisture (SM) is a fundamental state variable of the terrestrial water and energy cycles, governing the partitioning of precipitation into infiltration and runoff [
1], modulating land–atmosphere exchanges of latent and sensible heat [
2], and directly constraining crop water availability and root-zone replenishment [
3,
4,
5]. At the field scale (~10–100 m), accurate knowledge of SM distribution underpins precision irrigation scheduling, drought early warning, and the parameterisation of land surface models used in numerical weather prediction. Yet, sustained, spatially continuous SM monitoring at this resolution remains a long-standing observational challenge: in situ sensor networks are sparse, expensive to maintain, and subject to spatial representativeness errors arising from sub-grid heterogeneity in soil texture, topography, and vegetation cover.
Microwave remote sensing provides the only operational pathway to global, all-weather SM retrieval at satellite overpass frequencies [
6,
7,
8]. Passive L-band missions—including the NASA Soil Moisture Active Passive (SMAP) [
9] and ESA Soil Moisture and Ocean Salinity (SMOS) [
10]—exploit the strong sensitivity of the soil dielectric constant to liquid water content in the 1–3 GHz range, achieving retrieval accuracies of approximately 0.04 m
3 m
−3 ubRMSE under the SMAP mission requirement. However, the coarse spatial resolution inherent to passive radiometry (~9–40 km) is fundamentally incompatible with field-scale applications, where SM can vary by more than 0.10 m
3 m
−3 over distances of tens of metres [
11,
12]. Disaggregation approaches that downscale passive products using optical or radar auxiliary data partially address this limitation [
13,
14], but introduce additional assumptions and error propagation that degrade accuracy at fine scales.
Active C-band SAR, as provided by the ESA Sentinel-1 constellation since 2014, offers a compelling alternative: 10-metre spatial resolution, 12-day repeat cycle, and free global data access [
15]. The sensitivity of SAR backscatter to SM arises from the dielectric contrast between dry soil (ε ≈ 3–5) and wet soil (ε ≈ 20–30) [
16], which governs the magnitude of VV (vertical–vertical co-polarisation)- and VH (vertical–horizontal cross-polarisation)-polarised surface scattering. Semi-empirical models—including the Dubois model [
17], and Water Cloud Model (WCM) [
18]—describe the relationship between sigma-naught (σ
0) and SM as a function of surface roughness and vegetation water content. While physically interpretable, these models rely on ancillary roughness parameters that are difficult to retrieve operationally, and their performance degrades markedly under moderate-to-dense vegetation cover where volume scattering from the canopy attenuates the soil signal [
19,
20,
21].
Machine learning (ML) methods have gained substantial traction as data-driven alternatives that circumvent explicit roughness parameterisation by directly learning the mapping from multi-source input features to SM from collocated training observations. Random Forest (RF), Support Vector Regression (SVR) [
22], and gradient boosting methods (XGBoost, LightGBM) have demonstrated competitive accuracy in field-scale SAR SM retrieval, with reported R
2 values in the range 0.67–0.82 and RMSE of 0.025–0.048 m
3 m
−3 across diverse study sites [
23,
24,
25,
26,
27,
28]. Ensemble stacking—wherein the predictions of diverse base learners are combined via a meta-learner trained on out-of-fold outputs—further reduces prediction variance while retaining interpretability through feature attribution tools [
29]. Deep learning architectures, particularly Long Short-Term Memory (LSTM) networks [
30], have been explored for temporal SM prediction, but typically require larger training datasets and exhibit inferior generalisation to spatially novel locations compared to gradient-boosted tree ensembles at the station-observation scale [
12].
Despite these advances, a critical source of predictive information has been systematically overlooked in the SAR SM retrieval literature: the strong temporal autocorrelation of surface soil moisture. At 12-day intervals—matching the Sentinel-1 repeat cycle—the lag-1 autocorrelation of SM typically exceeds r = 0.75, reflecting the physical persistence of soil water controlled by capillary retention, drainage, and evapotranspiration [
31]. This ‘soil moisture memory’ is well-established in hydrological theory and exploited in data assimilation frameworks [
32,
33], yet the remote-sensing retrieval community has largely treated each satellite acquisition as an independent, memoryless observation. Including the most recent observed SM as a model input—what we term a temporal lag feature—provides a physically motivated prior that reduces the effective dimensionality of the SAR–SM inversion problem. To the authors’ knowledge, no prior study has explicitly incorporated temporally lagged in situ SM observations as predictive features within a SAR-based ML retrieval framework: the SM memory signal, well-exploited in data assimilation, has been systematically absent from SAR SM retrieval pipelines, representing an untapped and physically grounded source of predictive information.
A second unresolved challenge is spatial generalisation: the ability of a trained model to produce accurate SM estimates at monitoring stations or grid cells whose spatial characteristics were not represented in the training data. This is a prerequisite for the operational deployment of SM retrieval algorithms across landscapes that extend beyond the calibration domain. Existing ML-based SAR SM studies predominantly evaluate performance via temporal cross-validation—splitting records by time at the same stations—which conflates in-distribution prediction accuracy with true generalisation ability. The few studies employing station-level hold-out spatial cross-validation have reported pronounced accuracy degradation, with R
2 frequently degrading substantially, in some cases becoming negative, at unseen locations [
34,
35,
36], indicating that models learn station-specific characteristics rather than transferable SAR–SM relationships. Whether temporal lag features can simultaneously improve accuracy and spatial generalisation has not, to the authors’ knowledge, been systematically investigated.
A third limitation of current ML-based SM retrieval studies is the limited deployment of explainability methods to verify that model decisions are physically coherent. High test-set accuracy is a necessary but insufficient condition for trustworthy operational use: a model that achieves R
2 = 0.78 by exploiting spurious correlations between SM and spatially varying soil properties may fail catastrophically when applied to a new region with different soil characteristics. SHapley Additive exPlanations (SHAP) [
37] provide a rigorous, game-theoretic framework for decomposing individual predictions into additive feature contributions, enabling the identification of non-linear feature dependencies and interaction effects. Combined with the Boruta permutation-based feature selection algorithm [
38], SHAP attribution constitutes a robust explainable artificial intelligence (XAI) pipeline that bridges predictive performance and physical interpretability.
These three gaps—unexploited SM temporal memory, overestimated spatial generalisation, and insufficient physical interpretability—point to a common underlying deficiency: existing SAR SM retrieval frameworks treat each satellite acquisition as an isolated, memoryless event, without leveraging the strong temporal structure of the SM time series or subjecting model decisions to rigorous physical scrutiny. None of the reviewed studies has simultaneously addressed all three gaps within a unified stacking ensemble framework augmented with explicit temporal lag features and validated under both temporal and spatial cross-validation protocols using scale-independent error decomposition. This constitutes the central methodological contribution of the present study.
To address the above gaps, this study develops and evaluates a Stacked Additive Boosting-based Model (SABM) that combines Sentinel-1 C-band SAR, Sentinel-2 optical, and ancillary geophysical features with temporal lag SM features derived from a site’s own antecedent in situ record, evaluated at the Little Washita River Watershed, Oklahoma—one of the most comprehensively instrumented and benchmarked SM validation sites globally, selected specifically because the framework requires historical in situ SM observations at the target location and is therefore intended for instrumented, monitored field sites rather than as a stand-alone SAR retrieval method for ungauged locations. The Little Washita Micronet, comprising 21 continuous in situ stations spanning 2016–2023, provides the observational foundation for both model training and error characterisation via Extended Triple Collocation (ETC), which jointly estimates the random error of the SAR-based model, SMAP microwave, and the ground sensor network without designating any system as a noise-free reference. Because the same ground records also underpin SABM’s training and lag features, this characterisation does not constitute a fully independent validation in the classical ETC sense (see
Section 2.6 and
Section 4.4).
The specific objectives of this study are: (1) to quantify the accuracy improvement achieved by incorporating temporal lag SM features into the SABM stacking framework under both temporal and spatial cross-validation settings; (2) to characterise SABM’s point-scale error using ETC against SMAP L3 Enhanced as a benchmark, while examining the extent to which this characterisation satisfies ETC’s mutual-independence assumption; (3) to characterise the relative contributions of SAR backscatter, optical spectral indices, terrain, soil physicochemical, and lag SM features using Boruta–SHAP; (4) to diagnose the physical mechanisms underlying key SHAP dependence patterns, with particular emphasis on the non-linear roles of soil pH, seasonal forcing, and vegetation density in mediating the SAR–SM relationship; and (5) to conduct a preliminary assessment of cross-site transferability by applying the trained model to the REMEDHUS network (Spain) without site-specific retraining.
4. Discussion
4.1. Physical Basis and Operational Implications of Temporal Lag Features
The substantial accuracy improvement achieved by incorporating SM
lag1 and SM
lag2 (temporal hold-out validation R
2: 0.6255 → 0.7771; RMSE: 0.0535 → 0.0413 m
3 m
−3) is not merely a statistical artefact of additional predictors, but reflects a physically well-grounded mechanism. Surface soil moisture at the 0–5 cm depth responds to precipitation events on timescales of hours, but its subsequent recession—governed by drainage, capillary redistribution, and evapotranspiration—is a continuous deterministic process describable by the Richards equation [
57]. For a sub-humid site such as Little Washita, where precipitation events are episodic and interspersed by extended dry-down periods, the SM state at time
t provides a strong constraint on the achievable range of SM at time
t + 12 days: a very dry observation (SM < 0.10 m
3 m
−3) effectively excludes saturated conditions 12 days later in the absence of intervening precipitation. This prior reduces the effective search space the model must navigate from the full climatological SM distribution (~0.35 m
3 m
−3 range) to a conditional distribution substantially narrower around the prior SM value. The empirical lag-1 autocorrelation of r = 0.8168 at the 12-day Sentinel-1 repeat cycle quantifies this constraint, and the dominant SHAP importance of SM
lag1 (mean |SHAP| = 0.0540 m
3 m
−3, exceeding post-lag soil pH at 0.0058 m
3 m
−3) confirms that the model learns to exploit this information preferentially over instantaneous SAR signals.
This formulation is analogous to a one-step-ahead data assimilation scheme in which the previous observation serves as the background state and the SAR-derived features provide the innovation. Unlike ensemble Kalman filter approaches [
58] that require an explicit land surface model, the SABM framework learns the effective soil moisture decay function implicitly from data, without imposing prior assumptions about soil hydraulic parameters. The proposed framework is therefore directly applicable to any monitoring network where in situ observations are collected at intervals comparable to the satellite repeat cycle—a condition satisfied by the majority of operational SM networks worldwide (USDA Micronet, OzNet, SCAN [
59], COSMOS [
60]).
A practical consideration for operational deployment is the requirement for antecedent SM observations. Unlike a pure SAR retrieval model that can be applied globally at any unmonitored location, the SABM + Lag SM framework requires at least two prior SM observations per prediction site. This positions the method as a monitoring enhancement tool rather than a spatial mapping algorithm—the appropriate comparison class is operational soil moisture forecasting or data assimilation at instrumented sites, rather than satellite-only global SM mapping products such as SMAP. Within this framing, the lag-based SABM represents an attractive complement to sparse monitoring networks: the SAR component provides the spatial and temporal resolution unattainable by point sensors alone, while the lag features anchor predictions to locally observed conditions, suppressing the systematic errors that arise when a globally trained model encounters site-specific soil characteristics outside its training distribution.
In practice, however, the cold-start requirement is less prohibitive than it may initially appear. Only two Sentinel-1-collocated SM observations—separated by approximately 12 and 24 days—are needed before the full SABM + Lag SM framework becomes operational. For any new monitoring station, this warm-up period is completed within the first month of deployment. Furthermore, for truly ungauged locations where no in situ record exists, satellite-derived SM products offer a viable surrogate pathway: SMAP L3 Enhanced (9 km, daily) or ERA5-Land reanalysis (9 km, hourly) SM estimates can be spatially interpolated to the target site and used as initialisation values for SM
lag1 and SM
lag2. While such satellite-derived priors introduce additional uncertainty compared to direct in situ measurements—SMAP carries its own ETC RMSE of 0.0419 m
3 m
−3 at Little Washita—they reduce the operational barrier from ‘requires a monitoring station’ to ‘requires any SM data source at 12-day intervals’, substantially broadening applicability. This performance degradation associated with satellite-derived lag proxies, relative to in situ lag inputs, is quantified directly in
Section 3.10 and discussed further in
Section 4.6.
4.2. Mechanisms of Generalisation Improvement at Unseen, Instrumented Stations
The recovery of spatial hold-out validation R
2 from −0.6024 (original SABM, all folds negative) to +0.5134 (SABM + Lag SM) represents the most consequential finding of this study from a methodological perspective, understood as generalisation to stations unseen during training but retaining their own historical SM record (
Section 2.7.2), rather than to fully ungauged locations. The failure of the original SABM at spatially novel locations can be explained by examining what information the model exploits during training. Under temporal hold-out validation—where the same stations appear in both training and test splits—the model can learn arbitrary station-specific mappings between feature values and SM: for instance, a particular combination of soil pH and elevation that characterises a specific station may become a reliable predictor simply because that station always exhibits certain SM conditions, regardless of whether that relationship generalises across space. This form of spatial memorisation inflates temporal hold-out validation accuracy while producing negative R
2 when the model is applied to stations with different soil-topographic configurations [
61].
The lag SM feature disrupts this memorisation mechanism directly: it replaces the need to recall station-specific SM climatology from the static feature set with a direct observation of current SM conditions at the target site. When predicting SM at a spatially novel station, the static features (soil pH, elevation, coordinates) may lie outside the training distribution, causing the model to extrapolate unreliably. However, SMlag1 provides a site-specific anchor that reduces reliance on static spatial features as proxies for climatological SM. In information-theoretic terms, the lag feature reduces the conditional entropy H(SM_t|SAR, static features) at novel locations by providing a strong additional signal H(SM_t|SMlag1) that is largely location-invariant in its information structure. The model needs only learn how much SM typically changes between consecutive overpasses given the prevailing SAR conditions—a relationship that is more transferable across space than absolute SM levels.
The residual failure of Fold 2 (R2 = −0.0260; stations 132, 136, 234, 256) is consistent with this interpretation. These stations exhibit mean SM signal standard deviation of 0.0380 m3 m−3—among the lowest in the network—indicating that SM varies little over time at these locations. Low temporal variability reduces the signal content of both the lag features and the SAR backscatter anomalies, leaving the model with insufficient information to reconstruct the subtle SM dynamics. Physically, these stations may correspond to locations with high soil clay content that buffers SM fluctuations, or to topographic positions with persistent shallow water tables that maintain near-constant SM regardless of precipitation inputs. Future work should investigate whether the inclusion of soil textural priors or topographic wetness indices can improve performance at such low-variability sites.
4.3. XAI Insights into SAR–Soil Moisture Dynamics
The Boruta–SHAP analysis reveals a feature importance hierarchy that is physically coherent but challenges some assumptions implicit in conventional semi-empirical SAR SM retrieval models. Most notably, soil pH emerges as the single most important predictor (mean |SHAP| = 0.0345 m
3 m
−3), outranking VV backscatter (0.0162 m
3 m
−3) by a factor of approximately 2. This ranking does not imply that soil pH directly controls SAR backscatter—pH has no direct microwave radiative effect. A feature-group ablation (
Section 3.9) indicates that this ranking should be interpreted with caution: a model trained exclusively on time-invariant static features (including soil pH) achieves R
2 = 0.2350 under temporal hold-out validation despite containing no information that varies between repeat visits to the same station, while removing soil pH alone from the full feature set changes performance negligibly (ΔR
2 = −0.0062). Together these results suggest that soil pH’s high SHAP ranking reflects, at least in part, its role as one of several correlated static features that the no-lag model exploits as an implicit station identifier—learning station-specific SM regimes from static feature combinations—rather than a uniquely indispensable physical driver. This does not preclude a genuine physical association: in the Little Washita watershed, soil pH correlates with clay mineralogy (acidic soils in topographic lows tend toward smectitic 2:1 clays with higher water retention, while alkaline upland soils are more montmorillonite-poor), consistent with the sigmoidal pH–SHAP dependence crossing zero near pH = 6.8. However, given the ablation evidence, we present this clay–mineralogy association as a plausible contributing explanation rather than a confirmed mechanism, and caution against using global soil pH maps (e.g., SoilGrids [
45]) as a stand-alone proxy for SAR–SM calibration without independent validation of the underlying soil-property relationship at the deployment site.
The DOY SHAP pattern (
Figure 6b) reveals that seasonal forcing contributes as much to SM prediction as VV backscatter under the study conditions. The bimodal seasonal SHAP structure—positive in late winter/spring, negative in summer—is consistent with the regional climate: winter precipitation and low evapotranspiration produce persistently elevated SM (0.20–0.35 m
3 m
−3) from January to April, while summer evapotranspiration from winter wheat and grassland drives SM below 0.10 m
3 m
−3 by July. The model has learned to use DOY as a seasonal prior that adjusts predictions before the SAR signal is interpreted: the same VV backscatter value of −14 dB implies very different SM depending on the season. This season-dependent SAR sensitivity is a known phenomenon in the semi-empirical literature [
16], and the SHAP framework makes its magnitude and directionality explicit without requiring explicit seasonal parameterisation.
The vegetation stratification of VV SHAP values (
Figure 6c) provides empirical confirmation of the theoretical vegetation attenuation effect. The decrease in VV–SHAP slope from sparse (NDVI < 0.2) to dense (NDVI > 0.4) vegetation classes is consistent with the Water Cloud Model [
18,
62], in which vegetation water content attenuates the soil backscatter contribution in proportion to canopy optical depth. Notably, the SABM framework recovers physically sensible vegetation-stratified relationships without explicit WCM parameterisation, suggesting that the machine learning model has implicitly learned a functional equivalent of the WCM attenuation term from the joint distribution of VV, NDVI, and SM in the training data. This provides post hoc physical validation of the model and increases confidence in its behaviour under distribution-shift conditions.
4.4. ETC Benchmarking Against SMAP and Methodological Considerations
The ETC analysis places the SABM in a rigorous multi-system error framework that avoids the circular dependence of direct validation on a nominally error-free reference [
63]. The SABM ETC RMSE of 0.0347 m
3 m
−3 compares favourably with values reported in the literature for C-band SAR SM retrievals: Coupling SAR and optical data via Water Cloud Model approaches typically yield ETC RMSE in the range 0.040–0.055 m
3 m
−3 [
64,
65], while the SMAP L3 Enhanced product achieves approximately 0.040 m
3 m
−3 at core validation sites globally [
66]. This resolution difference means the comparison is not a fully equivalent accuracy benchmark across support scales: the point-support ETC reference used here favours SABM, which shares the same point support as the ground sensors, whereas SMAP’s ~9 km footprint necessarily aggregates sub-pixel heterogeneity that a point-scale comparison cannot credit. The results should therefore be interpreted as evidence that SABM provides accurate, high-resolution point estimates at instrumented sites, rather than as a demonstration that a point-anchored, lag-assisted model is unconditionally superior to an independent, globally consistent satellite product across all operational contexts, for which SMAP remains uniquely suited.
The direct comparison between SABM (ETC RMSE = 0.0347 m3 m−3) and SMAP (ETC RMSE = 0.0419 m3 m−3) at the same validation site reveals a 17% error reduction, consistent with the expected improvement from resolving sub-pixel SM heterogeneity that is smoothed in the SMAP footprint average. It should be noted, however, that the ETC assumption of mutually uncorrelated errors may be partially violated at the station level: both SABM and SMAP are sensitive to canopy water content through vegetation indices and the WCM, respectively, which could introduce a common-mode error component correlated across systems. Future work applying ETC over longer records and multiple sites would provide more robust uncertainty bounds on the inter-system error comparisons reported here.
A more fundamental independence concern applies specifically to the Ground–SABM pair. Because the lag-augmented SABM configuration uses each station’s own antecedent ground SM (SMlag1, SMlag2) as direct predictors, and because SABM is trained against ground SM as its regression target, SABM’s prediction error cannot be assumed statistically independent of the ground sensor’s own error, as ETC formally requires. This is not merely a theoretical concern: the raw, pre-clipping ETC error variance for the ground reference was negative—a mathematically inadmissible result under the ETC model, and a recognised diagnostic of correlated errors among collocated systems—at 18 of the 20 stations, precluding a stable, physically interpretable Ground ETC RMSE at the network level. Consistent with this, the point-scale correlation between ground observations and SABM predictions (r = 0.88) was substantially higher than either the Ground–SMAP (r = 0.53) or SABM–SMAP (r = 0.33) correlations, as expected if a component of SABM’s output reproduces the ground record’s own persistence structure rather than representing an independent estimate of it. SABM’s cross-station variability (SD = 0.0112 m3 m−3) was also roughly twice that of SMAP (SD = 0.0054 m3 m−3), indicating that SABM’s point-scale error is less stable across the network, plausibly reflecting station-to-station differences in lag-feature informativeness (e.g., local rainfall persistence patterns) rather than a spatially uniform error source as SMAP’s coarser footprint tends to produce. The SABM and SMAP error variances themselves remained non-negative at every station; however, non-negativity is a necessary admissibility condition, not evidence of unbiasedness, and does not rule out that the same Ground–SABM error correlation responsible for the negative Ground variance that also biases the SABM (and potentially SMAP) variance estimates—most plausibly toward underestimating SABM’s true error, since correlated errors are known to deflate estimated variances in triple collocation analysis. The ETC results as a whole should therefore be read as a relative benchmark of SABM’s and SMAP’s error magnitudes against a ground reference whose own error term cannot be cleanly isolated, rather than as a fully independent, assumption-satisfying validation of SABM in the classical triple collocation sense. Establishing a genuinely independent estimate of SABM’s absolute accuracy would require evaluation against a reference system sharing no information with SABM’s training or input data, such as a densification network or field campaign not used in model development.
4.5. Preliminary Cross-Site Transferability: Evidence from REMEDHUS
As a preliminary assessment of cross-site generalizability, the REMEDHUS validation provides encouraging evidence that the SABM + Lag SM framework extends beyond the Little Washita calibration domain, although the ablation presented in
Section 3.11 indicates that this transferability is attributable overwhelmingly to antecedent soil moisture persistence rather than to the transferable SAR–SM relationships. The zero-shot transfer result (R
2 = 0.6729, RMSE = 0.0518 m
3 m
−3) is exceeded by a lag-only model trained on Little Washita using no SAR, optical, or soil information at all (R
2 = 0.8783), and a REMEDHUS persistence baseline using only the site’s own antecedent SM record performs better still (R
2 = 0.9349). Adding the SAR/optical/soil/terrain features learned at Little Washita to the lag-only model reduces zero-shot R
2 by 0.2054 rather than improving it, indicating that the specific SAR–SM relationships calibrated at Little Washita do not transfer cleanly to REMEDHUS’s different pedoclimatic setting.
This transferability is physically grounded predominantly in the consistent temporal persistence of soil moisture across climates. The lag-1 SM autocorrelation at REMEDHUS (r > 0.80 at the 12-day Sentinel-1 cycle) is comparable to Little Washita (r = 0.8168), confirming that the temporal persistence of surface SM is a climate-independent signal [
67] that the model can exploit at both sites; this single mechanism accounts for the large majority of the observed zero-shot skill (
Section 3.11). While SAR backscatter sensitivity to SM is in principle governed by a universal dielectric mixing model [
16], the present results suggest that the specific, site-calibrated form of this relationship learned at Little Washita does not transfer as readily as the persistence signal: the full model’s zero-shot R
2 is measurably lower than that of a lag-only model excluding all SAR, optical, and soil information. We therefore attribute the positive zero-shot R
2 primarily to antecedent SM persistence, with the SAR/optical/soil-property component contributing comparatively little transferable skill in this cross-continental setting.
The observed positive bias of +0.0277 m
3 m
−3 in zero-shot transfer may be partly associated with the soil pH distribution mismatch between the two sites: REMEDHUS soils have a mean pH of approximately 7.2 (alkaline, smectite-poor) [
42], which lies near the upper boundary of the Little Washita training distribution (pH 6.4–7.2), and the model’s learned pH–SM dependence, if extrapolated beyond this boundary, would predict an overestimate in the observed direction. However, this account remains directional and unverified rather than quantitatively established, and it should be weighed against the ablation evidence in
Section 3.9 indicating that soil pH’s SHAP importance is substantially attributable to station-identity memorisation within Little Washita—a mechanism that does not straightforwardly transfer to REMEDHUS’s entirely distinct station network. Alternative or additional contributors to the bias (e.g., other static-feature distribution shifts, sensor differences, or SAR calibration offsets) cannot be ruled out. A simple local bias correction (subtracting the mean residual from a small calibration subset) would likely reduce this bias substantially; however, such post-processing is intentionally omitted here to preserve a conservative, worst-case assessment of zero-shot transferability.
The local leave-one-station-out CV result at REMEDHUS (R
2 = 0.9175, RMSE = 0.0260 m
3 m
−3) exceeds Little Washita performance and indicates that the SABM + Lag SM architecture generalises effectively to Mediterranean semi-arid conditions. The larger SM dynamic range at REMEDHUS (std = 0.0978 vs. 0.0932 m
3 m
−3) is only a ~5% difference and cannot by itself account for the full R
2 improvement; the more substantial driver is the absolute RMSE reduction (0.0413 → 0.0260 m
3 m
−3), which may reflect differences between the leave-one-station-out and temporal hold-out validation evaluation protocols, a stronger local antecedent-persistence signal at REMEDHUS stations, or REMEDHUS’s smaller station count reducing between-station heterogeneity. Distinguishing between these explanations would require further investigation. Taken together, the cross-site results suggest that the method is applicable across both sub-humid (Little Washita) and semi-arid Mediterranean (REMEDHUS) regions with in situ SM monitoring infrastructure, and that zero-shot transfer achieves operationally useful accuracy without local retraining, provided a local antecedent SM record is available to compute the lag features (
Section 3.11); the SAR/optical/soil-based component of this accuracy does not itself constitute demonstrated cross-site transferability.
4.6. Limitations and Future Research Directions
Despite the encouraging cross-site results, several limitations warrant explicit acknowledgement. First, the two validation sites span different climate classifications (sub-humid continental at Little Washita vs. semi-arid Mediterranean at REMEDHUS), which if anything strengthens the transferability evidence presented here; nonetheless, transferability to humid tropical, boreal, or monsoon-dominated climates—where SM dynamics are driven by fundamentally different processes—remains to be demonstrated. Second, the zero-shot transfer bias (+0.0277 m3 m−3) highlights the sensitivity of SABM predictions to soil pH distribution shifts across sites, suggesting that a small local calibration dataset (as few as 50–100 observations) could be used to fine-tune the model before operational deployment at a new location.
Third, REMEDHUS Sentinel-2 optical coverage was 71.6% (vs. ~100% at Little Washita), introducing a moderate proportion of missing optical features. Records with incomplete optical coverage were excluded via listwise deletion rather than imputed, reducing the effective REMEDHUS evaluation sample; a principled treatment of missing optical inputs—for instance, through imputation or uncertainty-weighted feature integration—would improve robustness and sample retention in persistently cloudy regions. Fourth, the lag SM framework requires antecedent in situ observations, limiting applicability to ungauged locations without historical SM records.
Future research directions include: (i) validation across humid and tropical sites using ISMN networks (e.g., HOBE in Denmark, OzNet in Australia); (ii) development of explicit bias correction (e.g., CDF matching), spatial downscaling, or data-assimilation techniques for satellite-derived lag proxies—motivated by the direct test in
Section 3.10, which found that naive SMAP substitution at held-out stations does not recover the accuracy benefit of true in situ lag (mean R
2 = −0.5520 vs. +0.5134); (iii) incorporation of local bias correction mechanisms for improved zero-shot transferability; and (iv) extension of the temporal lag framework to root-zone SM through empirical transfer functions or coupled vadose-zone models.