Next Article in Journal
Anomalies on Ionospheric Electron Density Before the 2024 Noto Peninsula Earthquake Using Oblique Ionosondes
Previous Article in Journal
Evaluation of Precipitation and Temperature from Multiple Products and CMIP6 Simulations over the Qinghai–Tibet Plateau
Previous Article in Special Issue
Assessing Climate Efficiency with Random Forest, DEA, and SHAP in the Eastern Black Sea Region, Türkiye
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Machine Learning-Based Hydrological Drought Prediction Integrating Teleconnections and Hydrological Memory in a Semi-Arid Basin, Algeria

1
Department of Civil Engineering, Erzincan Binali Yildirim University, Erzincan 24060, Türkiye
2
Laboratory of Water & Environment, Faculty of Nature and Life Sciences, University Hassiba Benbouali of Chlef, B.P 78C, Ouled Fares, Chlef 02180, Algeria
3
Department of Civil Engineering, Siirt University, Siirt 56100, Türkiye
4
Geography Department, Faculty of Arts and Sciences, Igdir University, Igdir 76000, Türkiye
5
Department of Geography, Nakhchivan State University, AZ 7012 Nakhchivan, Azerbaijan
6
G.B. Pant National Institute of Himalayan Environment, Garhwal Regional Centre, Srinagar 246174, Uttarakhand, India
*
Authors to whom correspondence should be addressed.
Atmosphere 2026, 17(7), 670; https://doi.org/10.3390/atmos17070670
Submission received: 3 June 2026 / Revised: 1 July 2026 / Accepted: 1 July 2026 / Published: 4 July 2026
(This article belongs to the Special Issue Machine Learning for Hydrological Prediction and Water Management)

Abstract

Hydrological drought forecasting in semi-arid basins is challenging due to the combined influence of meteorological forcing, large-scale atmospheric teleconnections, and basin memory processes, which are rarely jointly analysed within a leakage-free predictive framework. This study addresses this gap by evaluating gradient-boosted trees and neural forecasting models for one-month-ahead prediction of the Standardized Runoff Index (SRI) in two sub-basins of the Wadi Sahaouat Basin, Algeria. The models include gradient-boosted regression trees (GBRT), A-N-BEATS, A-N-HiTS, and TiDE, representing distinct forecasting architectures. Predictors consist of the Standardised Precipitation Index (SPI), seven teleconnection indices (NAO, AO, EAWR, SCAND, MEI, SOI, WeMO), and their one- to three-month lags. Two scenarios are tested: Scenario 1 uses SPI and teleconnection lags only, while Scenario 2 additionally includes lagged SRI values (SRI_lag1–3) to represent hydrological memory. A train-only Variance Inflation Factor (VIF > 10) procedure is applied to remove multicollinearity without data leakage. In Basin 1, SRI lags were excluded due to strong collinearity with SPI lags (r = 0.984), resulting in identical inputs for both scenarios. In Basin 2, SRI lags were retained to assess their predictive contribution. GBRT achieved the best overall performance across both basins and scenarios, with mean RMSE, NSE, and KGE values of 0.0682, 0.9907, and 0.8945, respectively. TiDE ranked second overall, with a mean RMSE of 0.1166, followed by A-N-HiTS in third place with a mean RMSE of 0.1203 and A-N-BEATS with the weakest overall performance, with a mean RMSE of 0.2159. These results indicate that gradient-boosted trees remain highly competitive with neural models for small monthly hydrological datasets and that the value of hydrological memory is basin-dependent and varies according to its independence from concurrent meteorological forcing.

1. Introduction

Drought is a complex and recurrent hydro-meteorological phenomenon characterized by a prolonged deficiency of precipitation that leads to significant reductions in water availability across atmospheric, terrestrial, and hydrological systems [1]. It manifests in multiple forms, including meteorological, agricultural, and hydrological droughts. Each form represents a different stage of water deficit and collectively they exert profound impacts on ecosystems, agriculture, and socio-economic systems [2,3]. In semi-arid regions, where climatic conditions are inherently variable and water resources are already limited, droughts occur more frequently and with greater severity [4]. The combined influence of erratic rainfall patterns, high evapotranspiration rates, and increasing climate variability has further intensified regional vulnerability, making accurate drought assessment and timely prediction a critical priority [5,6].
Traditionally, drought assessment has relied on standardized indices such as SPI, SPEI, and PDSI, along with classical statistical approaches, to quantify drought severity and duration. While these methods provide valuable insights into drought characterization, they are inherently limited in capturing the complex, nonlinear interactions among hydro-meteorological variables [7,8]. These limitations become more pronounced under changing climate conditions, where non-stationarity and the increasing frequency of extreme events challenge the reliability of conventional statistical frameworks [9]. Empirical evidence further highlights these constraints; for instance, Yang et al. (2020) reported that PDSI overestimates global drought-affected areas by approximately 10–20%, despite showing moderate to strong agreement (r ≈ 0.6–0.8) with climate model estimates, with notable regional inconsistencies [10]. Similarly, Jain et al. (2015) demonstrated that although indices such as SPI, EDI, and CZI exhibit strong inter-correlation (r ≈ 0.85–0.95), simpler approaches like rainfall departure fail to adequately represent drought severity, while SPI-12 effectively captures major drought events (e.g., 2002 and 2007) in the Ken River Basin [11]. These findings underscore the need for more robust and adaptive methodologies.
Building upon these advancements, recent years have witnessed a rapid shift toward machine learning (ML) techniques for hydro-meteorological drought assessment and forecasting, owing to their ability to model complex, nonlinear relationships [12,13,14]. Models such as Artificial Neural Networks (ANN), Fuzzy Logic, Coactive Neuro-Fuzzy Inference System (CANFIS), Support Vector Machines (SVM), Adaptive Neuro-Fuzzy Inference System (ANFIS), Random Forest (RF), and Extreme Learning Machines (ELM), along with hybrid and deep learning approaches, have demonstrated strong capabilities in capturing intricate interactions among climatic and hydrological variables. By integrating multi-source datasets including precipitation, temperature, evapotranspiration, soil moisture, and remote sensing variables, ML models significantly enhance drought prediction accuracy and robustness [15,16,17,18,19,20]. For instance, ANN achieved the highest prediction accuracy (NSE ≈ 0.92–0.96; RMSE ≈ 0.10–0.18), outperforming conventional approaches, while SVM and RF also exhibited strong performance (NSE ≈ 0.85–0.93). Further improvements were achieved using hybrid ELM models in the Wadi Mina Basin, with NSE values of approximately 0.95–0.98 and prediction errors reduced by 15–25% [21]. A GAM-based hybrid statistical–machine learning model for meteorological drought forecasting also demonstrated strong predictive performance, with NSE ≈ 0.88–0.94 and RMSE ≈ 0.10–0.20, outperforming traditional statistical models (NSE ≈ 0.75–0.85) and reducing forecast error by approximately 10–25% across different lead times [22]. More recently, Random Forest (RF), Support Vector Machine (SVM), Gradient Boosting Machine (GBM), and Artificial Neural Network (ANN) achieved high predictive accuracy for meteorological drought forecasting, with NSE ≈ 0.90–0.97, RMSE ≈ 0.08–0.18, and MAE ≈ 0.05–0.12, reducing prediction error by approximately 15–25% compared with conventional models in semi-arid regions [23].
Despite these advancements, significant research gaps remain in integrating lagged multivariate inputs with emerging deep learning architectures for short-term hydrological drought prediction, particularly in data-scarce semi-arid basins. Moreover, limited attention has been given to disentangling the relative influence of atmospheric teleconnections, basin memory, and local meteorological drivers within unified modeling frameworks. The present study aims to develop a lagged multivariate forecasting framework for one-month-ahead prediction of the Standardized Runoff Index (SRI) in the semi-arid Wadi Sahaouat Basin (Algeria) by comparing the performance of gradient-boosted regression trees (GBRT) and advanced neural architectures (A-N-BEATS, A-N-HiTS, and TiDE), while disentangling the relative contributions of meteorological forcing, atmospheric teleconnections, and basin memory processes in hydrological drought prediction.

2. Data and Methodology

2.1. Study Area and Data

The Wadi Sahaouat basin (2140 km2) is a sub-basin of the Macta basin (14,390 km2), located approximately 400 km west of Algiers (Figure 1). The basin is bounded by longitudes 0°07′ E and 0°56′ E and latitudes 35°07′ N and 35°20′ N, and encompasses two main administrative provinces Saida (75% of the total area) and Mascara (25%). The Digital Elevation Model (DEM, 30 m resolution; downloaded from http://earthexplorer.usgs.gov/) indicates that the basin ranges in elevation from 424 m to 1327 m above sea level, with the highest-altitude class (1000–1200 m) covering 35.22% of the total area. The Wadi Sahaouat watercourse is formed by the confluence of Wadi Taria and Wadi Saida, with a total channel length of 98 km, supplying the Ouizert dam reservoir, which has an initial storage capacity of 100 hm3.
Monthly rainfall and runoff records were obtained from the National Agency for Hydraulic Resources (ANRH) as shown in Table 1. Although the available station archive extends over 1973/74–2014/15, the forecasting experiments were conducted using the common monthly period shared by the rainfall, runoff, teleconnection, SPI/SRI, and lagged predictor datasets. After temporal alignment and feature construction, the final modelling dataset covered 1979–2015, corresponding to 440 monthly records for each basin. All SPI-1 and SRI-1 calculations, lagged predictor construction, VIF screening, chronological train–validation–test splitting, and model evaluation were therefore performed on this common modelling period.
The climatology of the study area is typically Mediterranean continental. The analysis of the monthly temperatures (T) over 34 years (1977/2010) of the Saida station (ID: 536) showed that July and August are the hottest months of the year, with average temperatures of 27 °C, while January records low temperatures of up to 3.1 °C. The inter-annual average temperature is 16.7 °C. The mean annual relative humidity (RH) is 58%. Monthly RH exhibits a pronounced seasonal pattern, with the lowest values occurring in July (mean: 39%), when monthly averages can decrease to as low as 17% in some years. In contrast, the highest RH values are observed during the winter months (December–February), with monthly means ranging from 68% to 71% and reaching up to 90% in individual years. For the wind speed (Ws), Average monthly wind speeds generally range from 2.4 m/s in September and October to 3 m/s in April. Analysis of sunshine hours at the Saida station shows that the highest level was recorded in July, with an average of 325 h, and the lowest in January, with 183 h. Furthermore, the calculation of corrected potential evapotranspiration showed a maximum value of 139.70 mm and a minimum of 17.80 mm. Summer is the most dominant period of the year, due to the rise in temperature at this time of year. The aridity index, which stands at around 12.94 at the Saida station, reflects a semi-arid climate.

2.2. Drought Index Calculation

The quantitative assessment of drought in this study is conducted using a suite of standardized and multivariate indices, providing a multi-dimensional perspective on water deficits. To ensure computational consistency and transparency, all drought indices, including meteorological and hydrological, were calculated using the PyDRGHT Python 3.11.9 package [24].

2.2.1. Standardized Precipitation Index (SPI)

The Standardized Precipitation Index (SPI) (McKee et al. 1993) is the primary metric used to quantify meteorological drought by assessing precipitation deficits over user-defined timescales [25]. The calculation begins by fitting long-term monthly precipitation data to a probability density function, most commonly the Gamma distribution, given its effectiveness in representing skewed rainfall data [26]. The Gamma distribution is defined by the function:
f x = 1 β α Γ α x α 1 e x β
where α and β represent the shape and scale parameters, respectively, and Γ(α) is the Gamma function. Since precipitation data often includes zero values where the Gamma function is undefined, the cumulative probability H(x) is adjusted using the formula:
H x = q + 1 q G x
where q is the probability of zero precipitation and G(x) is the cumulative probability from the fitted Gamma distribution. Finally, the cumulative probability is transformed into a standard normal variable (Z) with a mean of zero and a variance of one. This standardization allows the SPI to indicate the number of standard deviations an observed precipitation event deviates from the long-term median, with values below −1.0 generally indicating the start of a drought event.

2.2.2. Standardized Runoff Index (SRI)

The SRI serves as a critical indicator for hydrological drought, focusing on surface water availability rather than atmospheric moisture. The computational procedure for the SRI is mathematically analogous to the SPI; however, the input variable is streamflow or runoff data instead of precipitation. This index is particularly valuable because it incorporates hydrological processes such as snowmelt, infiltration, and groundwater lag, which often cause hydrological drought to persist even after meteorological conditions have improved. Runoff data is typically fitted to an appropriate distribution, such as the Log-normal or Gamma distribution, and then transformed into a standard normal distribution. The resulting index provides a standardized measure of water volume within the system, where negative values indicate a deficit in runoff relative to historical norms [27]. The SRI-1 timescale was selected to maintain consistency with the one-month-ahead forecasting framework adopted in this study. While longer timescales (e.g., SRI-3, SRI-6, and SRI-12) are commonly used to characterize seasonal and long-term hydrological drought conditions, the objective of this work was to evaluate short-term drought forecasting performance. Therefore, SRI-1 was considered the most appropriate indicator for the selected prediction horizon. Prior to standardization, runoff observations were examined for low-flow and zero-flow conditions to ensure the suitability of the fitted probability distribution. This procedure improves the reliability of the resulting SRI values and supports robust short-term drought assessment.

2.3. Predictor Design and Input Scenarios

2.3.1. Teleconnection and Seasonal Predictors

All analyses were conducted at the monthly time step. The core drought variables were SPI-1 and SRI-1, representing meteorological and hydrological drought at the same timescale. The teleconnection block comprised monthly series for the North Atlantic Oscillation (NAO), Arctic Oscillation (AO), East Atlantic/West Russia pattern (EAWR), Scandinavian pattern (SCAND), Multivariate ENSO Index version 2 (MEI), Southern Oscillation Index (SOI), and Western Mediterranean Oscillation (WeMO), aligned to a common monthly calendar. Although ENSO-related indices (MEI and SOI) are primarily associated with tropical Pacific variability, they are included due to their ability to influence North Atlantic and Mediterranean climate conditions through large-scale atmospheric teleconnection pathways and delayed circulation responses, which may indirectly modulate precipitation and drought dynamics in the study region. Seasonal structure was encoded through sine and cosine transformations of the calendar month (month_sin, month_cos), which preserve continuity between December and January and are better aligned with the physical periodicity of drought and teleconnection signals than a raw integer month label.

2.3.2. Lagged Predictors and Hydrological Memory

Lags were constructed deliberately rather than exhaustively. For every atmospheric index, the contemporaneous value and the one-, two-, and three-month lags were retained. The same lag depth was applied to SPI and, in Scenario 2, to SRI. This window is physically motivated: meteorological drought propagates into hydrological drought over timescales of weeks to a few months, depending on soil moisture buffering, catchment storage, and runoff response.
Multicollinearity was screened using the Variance Inflation Factor (VIF), computed exclusively on the training partition to prevent information from the validation or test sets from influencing feature selection. The procedure is iterative: at each step, the predictor with the highest VIF above a threshold of 10 is identified [28]; unless protected, it is removed and VIF is recomputed on the reduced set. SPI and its lags (SPI lag1–3) were always protected from removal as the primary meteorological forcing. In Scenario 2 for Basin 2, the lagged SRI values (SRI_lag1–3) were additionally protected to allow the scientific question of whether hydrological memory improves forecasting to be directly tested. Despite this protection, the VIF values of the SRI lags in Basin 2 were elevated (SRI_lag1 = 46.6, SRI_lag2 = 47.0, SRI_lag3 = 43.6), reflecting substantial but not perfect collinearity with SPI lags. In Basin 1, SRI lags were not protected because their VIF exceeded 40 (r = 0.984 collinearity with SPI lags), making their inclusion genuinely redundant. Despite its elevated VIF, SRI was intentionally retained because it represents hydrological memory and was needed to assess its added predictive value.

2.4. Leakage-Free Preprocessing

2.4.1. Chronological Train–Validation–Test Split

Model development followed a strictly chronological split: approximately 70% of the monthly series was assigned to training, the next 15% to validation, and the final 15% to testing. Predictor scaling was estimated from the training subset alone and applied to validation and test data. The same rule governed target scaling in the deep-learning models, ensuring that test scores reflect forward-looking model skill rather than leakage artefacts.

2.4.2. Scaling and VIF Screening

Two input scenarios were designed to isolate a specific hydrological question: is monthly streamflow deficit (SRI) primarily explained by concurrent and antecedent precipitation deficit (SPI) and atmospheric teleconnection forcing, or does the basin’s own recent hydrological history add predictive value beyond what meteorological signals provide?
Scenario 1 used SPI-1 (contemporaneous), SPI_lag1–3, seven atmospheric teleconnection indices and their lags 1–3, plus the two seasonal features, giving a maximum of 34 predictors before VIF screening. No SRI-derived predictor was included. Scenario 2 added SRI_lag1, SRI_lag2, and SRI_lag3, raising the maximum feature count to 37. Both scenarios used identical chronological splits and VIF screening logic; the sole structural difference was whether past SRI values were protected from removal and allowed to enter as predictors (Table 2).
Atmospheric oscillations influence hydrological drought through both direct meteorological forcing and delayed hydrological responses, as reflected in the variable design and lag structure used in this study. The inclusion of contemporaneous and 1–3 month lagged values of teleconnection indices (NAO, AO, EAWR, SCAND, MEI, SOI, and WeMO) indicates that their effects are not instantaneous but propagate over time. This is consistent with the physical understanding that meteorological drought, represented by SPI, translates into hydrological drought (SRI) over timescales of weeks to several months, depending on soil moisture storage, catchment characteristics, and runoff processes [29]. Previous studies have also shown that large-scale atmospheric circulation patterns significantly influence regional precipitation regimes and, consequently, streamflow variability [30,31]. In this study, results from Scenario 1 suggest that meteorological drought and atmospheric forcing alone can explain a substantial portion of hydrological drought variability, while Scenario 2 demonstrates that incorporating lagged SRI values improves predictions, highlighting the role of hydrological memory. This finding aligns with the understanding that hydrological systems are influenced not only by current atmospheric conditions but also by antecedent states [32,33].
The VIF procedure was used as a leakage-free screening tool, but physically meaningful predictors were allowed to be protected when they were required to test a specific hydrological hypothesis. In Scenario 2, SRI_lag1–3 were introduced to evaluate whether antecedent hydrological conditions add predictive information beyond SPI, teleconnection indices, and seasonality. Therefore, the final predictor sets were not forced to be identical across basins. In Basin 1, SRI lags were removed because they were nearly redundant with SPI lags, indicating that hydrological memory was not separable from meteorological memory at the monthly scale. In Basin 2, SRI_lag1–3 were retained as a controlled sensitivity experiment to quantify the predictive effect of hydrological memory despite high VIF values. Because the study is predictive rather than coefficient-explanatory, high VIF was interpreted as a caution for feature attribution, not as an automatic reason to remove hydrologically meaningful lag variables.

2.5. Forecasting Models

In this study, one tree-based benchmark and three custom neural forecasting variants were evaluated within a one-step-ahead lagged multivariate regression framework. The tree-based benchmark was implemented as GBRT (Gradient-Boosted Regression Trees), while the neural models were constructed as custom implementations inspired by N-BEATS (Neural Basis Expansion Analysis for Interpretable Time Series Forecasting), and N-HiTS (Neural Hierarchical Interpolation for Time Series Forecasting). Unlike their original sequence-based formulations, these architectures were applied to VIF-screened multivariate lagged predictors representing hydroclimatic memory, teleconnection signals, and seasonal covariates. Accordingly, the N-BEATS-inspired model was used as a residual fully connected forecasting block, the N-HiTS-inspired model as a multi-branch hierarchical dense regressor, and the TiDE-inspired model as a dense encoder–decoder regressor with a skip connection. Therefore, the present implementations should be interpreted as architecture-inspired neural regressors rather than exact replications of the original sequence-based forecasting models (Figure 2). For a one-month-ahead forecast, SRI(t + 1) was predicted exclusively from variables available at time t or earlier. No contemporaneous or future information from month (t + 1) was included among the predictors.

2.5.1. Gradient-Boosted Regression Trees (GBRT)

GBRT is a powerful ensemble learning method in which decision trees are trained sequentially, with each new tree learning the residual errors of the previous model to improve overall performance [34]. Due to its ability to capture nonlinear relationships, it is widely used in hydroclimatic forecasting problems [35].
In this study, the monthly lagged predictor set obtained after VIF screening was directly used as input to the GBRT model. Model performance was optimized through a compact hyperparameter search over the number of boosting rounds, learning rate, and tree complexity (max depth) on the validation set. These parameters are also central to modern gradient boosting frameworks [36].

2.5.2. Adapted Neural Basis Expansion Analysis for Time Series (A-N-BEATS)

The N-BEATS-inspired architecture offers a clear advantage through its residual, stacked fully connected structure, which progressively refines the predictive signal via successive nonlinear transformations. This design is well suited to capturing informative patterns in lagged hydroclimatic predictors while preserving a relatively simple and stable feed-forward formulation [37]. The model functions by iteratively decomposing the input into components that enhance both backcast and forecast representations. In this study, a custom NumPy-based implementation was developed in line with the original design principles and applied in a monthly forecasting context using the VIF-screened lagged predictor vector. N-BEATS was selected because drought dynamics often combine smooth, low-frequency variability with short-term anomalies, a structure that aligns naturally with residual basis expansion.

2.5.3. Adapted Neural Hierarchical Interpolation for Time Series (A-N-HiTS)

N-HiTS extends the N-BEATS framework by incorporating hierarchical interpolation and multi-rate processing. The N-HiTS-inspired model introduces a multi-branch architecture that enables the representation of predictor information at different levels of abstraction. This enhances predictive flexibility by allowing the model to capture both dominant patterns and finer-scale variations within lagged hydroclimatic inputs [38]. The architecture comprises three parallel branches operating at sub-sampling rates of 1, 2, and 3, each learning features at a distinct effective resolution before their outputs are combined. This configuration facilitates the separation of seasonal variability from inter-annual teleconnection signals, providing a physically meaningful inductive bias for modelling SPI-to-SRI propagation.

2.5.4. Time-Series Dense Encoder (TiDE)

TiDE is a dense encoder–decoder architecture designed to model multivariate covariates and nonlinear dependencies without relying on attention mechanisms [39]. The model encodes the input through multilayer perceptron blocks and decodes it into the forecast horizon, while a skip connection preserves the direct contribution of exogenous variables. TiDE was included because recent benchmarking studies indicate that carefully structured dense architectures can achieve strong performance when forecasting problems are governed by short memory and compact lag structures. Although N-BEATS, N-HiTS, and TiDE were originally developed as sequence-based forecasting models, they were implemented here as supervised regression models using lagged multivariate input features. As a result, the comparison reflects their behaviour within a tabular, feature-based forecasting framework rather than their original sequence-based formulations. Although the TiDE-based model was also implemented within the same forecasting framework, the original TiDE name was retained because the implemented structure follows the core TiDE formulation more directly. In contrast, the prefix “A-” was used for A-N-BEATS and A-N-HiTS to indicate more substantial architectural adaptations from the original N-BEATS and N-HiTS designs.

2.6. Local Interpretability

To support the physical interpretation of model outputs, a local surrogate explanation module was implemented following the LIME framework [40]. A LIME-inspired approach was applied to interpret the best-performing model at the instance level. For a selected input vector, 300 perturbed samples were generated, and model predictions were obtained for each sample. The distances between perturbed samples and the target instance were calculated, and a Gaussian kernel was used to assign proximity-based weights. A weighted Ridge regression model was then fitted locally, and the resulting coefficients were interpreted as local feature-importance scores (LIME weight), indicating both the direction and magnitude of each predictor’s contribution around the selected prediction.

2.7. Evaluation Metrics and Statistical Tests

Model performance was evaluated using RMSE, MAE, NSE, and KGE. RMSE and MAE quantify absolute forecast error, while NSE represents the proportion of explained variance. NSE and KGE provide hydrologically meaningful performance measures and are more interpretable than generic machine learning metrics when the target variable is a hydroclimatic index [41,42,43]. Statistical differences between model error distributions were assessed using the Wilcoxon signed-rank test and the paired-sample t-test, both applied to the 66-month held-out test set, which was not used during model training or hyperparameter tuning.
R M S E = 1 n i = 1 n y i y ^ i 2
M A E = 1 n i = 1 n y i y ^ i
N S E = 1 i = 1 n y i y ^ i 2 i = 1 n y i y - i 2
K G E = 1 r 1 2 + α 1 2 + β 1 2
α = σ y ^ σ y , β = μ y ^ μ y
Here, yi represents the observed value, ŷi is the predicted value, and ȳ is the mean of the observed values, while n denotes the number of observations. In the KGE metric, r is the correlation between observed and predicted values, α is the ratio of standard deviations and β is the ratio of means. Note that because SRI is a standardized, zero-centred index, the KGE bias-ratio component β = μ y ^ / μ y may become unstable when the observed mean over the test period is close to zero. Therefore, KGE was interpreted together with its components r, α, and β rather than as a standalone performance score.

3. Results

3.1. Basin 1: Forecast Performance

The GBRT model shows the strongest alignment, with NSE values around 0.997–0.998 and a low MAE of approximately 0.029. The points are tightly clustered along the diagonal, indicating that this model reproduces the observed values with relatively small deviations. Based on these findings, GBRT appears to provide the most consistent fit among the models shown.
The A-N-BEATS model demonstrates a slightly weaker performance, with NSE values around 0.95–0.96 and MAE values between 0.12 and 0.14. Although the general trend is still captured, the scatter around the diagonal is more pronounced, suggesting higher variability in the predictions compared to GBRT.
The A-N-HiTS model performs better than A-N-BEATS, with NSE values close to 0.98 and MAE around 0.076–0.078. The points are more tightly distributed along the diagonal, indicating improved predictive accuracy, though still not reaching the level observed in GBRT.
Finally, the TiDE model shows performance comparable to A-N-HiTS, with NSE values around 0.975–0.98 and MAE values ranging from approximately 0.086 to 0.099 (Figure 3). The predictions generally follow the observed values closely, although minor deviations are visible.
All metrics are computed on the 66-month held-out test set, which was withheld from both training and hyperparameter optimisation. In Scenario 1, GBRT achieved the best performance (RMSE = 0.0377, MAE = 0.0288, NSE = 0.9978, NSE = 0.9978, KGE = 0.8588). N-HiTS ranked second (RMSE = 0.0979, NSE = 0.9849, KGE = 0.4470), followed by TiDE (RMSE = 0.1240, KGE = 0.3695) and A-N-BEATS (RMSE = 0.1531, KGE = 0.4801).
In Scenario 2, VIF screening eliminated all three SRI lags owing to r = 0.984 collinearity with SPI lags; GBRT was therefore trained on an identical feature set and produced identical predictions (RMSE = 0.0377). Among the neural models, N-HiTS slightly worsened to RMSE = 0.1027 (KGE = 0.7948) and TiDE improved to RMSE = 0.1116 (KGE = 0.6976). A-N-BEATS deteriorated substantially (KGE = −0.0338), suggesting that its residual decomposition was destabilised by the autoregressive SPI lag configuration in this basin (Table 3).
GBRT appears to align more closely with the observed values, showing smaller deviations during both low and high SRI conditions (Figure 4). A-N-HITS and TIDE also reproduce the general variability reasonably well, although some discrepancies are visible at peak events. A-N-BEATS shows comparatively larger deviations, particularly during sudden changes, which may suggest some limitations in capturing abrupt fluctuations.
Residual boxplots further illustrate differences in model performance across Basin 1 scenarios (Figure 5). GBRT exhibits the narrowest residual distributions and medians closest to zero, indicating more accurate and stable predictions. A-N-HiTS and TiDE show relatively moderate residual variability, while A-N-BEATS displays the widest spread of residuals and the largest interquartile ranges, suggesting less consistent predictive performance. Although all models produce residuals on both sides of zero, the smaller dispersion observed for GBRT indicates lower prediction errors and greater robustness across the evaluated scenarios.
Residual boxplots and performance metrics provide a quantitative comparison. The residual distributions show that GBRT has a narrower spread centered near zero, which may indicate more stable and less biased predictions (Figure 6). A-N-BEATS displays a wider range of residuals, including more extreme values, implying higher variability. A-N-HiTS and TiDE fall between these two, with moderate dispersion.
The metric plots in Figure 6 support these observations. GBRT consistently achieves the lowest RMSE, along with the highest NSE and KGE values across both scenarios, indicating a closer match to observed data. A-N-HiTS and TiDE demonstrate intermediate performance, with moderate error levels and relatively strong efficiency scores. A-N-BEATS shows higher RMSE and lower NSE and KGE, suggesting comparatively larger deviations from observations.
All models capture the general behavior of SRI-1, but differences in accuracy and error distribution are evident, with GBRT showing more consistent agreement across the presented evaluations.

3.2. Basin 2: Forecast Performance

Basin 2 was analysed using its own SPI-1 and SRI-1 series over the common modelling period January 1979–August 2015, comprising 440 monthly records after temporal alignment and lagged predictor construction. In Scenario 1, GBRT ranked first (RMSE = 0.0985, NSE = 0.9837, KGE = 0.9286). TiDE ranked second (RMSE = 0.1258, KGE = 0.9372), A-N-HiTS third (RMSE = 0.1379, KGE = 0.8980), and A-N-BEATS fourth (RMSE = 0.2918, KGE = 0.8205).
In Scenario 2, SRI_lag1–3 were explicitly retained despite high VIF values (SRI_lag1 = 46.6, SRI_lag2 = 47.0, SRI_lag3 = 43.6). GBRT performance was essentially unchanged (RMSE = 0.0988, KGE = 0.9318), confirming that gradient-boosted trees already exploit the SPI-SRI relationship efficiently. TiDE improved substantially to RMSE = 0.1049 (KGE = 0.9878), the highest KGE recorded across all models and configurations; its encoder–decoder architecture with skip connections appears particularly effective at integrating autoregressive SRI signals with atmospheric covariates. A-N-HiTS achieved RMSE = 0.1428 (KGE = 0.9009). A-N-BEATS improved markedly from Scenario 1 (KGE = 0.8205) to Scenario 2 (KGE = 0.9453), indicating that its residual decomposition blocks can exploit SRI memory even when the direct meteorological signal is the dominant forcing (Table 4).
GBRT appears to track the observed series more closely, with relatively small deviations across the test period (Figure 7). A-N-HiTS and TIDE also capture the overall dynamics, although some differences are visible at peak values. A-N-BEATS shows comparatively larger departures, particularly during rapid changes, suggesting some difficulty in representing sharp transitions.
In Figure 8, Scenario 2 includes lagged predictors (SRI_lag1–3), a slight improvement in alignment can be observed for most models. The predictions appear somewhat smoother and closer to the observed series, especially for A-N-HiTS and TIDE. A-N-BEATS still exhibits noticeable deviations, although there is some reduction in extreme mismatches compared to Scenario 1.
GBRT shows points tightly clustered around the 1:1 line, indicating strong agreement between observed and predicted values (Figure 9). A-N-HiTS and TiDE also demonstrate relatively good alignment, though with more dispersion than GBRT. A-N-BEATS displays a wider spread of points, which suggests higher variability and larger prediction errors. Points above and below the line indicate that both overestimation and underestimation occur across models, but these deviations appear more pronounced for A-N-BEATS.
The residual distributions indicate that GBRT has a relatively narrow spread centered near zero, suggesting more stable predictions (Figure 10). A-N-BEATS shows the widest spread, including more extreme residuals, indicating higher variability. A-N-HiTS and TiDE fall in between, with moderate dispersion.
The metric comparisons are consistent with these observations. GBRT achieves the lowest RMSE (0.099), along with high NSE (0.984) and KGE (0.93) values across both scenarios. A-N-HITS and TiDE demonstrate intermediate performance, with moderate error levels and relatively high efficiency scores, and some improvement is visible in Scenario 2. A-N-BEATS shows higher RMSE (up to 0.292), along with comparatively lower NSE and more variable KGE values. All models capture the general behavior of SRI-1 in Basin 2, but their accuracy and consistency differ. GBRT appears to provide a more stable and closer agreement with observations, while A-N-HiTS and TIDE show moderate performance. A-N-BEATS exhibits higher variability and larger deviations, although some improvements are observed when lagged predictors are included in Scenario 2 (Table 5).
The marginal contribution of SRI_lag1–3 was evaluated by comparing Scenario 2 against Scenario 1 on the same held-out test period. In Basin 1, this contribution could not be separately estimated because VIF screening removed the SRI lags, producing effectively the same final input structure. This indicates that, in Basin 1, the antecedent runoff deficit largely overlaps with antecedent precipitation deficit. In Basin 2, however, retaining SRI_lag1–3 revealed a model-dependent hydrological-memory effect. TiDE improved from RMSE = 0.1258 to 0.1049 and KGE = 0.9372 to 0.9878, while A-N-BEATS improved from RMSE = 0.2918 to 0.2473 and KGE = 0.8205 to 0.9453. In contrast, GBRT remained essentially unchanged, suggesting that the tree-based model already exploited the SPI–SRI relationship efficiently. These findings indicate that hydrological memory provides additional predictive value in Basin 2, but the magnitude of this value depends on model architecture and on the degree to which SRI lags contain information not already represented by SPI lags.
For Basin 1, because SRI_lag1–3 was removed by VIF screening, the marginal contribution of SRI lags is not separately estimable. This result is itself hydrologically meaningful: in Basin 1, antecedent meteorological drought and antecedent hydrological drought are almost indistinguishable at the monthly scale.
For Basin 2, the marginal contribution of SRI_lag1–3 was quantified as follows:
Table 5. Comparison of model performance between Scenario 1 and Scenario 2 in Basin 2 and the associated changes in RMSE and KGE.
Table 5. Comparison of model performance between Scenario 1 and Scenario 2 in Basin 2 and the associated changes in RMSE and KGE.
ModelRMSE Scenario 1RMSE Scenario 2ΔRMSEKGE Scenario 1KGE Scenario 2ΔKGEInterpretation
GBRT0.09850.0988+0.00030.92860.9318+0.0032Almost no marginal gain; GBRT already captures the SPI–SRI relation efficiently.
TiDE0.12580.1049−0.02090.93720.9878+0.0506Clear benefit from SRI memory; RMSE decreased by approximately 16.6%.
A-N-HiTS0.13790.1428+0.00490.89800.9009+0.0029No meaningful RMSE improvement; KGE changed only slightly.
A-N-BEATS0.29180.2473−0.04450.82050.9453+0.1248Strong improvement, although the model remains less accurate than GBRT and TiDE.
These results show that the contribution of SRI_lag1–3 is model-dependent. It is small for GBRT, substantial for TiDE and A-N-BEATS, and limited for A-N-HiTS. Therefore, the inclusion of SRI_lag1–3 does not uniformly improve all models, but it reveals that some architectures are more sensitive to autoregressive hydrological memory than others.

3.3. Cross-Basin Model Ranking

Averaging performance metrics across both basins and both scenarios, GBRT ranked first overall (mean RMSE = 0.0682, NSE = 0.9907, KGE = 0.8945). TiDE ranked second (mean RMSE = 0.1166, NSE = 0.9778, KGE = 0.7480), A-N-HiTS third (mean RMSE = 0.1203, NSE = 0.9755, KGE = 0.7602), and A-N-BEATS fourth (mean RMSE = 0.2159, NSE = 0.9177, KGE = 0.5530). TiDE and A-N-HiTS were closely matched on average RMSE and NSE, but TiDE achieved the highest single-configuration KGE of any model (0.9878 in Basin 2, Scenario 2), consistent with its architecture benefiting most when SRI memory is available as a predictor (Table 6).

3.4. Statistical Significance and Local Interpretability

Pairwise differences in model forecast errors were examined using the Wilcoxon signed-rank test and the paired-sample t-test (Table 7 and Table 8). These tests provide exploratory information on whether the error distributions differ between model pairs. However, because the test data consist of monthly SRI forecasts, the resulting residuals and loss differences may not be fully independent. Therefore, the reported p-values should be interpreted cautiously and should not be treated as standalone proof of model superiority.
Figure 11, Figure 12, Figure 13 and Figure 14 suggest that while model performances are often statistically distinguishable, some models (particularly A-N-HiTS and TiDE) show comparable behavior under certain scenarios. The LIME analysis indicates that both current and lagged hydro-climatic variables contribute to predictions, with differences in feature importance patterns between basins.
The statistical tests generally support the performance-based ranking, particularly the stronger performance of GBRT in terms of RMSE, MAE, NSE, and KGE. Nevertheless, the interpretation of model differences is based primarily on the consistency of the evaluation metrics across basins and scenarios rather than on raw p-values alone. Cases with marginal p-values, such as TiDE versus A-N-HiTS in some scenarios, are therefore discussed as indicative rather than conclusive.
A similar pattern is observed. GBRT again shows statistically significant differences compared to other models in most cases (Figure 12). A-N-HiTS and TiDE appear statistically similar in Scenario 1 (p = 0.282), while in Scenario 2 the difference between them becomes marginally significant (p = 0.039). Additionally, GBRT and TIDE show a small but statistically significant difference in Scenario 2 (p = 0.006), which may suggest some convergence in performance compared to Basin 1. Given the high VIF values of some retained lagged SRI variables in Basin 2, the LIME-based local attributions should be interpreted primarily at the feature-group level rather than as isolated per-variable effects. Therefore, the relative contributions of correlated predictors such as SPI, SPI_lag1, and SPI_lag3 are discussed with caution.
Hyperparameter tuning was performed using a compact grid-search strategy on the validation subset in Table 9. The search was intentionally kept small because the available monthly hydrological dataset was limited, and overly broad tuning could increase the risk of overfitting. For GBRT, two parameter combinations were evaluated: (i) n_estimators = 80, learning_rate = 0.05, max_depth = 4, and subsample = 0.8; and (ii) n_estimators = 100, learning_rate = 0.03, max_depth = 5, and subsample = 0.9. For the neural models, two configurations were tested: (i) hidden = 64, learning rate = 0.01, epochs = 150, patience = 20, dropout = 0.05, and batch size = 32; and (ii) hidden = 64, learning rate = 0.005, epochs = 150, patience = 20, dropout = 0.10, and batch size = 32.
The optimal configuration was selected according to the lowest validation RMSE. Early stopping was applied to the neural models by monitoring validation loss, with training stopped when no improvement was observed for 20 consecutive epochs. The maximum number of epochs was set to 150. For GBRT, the scikit-learn GradientBoostingRegressor internal early-stopping mechanism was used with n_iter_no_change = 15 and validation_fraction = 0.1. This internal validation subset was drawn only from the chronological training block and did not include observations from the external validation or test partitions. The external chronological validation set was used only to select the best GBRT candidate configuration based on validation RMSE. In contrast, early stopping for the neural models was performed directly by monitoring the external chronological validation loss. Therefore, the stopping mechanisms were model-specific, but the held-out test set remained completely unused during training, early stopping, and hyperparameter selection. Neural network weights were initialized using seed-controlled He-type random initialization. Therefore, the repeated appearance of identical optimal hyperparameters across basins and scenarios reflects the outcome of the validation-based selection process within a compact search space, rather than the use of arbitrarily fixed parameters.
The LIME-based local surrogate outputs were re-examined because the resulting feature weights were extremely small, approximately on the order of 10−14. At this numerical scale, the apparent ordering of predictors should not be interpreted as a robust local sensitivity ranking. Therefore, the LIME results were not used to draw strong conclusions about the dominance of individual predictors. Instead, Figure 13 and Figure 14 are interpreted only as diagnostic local surrogate outputs with limited explanatory value.
Consequently, the interpretation of model behaviour in this study relies primarily on the leakage-free performance evaluation, scenario comparison, and hydrologically meaningful metrics. The comparison between Scenario 1 and Scenario 2 remains useful for assessing the role of hydrological memory, whereas the LIME results are treated cautiously because the near-zero surrogate coefficients may reflect numerical instability or a degenerate local approximation.
Figure 11. Unadjusted Wilcoxon signed-rank test p-value heatmaps for exploratory pairwise comparisons of forecasting model errors in Basin 1 under Scenario 1 and Scenario 2. The p-values should be interpreted cautiously because monthly residuals may exhibit serial dependence.
Figure 11. Unadjusted Wilcoxon signed-rank test p-value heatmaps for exploratory pairwise comparisons of forecasting model errors in Basin 1 under Scenario 1 and Scenario 2. The p-values should be interpreted cautiously because monthly residuals may exhibit serial dependence.
Atmosphere 17 00670 g011
Figure 12. Unadjusted Wilcoxon signed-rank test p-value heatmaps for exploratory pairwise comparisons of forecasting model errors in Basin 2 under Scenario 1 and Scenario 2. These heatmaps are used as supplementary diagnostics rather than definitive evidence of model superiority.
Figure 12. Unadjusted Wilcoxon signed-rank test p-value heatmaps for exploratory pairwise comparisons of forecasting model errors in Basin 2 under Scenario 1 and Scenario 2. These heatmaps are used as supplementary diagnostics rather than definitive evidence of model superiority.
Atmosphere 17 00670 g012
Figure 13. LIME-based local surrogate output for Basin 1. Positive values indicate local contributions toward higher predicted SRI, whereas negative values indicate local contributions toward lower predicted SRI. Because the LIME weights are extremely small, approximately on the order of 10−14, the apparent feature ordering should be interpreted cautiously and should not be considered a robust predictor-importance ranking.
Figure 13. LIME-based local surrogate output for Basin 1. Positive values indicate local contributions toward higher predicted SRI, whereas negative values indicate local contributions toward lower predicted SRI. Because the LIME weights are extremely small, approximately on the order of 10−14, the apparent feature ordering should be interpreted cautiously and should not be considered a robust predictor-importance ranking.
Atmosphere 17 00670 g013
Figure 14. LIME-based local surrogate output for Basin 2. Blue bars indicate positive local contributions associated with higher predicted SRI, whereas orange bars indicate negative local contributions associated with lower predicted SRI. Because the LIME coefficients are close to numerical zero, the plot is retained only as a diagnostic local explanation and is not used to infer dominant predictor effects.
Figure 14. LIME-based local surrogate output for Basin 2. Blue bars indicate positive local contributions associated with higher predicted SRI, whereas orange bars indicate negative local contributions associated with lower predicted SRI. Because the LIME coefficients are close to numerical zero, the plot is retained only as a diagnostic local explanation and is not used to infer dominant predictor effects.
Atmosphere 17 00670 g014

4. Discussion

This study addresses a critical gap in hydrological drought forecasting by systematically evaluating both ensemble and deep learning approaches within a lagged multivariate supervised learning framework for one-month-ahead prediction of the SRI in two sub-basins of the Wadi Sahaouat Basin, Algeria. Four architecturally distinct models namely GBRT, A-N-BEATS, A-N-HiTS, and TiDE were compared to capture diverse learning paradigms. The predictor framework integrates both local and large-scale climatic drivers, including the SPI and key teleconnection indices such as the North Atlantic Oscillation, Arctic Oscillation, and El Niño–Southern Oscillation, along with their lagged effects. To explicitly assess hydrological memory, two scenarios were designed, with Scenario 2 incorporating lagged SRI, enabling a deeper understanding of antecedent basin conditions.
The dominance of GBRT in Basin 1 is strongly supported by existing literature emphasizing the robustness of ensemble tree-based methods under nonlinear and multicollinear conditions. For instance, Chen and Guestrin (2016) demonstrated the effectiveness of boosting algorithms in handling feature redundancy, while Mosavi et al. (2018) highlighted their superior performance in hydrological applications [44,45]. These findings are further reinforced by regional case studies: Achite et al. (2022) reported that ANN, SVM, and RF achieved high predictive accuracy (NSE up to 0.96), with hybrid ELM models further improving performance (NSE up to 0.98 and error reductions of 15–25%) in Algerian basins, demonstrating the effectiveness of ML approaches under semi-arid conditions [21]. Similarly, Pande et al. (2025) showed that models such as MLP and RF achieved NSE up to 0.96 and NSE > 0.90 with improved lead times, while Basak et al. (2022) highlighted the suitability of data-driven models like Prophet (NSE up to 0.92) in data-scarce regions [23,46]. Furthermore, Li et al. (2023) demonstrated that hybrid approaches (e.g., VMD-MLP, EEMD-RF) significantly enhance drought prediction accuracy (NSE up to 0.97), reinforcing the reliability of advanced ML frameworks across diverse climatic settings [47].
In contrast, the improved performance of deep learning models in Basin 2 under Scenario 2 highlights the critical role of hydrological memory. The exceptional performance of TiDE (KGE = 0.9878) is consistent with findings from Lim and Zohren (2021), who showed that encoder–decoder architectures excel when combining exogenous and autoregressive inputs [48]. This is further supported by the classical work of Hundecha and Bárdossy (2004), which established that hydrological systems exhibit strong temporal persistence [49]. Together with recent hybrid and deep learning studies, these results confirm that while ensemble models provide stable baseline performance, deep learning architectures can outperform them when basin memory is explicitly incorporated.
Despite its robust findings, this study has some limitations. The relatively short temporal dataset may not fully capture long-term variability associated with large-scale climate drivers such as the El Niño–Southern Oscillation and North Atlantic Oscillation. Additionally, the analysis is limited to two semi-arid basins, which may restrict broader applicability. Model performance also depends on feature selection and data quality, and the study does not fully explore global interpretability or causal mechanisms, suggesting avenues for future research.
The very high GBRT NSE values indicate that the monthly SRI series contains a strong predictable component under the selected feature design. However, this should not be interpreted as evidence that hydrological drought in a flashy semi-arid basin is nearly deterministic. Instead, the result likely reflects the combined effects of monthly aggregation, strong SPI–SRI coupling, short-term hydrological memory, and the capacity of GBRT to approximate nonlinear relationships among lagged hydroclimatic predictors. GBRT was included as a strong tree-based machine-learning benchmark against the adapted neural architectures, not as a naive hydrological baseline. Because persistence, climatology, AR(1), and ARIMA-type references were not included in the present experimental design, the absolute NSE values should be interpreted as model-set-specific performance rather than as a complete measure of practical forecasting improvement. Future studies should explicitly quantify the incremental value of the proposed framework relative to persistence and autoregressive baselines using rolling-origin validation and skill scores relative to persistence.
The paired t-test and Wilcoxon signed-rank test results were retained as exploratory diagnostics for pairwise model comparison. However, because the test period consists of monthly SRI forecasts, residuals and loss differences may exhibit serial dependence. Therefore, the raw p-values should not be interpreted as definitive evidence of model superiority.
The present modelling design should be interpreted as a targeted hydrological-memory experiment rather than as an exhaustive algorithmic benchmark. GBRT was included as the main strong tree-based benchmark against the adapted neural architectures. Scenario 2 intentionally incorporated SRI_lag1–3 to test whether antecedent hydrological conditions provide additional information beyond SPI, teleconnection indices, and seasonality. Because monthly SRI contains inherent temporal persistence, the high performance obtained in Scenario 2 should be interpreted with caution: it may partly reflect the exploitation of hydrological autocorrelation rather than a purely independent gain from exogenous predictors. Future work should extend this framework by including persistence, climatology, linear regression, AR/ARIMAX, random forest, XGBoost, and LightGBM under rolling-origin validation to quantify the incremental value of complex models over simple but strong reference forecasts.
A more rigorous time-series-aware significance assessment would require explicit reporting of loss-difference autocorrelation, effective sample size, moving-block bootstrap confidence intervals, and multiple-comparison-adjusted p-values. Since these adjusted quantities are not reported in the present version, they are discussed as a recommended extension rather than as evidence supporting the current conclusions.
The neural forecasting models evaluated in this study were adapted implementations inspired by the original N-BEATS, N-HiTS, and TiDE architectures. Therefore, the reported results should be interpreted as comparisons among the implemented model variants under a common experimental framework rather than definitive assessments of the original architectures. Consequently, conclusions regarding the relative performance of neural and boosting-based approaches are limited to the specific implementations, dataset, and experimental settings considered in this study. Although pairwise statistical tests were used to compare model performance, the results should be interpreted in conjunction with the magnitude of performance differences and their practical relevance. Statistical significance does not necessarily imply hydrologically meaningful improvements, particularly when multiple comparisons are performed and model performance metrics are already near their optimal values. Model evaluation was based on a chronological train–test split designed to preserve temporal consistency and avoid information leakage. While this approach reflects a realistic forecasting framework, additional validation strategies may provide further insight into model robustness. Therefore, the reported findings should be interpreted within the scope of the adopted validation design.

5. Conclusions

This study establishes an operationally relevant framework for monthly hydrological drought forecasting by integrating meteorological drought indicators (SPI), atmospheric teleconnection indices, and lagged SRI variables within a unified, leakage-free modelling pipeline. The results consistently show that gradient-boosted regression trees (GBRT) deliver the highest predictive accuracy across both basins and scenarios (mean RMSE = 0.0682; NSE = 0.9907), confirming their robustness for small-sample hydroclimatic applications with limited monthly observations.
A key finding is that the contribution of hydrological memory is strongly basin-dependent. In Basin 1, SRI lag variables were effectively redundant due to near-perfect collinearity with SPI (r = 0.984), indicating that meteorological forcing alone adequately captures hydrological drought dynamics. In contrast, in Basin 2, the inclusion of SRI lags led to clear performance improvements, most notably for the TiDE-inspired model, where KGE increased from 0.9372 to 0.9878. This result demonstrates that autoregressive runoff information can provide additional predictive value when it is not fully explained by precipitation-driven processes.
From a methodological perspective, the study further demonstrates that neural architectures adapted to lagged multivariate inputs can achieve competitive performance, although their advantages remain constrained under limited data availability and short memory structures. Overall, the findings underscore that effective model selection in hydrological drought forecasting depends not only on model complexity but also on data scale, basin-specific characteristics, and the degree to which hydrological memory contains independent information beyond meteorological forcing.
Although GBRT achieved the highest absolute accuracy among the tested models, its performance should be interpreted within the limits of the selected experimental design. The results suggest that gradient-boosted trees are effective in exploiting the short-term memory and SPI–SRI coupling present in monthly hydrological drought series. However, because persistence, climatology, AR(1), and ARIMA-type baselines were not included, the high NSE values should not be interpreted as standalone evidence of near-deterministic one-month-ahead drought predictability. The practical improvement over simple memory-based forecasts remains an important issue for future work.

Author Contributions

O.M.K. and M.A.: Study concept and design, data collection, analyses, interpretation of results. M.A., V.K., M.A.Ç. and K.P.: Writing, Manuscript preparation, review and approved the final version of the manuscript. All authors have read and agreed to the published version of the manuscript.

Funding

This research did not receive any specific grant from public, commercial, or not-for-profit funding agencies.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The datasets used and/or analysed during the current study are available from the corresponding author upon reasonable request.

Acknowledgments

The authors thank the National Agency of Water Resources (ANRH) for providing the hydrological data used in this study, and the General Directorate of Scientific Research and Technological Development of Algeria (DGRSDT) for institutional support. The authors also acknowledge the National Oceanic and Atmospheric Administration (NOAA) Physical Sciences Laboratory (PSL) and Climate Prediction Center (CPC) for maintaining and disseminating the monthly climate-index datasets that underpinned the teleconnection analysis.

Conflicts of Interest

The authors declare no known competing financial interests or personal relationships that could have influenced the work reported in this paper.

Correction Statement

This article has been republished with a minor correction of the information included in the Institutional Review Board Statement and Informed Consent Statement. This change does not affect the scientific content of the article.

References

  1. Rahman, G.; Jung, M.K.; Kim, T.W.; Kwon, H.H. Drought Impact, Vulnerability, Risk Assessment, Management and Mitigation under Climate Change: A Comprehensive Review. KSCE J. Civ. Eng. 2025, 29, 100120. [Google Scholar] [CrossRef]
  2. Giri, S.; Mishra, A.; Zhang, Z.; Lathrop, R.G.; Alnahit, A.O. Meteorological and Hydrological Drought Analysis and Its Impact on Water Quality and Stream Integrity. Sustainability 2021, 13, 8175. [Google Scholar] [CrossRef]
  3. Barker, L.J.; Hannaford, J.; Chiverton, A.; Svensson, C. From Meteorological to Hydrological Drought Using Standardised Indicators. Hydrol. Earth Syst. Sci. 2016, 20, 2483–2505. [Google Scholar] [CrossRef]
  4. Schwabe, K.; Albiac, J.; Connor, J.D.; Hassan, R.M.; Meza González, L. (Eds.) Drought in Arid and Semi-Arid Regions; Springer: Dordrecht, The Netherlands, 2013; ISBN 978-94-007-6635-8. [Google Scholar]
  5. Abraha, T.; Tibebu, A.; Ephrem, G. Rapid Urbanization and the Growing Water Risk Challenges in Ethiopia: The Need for Water Sensitive Thinking. Front. Water 2022, 4, 890229. [Google Scholar] [CrossRef]
  6. Mmbando, G.S. Recent Possible Agriculture Techniques for Supporting Climate Resilient Crops in Africa. Cogent Food Agric. 2025, 11, 2557341. [Google Scholar] [CrossRef]
  7. Hao, Z.; Yuan, X.; Xia, Y.; Hao, F.; Singh, V.P. An Overview of Drought Monitoring and Prediction Systems at Regional and Global Scales. Bull. Am. Meteorol. Soc. 2017, 98, 1879–1896. [Google Scholar] [CrossRef]
  8. Raible, C.C.; Bärenbold, O.; Gómez-navarro, J.J. Drought Indices Revisited–Improving and Testing of Drought Indices in a Simulation of the Last Two Millennia for Europe. Tellus Ser. A Dyn. Meteorol. Oceanogr. 2017, 69, 1287492. [Google Scholar] [CrossRef]
  9. Gholinia, A.; Abbaszadeh, P. Agricultural Drought Monitoring: A Comparative Review of Conventional and Satellite-Based Indices. Atmosphere 2024, 15, 1129. [Google Scholar] [CrossRef]
  10. Yang, Y.; Zhang, S.; Roderick, M.L.; McVicar, T.R.; Yang, D.; Liu, W.; Li, X. Comparing Palmer Drought Severity Index Drought Assessments Using the Traditional Offline Approach with Direct Climate Model Outputs. Hydrol. Earth Syst. Sci. 2020, 24, 2921–2930. [Google Scholar] [CrossRef]
  11. Jain, V.K.; Pandey, R.P.; Jain, M.K.; Byun, H.R. Comparison of Drought Indices for Appraisal of Drought Characteristics in the Ken River Basin. Weather Clim. Extrem. 2015, 8, 1–11. [Google Scholar] [CrossRef]
  12. Zhang, B.; Abu Salem, F.K.; Hayes, M.J.; Smith, K.H.; Tadesse, T.; Wardlow, B.D. Explainable Machine Learning for the Prediction and Assessment of Complex Drought Impacts. Sci. Total Environ. 2023, 898, 165509. [Google Scholar] [CrossRef] [PubMed]
  13. En-Nagre, K.; Aqnouy, M.; Ouarka, A.; Ali Asad Naqvi, S.; Bouizrou, I.; Eddine Stitou El Messari, J.; Tariq, A.; Soufan, W.; Li, W.; El-Askary, H. Assessment and Prediction of Meteorological Drought Using Machine Learning Algorithms and Climate Data. Clim. Risk Manag. 2024, 45, 100630. [Google Scholar] [CrossRef]
  14. Mumtaz, I.; Niaz, R.; Sajid, Z.; Alameri, A.Q.; Ali, Z.; Gepreel, K.A. Utilising Machine Learning Classification Models for Meteorological Drought Monitoring and Analysis. All Earth 2025, 37, 1–21. [Google Scholar] [CrossRef]
  15. Zhao, Y.; Zhang, J.; Bai, Y.; Zhang, S.; Yang, S.; Henchiri, M.; Seka, A.M.; Nanzad, L. Drought Monitoring and Performance Evaluation Based on Machine Learning Fusion of Multi-Source Remote Sensing Drought Factors. Remote Sens. 2022, 14, 6398. [Google Scholar] [CrossRef]
  16. Lalika, C.; Mujahid, A.U.H.; James, M.; Lalika, M.C.S. Machine Learning Algorithms for the Prediction of Drought Conditions in the Wami River Sub-Catchment, Tanzania. J. Hydrol. Reg. Stud. 2024, 53, 101794. [Google Scholar] [CrossRef]
  17. Nikdad, P.; Mohammadi Ghaleni, M.; Moghaddasi, M.; Pradhan, B. Enhancing a Machine Learning Model for Predicting Agricultural Drought through Feature Selection Techniques. Appl. Water Sci. 2024, 14, 125. [Google Scholar] [CrossRef]
  18. Tamrakar, Y.; Das, I.C.; Sharma, S. Machine Learning for Improved Drought Forecasting in Chhattisgarh India: A Statistical Evaluation. Discov. Geosci. 2024, 2, 84. [Google Scholar] [CrossRef]
  19. Obsie, E.Y.; Liu, Y. Enhancing Drought Prediction through Machine Learning: Advanced Techniques Combining Phenotypic and Agrometeorological Data. Smart Agric. Technol. 2025, 12, 101227. [Google Scholar] [CrossRef]
  20. Melese, T.; Assefa, G.; Terefe, B.; Belay, T.; Bayable, G.; Senamew, A. Machine Learning-Based Drought Prediction Using Palmer Drought Severity Index and TerraClimate Data in Ethiopia. PLoS ONE 2025, 20, e0326174. [Google Scholar] [CrossRef] [PubMed]
  21. Achite, M.; Jehanzaib, M.; Elshaboury, N.; Kim, T.W. Evaluation of Machine Learning Techniques for Hydrological Drought Modeling: A Case Study of the Wadi Ouahrane Basin in Algeria. Water 2022, 14, 431. [Google Scholar] [CrossRef]
  22. Zhang, R.; Chen, Z.Y.; Xu, L.J.; Ou, C.Q. Meteorological Drought Forecasting Based on a Statistical Model with Machine Learning Techniques in Shaanxi Province, China. Sci. Total Environ. 2019, 665, 338–346. [Google Scholar] [CrossRef] [PubMed]
  23. Pande, C.B.; Vishwakarma, D.K.; Srivastava, A.; Moharir, K.N.; Alshehri, F.; Din, N.M.; Sidek, L.M.; Đurin, B.; Tolche, A.D. A Novel Machine Learning Models for Meteorological Drought Forecasting in the Semi-Arid Climate Region. Appl. Water Sci. 2025, 15, 121. [Google Scholar] [CrossRef]
  24. Terzi, T.B. PyDRGHT: A Comprehensive Python Package for Drought Analysis. Environ. Model. Softw. 2026, 197, 106847. [Google Scholar] [CrossRef]
  25. McKee, T.B.; Doesken, N.J.; Kleist, J. The relationship of drought frequency and duration to time scales. In Proceedings of the 8th Conference on Applied Climatology, Memphis, TN, USA, 19–23 May 1993; Volume 17, pp. 179–183. [Google Scholar]
  26. Wang, H.; Chen, Y.; Pan, Y.; Li, W. Spatial and Temporal Variability of Drought in the Arid Region of China and Its Relationships to Teleconnection Indices. J. Hydrol. 2015, 523, 283–296. [Google Scholar] [CrossRef]
  27. Shukla, S.; Wood, A.W. Use of a Standardized Runoff Index for Characterizing Hydrologic Drought. Geophys. Res. Lett. 2008, 35, L02405. [Google Scholar] [CrossRef]
  28. O’brien, R.M. A Caution Regarding Rules of Thumb for Variance Inflation Factors. Qual. Quant. 2007, 41, 673–690. [Google Scholar] [CrossRef]
  29. Van Loon, A.F. Hydrological Drought Explained. WIREs Water 2015, 2, 359–392. [Google Scholar] [CrossRef]
  30. Hurrell, J.W. Decadal Trends in the North Atlantic Oscillation: Regional Temperatures and Precipitation. Science 1995, 269, 676–679. [Google Scholar] [CrossRef] [PubMed]
  31. Kingston, D.G.; Lawler, D.M.; McGregor, G.R. Linkages between Atmospheric Circulation, Climate and Streamflow in the Northern North Atlantic: Research Prospects. Prog. Phys. Geogr. 2006, 30, 143–174. [Google Scholar] [CrossRef]
  32. Stahl, K.; Kohn, I.; Blauhut, V.; Urquijo, J.; De Stefano, L.; Acácio, V.; Dias, S.; Stagge, J.H.; Tallaksen, L.M.; Kampragou, E.; et al. Impacts of European Drought Events: Insights from an International Database of Text-Based Reports. Nat. Hazards Earth Syst. Sci. 2016, 16, 801–819. [Google Scholar] [CrossRef]
  33. Van Loon, A.F.; Kchouk, S.; Matanó, A.; Tootoonchi, F.; Alvarez-Garreton, C.; Hassaballah, K.E.A.; Wu, M.; Wens, M.L.K.; Shyrokaya, A.; Ridolfi, E.; et al. Review Article: Drought as a Continuum—Memory Effects in Interlinked Hydrological, Ecological, and Social Systems. Nat. Hazards Earth Syst. Sci. 2024, 24, 3173–3205. [Google Scholar]
  34. Friedman, J.H. Greedy Function Approximation: A Gradient Boosting Machine. Ann. Stat. 2001, 29, 1189–1232. [Google Scholar] [CrossRef]
  35. Liao, S.; Liu, Z.; Liu, B.; Cheng, C.; Jin, X.; Zhao, Z. Multistep-ahead daily inflow forecasting using the ERA-Interim reanalysis data set based on gradient-boosting regression trees. Hydrol. Earth Syst. Sci. 2020, 24, 2343–2363. [Google Scholar] [CrossRef]
  36. Ke, G.; Meng, Q.; Finley, T.; Wang, T.; Chen, W.; Ma, W.; Ye, Q.; Liu, T.Y. LightGBM: A Highly Efficient Gradient Boosting Decision Tree. In Proceedings of the Advances in Neural Information Processing Systems; NeurIPS: Sydney, Australia, 2017; Volume 2017. [Google Scholar]
  37. Oreshkin, B.N.; Carpov, D.; Chapados, N.; Bengio, Y. N-Beats: Neural Basis Expansion Analysis For Interpretable Time Series Forecasting. In Proceedings of the 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, 26–30 April 2020. [Google Scholar]
  38. Challu, C.; Olivares, K.G.; Oreshkin, B.N.; Ramirez, F.G.; Mergenthaler-Canseco, M.; Dubrawski, A. NHITS: Neural Hierarchical Interpolation for Time Series Forecasting. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, AAAI 2023, Washington, DC, USA, 7–14 February 2023; Volume 37. [Google Scholar]
  39. Das, A.; Kong, W.; Leach, A.; Mathur, S.; Sen, R.; Yu, R. Long-Term Forecasting with TiDE: Time-Series Dense Encoder. arXiv 2023, arXiv:2304.08424. [Google Scholar] [CrossRef]
  40. Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining; ACM: New York, NY, USA, 2016; pp. 1135–1144. [Google Scholar]
  41. Gupta, H.V.; Kling, H.; Yilmaz, K.K.; Martinez, G.F. Decomposition of the Mean Squared Error and NSE Performance Criteria: Implications for Improving Hydrological Modelling. J. Hydrol. 2009, 377, 80–91. [Google Scholar] [CrossRef]
  42. Nash, J.E.; Sutcliffe, J.V. River Flow Forecasting through Conceptual Models Part I—A Discussion of Principles. J. Hydrol. 1970, 10, 282–290. [Google Scholar] [CrossRef]
  43. Onyutha, C. A Hydrological Model Skill Score and Revised R-Squared. Hydrol. Res. 2022, 53, 51–64. [Google Scholar] [CrossRef]
  44. Chen, T.; Guestrin, C. XGBoost: A Scalable Tree Boosting System. In Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016. [Google Scholar]
  45. Mosavi, A.; Ozturk, P.; Chau, K.W. Flood Prediction Using Machine Learning Models: Literature Review. Water 2018, 10, 1536. [Google Scholar] [CrossRef]
  46. Basak, A.; Rahman, A.T.M.S.; Das, J.; Hosono, T.; Kisi, O. Drought Forecasting Using the Prophet Model in a Semi-Arid Climate Region of Western India. Hydrol. Sci. J. 2022, 67, 1397–1417. [Google Scholar] [CrossRef]
  47. Lu, Y.; Li, T.; Hu, H.; Zeng, X. Short-Term Prediction of Reference Crop Evapotranspiration Based on Machine Learning with Different Decomposition Methods in Arid Areas of China. Agric. Water Manag. 2023, 279, 108175. [Google Scholar] [CrossRef]
  48. Lim, B.; Zohren, S. Time-Series Forecasting with Deep Learning: A Survey. Philos. Trans. R. Soc. A Math. Phys. Eng. Sci. 2021, 379, 20200209. [Google Scholar] [CrossRef]
  49. Hundecha, Y.; Bárdossy, A. Modeling of the Effect of Land Use Changes on the Runoff Generation of a River Basin through Parameter Regionalization of a Watershed Model. J. Hydrol. 2004, 292, 281–295. [Google Scholar] [CrossRef]
Figure 1. Location of the Wadi Sahaouat study basin within the Macta watershed, Algeria, showing the positions of the two pluviometric stations (S1, S2) and two hydrometric stations (H1, H2) used in this study.
Figure 1. Location of the Wadi Sahaouat study basin within the Macta watershed, Algeria, showing the positions of the two pluviometric stations (S1, S2) and two hydrometric stations (H1, H2) used in this study.
Atmosphere 17 00670 g001
Figure 2. Flowchart of the present study.
Figure 2. Flowchart of the present study.
Atmosphere 17 00670 g002
Figure 3. Observed versus predicted SRI scatter plots for Basin 1 under Scenario 1 (top row) and Scenario 2 (bottom row). Results are shown for GBRT, A-N-BEATS, A-N-HiTS, and TiDE during the test period. The dashed line represents the 1:1 agreement line, while NSE and MAE values summarize predictive performance.
Figure 3. Observed versus predicted SRI scatter plots for Basin 1 under Scenario 1 (top row) and Scenario 2 (bottom row). Results are shown for GBRT, A-N-BEATS, A-N-HiTS, and TiDE during the test period. The dashed line represents the 1:1 agreement line, while NSE and MAE values summarize predictive performance.
Atmosphere 17 00670 g003
Figure 4. Comparison of observed and predicted SRI time series for Basin 1 during the test period under (a) Scenario 1 and (b) Scenario 2. Predictions generated by GBRT, A-N-BEATS, A-N-HiTS, and TiDE are shown alongside the observed SRI series to evaluate the ability of each model to reproduce temporal hydrological drought dynamics.
Figure 4. Comparison of observed and predicted SRI time series for Basin 1 during the test period under (a) Scenario 1 and (b) Scenario 2. Predictions generated by GBRT, A-N-BEATS, A-N-HiTS, and TiDE are shown alongside the observed SRI series to evaluate the ability of each model to reproduce temporal hydrological drought dynamics.
Atmosphere 17 00670 g004
Figure 5. Residual boxplots for Basin 1 under (a) Scenario 1 and (b) Scenario 2. The distributions summarize prediction errors for GBRT, A-N-BEATS, A-N-HiTS, and TiDE. The dashed horizontal line represents zero residual, while the spread of each boxplot reflects model error variability. Circles indicate outlier observations.
Figure 5. Residual boxplots for Basin 1 under (a) Scenario 1 and (b) Scenario 2. The distributions summarize prediction errors for GBRT, A-N-BEATS, A-N-HiTS, and TiDE. The dashed horizontal line represents zero residual, while the spread of each boxplot reflects model error variability. Circles indicate outlier observations.
Atmosphere 17 00670 g005
Figure 6. Comparison of model performance metrics (RMSE, NSE, and KGE) for Basin 1 under Scenarios 1 and 2. Bars represent the performance of GBRT, A-N-BEATS, A-N-HiTS, and TiDE models.
Figure 6. Comparison of model performance metrics (RMSE, NSE, and KGE) for Basin 1 under Scenarios 1 and 2. Bars represent the performance of GBRT, A-N-BEATS, A-N-HiTS, and TiDE models.
Atmosphere 17 00670 g006
Figure 7. Observed versus predicted SRI scatter plots for Basin 2 under Scenario 1 and Scenario 2. The dashed line indicates perfect agreement between observations and predictions.
Figure 7. Observed versus predicted SRI scatter plots for Basin 2 under Scenario 1 and Scenario 2. The dashed line indicates perfect agreement between observations and predictions.
Atmosphere 17 00670 g007
Figure 8. Observed and predicted SRI values during the test period for Basin 2 under (a) Scenario 1 and (b) Scenario 2. Predictions from GBRT, A-N-BEATS, A-N-HiTS, and TiDE are compared against observed SRI values to evaluate the ability of each model to reproduce temporal drought dynamics.
Figure 8. Observed and predicted SRI values during the test period for Basin 2 under (a) Scenario 1 and (b) Scenario 2. Predictions from GBRT, A-N-BEATS, A-N-HiTS, and TiDE are compared against observed SRI values to evaluate the ability of each model to reproduce temporal drought dynamics.
Atmosphere 17 00670 g008
Figure 9. Distribution of residuals for the four forecasting models in Basin 2 under Scenario 1 and Scenario 2. The dashed line indicates zero residual.
Figure 9. Distribution of residuals for the four forecasting models in Basin 2 under Scenario 1 and Scenario 2. The dashed line indicates zero residual.
Atmosphere 17 00670 g009
Figure 10. Comparison of forecasting performance metrics (RMSE, NSE, and KGE,) for the four models in Basin 2 under Scenario 1 and Scenario 2. Lower RMSE values indicate better performance, whereas higher NSE and KGE values indicate superior predictive skill.
Figure 10. Comparison of forecasting performance metrics (RMSE, NSE, and KGE,) for the four models in Basin 2 under Scenario 1 and Scenario 2. Lower RMSE values indicate better performance, whereas higher NSE and KGE values indicate superior predictive skill.
Atmosphere 17 00670 g010
Table 1. Characteristics of the meteorological (S1–S2) and hydrological (H1–H2) stations used in the study. The table distinguishes the available station archive from the common monthly period used for modelling.
Table 1. Characteristics of the meteorological (S1–S2) and hydrological (H1–H2) stations used in the study. The table distinguishes the available station archive from the common monthly period used for modelling.
No.Station NameStation IDSub-BasinLongitude (°)Latitude (°)Altitude (m)Available Station ArchiveCommon Modelling Period
S1Oued Taria111201Taria0°5′42″ E35°6′ N4981973/74–2014/15January 1979–August 2015, n = 440
S2Sidi Boubekeur111102Saida0°3′38″ E35°1′ N525
H1Oued Taria111201Taria0°07′ E35°10′ N501
H2Sidi Boubekeur111129Saida0°02′ E35°05′ N540
Table 2. Summary of the forecasting scenarios and corresponding predictor sets used for one-month-ahead SRI prediction. Scenario 2 includes lagged SRI variables in addition to the meteorological, teleconnection, and seasonal predictors considered in Scenario 1.
Table 2. Summary of the forecasting scenarios and corresponding predictor sets used for one-month-ahead SRI prediction. Scenario 2 includes lagged SRI variables in addition to the meteorological, teleconnection, and seasonal predictors considered in Scenario 1.
ScenarioPurposePredictorsTarget
Scenario 1Assess whether monthly meteorological drought and atmospheric forcing alone can explain hydrological drought.NAO, AO, EAWR, SCAND, MEI, SOI, WeMO; lag1–lag3 for all teleconnections; SPI; SPI_lag1–lag3; month_sin; month_cosSRI
Scenario 2Assess whether adding lagged SRI values (hydrological memory) improves forecasting beyond Scenario 1.All Scenario 1 predictors + SRI_lag1, SRI_lag2, SRI_lag3SRI
Table 3. Test-set performance metrics and ranking of the forecasting models in Basin 1 under both experimental scenarios.
Table 3. Test-set performance metrics and ranking of the forecasting models in Basin 1 under both experimental scenarios.
Scenario Model RMSE MAE NSE KGE Features Rank
Scenario 1GBRT0.03770.02880.99780.8588321
Scenario 1A-N-HiTS0.09790.07800.98490.4470322
Scenario 1TiDE0.12400.09890.97580.3695323
Scenario 1A-N-BEATS0.15310.12250.96310.4801324
Scenario 2GBRT0.03770.02880.99780.8588321
Scenario 2A-N-HiTS0.10270.07570.98340.7948322
Scenario 2TiDE0.11160.08640.98040.6976323
Scenario 2A-N-BEATS0.17130.14040.9538−0.0338324
Table 4. Test-set forecasting performance of the four models in Basin 2 under Scenario 1 and Scenario 2.
Table 4. Test-set forecasting performance of the four models in Basin 2 under Scenario 1 and Scenario 2.
ScenarioModelRMSEMAENSEKGEFeaturesRank
Scenario 1GBRT0.09850.06370.98370.9286321
Scenario 1TiDE0.12580.10300.97340.9372322
Scenario 1A-N-HiTS0.13790.11520.96800.8980323
Scenario 1A-N-BEATS0.29180.23620.85680.8205324
Scenario 2GBRT0.09880.06280.98360.9318351
Scenario 2TiDE0.10490.08810.98150.9878352
Scenario 2A-N-HiTS0.14280.11450.96570.9009353
Scenario 2A-N-BEATS0.24730.20290.89720.9453354
Table 6. Overall performance ranking of GBRT, A-N-BEATS, A-N-HiTS, and TiDE based on average evaluation metrics computed across Basin 1, Basin 2, and both forecasting scenarios.
Table 6. Overall performance ranking of GBRT, A-N-BEATS, A-N-HiTS, and TiDE based on average evaluation metrics computed across Basin 1, Basin 2, and both forecasting scenarios.
RankModelMean RMSEMean MAEMean NSEMean KGE
1GBRT0.06820.04600.99070.8945
2TiDE0.11660.09410.97780.7480
3A-N-HiTS0.12030.09590.97550.7602
4A-N-BEATS0.21590.17550.91770.5530
Table 7. Exploratory pairwise comparison of forecasting model errors in Basin 1 using the Wilcoxon signed-rank test and paired t-test. The reported p-values are unadjusted and should be interpreted cautiously because monthly residuals may exhibit serial dependence.
Table 7. Exploratory pairwise comparison of forecasting model errors in Basin 1 using the Wilcoxon signed-rank test and paired t-test. The reported p-values are unadjusted and should be interpreted cautiously because monthly residuals may exhibit serial dependence.
ScenarioModel AModel BW StatW p-Valuet Statt p-ValueSig. (W)Sig. (t)
Sc. 1GBRTA-N-BEATS1458.48 × 10−10−7.7548.00 × 10−11YesYes
Sc. 1GBRTA-N-HiTS2443.73 × 10−8−6.0328.52 × 10−8YesYes
Sc. 1GBRTTiDE1541.21 × 10−9−7.4023.37 × 10−10YesYes
Sc. 1A-N-BEATSA-N-HiTS4991.07 × 10−44.0961.19 × 10−4YesYes
Sc. 1A-N-BEATSTiDE8510.1042.1360.036NoYes
Sc. 1A-N-HiTSTiDE7370.019−2.2150.030YesYes
Sc. 2GBRTA-N-BEATS744.42 × 10−11−9.1233.01 × 10−13YesYes
Sc. 2GBRTA-N-HiTS3491.35 × 10−6−5.1842.31 × 10−6YesYes
Sc. 2GBRTTiDE2474.15 × 10−8−6.3192.72 × 10−8YesYes
Sc. 2A-N-BEATSA-N-HiTS3742.97 × 10−64.9505.55 × 10−6YesYes
Sc. 2A-N-BEATSTiDE4644.17 × 10−54.8687.54 × 10−6YesYes
Sc. 2A-N-HiTSTiDE9050.200−1.1230.266NoNo
Note: Significance at alpha = 0.05 is denoted “Yes”.
Table 8. Exploratory pairwise comparison of forecasting model errors in Basin 2 using the Wilcoxon signed-rank test and paired t-test. The reported p-values are unadjusted and are used only as supplementary diagnostics.
Table 8. Exploratory pairwise comparison of forecasting model errors in Basin 2 using the Wilcoxon signed-rank test and paired t-test. The reported p-values are unadjusted and are used only as supplementary diagnostics.
ScenarioModel AModel BW StatW p-Valuet Statt p-ValueSig. (W)Sig. (t)
Sc. 1GBRTA-N-BEATS1417.21 × 10−10−7.7986.66 × 10−11YesYes
Sc. 1GBRTA-N-HiTS4361.90 × 10−5−3.8652.59 × 10−4YesYes
Sc. 1GBRTTiDE5201.84 × 10−4−3.2291.95 × 10−3YesYes
Sc. 1A-N-BEATSA-N-HiTS4231.30 × 10−55.3021.47 × 10−6YesYes
Sc. 1A-N-BEATSTiDE2719.78 × 10−86.1126.20 × 10−8YesYes
Sc. 1A-N-HiTSTiDE9370.2820.9160.363NoNo
Sc. 2GBRTA-N-BEATS1894.78 × 10−9−6.7983.93 × 10−9YesYes
Sc. 2GBRTA-N-HiTS3925.17 × 10−6−4.1489.92 × 10−5YesYes
Sc. 2GBRTTiDE6725.62 × 10−3−2.2550.028YesYes
Sc. 2A-N-BEATSA-N-HiTS4971.01 × 10−44.3544.84 × 10−5YesYes
Sc. 2A-N-BEATSTiDE2616.86 × 10−86.4541.58 × 10−8YesYes
Sc. 2A-N-HiTSTiDE7830.0392.2440.028YesYes
Table 9. Optimal hyperparameters selected using the external chronological validation subset. For GBRT, the candidate configuration was selected according to external validation RMSE, while scikit-learn’s internal early-stopping mechanism used only a 10% subset of the training block. For the neural models, early stopping was based on the external validation loss.
Table 9. Optimal hyperparameters selected using the external chronological validation subset. For GBRT, the candidate configuration was selected according to external validation RMSE, while scikit-learn’s internal early-stopping mechanism used only a 10% subset of the training block. For the neural models, early stopping was based on the external validation loss.
BasinScenarioModelOptimal Hyperparameters
Basin 1 Scenario 1GBRTn_estimators = 80; learning_rate = 0.05; max_depth = 4; subsample = 0.8
Scenario 1A-N-BEATShidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 1A-N-HiTShidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 1TiDEhidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 2GBRTn_estimators = 80; learning_rate = 0.05; max_depth = 4; subsample = 0.8
Scenario 2A-N-BEATShidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 2A-N-HiTShidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 2TiDEhidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Basin 2Scenario 1GBRTn_estimators = 80; learning_rate = 0.05; max_depth = 4; subsample = 0.8
Scenario 1A-N-BEATShidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 1A-N-HiTShidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 1TiDEhidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 2GBRTn_estimators = 80; learning_rate = 0.05; max_depth = 4; subsample = 0.8
Scenario 2A-N-BEATShidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 2A-N-HiTShidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Scenario 2TiDEhidden = 64; lr = 0.01; epochs = 150; patience = 20; dropout = 0.05; batch = 32
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Katipoğlu, O.M.; Achite, M.; Kartal, V.; Çelik, M.A.; Pandey, K. Machine Learning-Based Hydrological Drought Prediction Integrating Teleconnections and Hydrological Memory in a Semi-Arid Basin, Algeria. Atmosphere 2026, 17, 670. https://doi.org/10.3390/atmos17070670

AMA Style

Katipoğlu OM, Achite M, Kartal V, Çelik MA, Pandey K. Machine Learning-Based Hydrological Drought Prediction Integrating Teleconnections and Hydrological Memory in a Semi-Arid Basin, Algeria. Atmosphere. 2026; 17(7):670. https://doi.org/10.3390/atmos17070670

Chicago/Turabian Style

Katipoğlu, Okan Mert, Mohammed Achite, Veysi Kartal, Mehmet Ali Çelik, and Kusum Pandey. 2026. "Machine Learning-Based Hydrological Drought Prediction Integrating Teleconnections and Hydrological Memory in a Semi-Arid Basin, Algeria" Atmosphere 17, no. 7: 670. https://doi.org/10.3390/atmos17070670

APA Style

Katipoğlu, O. M., Achite, M., Kartal, V., Çelik, M. A., & Pandey, K. (2026). Machine Learning-Based Hydrological Drought Prediction Integrating Teleconnections and Hydrological Memory in a Semi-Arid Basin, Algeria. Atmosphere, 17(7), 670. https://doi.org/10.3390/atmos17070670

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop