Previous Article in Journal
Neural Signed Distance Surrogates for Cut-Cell Finite-Volume Solvers
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Robust Short-Term Multivariate Water-Level Forecasting Using a Hybrid LSTM–EnKF Model Under White-Noise Disturbances

by
Jackson B. Renteria-Mena
1 and
Eduardo Giraldo
2,*
1
Facultad de Ingeniería, Universidad Tecnológica del Chocó, Quibdo 270001, Colombia
2
Electrical Engineering Department, Research Group in Automatic Control, Universidad Tecnológica de Pereira, Pereira 660003, Colombia
*
Author to whom correspondence should be addressed.
Computation 2026, 14(9), 199; https://doi.org/10.3390/computation14090199
Submission received: 21 July 2026 / Revised: 25 August 2026 / Accepted: 26 August 2026 / Published: 28 August 2026
(This article belongs to the Section Computational Intelligence)

Abstract

Short-term multivariate forecasting of hydrological variables remains challenging because river systems exhibit nonlinear and time-dependent dynamics, complex relationships among water level, flow, and precipitation, and uncertainty arising from measurement errors and external disturbances. Although neural network models can learn nonlinear relationships among hydrological variables, their predictive performance often deteriorates in the presence of noise. Moreover, existing approaches rarely integrate the learning of long-term temporal dependencies and cross-variable relationships with a data assimilation mechanism capable of recursively updating state estimates and reducing forecast uncertainty. This limitation reveals the need for a robust forecasting framework that combines both capabilities. Therefore, this study aimed to develop and evaluate a hybrid Long Short-Term Memory–Ensemble Kalman Filter (LSTM–EnKF) model for short-term multivariate water-level forecasting under noisy conditions. The proposed framework extends a previously developed NARX–EnKF approach by replacing the NARX network with an LSTM architecture capable of learning nonlinear temporal patterns and relationships among water level, flow, and precipitation. The model was implemented using data from two hydrological stations located along the Atrato River in Colombia and configured to generate water-level forecasts with a two-day prediction horizon. The LSTM network generated the initial forecasts, whereas the EnKF assimilated the available observations to recursively update the estimated states and reduce forecast uncertainty. Model robustness was examined by introducing Gaussian white noise with variance levels of 0.001, 0.05, 0.10, and 0.20 to represent measurement uncertainty and external disturbances. Performance was evaluated using the root mean square error (RMSE), mean absolute error (MAE), and Nash–Sutcliffe efficiency (NSE). Across all evaluated noise levels, the LSTM–EnKF model outperformed the standalone LSTM model. Its total RMSE ranged from 0.1009 to 0.1929 m, compared with 0.1740 to 0.2356 m for the standalone LSTM, representing reductions of approximately 18.1–42%. The hybrid model achieved NSE values ranging from 0.9760 to 0.9992, whereas the standalone LSTM produced values between 0.9200 and 0.9776. Furthermore, the LSTM–EnKF reduced the MAE by approximately 51.9–56.2% across both outputs. These results indicate that integrating LSTM-based temporal learning with EnKF-based data assimilation improves short-term forecasting accuracy and robustness under noisy conditions. The developed framework provides a promising tool for supporting flood early-warning systems, flood-risk management, and the protection of riverine communities.

1. Introduction

Accurate modeling and forecasting of hydrological variables are essential for reducing flood risk and for the effective planning, operation, and management of water-resources systems [1]. Nevertheless, short-term water-level forecasting remains challenging because river systems exhibit nonlinear and time-dependent dynamics, complex interactions among hydrological variables, spatial dependencies between monitoring stations, and uncertainty arising from measurement errors and external disturbances. These factors can reduce forecasting accuracy, particularly when models are applied under conditions that differ from those represented in the training data.
Long Short-Term Memory (LSTM) networks provide a suitable approach for modeling these temporal dynamics. Their gated architecture mitigates the vanishing-gradient problem and enables relevant information to be retained over long sequences. Consequently, LSTM networks have been applied to multivariate hydrological forecasting to learn temporal patterns and cross-variable relationships from observations collected at different hydrological stations [2,3]. Despite these advantages, a standalone LSTM generates predictions exclusively from patterns learned during training and does not inherently include a sequential mechanism for assimilating newly available observations. Its performance may therefore deteriorate when the input data are affected by measurement noise, unmodeled dynamics, or disturbances that are not adequately represented in the training dataset. Thus, the robustness of multivariate LSTM forecasting under controlled levels of measurement disturbance requires further investigation.
Data assimilation offers a potential means of addressing this limitation by combining model predictions with available observations. Different variants of the Kalman filter have been used for this purpose, including the Kalman Filter (KF) [4,5], the Unscented Kalman Filter (UKF) [6,7], and the Ensemble Kalman Filter (EnKF) [8,9]. The EnKF represents forecast uncertainty through an ensemble of possible system states and recursively updates this ensemble as observations become available [10]. This formulation makes the EnKF suitable for high-dimensional and nonlinear dynamic systems in meteorology, oceanography, engineering, and water-resources management [11]. Nevertheless, its performance depends on the quality of the forecasting model, the ensemble representation, and the specification of process and observation errors.
Previous studies have combined machine-learning methods with the EnKF or related data-assimilation strategies. Machine learning has been employed to approximate covariance matrices and reduce the computational cost of the EnKF [12]. Recurrent neural networks have also been combined with the EnKF for dynamic-system modeling [13] and data-driven oil-reservoir production forecasting [14]. Other applications include a matrix EnKF-based multi-arm neural-network architecture for approximating deep neural networks [15] and an LSTM-embedded nudging strategy for nonlinear data assimilation in geophysical flows [16]. These studies demonstrate the broader potential of integrating sequential learning with data assimilation; however, they do not directly address short-term multivariate water-level forecasting under different levels of observational disturbance.
More recent approaches have integrated the EnKF with LSTM-based architectures in other application domains. An EnKF–LSTM assimilation algorithm has been used to improve crop-growth modeling by incorporating observations into model-state estimation [17]. Similarly, the EnKF has been combined with Generative Adversarial Networks and Convolutional LSTM models for long-lead forecasting [18], while multi-source spatiotemporal information has been incorporated into related architectures to improve lightning-detection probability [19]. Although these methods provide relevant alternatives, differences in their objectives, data structures, forecasting horizons, and application domains limit their direct transfer to multivariate river-level forecasting.
In previous hydrological research, a hybrid Nonlinear AutoRegressive Neural Network with exogenous inputs and Ensemble Kalman Filter (NARX–EnKF) framework was developed for short-term water-level forecasting [20]. The NARX component represented nonlinear relationships among historical hydrological variables, whereas the EnKF assimilated available observations to update the estimated states. However, the dependence of NARX models on a predefined set of lagged inputs may constrain their ability to represent longer temporal dependencies. LSTM networks provide a potentially suitable alternative because their internal memory mechanisms can learn relevant dependencies directly from sequential data. However, it remains unclear whether coupling an LSTM network with the EnKF can consistently improve multivariate short-term water-level forecasts when observations are subjected to increasing levels of noise.
Accordingly, the research gap addressed in this study concerns the limited evaluation of hybrid LSTM–EnKF models for multivariate short-term water-level forecasting under explicitly controlled measurement-noise conditions. In particular, further evidence is required to determine whether the sequential assimilation of observations can reduce the sensitivity of LSTM forecasts to noise while preserving the temporal and cross-variable information contained in water-level, flow, and precipitation records.
Therefore, this study aims to develop and evaluate a hybrid LSTM–EnKF framework for two-day-ahead water-level forecasting at two hydrological stations along the Atrato River in Colombia. The analysis uses water-level, flow, and precipitation records provided by the Colombian Institute of Hydrology, Meteorology and Environmental Studies (IDEAM). The specific objectives are to: (i) construct an LSTM model capable of representing the temporal and cross-variable relationships within the multivariate hydrological records; (ii) integrate the LSTM forecasts with an EnKF-based sequential data assimilation procedure; and (iii) compare the coupled LSTM–EnKF framework with a standalone LSTM model under Gaussian white-noise variance levels of 0.001, 0.05, 0.10, and 0.20 using the root mean square error (RMSE), mean absolute error (MAE), and Nash–Sutcliffe efficiency (NSE).
The study addresses the following research questions: (1) Does the assimilation of observations through the EnKF improve the predictive performance of an LSTM model for two-day-ahead water-level forecasting? (2) How does the performance of the coupled and standalone models change as the variance of Gaussian white noise increases? The research hypothesis is that sequential EnKF assimilation reduces forecasting errors and maintains higher model efficiency than the standalone LSTM across the evaluated noise conditions. This hypothesis is examined through a comparative experimental analysis rather than assumed as an outcome.
The methodological contribution of this study is the adaptation and evaluation of an LSTM–EnKF architecture for multivariate hydrological forecasting under controlled noise conditions. In contrast to the previous NARX–EnKF formulation, the present framework uses an LSTM network to model temporal dependencies while retaining the EnKF as a mechanism for observation-based state updating. The contribution lies in evaluating this integration for multivariate water-level forecasting and quantifying its performance across several disturbance levels without assuming beforehand that the hybrid model is superior.
The remainder of this article is organized as follows. Section 2 describes the LSTM architecture, its integration with the EnKF, the study area, and the hydrological dataset, which comprises measurements sampled every 12 h over 789 days. Section 3 presents the forecasting results and the comparative evaluation of the coupled and standalone models. The subsequent sections discuss the findings, methodological uncertainties, principal sources of error, and conclusions.

2. Theoretical Framework

The Department of Chocó, situated on Colombia’s Pacific coast, boasts a unique blend of hydrology, geology, and climate dynamics, with a vast river network dominated by the Atrato River, the region enjoys ecological richness but also faces flood risks during the rainy season due to rapid flow and heavy precipitation. Geologically diverse, Chocó experiences seismic and volcanic influences, featuring mountains, valleys, and fertile plains. Its tropical climate nurtures exceptional biodiversity while presenting challenges for infrastructure and natural hazard management, such as floods and landslides. In [21], the intricate interplay of these natural elements underscores Chocó’s distinctiveness and the importance of holistic and sustainable approaches to its governance and development.
The Atrato River is the nation’s third most navigable waterway, trailing only the Magdalena and Cauca Rivers. Originating from the Cerro del Plateado in El Carmen de Atrato, within the Andean mountain range, it eventually converges into the Gulf of Urabá near the Caribbean Sea, marking the border with Panama. Serving as a vital transportation route in the Chocó region, the river also delineates segments of the Chocó and Antioquia departments. Renowned for its navigability, the Atrato River is crucial in local transportation networks. Moreover, it traverses the biogeographically significant Chocó region, renowned for its unparalleled biodiversity and abundant rainfall, contributing to its consistently high flow rate.
Stretching 750 km, the Atrato River boasts a width fluctuating between 150 and 500 m and depths ranging from 31 to 38 m. Its mouth, forming a delta with 18 channels, empties into the Gulf of Urabá. Throughout its course, the river is fed by approximately 150 tributaries and over 3000 streams, rendering it a vital repository of genetic diversity, as recognized by the World Wildlife Fund.
The meteorological dataset primarily comprises three key hydrological parameters: flow rates, water levels, and precipitation. These data are sourced from two hydrological stations, each providing comprehensive measurements of the aforementioned variables. The sampling period is 789 days, and the hydrological stations measure each variable every 12 h, resulting in approximately 1578 data samples for each measured variable and its corresponding hydrological station. Figure 1 shows the location of the hydrological stations and the exact place where the data collection of the study is performed in satellite form in the department of Chocó, Colombia, Atrato River.
In Table 1, the geographical positions of the hydrological stations are shown.
Figure 2 shows the prototype station considered for the case study of this article, which consist of an ultrasonic level sensor, a water flow sensor and a precipitation sensor.
Numerous strategically positioned measurement stations along the river enable tracking of variable dynamics in various locations. This inherent coupling among measurement variables facilitates the definition of a multivariate coupled model, allowing a comprehensive understanding of hydrological processes and interactions throughout the river system.
In addition, in Table 2 the information period of the data sets used and the specified hydrological variables of the system can be observed.
The water flow and precipitation in the Atrato River from the hydrological data of the two stations are used as input data, and the water level of the two hydrological stations is used as a prediction for model training. Thus, the system has a first-order structure with a training model structure of 4 inputs and 2 outputs.
A short-term LSTM network is a type of RNN designed to capture patterns in data sequences over immediate time periods. Unlike traditional RNNs, LSTMs feature memory units capable of retaining information over longer spans, but in this context, they focus on predicting events or behaviors in the short term without engaging in plagiarism practices [22].
The weights that can be learned from the LSTM layer are input weights W (InputWeights), recurrent weights R (RecurrentWeights) and bias b (Bias). The matrices W, R and b are the combination of input weights, cycle weights and offsets for each component, respectively. This layer connects matrices according to the following Equation (1):
W = W i W f W g W 0 , R = R i R f R g R 0 , b = b i b f b g b o
where i, f, g and o determine the entry gate, forgetting gate, cell candidate and exit gate, respectively.
The cell state at time unit t is given by (2)
C t = f t C t 1 + i t g t ,
where ⊙ determines the Hadamard product (element-level vector multiplication).
The hidden state at time unit t is given by (3)
h t = o t σ c + C t ,
where σ c determines the state activation function. By default, the lstmLayer function uses the hyperbolic tangent function (tanh) to calculate the state activation function.
The equations of the components in time unit t are described below. In (4), we describe the gateway equation entrance door, in (5) oblivion door, in (6) cell candidate and in (7) open door.
i t = σ g ( W i u t + R i h t 1 + b i )
f t = σ g ( W f u t + R f h t 1 + b f )
g t = σ c ( W g u t + R g h t 1 + b g )
o t = σ g ( W o u t + R o h t 1 + b o )

2.1. EnKF Forecasting

The state estimation x k can be obtained by applying the EnKF in two stages [23,24,25]. First, the forecast stage, where an ensemble of q forecasted state estimates is computed
x k f i = F ( x ¯ k a , u k ) + ω k i
being i = 1 , , q , ω k i R n × 1 is a zero mean random variable with normal distribution an covariance Ω k R n × n . The sample error covariance matrix computed from ω k i converges to Ω k as q . The forecasted measurement is given by
y k f i = f ( x k f i )
where f i denotes the i-th forecast ensemble member. Then, the ensemble mean for states x ¯ k f R n × 1 is defined by
x ¯ k f = 1 q i = 1 q x k f i
and the ensemble mean for measurements y ¯ k f is defined by
y ¯ k f = 1 q i = 1 q y k f i
The ensemble error matrix around the ensemble mean is defined as
E k f = x k f 1 x ¯ k f x k f q x ¯ k f
being E k f R n × q , and the ensemble output error
E y k f = y k f 1 y ¯ k f y k f q y ¯ k f
being E y k f R d × q , and where the forecast covariances are approximated as
P k f = 1 q 1 E k f E k f P x y k f = 1 q 1 E k f E y k f P y y k f = 1 q 1 E y k f E y k f
being P k f R n × n , P x y k f R n × d and P y y k f R d × d .
The second stage is the analysis stage, where an ensemble of perturbed observations y k i is obtained as follows
y k i = y k + υ k i
where υ k i R d × 1 is a zero mean random variable with normal distribution an covariance Υ k R d × d . The sample error covariance matrix computed from υ k i converges to Υ k as q . Also, the EnKF gain matrix is approximated as
K k = P x y k f P y y k f 1
where it is worth noting that P y y k f 1 is an inverse of d × d , being d < < n .
Therefore, an ensemble of data assimilation cycles is obtained as follows:
x k a i = x k f i + K k y k i f ( x k f i )
where a i denotes the i-th analysis ensemble member. And finally, the ensemble mean of the analysis stage is computed as
x ¯ k a = 1 q i = 1 q x k a i
where x ¯ k a is the estimated activity at sample k.
In order to validate the proposed approach, a comparative analysis of the LSTM model coupled and decoupled with a nonlinear recursive EnKF is developed. This evaluation is conducted using a state space model of the LSTM short-term recurrent neural network, aiming to validate the prediction capabilities of the nonlinear model for a system of fourth order. The dataset comprises data from two hydrological stations featuring four inputs and two outputs, and it is trained by using online training. We introduce four noise perturbations as variance parameters to assess the model’s robustness in level prediction. Visual comparisons between real and estimated signals for each noise parameter are presented for the LSTM model with the EnKF filter and the LSTM model without it. Additionally, a quantitative evaluation based on machine learning regression metrics is performed.
For a better description in Table 3 presents the parameters used for the recurrent LSTM neural network and the EnKF algorithm, including the number of input, hidden, and output nodes, as well as the learning rate and number of training epochs.
Table 3 summarizes the configuration and training parameters of the proposed LSTM–EnKF model. The LSTM uses a sequence length of two time steps, corresponding to the previous 24 h of observations. The network comprises four input nodes, one hidden layer with 40 nodes, and two output nodes. It was trained for 20 epochs using the Adam optimizer with a learning rate of 0.02. The EnKF uses an ensemble size of q = 80 .

2.2. Hydrological Variables

Using the hydrological records from the two stations described previously, the forecasting problem is formulated as a nonlinear multivariate dynamic system. This formulation accounts for the temporal dependencies and cross-variable relationships among water level, flow, and precipitation. The input and output variables considered in the forecasting model are defined in Equation (18).
y k = y L 1 [ k ] y L 2 [ k ] , u k = u F 1 [ k ] u P T 1 [ k ] u F 2 [ k ] u P T 2 [ k ]
where y L 1 [ k ] and y L 2 [ k ] correspond to the 2 outputs of the level variable of the two stations, u F 1 [ k ] , u F 2 [ k ] , u P T 1 [ k ] , and u P T 2 [ k ] are the 4 inputs of the neural network system corresponding to the two stations of the multivariable system, i.e., the u F j [ k ] represents the j-th 2 inputs of the flow variable of the two stations, and u P T j [ k ] represents the j-th two more inputs of the rainfall variable of the two stations already mentioned, thus obtaining a multivariable system with 4 inputs and two outputs.

2.3. Regression Metrics for the Estimation of the Quadratic Error

In this study, the evaluation of predictive performance is essential to assess the quality of the model and to optimize its efficiency. Key indicators, such as Root Mean Square Error (RMSE), Mean Absolute Error (MAE) and Nash-Sutcliffe Model Efficiency (NSE) coefficient, are used to evaluate the prediction performance of the LSTM recurrent neural network model. The model includes noise in both input and output data for a configuration of 4 inputs and 2 outputs, designed for a fourth-order system. In addition, the LSTM model is combined with a nonlinear Ensemble Kalman filter (EnKF) to improve the estimation of the LSTM model by four noise variance parameters ( 0.001 , 0.05 , 0.1 , 0.2 ) .
R M S E = 1 n Σ i = 1 n ( y 0 i y i ) 2
M A E = 1 n Σ i = 1 n [ y 0 i y i ]
N S E = 1 Σ i = 1 n ( y 0 i y i ) 2 Σ i = 1 n ( y 0 i y 0 ¯ ) 2

3. Results

Estimation Results

In this subsection are presented the forecasting results of the two methods considered: the multivariate LSTM approach coupled to an EnKF filter and the multivariate LSTM, where we obtain the responses of the real signal with noise in the images described as Y-real noise, the signal estimated with the LSTM-EnKF model, where it is described as Y-Estim EnKF, and the signal without noise.
Figure 3 and Figure 4 present the estimation results of the LSTM method coupled with the multivariate nonlinear recursive EnKF filter for two water level outputs. It uses 40 nodes in the LSTM hidden layer with white noise variance parameters of 0.01, 0.1, and 0.2, according to the results of the referenced tables. Figure 3 and Figure 4 the subfigures (a) and (b) show the estimation of the first and second water level outputs of hydrological stations, respectively. It is worth noting that Figure 3 and Figure 4 shows a comparison by using a fourth-order system of the measured real signal with noise, the estimated signal, and the real signal without noise.
In Figure 5, Figure 6 and Figure 7 are shown the results of the estimation carried out using a multivariate LSTM neural network model, excluding the recursive nonlinear EnKF filter, for two water level outputs. The model comprises 40 nodes in the hidden layer and utilizes noise variance parameters set at 0.01, 0.1, and 0.2. In Figure 5, Figure 6 and Figure 7 the subfigures (a) and (b) are shown the estimations of the first and second water level outputs from hydrological station, respectively, considering a fourth-order system. The comparison includes the real signal with noise, the estimated signal, and the real signal without noise.
In addition, Table 4 shows the RMSE estimates of the LSTM model without the EnKF filter and with the EnKF filter for the two outputs corresponding to the level of the hydrological variables for the two stations already described. It is worth mentioning that the EnKF implements the Bayesian Monte Carlo update problem: it updates the probability density function of the modeled system state.
In Table 4, a metric regression analysis of the data for the root mean square error (RMSE) estimates of an LSTM neural network with noise coupled to an EnKF and with decoupled noise without an EnKF is shown.
From Table 4, it is observed that the total estimation error of the Feedforward LSTM neural network is greater than the estimation error of the LSTM neural network coupled to the nonlinear recursive filter EnKF. It can be observed that the estimation of the two-station hydrological model used in this research presents an estimation of the model outputs. This can be observed in the calculated estimation error, which is quite low and well below 1. Also, the total estimation error of the recurrent neural network is observed to be larger than that of the LSTM neural network decoupled to the EnKF nonlinear recursive filter. It can be observed that the estimation of the two-station hydrological model used in this research presents an estimation. However, it can be observed that it does not have an optimal response as the LSTM neural network coupled to the EnKF.
A metric regression data analysis is shown for the Nash-Sutcliffe model efficiency coefficient (NSE) and Mean Absolute Error (MAE) for the LSTM neural network with noise coupled and decoupled to the EnKF filter.
A highlight of the results obtained in this study is the analysis of the Nash-Sutcliffe efficiency coefficient (NSE), a key indicator for assessing the quality of the predictions. A value of NSE close to 1 suggests that the predictions are optimal and that the model adequately describes the variability of the observed data. In the study, the performance of the LSTM model coupled and decoupled to the nonlinear recursive EnKF filter is analyzed by observing the impact of noise in the variance parameter on the NSE coefficient.
For the LSTM model without the EnKF filter, the NSE coefficient tends to move away from 1 as the noise in the variance parameter increases, indicating that the estimation quality decreases. Table 5 shows that for a noise in the variance parameter of 0.2, the NSE value is 0.9325 for output 2 of the LSTM model, suggesting that the prediction accuracy is lower with additional noise.
In contrast, for the LSTM model coupled to the EnKF filter, the NSE coefficient remains close to 1, even with noise in the variance parameter. In the same Table, for a variance noise of 0.2, the NSE value is 0.9799, indicating that the combination of LSTM and EnKF provides more reliable predictions. These results demonstrate the effectiveness of the EnKF filter in improving the stability and accuracy of LSTM model predictions, especially in noisy environments.
Another outstanding result is the mean absolute error (MAE), which shows that the LSTM recurrent neural network coupled to the EnKF filter shows an estimation error percentage of 10 % compared to the LSTM neural network without the EnKF filter, which shows an error percentage of approximately 20 % according to the MAE data obtained, this means that the EnKF filter coupled to the LSTM shows a decrease of 10 % of the percentage of error.
The LSTM (Long Short-Term Memory) model is an efficient recurrent neural network for handling long-term dependencies, which makes it suitable for river water level prediction. Configured with four inputs (precipitation, flow, upstream water level, and two outputs (water levels at two stations on the Atrato River), the LSTM offers some robustness to noise. However, its accuracy may decrease when faced with significant external disturbances.
The LSTM-ENKF model combines LSTM with the Ensemble Kalman Filter (ENKF), improving prediction capability and robustness to noise. ENKF adjusts LSTM predictions using new observations, which mitigates the impact of noise and results in more accurate estimates. This approach, although more accurate, involves greater complexity and demand for computational resources and processing time. In addition, the LSTM-ENKF is ideal when high accuracy is needed and sufficient resources are available due to its superior performance in the face of noise. However, in resource-constrained situations or where a simpler solution is required, pure LSTM may be sufficient, albeit with a possible reduction in accuracy under noisy conditions. The choice between the two models will depend on specific project objectives and operational constraints.

4. Discussion

The results show that coupling the LSTM network with the EnKF improves two-day-ahead water-level forecasting under all evaluated Gaussian white-noise conditions. The LSTM–EnKF achieved total RMSE values ranging from 0.1009 to 0.1929 m, compared with 0.1740 to 0.2356 m for the standalone LSTM, corresponding to reductions of approximately 18.1–42%. The hybrid model also produced higher NSE values (0.9760–0.9992) and reduced the MAE by approximately 51.9–56.2% across both outputs. These results indicate that assimilating observations through the EnKF reduces the forecasting errors generated by the standalone LSTM.
The improvement can be explained by the complementary functions of the two components. The LSTM learns nonlinear temporal dependencies and cross-variable relationships among water level, flow, and precipitation, whereas the EnKF updates the predicted states as observations become available. This sequential correction limits the propagation of errors caused by measurement disturbances. However, the RMSE reduction decreased from approximately 42 % at a noise variance of 0.001 to 18.1 % at a variance of 0.20 . Thus, although the EnKF mitigates the effect of observational uncertainty, it does not completely eliminate the influence of increasingly noisy observations.
These findings are consistent with previous multivariate LSTM studies that demonstrated the ability of recurrent architectures to represent temporal and cross-station hydrological relationships [2,3]. The present study extends this line of research by incorporating sequential data assimilation and evaluating the coupled model under controlled noise conditions. It also builds on the previous NARX–EnKF framework [20] by replacing the predefined lag structure of NARX with the internal memory mechanism of the LSTM. Nevertheless, direct numerical comparisons should be interpreted cautiously because the datasets, prediction horizons, model configurations, and experimental conditions used in these studies may differ.
The two-day prediction horizon may provide useful information for flood preparedness and water-resources management along the Atrato River. However, the present findings represent an experimental evaluation rather than the validation of a complete operational early-warning system. Practical deployment would require real-time data transmission, automated data-quality control, operational warning thresholds, communication protocols, and validation during extreme hydrological events.

Methodological Uncertainties and Main Sources of Error

The results are subject to several sources of uncertainty. The hydrological records may contain measurement errors, missing observations, outliers, or inconsistencies associated with sensor operation, calibration, and data transmission. These errors can affect both the LSTM input sequences and the observations assimilated by the EnKF, thereby influencing the estimated states and the resulting forecasts.
The additive Gaussian white noise used in the experiments provides a controlled representation of measurement uncertainty and enables the models to be compared under reproducible disturbance levels. However, this assumption may not fully represent errors observed in operational hydrological monitoring systems. Real disturbances may be biased, temporally correlated, heteroscedastic, nonstationary, or associated with sensor malfunction and communication failures. Therefore, the reported robustness should be interpreted in relation to the specific noise assumptions evaluated in this study.
Model performance is also influenced by the LSTM architecture, hyperparameter configuration, training procedure, and random initialization. Variations in sequence length, number of hidden units, learning rate, number of epochs, and training-data composition may produce different forecasting results. Similarly, the EnKF is sensitive to the ensemble size, initial ensemble distribution, and specification of the process- and observation-error covariance matrices. Inadequate parameter or covariance selection may underestimate forecast uncertainty, reduce the effectiveness of state updates, or allow errors to propagate across the two-day forecasting horizon.
Finally, the analysis is limited to 789 days of observations from the Atrato River. This observational period may not fully represent interannual variability, seasonal transitions, or rare extreme-flood events. Consequently, the generalizability of the framework to other river basins, climatic periods, monitoring configurations, and hydrological conditions remains to be established. Further validation should therefore consider longer records, independent basins, additional monitoring locations, extreme-event data, and more realistic disturbance scenarios.

5. Conclusions

This study developed and evaluated a hybrid LSTM–EnKF framework for two-day-ahead multivariate water-level forecasting along the Atrato River. Combining the temporal-learning capability of the LSTM with the observation-assimilation mechanism of the EnKF improved forecasting performance relative to the standalone LSTM under all four Gaussian white-noise variance levels evaluated.
The LSTM–EnKF achieved total RMSE values ranging from 0.1009 to 0.1929 m, representing reductions of approximately 18.1–42% relative to the standalone LSTM. It also produced NSE values ranging from 0.9760 to 0.9992 and reduced the MAE by approximately 51.9–56.2% across both outputs. These results support the research hypothesis that sequential EnKF assimilation can reduce forecasting errors and maintain higher model efficiency under noisy observational conditions.
The principal methodological contribution is the integration of an LSTM network, which represents nonlinear temporal and cross-variable relationships, with an EnKF that recursively updates the estimated states using available observations. In this respect, the framework extends the previous NARX–EnKF formulation by replacing the predefined lag structure of NARX with the memory-based sequential representation provided by the LSTM.
Although the framework provides promising support for flood-risk management and early-warning applications, the conclusions are limited to the observational period, noise assumptions, hydrological setting, and model configuration evaluated in this study. Future research should validate the approach using longer records, independent river basins, additional monitoring locations, temporally correlated and non-Gaussian disturbances, and a broader representation of extreme-flood events. Further studies should also examine ensemble-size sensitivity, covariance specification, hyperparameter uncertainty, computational requirements, and operational implementation using real-time hydrological observations.

Author Contributions

Conceptualization, J.B.R.-M. and E.G.; methodology, J.B.R.-M. and E.G.; software, J.B.R.-M.; validation, J.B.R.-M. and E.G.; formal analysis, J.B.R.-M. and E.G.; investigation, J.B.R.-M.; resources, E.G.; writing—original draft preparation, J.B.R.-M.; writing—review and editing, J.B.R.-M. and E.G.; supervision, E.G. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Ciaburro, G.; Venkateswaran, B. Neural Networks with R: Smart Models Using CNN, RNN, Deep Learning, and Artificial Intelligence Principles; Packt Publishing Ltd.: Birmingham, UK, 2017. [Google Scholar]
  2. Renteria-Mena, J.B.; Plaza, D.; Giraldo, E. Multivariate Hydrological Modeling Based on Long Short-Term Memory Networks for Water Level Forecasting. Information 2024, 15, 358. [Google Scholar] [CrossRef] [Scilit]
  3. Renteria-Mena, J.B.; Plaza, D.; Giraldo, E. Multivariable NARX Based Neural Networks Models for Short-Term Water Level Forecasting. Eng. Proc. 2023, 39, 60. [Google Scholar] [CrossRef] [Scilit]
  4. Zhong, C.; Guo, T.; Jiang, Z.; Liu, X.; Chu, X. A hybrid model for water level forecasting: A case study of wuhan station. In Proceedings of the 2017 4th International Conference on Transportation Information and Safety (ICTIS); IEEE: Piscataway, NJ, USA, 2017; pp. 247–251. [Google Scholar]
  5. Zhong, C.; Jiang, Z.; Chu, X.; Guo, T.; Wen, Q. Water level forecasting using a hybrid algorithm of artificial neural networks and local Kalman filtering. Proc. Inst. Mech. Eng. Part M J. Eng. Marit. Environ. 2019, 233, 174–185. [Google Scholar] [CrossRef] [Scilit]
  6. Chen, Y.; Cao, F.; Meng, X.; Cheng, W. Water Level Simulation in River Network by Data Assimilation Using Ensemble Kalman Filter. Appl. Sci. 2023, 13, 3043. [Google Scholar] [CrossRef] [Scilit]
  7. Wang, K.; Hu, T.; Zhang, P.; Huang, W.; Mao, J.; Xu, Y.; Shi, Y. Improving Lake Level Prediction by Embedding Support Vector Regression in a Data Assimilation Framework. Water 2022, 14, 3718. [Google Scholar] [CrossRef] [Scilit]
  8. Kong, L.; Li, Y.; Yuan, S.; Li, J.; Tang, H.; Yang, Q.; Fu, X. Research on water level forecasting and hydraulic parameter calibration in the 1D open channel hydrodynamic model using data assimilation. J. Hydrol. 2023, 625, 129997. [Google Scholar] [CrossRef] [Scilit]
  9. Zhou, Y.; Guo, S.; Xu, C.Y.; Chang, F.J.; Yin, J. Improving the reliability of probabilistic multi-step-ahead flood forecasting by fusing unscented Kalman filter with recurrent neural network. Water 2020, 12, 578. [Google Scholar] [CrossRef] [Scilit]
  10. Sun, Y.; Bao, W.; Valk, K.; Brauer, C.; Sumihar, J.; Weerts, A. Improving forecast skill of lowland hydrological models using ensemble Kalman filter and unscented Kalman filter. Water Resour. Res. 2020, 56, e2020WR027468. [Google Scholar] [CrossRef] [Scilit]
  11. Ziliani, M.G.; Ghostine, R.; Ait-El-Fquih, B.; McCabe, M.F.; Hoteit, I. Enhanced flood forecasting through ensemble data assimilation and joint state-parameter estimation. J. Hydrol. 2019, 577, 123924. [Google Scholar] [CrossRef] [Scilit]
  12. Liu, C.; Fu, R.; Xiao, D.; Stefanescu, R.; Sharma, P.; Zhu, C.; Sun, S.; Wang, C. Enkf data-driven reduced order assimilation system. Eng. Anal. Bound. Elem. 2022, 139, 46–55. [Google Scholar] [CrossRef] [Scilit]
  13. Mirikitani, D.T.; Nikolaev, N. Dynamic modeling with ensemble Kalman filter trained recurrent neural networks. In Proceedings of the 2008 Seventh International Conference on Machine Learning and Applications; IEEE: Piscataway, NJ, USA, 2008; pp. 843–848. [Google Scholar]
  14. Bao, A.; Gildin, E.; Huang, J.; Coutinho, E.J. Data-driven end-to-end production prediction of oil reservoirs by enkf-enhanced recurrent neural networks. In Proceedings of the SPE Latin America and Caribbean Petroleum Engineering Conference; SPE: Richardson, TX, USA, 2020; p. D021S004R001. [Google Scholar]
  15. Piyush, V.; Yan, Y.; Zhou, Y.; Yin, Y.; Ghosh, S. A Matrix Ensemble Kalman Filter-based Multi-arm Neural Network to Adequately Approximate Deep Neural Networks. arXiv 2023, arXiv:2307.10436. [Google Scholar]
  16. Pawar, S.; Ahmed, S.E.; San, O.; Rasheed, A.; Navon, I.M. Long short-term memory embedded nudging schemes for nonlinear data assimilation of geophysical flows. Phys. Fluids 2020, 32, 076606. [Google Scholar] [CrossRef] [Scilit]
  17. Zhou, S.; Wang, L.; Liu, J.; Tang, J. An EnKF-LSTM Assimilation Algorithm for Crop Growth Model. arXiv 2024, arXiv:2403.03406. [Google Scholar]
  18. Cheng, M.; Fang, F.; Navon, I.M.; Pain, C. Ensemble Kalman filter for GAN-ConvLSTM based long lead-time forecasting. J. Comput. Sci. 2023, 69, 102024. [Google Scholar] [CrossRef] [Scilit]
  19. Lu, M.; Jin, C.; Yu, M.; Zhang, Q.; Liu, H.; Huang, Z.; Dong, T. MCGLN: A multimodal ConvLSTM-GAN framework for lightning nowcasting utilizing multi-source spatiotemporal data. Atmos. Res. 2024, 297, 107093. [Google Scholar] [CrossRef] [Scilit]
  20. Renteria-Mena, J.B.; Plaza, D.; Giraldo, E. Water-Level Forecasting Based on an Ensemble Kalman Filter with a NARX Neural Network Model. Eng. Proc. 2025, 101, 2. [Google Scholar] [CrossRef] [Scilit]
  21. Angulo, C.D.; Viviana, Y.; Oviedo-Barrero, F. Modelación numérica para la determinación de la cota máxima de inundación, en la Ensenada de Utría desde Playa de Diego hasta Ciudad El Valle-Chocó. Ing. Investig. Tecnol. 2023, 24, 1–12. [Google Scholar] [CrossRef] [Scilit]
  22. Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Gillijns, S.; Mendoza, O.B.; Chandrasekar, J.; Moor, B.L.R.D.; Bernstein, D.S.; Ridley, A. What is the ensemble Kalman filter and how well does it work? In Proceedings of the 2006 American Control Conference; IEEE: Piscataway, NJ, USA, 2006; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  24. Muñoz-Gutiérrez, P.A.; Giraldo, E. Location of brain areas using ensemble Kalman filter and considering non-homogeneous and no linear dynamical model. In Proceedings of the 2017 IEEE 37th Central America and Panama Convention (CONCAPAN XXXVII); IEEE: Piscataway, NJ, USA, 2017; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  25. Munõz-Gutiérrez, P.A.; Giraldo, E. Ensemble Kalman filter for state estimation of brain activity by considering a large scale nonlinear dynamical model. In Proceedings of the VII Latin American Congress on Biomedical Engineering CLAIB 2016, Bucaramanga, Santander, Colombia, 26–28 October 2016; Torres, I., Bustamante, J., Sierra, D.A., Eds.; Springer: Singapore, 2017; pp. 445–448. [Google Scholar]
Figure 1. The location of hydrological stations in the Atrato River.
Figure 1. The location of hydrological stations in the Atrato River.
Computation 14 00199 g001
Figure 2. Description of the Hydrological Station.
Figure 2. Description of the Hydrological Station.
Computation 14 00199 g002
Figure 3. Water-level forecasts for both outputs obtained with the fourth-order multivariate LSTM–EnKF model at a noise variance of 0.01 : (a) Output 1 water-level forecast obtained with the fourth-order multivariate LSTM–EnKF model at a noise variance of 0.01 and (b) Output 2 water-level forecast obtained with the fourth-order multivariate LSTM–EnKF model at a noise variance of 0.01 .
Figure 3. Water-level forecasts for both outputs obtained with the fourth-order multivariate LSTM–EnKF model at a noise variance of 0.01 : (a) Output 1 water-level forecast obtained with the fourth-order multivariate LSTM–EnKF model at a noise variance of 0.01 and (b) Output 2 water-level forecast obtained with the fourth-order multivariate LSTM–EnKF model at a noise variance of 0.01 .
Computation 14 00199 g003
Figure 4. Results for two water-level outputs for a system order 4 Multivariable LSTM forecasting model coupled to an EnKF filter with noise variance parameter of 0.2 . (a) Output 1 water level forecasts for an order 4 system of a multivariate LSTM model coupled to an EnKF filter with noise variance parameter of 0.2 . (b) Output 2 water level forecast for an order 4 system of a multivariate LSTM model coupled to an EnKF filter with noise variance parameter of 0.2 .
Figure 4. Results for two water-level outputs for a system order 4 Multivariable LSTM forecasting model coupled to an EnKF filter with noise variance parameter of 0.2 . (a) Output 1 water level forecasts for an order 4 system of a multivariate LSTM model coupled to an EnKF filter with noise variance parameter of 0.2 . (b) Output 2 water level forecast for an order 4 system of a multivariate LSTM model coupled to an EnKF filter with noise variance parameter of 0.2 .
Computation 14 00199 g004
Figure 5. Results for two water-level outputs for a system order 4 multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01 . (a) Output 1 water level forecast for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01 . (b) Output 2 water level forecasting for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01 .
Figure 5. Results for two water-level outputs for a system order 4 multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01 . (a) Output 1 water level forecast for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01 . (b) Output 2 water level forecasting for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.01 .
Computation 14 00199 g005
Figure 6. Results for two water-level outputs for a system order 4 multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1 . (a) Output 1 water level forecast for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1 . (b) Output 2 water level forecast for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1 .
Figure 6. Results for two water-level outputs for a system order 4 multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1 . (a) Output 1 water level forecast for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1 . (b) Output 2 water level forecast for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.1 .
Computation 14 00199 g006
Figure 7. Results for two water-level outputs for a system order 4 multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.2 . (a) Output 1 water level forecast for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.2 . (b) Output 2 water level forecasting for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.2 .
Figure 7. Results for two water-level outputs for a system order 4 multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.2 . (a) Output 1 water level forecast for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.2 . (b) Output 2 water level forecasting for an order 4 system of a multivariate LSTM recurrent neural network model without recursive nonlinear EnKF with noise variance parameter of 0.2 .
Computation 14 00199 g007
Table 1. Description of the hydrological station locations.
Table 1. Description of the hydrological station locations.
Station 1 (E1)Station 2 (E2)
Longitude76°40′10.75″ W76°39′44.13″ W
Latitude5°45′53.38″ N5°41′52.77″ N
Altitude20.579 MASL20.83 MASL
CityBelén de BajiráQuibdó
Table 2. Information of the used data sets.
Table 2. Information of the used data sets.
TypeOrganizationAvailable DataDataset
Station 1 (S1)IDEAM2021–2023Flow, Level, Precipitation.
Station 2 (S2)IDEAM2021–2023Flow, Level, Precipitation.
Table 3. Parameters of the LSTM model, LSTM-EnKF.
Table 3. Parameters of the LSTM model, LSTM-EnKF.
Non-Linear Models# Input Nodes# Hidden Nodes# Output NodesLearning RateEpochs
ENKF480220
LSTM44020.0220
Table 4. Root mean square estimation (RMSE) with Gaussian noise for LSTM neural network coupled with EnKF and without EnKF.
Table 4. Root mean square estimation (RMSE) with Gaussian noise for LSTM neural network coupled with EnKF and without EnKF.
LSTM coupled with EnKF
Variance parameterRMSE Output 1 (m)RMSE Output 2 (m)Total (m)
0.0010.05330.04760.1009
0.050.06540.05330.1187
0.10.08230.06150.1438
0.20.10280.09010.1929
LSTM without EnKF
0.0010.09700.07700.1740
0.050.09880.08710.1859
0.10.11240.09230.2047
0.20.12330.11230.2356
Table 5. Nash-Sutcliffe model efficiency coefficient (NSE) and Mean Absolute Error (MAE) with Gaussian noise for the LSTM neural network coupled with EnKF and decoupled without EnKF.
Table 5. Nash-Sutcliffe model efficiency coefficient (NSE) and Mean Absolute Error (MAE) with Gaussian noise for the LSTM neural network coupled with EnKF and decoupled without EnKF.
LSTM coupled with EnKF
Variance parameterNSE Output 1 (m)NSE Output 2 (m)MAE Output 1MAE Output 2
0.0010.99160.99550.10420.1010
0.050.99170.99920.11900.1070
0.10.98660.99100.12330.1180
0.20.97600.97990.13220.1244
LSTM without EnKF
0.0010.96830.97760.23800.2300
0.050.94320.95620.24750.2412
0.10.93010.94270.26670.2511
0.20.92000.93250.28630.2612
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Renteria-Mena, J.B.; Giraldo, E. Robust Short-Term Multivariate Water-Level Forecasting Using a Hybrid LSTM–EnKF Model Under White-Noise Disturbances. Computation 2026, 14, 199. https://doi.org/10.3390/computation14090199

AMA Style

Renteria-Mena JB, Giraldo E. Robust Short-Term Multivariate Water-Level Forecasting Using a Hybrid LSTM–EnKF Model Under White-Noise Disturbances. Computation. 2026; 14(9):199. https://doi.org/10.3390/computation14090199

Chicago/Turabian Style

Renteria-Mena, Jackson B., and Eduardo Giraldo. 2026. "Robust Short-Term Multivariate Water-Level Forecasting Using a Hybrid LSTM–EnKF Model Under White-Noise Disturbances" Computation 14, no. 9: 199. https://doi.org/10.3390/computation14090199

APA Style

Renteria-Mena, J. B., & Giraldo, E. (2026). Robust Short-Term Multivariate Water-Level Forecasting Using a Hybrid LSTM–EnKF Model Under White-Noise Disturbances. Computation, 14(9), 199. https://doi.org/10.3390/computation14090199

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop