Abstract
Forecasting of water quality is crucial for safeguarding public health, optimizing resource management, and promoting environmental sustainability. However, predicting water quality dynamics is challenging due to nonlinear relationships among environmental variables, temporal fluctuations, and data scarcity. Recent advances in deep learning have significantly enhanced the predictive accuracy of water quality indices (WQIs) and associated parameters. This review critically examines deep learning approaches for water quality forecasting, focusing on surface water. Following a systematic literature search covering publications from 2015 to June 2026 and a detailed eligibility audit of all screened references, 37 primary studies were included in the systematic synthesis, supplemented by additional review, methodological, and contextual references cited for background. The architecture reviewed includes convolutional and recurrent neural networks, hybrid models, attention mechanisms, transformers, and graph neural networks for spatial and temporal modeling. Data sources examined include in situ sensors, laboratory analyses, remote sensing, and hydro-meteorological variables. Prevalent challenges such as missing data, non-stationarity, and spatial limitations are addressed. Mitigation strategies, including transfer learning, uncertainty quantification, and physics-informed approaches are synthesized. The key finding is that hybrid and attention-enhanced deep learning architectures are frequently reported to outperform single architecture baselines in the reviewed literature and offer the greatest promise for robust, interpretable, and scalable systems for predicting water quality.
1. Introduction
Ensuring access to clean and safe water is vital for public health, ecosystem resilience, and sustainable socioeconomic growth [1,2]. Globally, water resources face increasing threats from industrial effluents, agricultural runoff, and inadequate treatment of municipal wastewater [1,2]. As these pressures grow, the demand for effective water quality monitoring and forecasting systems becomes more urgent to enable timely actions and safeguard both people and aquatic ecosystems [3,4]. Water quality forecasting, which involves predicting future water conditions from hours to several weeks ahead, supports proactive management, enabling authorities to adjust treatment processes, issue public advisories, and implement pollution control measures before critical thresholds are surpassed [2,3]. However, accurate water quality forecasting remains a formidable challenge due to complex, nonlinear interactions among physical, chemical, and biological processes, further complicated by variability arising from meteorological, land-use, and human activities [2,5].
To interpret complex water quality data, water quality indices (WQIs) combine multiple parameters, including dissolved oxygen (DO), biochemical oxygen demand (BOD), chemical oxygen demand (COD), ammonia, suspended solids, and pH into a single composite score [1]. In Malaysia, the WQI is widely used for regulation and public communication, as documented by the Department of Environment Malaysia [3]. Precise prediction of the WQI and its components is thus essential for timely intervention, effective pollution management, and optimal resource allocation [1,4]. Conventional statistical and process-based models often struggle to capture the complexity and dynamic changes in natural water systems because they depend on simplified assumptions or require extensive, high-quality datasets, and their accuracy can suffer when data are limited or noisy [2,5]. Much research has consequently focused on forecasting water quality using machine learning and deep learning models [6,7].
Deep learning (DL) models excel at recognizing complex patterns and nonlinear relationships within large, high-dimensional datasets, often outperforming traditional approaches in predicting water quality parameters [2,5]. Advanced architectures such as convolutional neural networks (CNNs), long short-term memory (LSTM) networks, transformers, and hybrid models can learn both short-term and long-term trends by integrating data from sensors, satellite imagery, and weather records [5,8]. Unlike traditional models, deep learning methods extract features directly from raw data, reducing the need for manual feature engineering [2,6]. For example, CNNs can identify spatial anomalies, LSTMs capture long-range temporal dependencies [8,9], attention mechanisms emphasize the most relevant time segments [10,11], and graph neural networks (GNNs) model spatial relationships across monitoring sites [2,12,13]. This versatility allows deep learning models to more effectively handle the nonlinearities and heterogeneity present in water quality data [2,5].
The number of relevant publications on deep learning for water quality forecasting has grown significantly in recent years, reflecting broader advances in environmental AI and the growing accessibility of high-frequency sensor data [2,7]. Sit et al. [7] provide a comprehensive review documenting this growth across hydrology and water resources. Despite this growth, the literature remains fragmented: studies differ in their choice of model architectures, evaluation methods, and targeted water quality parameters, and few perform systematic comparisons across different model types [5,6]. Although several reviews have explored machine learning applications in water quality [2,5,14], none provides a comprehensive, architecture-focused synthesis of deep learning techniques, especially regarding hybrid models and attention mechanisms that have gained prominence in recent research [6].
This review pursues three main objectives. First, it outlines the fundamentals of water quality forecasting, including key indicators, data sources, major challenges, and the development of water quality indices [1,3]. Second, it examines how deep learning models are used for water quality prediction, providing a comparative analysis of single-model and hybrid-model performance [2,5,6]. Third, it reviews standard evaluation metrics, summarizes key findings and ongoing challenges, and highlights emerging research directions to enhance model interpretability, reliability and scalability [6,15,16].
1.1. Previous Review Studies
Recent review studies increasingly investigate how artificial intelligence (AI) and deep learning enhance water quality assessment and forecasting. These studies consistently report that models such as CNNs, LSTMs, and hybrid architectures effectively capture nonlinear relationships in water quality data [7,15]. Several review studies specifically analyse deep learning methods for predicting individual water quality parameters, often using WQI as an aggregated evaluation metric [5,15]. These studies emphasize the importance of data preprocessing techniques, including normalization, feature selection, and handling missing values for improving model performance [6,17]. Despite growing research, critical gaps remain in current reviews of AI for water quality [7,15]. Most reviews primarily focus on model performance comparison, with limited attention given to architectural innovations such as attention mechanisms and hybrid deep learning models. Additionally, the integration of explainable artificial intelligence (XAI) remains underexplored, despite its importance for improving model transparency and supporting decision-making.
As shown in Table 1, previous reviews have made important contributions from complementary perspectives, including broad coverage of deep learning applications, hybrid forecasting approaches, practical model development, and systematic or bibliometric synthesis [2,5,7,17]. However, their analytical priorities differ, and several dimensions relevant to multi-indicator water quality forecasting have not been consistently integrated within a single review framework. In particular, attention mechanisms have generally received limited or non-systematic treatment [2,7,17], while XAI and interpretability have not been consistently organized as a dedicated analytical dimension [18,19]. Regional and dataset characteristics have also not generally been synthesized alongside model architectures and forecasting evidence, making it difficult to assess how differences in monitoring settings, indicators, and data availability influence the reported results [2,7,17]. Furthermore, although hybrid models are increasingly discussed, their reported performance is difficult to compare horizontally because studies employ different datasets, target indicators, temporal resolutions, forecast horizons, input variables, and evaluation metrics [2,6,7]. The present review addresses these gaps by integrating architectural analysis, attention mechanisms, XAI and interpretability, regional and dataset characteristics, and comparative synthesis of hybrid-model performance within a single water quality forecasting framework. Section 5.4 provides a structured quantitative synthesis of hybrid-model performance across the included studies, addressing this heterogeneity directly.
Table 1.
Comparison of Previous Review Studies with the Present Study.
1.2. Scope and Aim of the Review
The aim of this review is to examine recent developments in deep learning approaches for water quality forecasting, with particular emphasis on model architectures for multi-indicator time-series data [2,6,15]. The review synthesizes key characteristics of water quality indicators, data sources, and associated challenges, and critically evaluates single-model, hybrid, and attention-enhanced architectures focused on surface water. In addition to architectural performance, the review considers attention mechanisms, explainable artificial intelligence (XAI), and the influence of regional and dataset characteristics on reported forecasting results. Particular attention is given to the comparative evidence for hybrid architectures while recognizing the heterogeneity of datasets, target indicators, forecast horizons, and evaluation protocols across studies.
In addition, this review discusses the role of attention mechanisms, which have been reported to improve predictive performance in several recent studies [10,20]. The main problems that need to be solved are data sparsity, non-stationarity, model interpretability, and practical deployment [5,6], along with potential future research directions encompassing standardized benchmark datasets [5,6], explainable AI (XAI) tools [18], physics-informed neural networks (PINNs) [21,22], transfer and few-shot learning [23,24], and edge computing with IoT integration [25,26].
1.3. Review Methodology
This review was conducted through a structured literature search, in compliance with the PRISMA guidelines to identify relevant studies on deep learning for water quality forecasting. The literature search was conducted across four major academic databases, Scopus, Web of Science, IEEE Xplore, and PubMed supplemented by the Google Scholar search engine, covering publications from the literature search and studies published from January 2015 through the final search date of 30 June 2026. The 2026 endpoint was defined by the final search date rather than by treating 2026 as a complete publication year. Therefore, only records available and indexed in the searched databases up to 30 June 2026 were eligible for inclusion. The scope of the review focuses on key areas, including deep learning architectures (CNNs, RNNs/LSTM/GRU, Transformers, GNNs, autoencoders), hybrid models, attention mechanisms, and XAI methods. Foundational references predating the search window such as seminal architecture papers, are cited for methodological context but were not identified through the systematic search.
Search Strategy: The following Boolean search string was applied:
(water quality OR water quality index) AND (deep learning OR machine learning OR neural network OR LSTM OR CNN OR transformer OR graph neural network OR GNN OR GRU OR autoencoder OR attention mechanism OR explainable) AND (prediction OR forecasting OR estimation OR monitoring).
Inclusion criteria: (1) Peer-reviewed journal articles or conference papers published from 2015 to 2026; (2) written in English; (3) applying at least one deep learning or machine learning architecture to water quality prediction or forecasting; (4) reporting at least one quantitative performance metric (RMSE, MAE, R2, NSE, accuracy, or F1-score).
Exclusion criteria: (1) Review papers were excluded from the primary-study pool. However, relevant review papers were retained in the reference list solely for contextual comparison, methodological background, or discussion of the broader literature and were not counted among the primary studies included in the systematic synthesis; (2) grey literature, reports, and theses; (3) studies focused solely on groundwater with no surface water quality relevance; and (4) articles with insufficient methodological reporting to extract performance data.
Study selection process: The study selection process was carried out using a systematic approach, and the overall procedure is illustrated using a PRISMA flow diagram (Figure 1). A total of 422 records were identified through database searching. After removing 75 duplicates, 347 records were screened based on titles and abstracts, of which 225 were excluded (irrelevant topics: 140; non-deep-learning studies: 85). A total of 122 full-text articles were assessed for eligibility. Following exclusion of 42 articles at this stage (18 did not apply deep learning; 12 focused solely on groundwater; 6 lacked quantitative performance metrics; and 6 were duplicate records identified at the full-text stage), 80 references were retained in the working bibliography. A subsequent detailed eligibility audit of these 80 references, cross-checked against the inclusion and exclusion criteria above, identified that 41 did not meet the full criteria for a primary included study: 12 were review papers, 12 were foundational or methodological references not specific to water quality, 5 employed traditional machine learning without a deep learning component, 2 addressed groundwater-only systems, 2 addressed engineered water distribution networks rather than natural surface water, 1 had unconfirmed surface-water scope, and 7 were grey literature.government reports, or other contextual references. These 41 references remain cited in the manuscript for background, methodological, or comparative purposes but are not counted among the included primary studies. The remaining 37 references constitute the final set of primary studies included in this systematic review.
Figure 1.
A PRISMA Flow Diagram of Study Selection.
1.4. Structure of the Review
This section outlines the overall structure of the review and the relationships between its key components. The review is organized into several interconnected parts, including water quality fundamentals, deep learning architectures, hybrid and attention-enhanced models, evaluation metrics, challenges, and future directions. Figure 2 illustrates how these components are integrated within the context of deep learning-based water quality forecasting.
Figure 2.
Structure of Review.
Section 2 presents the fundamentals of water quality forecasting, including key water quality indicators (DO, BOD, COD, pH, ammonia, TSS), data sources (in situ sensors, remote sensing, IoT), data challenges (missing values, non-stationarity, extreme events), and WQI formulations. Section 3 provides a comprehensive review of deep learning models for water quality prediction, including CNNs, RNNs (LSTM and GRU), Transformers, GNNs, and autoencoders/GANs. Section 4 focuses on hybrid and attention-enhanced models. Section 5 presents performance evaluation and comparative analysis. Section 6 discusses the challenges of current deep learning models. Section 7 outlines future research directions, including standardized benchmark datasets, explainable AI (XAI) tools, physics-informed neural networks (PINNs), transfer and few-shot learning, and edge computing with IoT integration. Section 8 concludes the review.
Key contributions of this review:
(1) A structured analysis of deep learning architectures for water quality forecasting, including single, hybrid, and attention-enhanced models.
(2) A systematic comparison of attention mechanisms (spatial, temporal, and spatiotemporal) for improving multi-indicator time-series prediction.
(3) A critical examination of explainable AI (XAI) techniques to enhance transparency, interpretability, and decision support in water quality assessment.
(4) Identification of key challenges related to data sparsity, non-stationarity, model interpretability, and real-world deployment, along with future research directions.
2. Fundamentals of Forecasting Water Quality
Effective application of deep learning to water quality forecasting requires a comprehensive understanding of the environmental context in which these models operate [2,3]. Water quality assessment involves various physical, chemical, and biological indicators [1]. Monitoring relies on diverse data sources with different temporal and spatial resolutions [4,25]. Datasets often present systematic challenges that impact model training and validation [6,27]. Composite indices combine multiple parameters into a single score, simplifying interpretation but complicating prediction [1,28]. The key fundamentals covered include: (i) Key water quality indicators and their typical predictive performance; (ii) Data sources for monitoring; (iii) Common data challenges; and (iv) Water Quality Index (WQI) formulation steps [1,2,28].
2.1. Key Water Quality Indicators
Water quality assessment involves analyzing physical, chemical, and biological parameters to determine suitability for various uses, including ecological, domestic, agricultural, and industrial purposes [1]. Basic physical parameters such as temperature and turbidity provide initial insights into water conditions [1]. Chemical indicators, including dissolved oxygen (DO), biochemical oxygen demand (BOD), chemical oxygen demand (COD), ammonia, nutrients (nitrogen and phosphorus), and pH, are regularly measured to evaluate pollution levels, nutrient loading, and buffering capacity [1]. Biological metrics, such as microorganisms and indicator species, help assess ecosystem health and potential human health risks [1].
Among these, DO is especially important because it directly impacts aquatic life; sustained low DO levels can lead to hypoxia, biodiversity loss, and ecosystem decline [6]. BOD and COD are complementary measures of organic pollution: BOD reflects oxygen consumption during microbial breakdown of organic matter, while COD measures the oxygen required for chemical oxidation of both biodegradable and nonbiodegradable organics [1]. Toxic substances, including heavy metals such as lead, mercury, and cadmium, along with pesticides and industrial pollutants, are monitored due to their persistence, potential for bioaccumulation, and long-term environmental impacts [1].
To simplify interpretation, composite indices such as the Water Quality Index (WQI) aggregate individual parameters into subindices and combine them into a single score [1]. In Malaysia, the WQI assigns the greatest weight to DO, with significant contributions from BOD, COD, ammonia, total suspended solids, and pH [3,4,29].
Based on the reviewed literature, microbiological and toxicological indicators were reported relatively infrequently compared with physical and chemical parameters, primarily because of low sampling frequency, high variability, and complex dependencies [2,6]. Prediction accuracy varies across indicators and is generally lower for parameters with infrequent and variable measurements, as summarized in Table 2 [2,6]. Temperature predictions consistently achieve the highest accuracy, with R2 values commonly reported above 0.90, owing to regular diurnal patterns and continuous monitoring coverage [2,7]. DO models also tend to perform strongly, with many reviewed studies reporting R2 values above 0.85 [2,6]. Chemical parameters such as BOD, COD, and nutrients show moderate predictive accuracy, typically with R2 values below 0.90 [2,6]. It should be noted that these ranges represent a synthesis across heterogeneous studies and should not be interpreted as direct quantitative comparisons or universal performance benchmarks, because they originate from studies with different datasets, target variables, temporal resolutions, forecast horizons, and evaluation settings.
Geographic differences are evident in the literature, with variations in water body types, dataset durations, parameter sets, and sampling frequencies [7,30]. Studies from China and North America typically use large, high-frequency datasets from rivers and reservoirs, focusing on DO, COD, and WQI [2,31]. Conversely, research from Africa and South America is less represented in the reviewed literature, often drawing on shorter datasets and emphasising parameters such as chlorophyll and nutrients [7,30].
Table 2.
Typical ranges of predictive performance and modeling features for major water quality parameters in recent studies.
2.2. Data Sources for Monitoring
Modern water quality monitoring integrates in situ measurements, laboratory analyses, remote sensing and real time interface [4,29]. Enhanced real time water quality index monitoring has been developed using interactive fuzzy logic. The interface integrates WQI computation with advanced visualization tools such as bar charts, fuzzy membership graphs, parameter tables, dynamic gauges, heatmaps and scatter plots [29]. Meanwhile, in situ sensors deployed in rivers, lakes, and reservoirs provide high-frequency data on temperature, pH, dissolved oxygen, turbidity, and conductivity [25,43]. Internet of Things (IoT) technologies enable continuous data transmission for near-real-time monitoring and early anomaly detection [4,25,44]. Laboratory analyses are essential for parameters that cannot be measured in situ, such as nutrients, trace metals, and microbiological indicators; these analyses provide reference data for sensor calibration, although at lower temporal resolution [1].
Satellite remote sensing provides broad spatial coverage of surface water bodies and enables estimation of proxies such as chlorophyll-a and turbidity to support large-scale assessments [45]. Meteorological and hydrological inputs including precipitation, air temperature, solar radiation, wind speed, river discharge, and reservoir releases serve as additional predictors, as they influence mixing, stratification, runoff loading, and dilution processes [2,36]. Several reviewed studies reported improvements in predictive performance when exogenous variables were incorporated as model inputs. However, the magnitude of improvement varied across datasets, target parameters, and forecasting scenarios.
Citizen science offers supplementary observations, such as Secchi depth and algal bloom reports, through mobile platforms [46]. While these data are less precise than sensor measurements, they expand spatial coverage and can provide early warnings of deteriorating conditions [25,46]. Emerging sources such as unmanned aerial vehicle (UAV) imagery, environmental DNA (eDNA), and smart-city sensors present new opportunities, though deployment remains limited [25]. Multi-source data integration requires architectures capable of handling diverse data types; graph neural networks, multi-input CNN–LSTM hybrids, and attention mechanisms show promising potential for this purpose [2,47].
2.3. Data Challenges
Water quality datasets frequently suffer from missing values due to sensor malfunctions, maintenance interruptions, or irregular sampling protocols [6,27]. These gaps can severely hinder the training of sequence-based deep learning models, which rely on continuous temporal inputs [6]. To maintain data integrity, many studies applied quality-control procedures in which observations with excessive missing data were excluded before model development. The specific threshold varied depending on dataset characteristics and preprocessing strategy. For smaller gaps, effective imputation uses statistical techniques such as mean, median, or mode calculations tailored to specific attribute types [6,27].
Another major challenge is non-stationarity. Seasonal fluctuations, climate variability, and human activities alter the statistical properties of water quality variables over time [2,6]. To address these challenges, researchers increasingly adopt decomposition-based hybrid models that separate complex nonstationary signals into simpler trend and seasonal components [48]. Including exogenous variables such as rainfall, solar radiation, and river discharge helps models capture changing environmental dynamics [2,36]. The reviewed literature indicates that incorporating exogenous variables, such as rainfall, solar radiation, and river discharge, can improve predictive performance by providing additional environmental context for water quality forecasting [2]. However, the magnitude of improvement varies across datasets, target variables, and forecasting scenarios.
Rare extreme events such as sudden pollution spills, algal blooms, or storm-driven runoff pose a disproportionate threat to water quality but are often underrepresented in training datasets [2]. Specialised techniques help address this problem, including log transformation and oversampling methods to handle data imbalance and skewness [6,27]. Advanced attention mechanisms, discussed in Section 4, can assign adaptive weights to significant parameters in real time, enabling models to prioritise critical indicators during sudden contamination events [19]. Table 3 summarizes the comparative strengths and evidence base of these solutions across the challenge categories.
Table 3.
A summary of the comparative strengths and evidence base of these solutions across the three challenge categories.
2.4. Water Quality Index Formulations
Water quality index (WQI) is commonly used to determine water quality. Constructing a water quality index involves four steps: selecting parameters, normalising them to create subindices, assigning weights, and aggregating using either additive or multiplicative methods [1,28]. WQI formulations are not standardized internationally and can differ in parameter selection, normalization procedures, and weighting or aggregation schemes. In Malaysia, the WQI uses six parameters—dissolved oxygen (which carries the greatest weight), biochemical oxygen demand, chemical oxygen demand, ammonia, total suspended solids, and pH—combined through a weighted arithmetic approach, as documented by the Department of Environment Malaysia [3,57]. Scores classify water quality as excellent, good, fair, poor, or very poor [1,3]. Modelling approaches may involve directly predicting the composite WQI score or estimating individual parameters first and then combining them [4,27]. Parameters with higher weights require more specific auxiliary inputs for effective forecasting [1].
3. Deep Learning Models for Water Quality Prediction
This section reviews the principal deep learning architectures applied to water quality prediction, covering convolutional neural networks (CNNs), recurrent neural networks (RNNs) including long short-term memory (LSTM) and gated recurrent unit (GRU) models, transformer architectures, graph neural networks (GNNs), and autoencoders and generative adversarial networks (GANs). Each subsection discusses the architectural principles, representative applications, reported performance, and known limitations, supported by evidence from the reviewed literature.
3.1. Convolutional Neural Networks (CNNs)
Convolutional neural networks (CNNs), initially developed to extract spatial features from grid-structured data, are increasingly applied in water quality research to model both temporal and spatial patterns [2,5]. Specifically, one-dimensional CNNs treat water quality observations as temporal sequences, applying convolutional kernels to identify localised patterns such as diurnal fluctuations in dissolved oxygen or rainfall-driven runoff dynamics [15]. This localised filtering improves short-term prediction accuracy and reduces noise through adaptive feature learning, rather than relying on predefined smoothing techniques [9]. Recent studies confirm that CNNs can capture subtle temporal dependencies often missed by traditional regression models. For instance, a practical comparative study reimplementing several CNN- and LSTM-based architectures on a new regional dataset found that model performance depended heavily on hyperparameter tuning and data imputation strategy rather than architecture choice alone, underscoring the sensitivity of automatic feature extraction to implementation details [15].
Two-dimensional CNNs are especially effective for spatial analysis using remote sensing data, as they identify patterns such as sediment plumes and algal blooms by processing imagery as spatial fields [46]. For integrated spatiotemporal modelling, CNNs are frequently combined with recurrent architectures, where convolutional layers extract spatial or local temporal features and recurrent layers capture longer-range temporal dependencies [15]. Pooling operations help reduce noise, while shared kernels enable the detection of recurring patterns across datasets [2]. Despite these advantages, CNNs have a limited receptive field that restricts their ability to capture long-range temporal dependencies unless improved through techniques such as dilated convolutions or hybrid model design [2,15]. In summary, CNNs remain highly effective for tasks such as sensor-based classification and image-driven parameter estimation, particularly when local feature patterns are dominant [2,46]. A 2026 study by Wang et al. [39] further extended this scope through a multimodal deep learning framework for water-quality forecasting and scenario assessment, integrating diverse data streams to shift water management from reactive monitoring toward proactive basin-level decision-making.
3.2. Recurrent Neural Networks (RNNs): LSTM and GRU
Recurrent neural networks (RNNs) maintain an internal state that propagates information across time steps, making them well-suited for modelling sequential data [9]. In water quality forecasting, long short-term memory (LSTM) and gated recurrent unit (GRU) architectures are preferred over standard RNNs because they mitigate vanishing-gradient issues that arise in long sequences [8,9]. LSTMs use input, output, and forget gates to regulate memory retention, enabling the preservation of relevant context over extended periods [8,9]. LSTM-based applications effectively capture delayed responses and seasonal persistence that are often not reproduced by traditional models [6]. These models leverage multi-day lagged inputs to forecast parameters such as dissolved oxygen, pH, and algal concentrations, supporting event-driven dynamics where rainfall impacts may emerge several days later [6,57].
According to the synthesis reported by Pyo et al. [6], several reviewed dissolved oxygen forecasting studies achieved R2 values exceeding 0.85, outperforming conventional regression methods. GRUs simplify the LSTM architecture by combining gating mechanisms into reset and update gates, thereby reducing the number of parameters and enhancing computational efficiency [58]. The GRU–N-Beats model proposed by Hao [34] demonstrated particularly strong dissolved oxygen prediction performance, achieving qR2 of 0.97 and RMSE of 0.171 mg/L, outperforming standalone LSTM, GRU, and TCN baselines by margins of 28.5% and 32.1% on RMSE and MAE respectively. Hybrid architectures combining temporal convolutional networks with GRU have similarly demonstrated strong performance in long-range prediction tasks, as evidenced in Yellow River source area studies achieving superior RMSE and MAE compared to standalone LSTM, GRU, and TCN baselines [42]. This simplification makes GRUs especially suitable for smaller datasets or applications with limited computational resources. Both LSTM and GRU architectures can process multivariate inputs, enabling the learning of complex interactions between variables [6]. Despite their strengths, gating mechanisms do not eliminate the risk of overfitting, and large architectures require careful regularisation, particularly in data-limited environments [2,6]. Overall, LSTM and GRU models remain dependable for daily-to-monthly forecasting tasks, despite recent advances in transformer-based and hybrid approaches [6,9,57].
3.3. Transformer Models
Transformer architectures, initially designed for natural language processing, have gained prominence in time-series modelling owing to their self-attention mechanism, which enables direct interactions between distant time steps rather than relying on sequential processing [59]. Self-attention allows each time step to assign adaptive weights to all other time steps, facilitating the learning of long-range dependencies without the gradient issues typical of recurrent models [59,60]. This feature is especially useful in water quality forecasting, where multiple seasonal patterns, long-term trends, and sudden events can simultaneously affect system dynamics [2,60].
Transformer models excel in long-horizon forecasting by integrating recent data, seasonal trends, and baseline shifts [11,45]. Informer and Aquaformer are two techniques that use sparse attention to focus on the most important time steps. Aquaformer is a transformer-based model that combines phase space reconstruction with multi-source transfer learning to deal with the lack of data in long-sequence water quality forecasting [23]. Autoformer employs an autocorrelation-based decomposition approach that separates trend and seasonal components [48], though its direct application to water quality datasets remains limited in the reviewed literature. Research supports these benefits: Zhi et al. [2] reviewed studies indicating that transformer models outperform CNNs and LSTMs in multi-step reservoir water temperature predictions, with performance improvements growing as the forecast horizon extends. Transformers also perform well in multivariate and multi-site contexts, capturing cross-variable and cross-location dependencies through their attention mechanisms [11,60]. However, they generally require large datasets and can overfit unless adequately regularised or tuned [2,15]. In conclusion, transformer architecture is poised to become increasingly important for water quality forecasting as monitoring networks expand and datasets grow [2,23,32,59].
3.4. Graph Neural Networks (GNNs)
Graph neural networks (GNNs) offer a natural framework for modelling hydrological connectivity because river systems and monitoring networks are inherently graph-structured rather than collections of independent points [12,61]. In water quality network graphs, monitoring stations are represented as nodes, with edges encoding flow connectivity or influence pathways, allowing models to directly capture upstream and downstream dependencies and overcoming the limitations of traditional spatial methods that rely on simplified distance assumptions [12,61].
GNN architectures work by iteratively aggregating information from neighbouring nodes, enabling each station to incorporate local observations and upstream context [12]. Common variants include graph convolutional networks (GCNs), which extend convolution operations to graph-structured data [12], and graph attention networks (GATs), which assign adaptive weights to neighbouring nodes based on learned importance [12,47]. Spatiotemporal trend-aware neural networks have been developed to capture both spatial dependencies across monitoring stations and temporal trends in river systems simultaneously [20].
Wan et al. [33] reported an accuracy improvement of 36.54–161.47% over baseline models, whereas Yuan et al. [38] specifically reported a 46.62% MAE reduction (with 37.68% and 45.67% reductions in RMSE and MAPE, respectively) compared with LSTM/GRU baselines. Wang et al. [41] applied a novel deep learning model for high-frequency water quality prediction across multi-station river networks, further demonstrating the scalability of spatiotemporal approaches to densely monitored systems. GNNs explicitly incorporate physical connectivity, providing more realistic spatial generalisation than grid-based or independent-station approaches [12,61]. However, their performance depends heavily on the quality of the underlying graph structure, which can be difficult to define in complex or partially observed systems, and they can be computationally demanding for large networks [12,61]. Reinforcement learning-based graph neural networks have also been explored as an adaptive extension, combining spatial relational modeling with reward-driven optimization to improve water quality prediction in dynamic environments [62].
3.5. Autoencoders and Generative Adversarial Networks (GANs)
Autoencoders and generative adversarial networks (GANs) primarily support water quality analysis pipelines rather than serving as standalone forecasting models [2,50]. Autoencoders learn compressed representations by reconstructing input data, making them useful for dimensionality reduction and denoising [2,50]. When trained on typical system behaviour, autoencoders accurately reproduce normal patterns, whereas anomalies such as contamination events or sensor malfunctions produce higher reconstruction errors, enabling unsupervised anomaly detection [50,51].
Seshan et al. [49] demonstrated the practical utility of LSTM-based autoencoders for real-time quality control of wastewater treatment sensor data, showing that both LSTM autoencoder and ARIMA-based approaches achieved high accuracy for reconciling single-point anomalies, though accuracy degraded quickly with increasing forecasting horizon for prolonged anomalous events [49]. GANs excel at generating realistic synthetic samples through adversarial training. In water quality datasets, rare events such as pollution spills or algal blooms are often underrepresented, limiting model exposure during training [2,17]. Although practical validation of GANs in water quality forecasting remains limited within the reviewed literature, they show potential for creating plausible scenarios for rare events [2].
4. Advanced Hybrid Deep Learning Architectures
Hybrid deep learning architectures have emerged as a dominant paradigm in water quality forecasting, combining the complementary strengths of multiple model types to address the multi-scale, nonlinear, and non-stationary dynamics of water quality systems [9,15]. By integrating architectures such as CNNs, LSTMs, GRUs, attention mechanisms, and decomposition strategies in the reviewed studies, hybrid models often outperform their single-architecture counterparts across a wide range of water quality prediction tasks [2,10,52]. This section reviews four principal categories of hybrid architectures: CNN–LSTM models, decomposition-based hybrids, attention-enhanced hybrids, and implementation considerations for operational deployment.
4.1. CNN–LSTM Architecture
Hybrid CNN-LSTM architectures integrate local pattern recognition with long-term sequence analysis, aligning with the multi-scale dynamics of water quality systems. Convolutional layers initially detect short-range patterns within input windows, followed by LSTM layers that capture temporal dependencies over extended periods [1,47,60]. Combining convolutional feature extraction with recurrent memory has proven effective for environmental time-series forecasting and hydrological prediction [36]. Empirical research supports the success of this hybrid approach: Barzegar et al. [9] reported that a CNN-LSTM hybrid model captured both low and high concentration ranges of dissolved oxygen and chlorophyll-a more effectively than standalone CNN or LSTM models in the Small Prespa Lake, Greece. Wai et al. [52] extended this concept with a CNN-coupled dual-path LSTM architecture, enhancing multi-step water quality index forecasts. Complementing this direction, Sabagh Torkan et al. [55] proposed a VMD–CNN–GRU framework that integrates Variational Mode Decomposition with convolutional and gated recurrent units, achieving multi-horizon forecasting of water quality dynamics with improved accuracy across short- to long-term prediction windows. Similarly, Meshram et al. [63] applied a CNN-LSTM-GRU triple-stage hybrid to predict total dissolved solids across three Hong Kong rivers using 24 years of monthly data (2000–2023), with the three-stage model outperforming standalone CNN, CNN-LSTM, and DNN baselines on all three evaluation metrics (RMSE, R2, NSE). Further extending hybrid architectures, the NGO-CNN-GRU model incorporates an optimization algorithm for hyperparameter tuning, achieving exceptional predictive performance (R2 > 0.986) for multiple water quality parameters in river systems [64]. These models also facilitate multimodal data integration by using separate convolutional encoders to process inputs such as meteorological data, water quality metrics, and spatial features, and then combining them via recurrent layers [65]. Operationally, CNN-LSTM models are well-suited for early warning systems that require accurate predictions of threshold exceedances with sufficient lead time [1,60]. However, the increased complexity increases the risk of overfitting, necessitating regularization techniques such as dropout, early stopping, and careful hyperparameter tuning.
4.2. Decomposition-Based Hybrid Models
Decomposition-based hybrid models address the nonstationarity and nonlinearity of water quality time series by splitting the signals into trend, seasonal, and residual components, which are then modeled separately [2,66]. This modular approach allows different modeling methods to be applied to components with unique statistical properties. For example, simple extrapolation can model long-term trends, periodic models capture seasonal patterns, and data-driven approaches handle residual variability. This strategy is popular in time series forecasting because it breaks down complex dynamics into simpler, more manageable sub-problems. This approach improves predictive accuracy, especially during periods of high variability, by isolating non-stationary behavior into easier-to-handle segments. It also enhances interpretability, since forecasts are linked to specific time scales. Additionally, decomposition-based methods facilitate hybrid modeling by combining traditional statistical techniques with modern deep learning architectures [2,10]. Nevertheless, this method requires additional preprocessing, such as selecting a decomposition technique and tuning its parameters, which can significantly affect performance.
4.3. Attention-Enhanced Hybrids
Attention-enhanced hybrid models include mechanisms that assign importance to relevant time steps and variables within the input data in real time. This feature is especially useful in event-driven water quality systems, where certain periods such as storm events have a much greater impact on system behavior [3,27]. Advanced attention mechanisms can assign adaptive weights to significant parameters in real time, enabling models to prioritize potentially informative indicators during sudden contamination events [27]. Models such as Autoformer use an autocorrelation-based mechanism rather than standard attention to identify key periodicities in long time series, enabling more efficient representation of complex temporal patterns [48]. By directing computational resources toward the most informative parts of the data, attention-based hybrids may enhance predictive performance and can provide information about which time steps or variables receive greater model weighting. However, attention weights should not automatically be interpreted as faithful explanations of model decisions. Studies have shown that attention distributions may not reliably correspond to feature importance, while subsequent work has argued that the interpretability of attention depends on how explanation and faithfulness are defined and evaluated [67,68]. Similarly, SHAP-based attributions should ideally be validated through perturbation-based methods. For example, occluding or altering features is identified as important and the resulting change in model output is assessed rather than being interpreted as definitive explanations without such validation.
4.4. Implementation Considerations
Deploying hybrid deep learning models successfully requires careful attention to data preprocessing, model design, and operational factors. Data preprocessing is vital, involving proper handling of missing data to prevent bias, feature scaling for numerical stability, and choosing optimal input window sizes [8]. Hyperparameter tuning is crucial for robust results, using methods such as grid search, random search, or Bayesian optimization to tune parameters such as convolutional kernel sizes, LSTM units, and dropout rates. To ensure models generalize well, appropriate validation techniques such as rolling-origin evaluation and consistent temporal data splits must be employed to avoid data leakage and overfitting. From a deployment standpoint, hybrid models are often implemented using containerized microservices, enabling scalable, flexible system integration. Customization for specific sites, including incorporating domain knowledge such as dam operations and designing loss functions that focus on extreme events, is essential for real-world success. Considering computational needs, especially GPU acceleration, is important [11]. Ultimately, local customization and adaptation to site-specific conditions often distinguish academic success from practical operational performance [8]. While Table 4 synthesizes reported performance patterns at the architecture-family level, Table 5 presents the underlying study-level evidence for specific named hybrid model implementations to preserve traceability between family-level synthesis and individual primary studies.
Table 4.
Previous deep learning models.
Table 5.
Comparative analysis of hybrid model architectures for water quality prediction.
5. Performance Evaluation and Comparative Analysis
This section presents the evaluation framework used to assess deep learning models for water quality prediction, including standard metrics, comparative results across architectures, statistical testing approaches, and a structured quantitative synthesis supported by geographic and architectural comparison tables.
5.1. Evaluation Metrics
The performance of deep learning models for water quality prediction is assessed using multiple complementary metrics, as no single metric fully captures model behaviour across all prediction contexts [5,15]. Error-based metrics, including root mean square error (RMSE) and mean absolute error (MAE), are most commonly reported [5,15]. RMSE is the square root of the mean squared prediction error and is especially sensitive to large deviations, making it suitable for evaluating extreme events such as pollution spikes or algal blooms. In contrast, MAE indicates the average magnitude of errors in the original measurement units, offering a more robust and interpretable measure for practical applications [15].
Across the reviewed studies, reported R2 values varied substantially depending on the target parameter, model architecture, forecast horizon, and data characteristics [6,15,52]. These findings represent a synthesis across heterogeneous studies and should not be interpreted as direct quantitative comparisons or universal performance benchmarks [6,15,52]. The Nash–Sutcliffe Efficiency (NSE) coefficient is widely used in hydrological and water quality modelling as a normalised measure of model skill relative to a simple mean baseline [30]. Recent studies have reported that attention-enhanced hybrid deep learning models achieved high predictive performance, with NSE values for dissolved oxygen ranging from 0.817 to 0.967 depending on the model configuration, at a representative monitoring site [10].
For classification tasks such as predicting whether water quality exceeds regulatory thresholds, evaluation metrics include accuracy, precision, recall, and F1 score [16]. In multi-step forecasting, model performance is assessed across different prediction horizons, as errors generally increase with longer lead times [10,52]. Reporting performance by forecast horizon provides a more thorough evaluation of model reliability and is recommended for studies targeting operational early warning systems [52].
Beyond these standard deterministic metrics, probabilistic evaluation provides an important complement to point-prediction metrics when uncertainty-aware forecasting is intended. However, calibrated uncertainty reporting remains relatively uncommon in the primary evidence examined in this review. Consequently, strong point-prediction performance should not by itself be interpreted as evidence of fully reliable operational deployment, particularly for early-warning and decision-support applications where the uncertainty associated with individual forecasts is consequential. Prediction Interval Coverage Probability (PICP) measures the proportion of observed values that fall within a model’s generated prediction intervals, with values approaching the nominal confidence level (95%) indicating well-calibrated uncertainty estimates [70,71].
5.2. Comparative Results
Comparative studies generally indicate that deep learning models can outperform traditional statistical methods when sufficient data are available [2,15]. Among these, hybrid architectures are among the top performers in predictive accuracy across many studies [2,16]. For example, Barzegar et al. [9] showed that a CNN–LSTM hybrid model was better at capturing both low and high concentration ranges of dissolved oxygen and chlorophyll-a than either a CNN or an LSTM model on its own.
Recurrent models such as LSTM and GRU remain strong baseline approaches due to their ability to capture temporal dependencies [6,8]. CNN-based models perform well in short-term forecasting and pattern recognition but tend to lose accuracy over longer prediction horizons [9,15]. Transformer-based models have shown excellent performance in long-horizon forecasting tasks, exhibiting slower error growth over extended lead times than LSTM models [11,48]. However, this benefit typically comes with higher computational costs and greater data requirements [2,15]. Hybrid deep learning frameworks have also outperformed traditional machine learning methods such as decision trees, random forests, and support vector machines [16,72]. Attention-enhanced hybrid models have achieved DO-prediction NSE values of 0.817–0.967 in high-frequency monitoring scenarios [10].
5.3. Statistical Analysis
While performance metrics measure how accurately models predict outcomes, statistical testing provides an important basis for determining whether observed differences between models are statistically significant [73]. Beyond significance testing, uncertainty quantification provides an important complement to point-prediction evaluation, particularly when forecasts are intended to support uncertainty-aware operational decisions.
Within the primary evidence examined in this review, statistical analysis is used for several complementary purposes, including testing model assumptions, assessing residual behaviour, evaluating probabilistic forecasts, and testing differences associated with experimental conditions. Lokman et al. [4,27] report the Breusch–Pagan test for assessing heteroscedasticity and the Shapiro–Wilk test for assessing residual normality, illustrating the use of statistical diagnostics alongside machine-learning modeling. These tests provide information about residual behaviour and model-assumption characteristics that cannot be inferred from predictive-error metrics alone. Zheng et al. [53] use a two-sided paired t-test to evaluate the significance of performance differences associated with different training-data quantities (100%, 50%, and 20%), with significance levels reported at 0.05, 0.01, and 0.001. This provides an example of using inferential testing to distinguish observed performance differences from differences that may arise from experimental variation. For probabilistic forecast assessment. Huan et al. [40] apply the Probability Integral Transform (PIT), together with Kolmogorov 5% significance bands, to assess whether test-set predictive distributions are consistent with the expected uniform distribution, providing a distributional diagnostic for probabilistic forecast reliability.
Taken together, these examples show that statistical analysis in water-quality forecasting extends beyond reporting RMSE, MAE, or related point-prediction metrics. Diagnostic tests can examine residual assumptions, inferential tests can assess whether observed differences are statistically meaningful, and distributional tests can evaluate aspects of probabilistic forecast reliability. However, the specific statistical procedures and reporting practices remain heterogeneous across the reviewed studies. Consequently, reported improvements in predictive metrics should not automatically be interpreted as statistically significant unless an appropriate inferential procedure is reported. This distinction is particularly important when relatively small differences between competing architectures are used to support claims of model superiority.
For uncertainty quantification, Wang et al. [39] provide a clear example of quantile-based uncertainty assessment, generating 0.1, 0.5, and 0.9 quantile forecasts and evaluating prediction-interval width and observed coverage against nominal levels. Importantly, ref. [39] also connects these quantile forecasts explicitly to early-warning decisions: when the 0.9 quantile forecast for a regulated pollutant exceeds a relevant threshold, the authors interpret this as indicating at least a 10% probability of exceedance and propose pre-emptive responses such as intensified sampling. This example illustrates the broader decision relevance of calibrated uncertainty for early-warning applications.
Other approaches include Monte Carlo (MC) dropout [70]. Conformal prediction and deep ensembles provide additional frameworks for constructing or evaluating predictive uncertainty; however, we did not identify a primary water-quality forecasting study in the reviewed corpus that empirically evaluated these approaches for uncertainty quantification. Within the reviewed evidence, robustness validation is more commonly addressed through ablation studies [34,62], noise-injection experiments [56], and data-scarcity experiments [53], which test model resilience under challenging conditions rather than providing probabilistic uncertainty estimates.
5.4. Quantitative Synthesis
To address variability among studies, this section synthesizes findings through a structured comparison. Table 6 summarises dataset characteristics across geographic regions, while Table 7 compares the key features of major deep learning architectures used in water quality forecasting.
Table 6.
Geographic distribution reflects coverage in the reviewed literature.
Table 7.
Architecture comparison of key deep learning architectures for water quality forecasting.
Table 6 summarizes dataset characteristics across geographic regions. The reviewed studies reveal a strong geographic concentration in East and Southeast Asia, particularly China and Malaysia, where high-frequency monitoring data and well-established water quality index frameworks are available. European studies contribute high-resolution IoT-based datasets, while global review studies provide broader synthesis across diverse water body types. Notably, coverage from Africa and South America remains sparse, reflecting limitations in monitoring infrastructure and data availability in these regions. These geographic disparities highlight the need for more inclusive benchmark datasets and international collaboration to support deep learning applications in underrepresented regions.
Analysis of Table 7 reveals a clear tradeoff between model complexity and performance. Hybrid models generally achieve the highest accuracy but require more computational resources and larger datasets. Simpler models are more interpretable but may have lower predictive accuracy. These findings highlight the importance of selecting models based on operational constraints rather than solely aiming for the highest accuracy. Additionally, Table 6 shows notable geographic biases, with East Asia dominating high-frequency data collection while Africa and South America rely on sparser, lower-frequency measurements.
6. Challenges of Current Deep Learning Models
Despite the demonstrated capabilities of deep learning architectures for water quality forecasting, significant challenges remain across three interconnected dimensions: data availability and quality, model design and generalisation, and operational deployment. This section systematically reviews these challenges, drawing on evidence from the reviewed literature.
6.1. Data-Related Challenges
Despite significant advances in environmental monitoring technologies, data-related limitations remain a major barrier to the effective use of deep learning models for water quality forecasting [2,5]. High-frequency in situ measurements are typically limited to a small set of core parameters and specific monitoring stations, whereas more comprehensive datasets covering nutrients, pathogens, and toxic substances are often collected at lower temporal resolutions or irregular intervals [4,25]. Missing data due to sensor malfunctions, maintenance interruptions, and transmission errors is a persistent problem in environmental monitoring systems [49,51].
Another major challenge is data non-stationarity. Water quality systems are affected by seasonal fluctuations, climate variability, and land-use changes, all of which alter the statistical properties of observed variables over time, violating the assumptions of many machine learning models [2,6]. Extreme environmental events, such as heavy rainfall or pollution discharge incidents, are often underrepresented in training datasets, limiting model exposure to the conditions that matter most for early warning applications [2]. In sub-Saharan Africa and other low-resource regions, data scarcity poses a particularly acute barrier to deep learning adoption, where monitoring infrastructure remains limited [7,75].
6.2. Model-Related Challenges
Although deep learning models such as LSTM, GRU, Transformers, and Graph Neural Networks have demonstrated strong predictive capabilities, several limitations remain [2,5]. One major challenge is overfitting, in which models may learn noise or spurious correlations from the training data, reducing their performance on unseen data; regularisation techniques and careful model design are required to mitigate this issue [2,15]. Additionally, many deep learning architectures operate as black-box models, making it difficult to interpret their internal decision-making processes [18]. Explainable AI (XAI) approaches, including SHAP-based feature attribution and attention-based analyses, have been explored to improve the interpretability of deep learning models [18,19]. However, attention-based analyses should be interpreted cautiously because attention weights do not necessarily constitute faithful explanations of model decisions [67,68]. For instance, an interpretable ANN-SHAP framework developed for the Poyang Lake Basin successfully integrated multiple data sources while providing transparent predictions of key water quality parameters [31].
Advanced architectures such as Transformers and Graph Neural Networks often require high computational resources and large datasets for effective training, which may limit their applicability in resource-constrained environments [2,15]. Furthermore, purely data-driven models may sometimes produce physically unrealistic outputs if environmental constraints are not explicitly incorporated into the learning process [21,22]. Physics-informed neural networks (PINNs) have been proposed as one approach to address this limitation by embedding physical laws directly into the model architecture [21,22].
6.3. Operational and Deployment Challenges
In practical applications, real-time water quality forecasting relies on robust data pipelines, dependable sensor networks, and continuous data-streaming infrastructure [25,37]. These systems are often challenging to maintain in real-world environmental monitoring, especially in developing regions with limited technical capacity [7,75]. Additionally, many models are tailored to specific sites and do not readily transfer across different regions due to variations in hydrological and environmental conditions, necessitating frequent retraining or recalibration in new locations. Transfer learning approaches, as discussed in Section 7.4, represent a promising mitigation strategy for this limitation [2,23]. Furthermore, ongoing model retraining is crucial to accommodate changing environmental conditions; without this, performance can decline over time due to evolving water system dynamics [2,15]. Handling non-stationarity requires evaluation and updating strategies that preserve temporal structure. Rolling-origin or blocked cross-validation can provide a more realistic assessment than random partitioning by repeatedly evaluating models on later times while restricting training to information available before each evaluation period. Concept-drift detection can further be used to identify sustained changes in the data-generating process and determine when model performance or input distributions have changed sufficiently to warrant intervention. Adaptive or continual retraining can then incorporate more recent observations without relying indefinitely on a static training distribution. Within the reviewed evidence base, the continual-learning framework reported by [76] provides an example of model adaptation to evolving water-quality observations, explicitly motivated by the dynamic nature of water-quality data and incorporating mechanisms to refine local models in response to new data. Wang et al. [39] applied a chronological (not random) split for their transfer-learning evaluation and separately compared chronological versus random partitioning for their main model, finding a modest 5% NSE increase under random partitioning. Direct evidence of the temporal-leakage risk that rolling-origin approaches are designed to avoid, though this does not constitute a rolling-origin or blocked cross-validation procedure in itself. We did not identify a primary study in the reviewed corpus that explicitly implemented rolling-origin or blocked cross-validation as a principal evaluation framework, nor a dedicated concept-drift detection procedure. These approaches therefore represent important methodological directions for strengthening future water-quality forecasting studies rather than established practices in the current evidence base. The integration of deep learning models with IoT sensor networks and edge computing platforms introduces additional engineering challenges related to latency, data synchronisation, and hardware constraints [25,26].
6.4. Limitations of the Review
Several methodological limitations of this review are acknowledged. First, as with many systematic reviews in applied machine learning, the available evidence may be influenced by publication bias, whereby studies reporting positive or statistically significant improvements are more likely to be published than studies reporting neutral or negative results. Consequently, the apparent superiority of hybrid and attention-enhanced models reported in the literature may partially reflect the characteristics of the published evidence rather than the complete body of conducted research.
Second, the heterogeneity of datasets and evaluation protocols across studies complicates direct quantitative comparisons. The studies reviewed vary widely in terms of water quality parameters, spatial and temporal resolutions, forecast horizons, environmental settings such as river basins, aquaculture systems, reservoirs, and performance metrics. While we have summarized these findings to illustrate the range of reported performance, the metrics should not be interpreted as universal benchmarks or direct comparisons. This heterogeneity also precludes the construction of a valid pooled performance benchmark or cross-study metric normalization; accordingly, this review synthesizes evidence by architecture family and reported context rather than through a unified numerical ranking. Formal meta-analytic synthesis such as pooled effect sizes or forest-plot summaries was considered but judged not methodologically defensible given the reported evidence. Meta-analytic pooling requires either a shared outcome metric or a validated cross-metric conversion, together with study-level variance or standard-error estimates for appropriate weighting. However, the 37 primary studies do not consistently report associated uncertainty estimates such as standard errors or confidence intervals, while RMSE and MAE values are not directly comparable across studies predicting different indicators with different natural scales and units. Consequently, this review adopts a structured narrative synthesis organized by architecture family and reported study context, rather than a quantitative pooled synthesis that would imply a level of cross-study comparability not supported by the underlying evidence.
Third, the geographic concentration of studies predominantly in East and Southeast Asia, particularly China and Malaysia, limits the generalizability of our conclusions to data-scarce regions such as Africa and South America (see Table 6). The geographically attributed studies in the reviewed evidence are concentrated in East Asia and Southeast Asia, with comparatively limited representation from Europe and Africa and no geographically attributed study from South America in Table 6. This geographic concentration may also compound the publication-bias limitation noted above, because the available evidence may not fully represent studies conducted in regions with more limited monitoring infrastructure and data availability. Then, the frequently reported strong performance of hybrid and attention-enhanced architectures should not be generalized to underrepresented regions without local validation and adaptation. While the findings indicate promising directions, they should be interpreted with caution when applied to new contexts without appropriate validation and local adaptation.
Fourth, the included studies did not consistently report comparisons with naive or simple baseline models, such as persistence or climatological forecasts. Because baseline selection varied across studies and was not systematically coded in the present review, the reported superiority of deep learning and hybrid architectures should not be interpreted as evidence of improvement over a common naive benchmark. This eligibility audit represents a refinement of the original screening process, ensuring that the final count of included primary studies reflects strict and consistent application of the stated inclusion and exclusion criteria.
Fifth, formal statistical significance testing of predictive differences between competing models is largely absent from the reviewed evidence base. The Diebold–Mariano (DM) test provides an established framework for comparing the predictive accuracy of two forecasts using their forecast-error sequences when competing models are evaluated on the same forecasting task [73]. However, none of the 37 primary studies included in our evidence base reported use of the DM test or an equivalent significance test. Consequently, many reported performance improvements over baseline models, including the percentage-based gains summarized throughout this review, reflect differences in point estimates rather than confirmed statistical significance. This represents a methodological gap in the current water-quality forecasting literature rather than a shortcoming specific to any individual study.
Sixth, reported performance differences between architectures may also reflect implementation-level factors, including hyperparameter tuning, training data volume, feature engineering, and preprocessing choices, rather than architecture alone. This concern is illustrated in the reviewed literature, where practical comparative implementations show that predictive performance can vary with model-configuration and data-imputation choices [15]. Because the reviewed primary studies do not consistently report sufficient experimental detail to isolate architectural contribution from these implementation factors, comparative performance statements throughout this review should be interpreted as literature-reported outcomes under study-specific experimental conditions rather than as controlled estimates of architectural superiority.
An additional limitation concerns the use of formal statistical inference in the primary literature. The reviewed studies do not provide a sufficiently consistent inferential framework for establishing statistically significant differences between architectures. Accordingly, comparisons among model families in this review should be interpreted as a structured descriptive synthesis of reported performance rather than as formal statistical evidence of architectural superiority. This limitation is particularly relevant because differences in reported predictive performance may also reflect differences in datasets, forecasting horizons, preprocessing, feature engineering, and model-tuning procedures.
Meaningful comparison of forecasting architectures requires evaluation conditions that minimize methodological confounding. Ideally, competing models should be evaluated using identical temporal splits and comparable hyperparameter-tuning budgets, with the test period kept strictly independent of model selection. Baseline selection is also important: persistence or climatological forecasts provide simple reference points, while regularized regression and gradient-boosting models can provide stronger classical machine-learning benchmarks. However, the reviewed literature exhibits substantial heterogeneity in baseline selection, data partitioning, and the reporting of model-selection procedures. As noted above, baseline reporting is not consistent across studies, while differences in hyperparameter tuning can also confound apparent architectural advantages. These limitations make it difficult to attribute reported performance improvements solely to the underlying deep-learning architecture. Because the final search was completed on 30 June 2026, publications indexed or made available in the searched databases after this cutoff are not represented. The 2026 evidence base should therefore be interpreted as incomplete for the full calendar year, and future updates of the review should extend the search beyond this cutoff.
An additional limitation concerns the screening process itself: eligibility assessment was conducted primarily by the lead author against the stated inclusion and exclusion criteria, with supervisory guidance at points during the process, rather than through a prospectively documented, independent dual-screening protocol with calculated inter-rater agreement. We mitigated this limitation through explicit inclusion/exclusion criteria and a subsequent detailed eligibility audit that reapplied these criteria across all retained references. Future systematic reviews in this area would benefit from independent dual screening with reported inter-rater reliability.
A formal appraisal of temporal-split validity and tuning-protocol reporting across the 37 primary studies (Supplementary Table, “Risk Classification” sheet) found 1 study at high risk of temporal data leakage (a random, non-chronological holdout split applied to temporal data), 5 studies at low risk (explicit chronological or rolling-forecast validation), and 31 studies where split methodology was not reported with sufficient detail to classify risk either way. This high proportion of unclassifiable studies is a notable finding, indicating that most primary studies in this literature do not report sufficient methodological detail for readers to independently assess the risk of temporal data leakage, a pervasive and under-scrutinized threat to validity in this field. Table 8 shows the challenges of current deep learning models for water quality forecasting.
Table 8.
Challenges of current deep learning models for water quality forecasting.
7. Future Research Directions
Despite the significant progress reviewed in preceding sections, several important research directions remain underexplored or insufficiently addressed in the current literature. This section identifies five priority areas for future investigation, drawing on evidence from the reviewed studies and newly assessed reference candidates.
7.1. Standardized Benchmark Datasets
Advancement in water quality forecasting depends on developing standardised benchmark datasets and evaluation protocols [1,2]. Unlike fields such as computer vision and natural language processing, where shared benchmarks allow direct comparison of methods, water quality research often relies on localised or proprietary datasets, which limit reproducibility and fair evaluation of models [1]. The absence of standardised datasets makes it difficult to distinguish methodological improvements from dataset-specific biases, leading to inconsistent conclusions across studies [1,2].
Well-defined benchmark datasets with comprehensive temporal and spatial coverage improve reproducibility and enable consistent comparisons between models trained under the same conditions [52]. Effective benchmark datasets should include multiple hydrological and meteorological variables alongside water quality parameters [1]. Standardised train–test splits and unified evaluation metrics are also essential for comparability across studies [1]. Data interoperability, achieved through standardised units, consistent formatting, and thorough metadata documentation, minimises preprocessing errors and enhances data usability across different models [2,25]. Additionally, open data initiatives led by national and international organisations support reproducibility and large-scale benchmarking [1,2]. Additionally, open data initiatives led by national and international organisations support reproducibility and large-scale benchmarking [1,2]. Alongside standardized datasets and evaluation protocols, future comparative studies should incorporate formal statistical significance testing, such as the Diebold–Mariano test, when evaluating competing models on a common forecasting task, to complement point-estimate performance metrics with a rigorous basis for claims of superiority.
7.2. Explainable Artificial Intelligence (XAI) Tools
As deep learning models become more complex, interpretability becomes crucial for widespread adoption and trust in environmental decision-making systems [18]. In water quality applications, model interpretability is especially important because predictions directly impact regulatory actions and public health decisions [18,71]. Attention-based architectures offer significant promise for explainability by assigning weights to input features and time steps, thereby highlighting influential variables in prediction processes [19,59]. Feature attribution methods such as SHAP values measure the contribution of each input variable to the model’s output, enabling identification of key environmental drivers such as temperature, flow rate, and nutrient levels [18,19,71].
A recent study by Karahan et al. employed an LSTM model used not primarily for predictive forecasting but as a data-driven framework together with SHAP to interpret the drivers of electrical conductivity and salinity in a Flemish river system. Although predictive performance was limited due to dynamics not represented in the training data, SHAP analysis revealed physically consistent feature influences, including neighbouring sensors, discharge, and temperature, illustrating how XAI can extract interpretable insight even from a model with constrained forecasting accuracy [19]. Although XAI methods improve transparency, they do not fully eliminate the black-box nature of deep learning models, and ongoing development of domain-specific interpretability tools remains necessary [18].
7.3. Physics-Informed Neural Networks (PINNs)
Physics-Informed Neural Networks (PINNs) incorporate governing physical laws directly into the neural network training process to ensure that predictions remain aligned with real-world hydrological processes [21,22]. This approach reduces dependence on purely data-driven learning and enhances model generalisation by guiding learning toward physically plausible solutions [21,22]. In water quality modelling, PINNs embed fundamental physical principles such as mass conservation, transport dynamics, and reaction kinetics [21,78]. By penalising violations of these constraints within the loss function, PINNs reduce the likelihood of physically implausible predictions, especially in extrapolation scenarios where observational data are limited [22,61]. PINNs are particularly effective in environments with limited data, where purely data-driven models are prone to overfitting [21,22]. Nonetheless, selecting suitable physical constraints and balancing them with data-driven objectives remains a key research challenge [21,22]. This difficulty is illustrated within the reviewed evidence by Frankel et al. [78], who developed a PINN framework for drinking-water disinfectant residuals in the presence of imperfect reaction models and partial data. Their study notes that laboratory-calibrated kinetic models for chlorine and monochloramine decay may not fully capture additional reactions occurring under actual environmental conditions, highlighting the difficulty of specifying a complete governing reaction system a priori. In such settings, PINN design requires careful consideration of how available physical knowledge should be incorporated alongside observational data when the underlying process is incompletely characterized. Recent studies highlight the potential of hybrid approaches that combine physics-based modelling with advanced deep learning architectures to enhance performance in complex environmental systems [22,78].
7.4. Transfer and Few-Shot Learning
Limited availability of labeled water quality data limits the performance of deep learning models trained from scratch [3,50]. Transfer learning overcomes this by pretraining models on large datasets and fine-tuning them on smaller local datasets, which improves generalization across different locations [46]. This approach allows models to learn broad hydrological representations from large-scale datasets and adapt to local environmental conditions with fewer training samples [3,46]. Studies indicate that transfer learning consistently outperforms models trained solely on small local datasets for predicting water quality [2]. Recent innovations include Aquaformer, a transformer-based model that integrates phase space reconstruction with multi-source transfer learning to address data scarcity in long-sequence water quality forecasting [23]. Few-shot learning builds on this idea by enabling model adaptation with only a small number of labeled samples, making it particularly useful in regions with limited monitoring [53,60]. Meta-learning-based methods further boost adaptability across various hydrological systems by learning how to learn from limited data [2].
7.5. Edge Computing and IoT Integration
The deployment of Internet of Things (IoT) sensors in water monitoring systems enables continuous collection of environmental data and supports near-real-time forecasting [25,44,79]. IoT-based smart monitoring systems have also demonstrated particular value in aquaculture environments, where automated machine learning pipelines support real-time water quality management and productivity optimization [44,80]. However, traditional centralised cloud-based processing can introduce latency and often relies on stable network connectivity [25]. Edge computing addresses these issues by running predictive models directly on sensor nodes or local gateways, enabling low-latency and immediate local decision-making [25]. This approach is especially important for early detection of water quality issues and emerging environmental threats [25]. To enable deployment on resource-limited devices, model optimisation techniques such as model compression, quantisation, and knowledge distillation are necessary [66]. Additionally, edge computing-enabled systems enhance overall resilience, as local models can continue functioning even during communication outages [25]. Federated learning frameworks, which allow collaborative model training across distributed sensor nodes without sharing raw data, represent a complementary approach that addresses both privacy and connectivity constraints [35,56,77]. Advances in federated continual learning further extend this paradigm, enabling IoT-based monitoring systems to adapt incrementally to new environmental conditions without retraining from scratch [76].
8. Conclusions
Recent advances in water quality monitoring demonstrate the potential of deep learning for precise and adaptable forecasting. Traditional statistical and process-based models, while useful, depend on simplified assumptions and might not fully capture nonlinear dynamics. In contrast, deep learning methods can learn spatiotemporal relationships directly from data, often providing competitive predictive accuracy and flexibility.
Various architectures, such as convolutional neural networks, recurrent networks, and transformers, excel in different applications. Convolutional models are effective for recognizing spatial patterns, while recurrent and attention-based models are better suited for understanding temporal dependencies. Hybrid models, like CNN-LSTM architectures, combine feature extraction with sequential modeling, addressing both short-term fluctuations and long-term trends, with strong performance reported in multiple studies under specific conditions.
However, increased model complexity presents challenges, including higher data and computational requirements, limited interpretability, and difficulties in generalization. Proper model evaluation using error metrics, appropriate statistical testing, and uncertainty analysis is important for assessing predictive reliability and distinguishing methodological improvements from dataset-specific effects.
Despite advances, persistent issues such as data sparsity, non-stationarity, extreme events, model opacity, and scalability remain. Tackling these problems will require standardized datasets, explainable and physics-informed models, and strategies to improve model transferability. In summary, deep learning offers significant promise for water quality forecasting and proactive environmental management. Ongoing research emphasizing robustness, transparency, and operational integration will be vital for transforming these models into dependable tools for policy decisions, public health, and ecosystem management.
Supplementary Materials
The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/environments13090515/s1, Table S1: Final Verified Primary Studies Included in the Systematic Review [78,79,80].
Author Contributions
Conceptualization, A.A., W.Z.W.I. and N.A.A.A.; methodology, W.Z.W.I. and A.A.; writing original draft preparation, W.Z.W.I. and A.A.; writing review and editing, W.Z.W.I., O.A.A., A.K.G. and N.A.A.A.; supervision, W.Z.W.I. and N.A.A.A.; funding acquisition W.Z.W.I. and N.A.A.A. All authors have read and agreed to the published version of the manuscript.
Funding
This research was funded by a grant from the Ministry of Higher Education (MOHE), Malaysia under the Fundamental of Research Grant Scheme (FRGS/1/2024/WAS02/USIM/02/1) and Universiti Sains Islam Malaysia (USIM/MG/AAU/FKAB/SEPADAN-A/70726). The Article Processing Charge (APC) is funded by Multimedia University.
Data Availability Statement
No new data were created or analyzed in this study. Data sharing is not applicable to this article.
Acknowledgments
We would like to acknowledge the support given by Universiti Sains Islam Malaysia and Multimedia University towards this project. The authors have reviewed and edited the output and take full responsibility for the content of this publication.
Conflicts of Interest
The authors declare no conflicts of interest. The funders had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results. All authors have read and agreed to the published version of the manuscript.
References
- Chidiac, S.; El Najjar, P.; Ouaini, N.; El Rayess, Y.; El Azzi, D. A comprehensive review of water quality indices (WQIs): History, models, attempts and perspectives. Rev. Environ. Sci. Biotechnol. 2023, 22, 349–395. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhi, W.; Appling, A.P.; Golden, H.E.; Podgorski, J.; Li, L. Deep learning for water quality. Nat. Water 2024, 2, 228–241. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Department of Environment Malaysia. Malaysia Environmental Quality Report 2020; Department of Environment Malaysia: Putrajaya, Malaysia, 2020. [Google Scholar]
- Lokman, A.; Ismail, W.Z.W.; Aziz, N.A.A.; Ghazali, A.K. Water quality index (WQI) forecasting and analysis based on neuro-fuzzy and statistical methods. Appl. Sci. 2025, 15, 9364. [Google Scholar] [CrossRef] [Scilit]
- Zhu, M.; Wang, J.; Yang, X.; Zhang, Y.; Zhang, L.; Ren, H.; Wu, B.; Ye, L. A review of the application of machine learning in water quality evaluation. Eco-Environ. Health 2022, 1, 107–116. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pyo, J.; Pachepsky, Y.; Kim, S.; Abbas, A.; Kim, M.; Kwon, Y.S.; Ligaray, M.; Cho, K.H. Long short-term memory models of water quality in inland water environments. Water Res. X 2023, 21, 100207. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sit, M.; Demiray, B.Z.; Xiang, Z.; Ewing, G.J.; Sermet, Y.; Demir, I. A comprehensive review of deep learning applications in hydrology and water resources. Water Sci. Technol. 2020, 82, 2635–2670. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Van Houdt, G.; Mosquera, C.; Nápoles, G. A review on the long short-term memory model. Artif. Intell. Rev. 2020, 53, 5929–5955. [Google Scholar] [CrossRef] [Scilit]
- Barzegar, R.; Aalami, M.T.; Adamowski, J. Short-term water quality variable prediction using a hybrid CNN–LSTM deep learning model. Stoch. Environ. Res. Risk Assess. 2020, 34, 415–433. [Google Scholar] [CrossRef] [Scilit]
- Zhang, M.; Zhang, Z.; Wang, X.; Liao, Z.; Wang, L. The use of attention-enhanced CNN-LSTM models for multi-indicator and time-series predictions of surface water quality. Water Resour. Manag. 2024, 38, 6103–6119. [Google Scholar] [CrossRef] [Scilit]
- Li, D.; Ji, X.; Liu, L. An accurate forecasting model for key water quality factors based on transformer with multi-scale attention mechanism. Environ. Model. Softw. 2025, 191, 106491. [Google Scholar] [CrossRef] [Scilit]
- Wu, Z.; Pan, S.; Chen, F.; Long, G.; Zhang, C.; Yu, P.S. A comprehensive survey on graph neural networks. IEEE Trans. Neural Netw. Learn. Syst. 2021, 32, 4–24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Salem, A.K.; Taha, A.F.; Abokifa, A.A. Graph neural networks-based dynamic water quality state estimation in water distribution networks. Eng. Appl. Artif. Intell. 2024, 138, 109426. [Google Scholar] [CrossRef] [Scilit]
- Nikoo, M.R.; Al Aamri, A.H.S.; Etri, T.; Al-Rawas, G.A.; Nazari, R. A review of machine learning, remote sensing, and statistical methods for reservoir water quality assessment. J. Hydrol. 2025, 659, 133323. [Google Scholar] [CrossRef] [Scilit]
- Helaly, M.A.; Rady, S.; Mabrouk, M.; Aref, M.M.; Villarroya, S.; Cotos, J.M.; Mera, D. Advancements in water quality prediction: A practical review of machine learning and deep learning approaches. Clust. Comput. 2025, 28, 598. [Google Scholar] [CrossRef] [Scilit]
- Utku, A.; Utku, E.D.; Kutlu, B. Deep learning based an effective hybrid model for water quality assessment. Water Environ. Res. 2023, 95, e10929. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Muñoz-Alegría, J.A.; Núñez, J.; Oyarzún, R.; Chávez, C.A.; Arumí, J.L.; Rodríguez-López, L. A bibliometric-systematic literature review (B-SLR) of machine learning-based water quality prediction: Trends, gaps, and future directions. Water 2025, 17, 2994. [Google Scholar] [CrossRef] [Scilit]
- Arrieta, A.B.; Díaz-Rodríguez, N.; Del Ser, J.; Bennetot, A.; Tabik, S.; Barbado, A.; García, S.; Gil-López, S.; Molina, D.; Benjamins, R.; et al. Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Inf. Fusion 2020, 58, 82–115. [Google Scholar] [CrossRef] [Scilit]
- Karahan, S.M.; Vandenbruwaene, W.; Alferes, J.; Verwaeren, J. Explainable AI for aquatic environmental intelligence: A SHAP-enhanced LSTM approach using high-frequency water quality data in a river system. J. Hydroinform. 2025, 27, 1918–1928. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Y.; Liu, X.; Wang, H.; Chen, Z.; Li, J. A spatial-temporal trend-aware neural network model for accurate water quality prediction in river. Water Res. 2025, 285, 124389. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Raissi, M.; Perdikaris, P.; Karniadakis, G.E. Physics-informed neural networks: A deep learning framework for solving forward and inverse problems involving nonlinear partial differential equations. J. Comput. Phys. 2019, 378, 686–707. [Google Scholar] [CrossRef] [Scilit]
- Karniadakis, G.E.; Kevrekidis, I.G.; Lu, L.; Perdikaris, P.; Yang, L. Physics-informed machine learning. Nat. Rev. Phys. 2021, 3, 422–440. [Google Scholar] [CrossRef] [Scilit]
- Sun, M.; Xu, C.; Jia, Q.; Jia, H. Aquaformer: Multi-source transfer learning model based on transformer and phase space reconstruction for long sequence water quality forecasting. J. Hydrol. 2025, 664, 134372. [Google Scholar] [CrossRef] [Scilit]
- Derdour, A.; Baz, M.; Alzaed, A.; Bojer, A.K.; Ghoneim, S.S.M. Groundwater quality assessment using few-shot learning with prototypical, Siamese, and matching networks. J. Water Process Eng. 2025, 75, 108003. [Google Scholar] [CrossRef] [Scilit]
- Bandara, R.M.P.N.S.; Jayasinghe, A.B.; Retscher, G. The integration of IoT sensors and location-based services for water quality monitoring: A systematic literature review. Sensors 2025, 25, 1918. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yadav, H.; Murugan, V.; Vijayakanthan, N.A.K.J.N.; Mathurkar, P. Integrating edge computing and IoT for real-time air and water quality monitoring systems. Int. J. Environ. Sci. 2025, 11, 1507–1511. [Google Scholar] [CrossRef] [Scilit]
- Lokman, A.; Ismail, W.Z.W.; Aziz, N.A.A. Water quality evaluation and analysis by integrating statistical and machine learning approaches. Algorithms 2025, 18, 494. [Google Scholar] [CrossRef] [Scilit]
- Uddin, M.G.; Nash, S.; Olbert, A.I. A sophisticated model for rating water quality. Sci. Total Environ. 2023, 868, 161614. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lokman, A.; Ismail, W.Z.W.; Aziz, N.A.A.; Ghazali, A.K. Interactive Fuzzy Logic Interface for Enhanced Real-Time Water Quality Index Monitoring. Algorithms 2025, 18, 591. [Google Scholar] [CrossRef] [Scilit]
- Baek, S.-S.; Pyo, J.; Chun, J.A. Prediction of water level and water quality using a CNN-LSTM combined deep learning approach. Water 2020, 12, 3399. [Google Scholar] [CrossRef] [Scilit]
- Yuan, Y.; Wu, J.; Deng, F.; Liu, W.; Sun, M.; Li, L. An interpretable deep learning framework for river water quality prediction—A case study of the Poyang Lake Basin. Water 2025, 17, 2496. [Google Scholar] [CrossRef] [Scilit]
- Liu, L.; Ji, X.; Li, Y.; Xing, H.; Wang, B.; Li, D. A long-term prediction model for key water quality based on transformer with parallel attention mechanism and adaptive spectral enhancement. Environ. Geochem. Health 2025, 47, 322. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wan, H.; Xiang, L.; Cai, Y.; Xie, Y.; Xu, R. Temporal and spatial feature extraction using graph neural networks for multi-point water quality prediction in river network areas. Water Res. 2025, 281, 123561. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hao, Z. A dissolved oxygen prediction model based on GRU–N-Beats. Front. Mar. Sci. 2024, 11, 1365047. [Google Scholar] [CrossRef] [Scilit]
- Das, B.; Adel, A.; Jan, T.; Wahiduzzaman, M. Water quality management using federated deep learning in developing southeastern Asian country. Water Resour. Manag. 2024, 39, 1893–1909. [Google Scholar] [CrossRef] [Scilit]
- Karbasi, M.; Ali, M.; Bateni, S.M.; Jun, C.; Jamei, M.; Farooque, A.A.; Yaseen, Z.M. Multi-step ahead forecasting of electrical conductivity in rivers by using a hybrid convolutional neural network-long short-term memory (CNN-LSTM) model enhanced by Boruta-XGBoost feature selection algorithm. Sci. Rep. 2024, 14, 15051. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mohammadi, M.; Assaf, G.; Assaad, R.H.; Chang, J. An intelligent cloud-based IoT-enabled multimodal edge sensing device for automated, real-time, comprehensive, and standardized water quality monitoring and assessment process using multisensor data fusion technologies. J. Comput. Civ. Eng. 2024, 38, 04024003. [Google Scholar] [CrossRef] [Scilit]
- Yuan, M.; Li, Y.; Zhang, L.; Zhao, W.; Li, J. Rapid prediction approach for water quality in plain river networks: A data-driven water quality prediction model based on graph neural networks. Water 2025, 17, 2543. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Shi, B.; Osmond, P.; Zhang, K. Multimodal deep-learning–driven water-quality forecasting and scenario assessment: Shifting from reactive monitoring to proactive basin management. J. Hydrol. 2026, 675, 135554. [Google Scholar] [CrossRef] [Scilit]
- Huan, J.; Zhang, C.; Qian, Y.; Zhang, H.; Fa, Y. River water quality forecasting: A novel LSTM-transformer approach enhanced by multi-source data fusion. Environ. Monit. Assess. 2025, 197, 336. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, Y.; Han, J.C.; Zhou, Y.; Wang, X.; Ikram, R.M.A.; Xu, Z.; Huang, Y. Application of novel deep learning model for high-frequency water quality prediction in multi-station river networks. J. Environ. Chem. Eng. 2025, 13, 119673. [Google Scholar] [CrossRef] [Scilit]
- Tian, Q.; Luo, W.; Guo, L. Water quality prediction in the Yellow River source area based on the DeepTCN-GRU model. J. Water Process Eng. 2024, 59, 105052. [Google Scholar] [CrossRef] [Scilit]
- Laha, S.R.; Pattanayak, B.K.; Kumar, S.; Ray, M.; Pattnaik, S. IoT-enabled machine learning for comprehensive water quality assessment in the Mahanadi River: A multibelt analysis of seasonal contamination and predictive modeling. J. Eng. 2025, 2025, 5549990. [Google Scholar] [CrossRef] [Scilit]
- Chandran, P.J.I.; Khalil, H.A.; Hashir, P.K.; Veerasingam, S. Smart technologies in aquaculture: An integrated IoT, AI, and blockchain framework for sustainable growth. Aquac. Eng. 2025, 111, 102584. [Google Scholar] [CrossRef] [Scilit]
- Pan, D.; Deng, Y.; Yang, S.X.; Gharabaghi, B. Recent advances in remote sensing and artificial intelligence for river water quality forecasting: A review. Environments 2025, 12, 158. [Google Scholar] [CrossRef] [Scilit]
- Biraghi, C.A.; Lotfian, M.; Carrion, D.; Brovelli, M.A. AI in support to water quality monitoring. Int. Arch. Photogramm. Remote Sens. Spat. Inf. Sci. 2021, XLIII-B4-2021, 167–174. [Google Scholar] [CrossRef] [Scilit]
- Liu, Y.; Zheng, H.; Zhao, J. Enhanced water quality prediction by LSTM and graph attention network (L-GAT): An analytical study of the Pearl River Basin. Water Res. X 2025, 28, 100383. [Google Scholar] [CrossRef] [Scilit]
- Wu, H.; Xu, J.; Wang, J.; Long, M. Autoformer: Decomposition transformers with auto-correlation for long-term series forecasting. Adv. Neural Inf. Process. Syst. 2021, 34, 22419–22430. [Google Scholar] [CrossRef] [Scilit]
- Seshan, S.; de Vries, D.; Immink, J.N.; van der Helm, A.; Poinapen, J. LSTM-based autoencoder models for real-time quality control of wastewater treatment sensor data. J. Hydroinform. 2024, 26, 441–458. [Google Scholar] [CrossRef] [Scilit]
- Sakurada, M.; Yairi, T. Anomaly detection using autoencoders with nonlinear dimensionality reduction. In Proceedings of the MLSDA 2014 2nd Workshop on Machine Learning for Sensory Data Analysis, Gold Coast, Australia, 2 December 2014; pp. 4–11. [Google Scholar] [CrossRef] [Scilit]
- Roukerd, F.R.; Rajabi, M.M. Anomaly detection in groundwater monitoring data using LSTM-autoencoder neural networks. Environ. Monit. Assess. 2024, 196, 692. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wai, K.P.; Koo, C.H.; Huang, Y.F.; Chong, W.C.; El-Shafie, A.; Sherif, M.; Ahmed, A.N. A practical temporal transfer learning model for multi-step water quality index forecasting using a CNN-coupled dual-path LSTM network. J. Hydrol. Reg. Stud. 2026, 60, 102553. [Google Scholar] [CrossRef] [Scilit]
- Zheng, Y.; Zhang, X.; Zhou, Y.; Zhang, Y.; Zhang, T.; Farmani, R. Deep representation learning enables cross-basin water quality prediction under data-scarce conditions. npj Clean Water 2025, 8, 33. [Google Scholar] [CrossRef] [Scilit]
- Xu, X.; Chen, G.; Xu, H.; Jia, F. WT-DSE-LSTM: A hybrid model for the multivariate prediction of dissolved oxygen. Alex. Eng. J. 2025, 124, 285–296. [Google Scholar] [CrossRef] [Scilit]
- Torkan, M.S.; Hekmatiyan, A.; Zamani, M.G. An integrated VMD–CNN–GRU framework for multi-horizon forecasting of water quality dynamics. Water Resour. Manag. 2026, 40, 104. [Google Scholar] [CrossRef] [Scilit]
- Rejula, M.A.; Minija, S.J.; Sophia, S.; Barakka, J.A. Decentralized water quality classification using federated learning with recurrent neural networks. Water Qual. Res. J. 2024, 60, 135–150. [Google Scholar] [CrossRef] [Scilit]
- Lokman, A.; Ismail, W.Z.W.; Aziz, N.A.A. Enhancing water quality index prediction accuracy in Mranti Lake and rivers in Malaysia using regression forest model. Appl. Water Sci. 2026, 16, 34. [Google Scholar] [CrossRef] [Scilit]
- Cho, K.; van Merrienboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning phrase representations using RNN encoder-decoder for statistical machine translation. arXiv 2014, arXiv:1406.1078. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar] [CrossRef] [Scilit]
- Bi, J.; Wang, Z.; Wu, X.; Wu, R.; Zhang, J.; Zhou, M.C. Long-term water quality prediction with transformer-based spatial-temporal graph fusion. IEEE Trans. Autom. Sci. Eng. 2025, 22, 11392–11404. [Google Scholar] [CrossRef] [Scilit]
- Mu, T.; Liu, X.; Wang, H.; Chen, Z.; Li, J. ST-GPINN: A spatiotemporal graph physics-informed neural network for enhanced water quality prediction in water distribution systems. npj Clean Water 2025, 8, 25. [Google Scholar] [CrossRef] [Scilit]
- Yan, M.; Wang, Z. Water quality prediction method based on reinforcement learning graph neural network. IEEE Access 2024, 12, 184421–184430. [Google Scholar] [CrossRef] [Scilit]
- Tu, J.; Nie, Z.; Neculita, M.; Fortea, C.; Antohi, V.M.; Meshram, S.G. Exploring advanced hybrid approaches in Hong Kong rivers for accurate prediction of surface water quality using CNN-LSTM-GRU model. Appl. Water Sci. 2026, 16, 124. [Google Scholar] [CrossRef] [Scilit]
- Ding, X.F.; Chen, Y.L.; Zeng, H.P.; Du, Y. Time series prediction of water quality based on NGO-CNN-GRU model—A case study of Xijiang River, China. Water 2025, 17, 2413. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Zhao, Y.; Zhu, Y.; Dong, Z.; Wang, F.; Huang, F. Research progress in water quality prediction based on deep learning technology: A review. Environ. Sci. Pollut. Res. 2024, 31, 26415–26431. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dahane, A.; Benameur, R.; Souihi, S.; Naloufi, M.; Lucas, F.; Mellouk, A. IoT urban river water quality system using federated learning via knowledge distillation. In Proceedings of the ICC 2024-IEEE International Conference on Communications, Denver, CO, USA, 9–13 June 2024; pp. 1515–1520. [Google Scholar] [CrossRef] [Scilit]
- Jain, S.; Wallace, B.C. Attention is not explanation. In Proceedings of the 2019 Conference North American Chapter Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), Minneapolis, MN, USA, 2–7 June 2019; pp. 3543–3556. [Google Scholar] [CrossRef] [Scilit]
- Wiegreffe, S.; Pinter, Y. Attention is not not explanation. In Proceedings of the 2019 Conference Empirical Methods in Natural Language Processing and 9th International Joint Conference Natural Language Processing (EMNLP-IJCNLP), Hong Kong, China, 3–7 November 2019; pp. 11–20. [Google Scholar] [CrossRef] [Scilit]
- Wu, M.; Blancaflor, E.B. A hybrid deep learning model for water quality prediction: GS-EHHO-CNN-BiLSTM applied to the Yellow River Basin. Int. J. Comput. Commun. Control 2025, 20, 6908. [Google Scholar] [CrossRef] [Scilit]
- Gal, Y.; Ghahramani, Z. Dropout as a Bayesian approximation: Representing model uncertainty in deep learning. Proc. Mach. Learn. Res. (PMLR) 2016, 48, 1050–1059. [Google Scholar] [CrossRef] [Scilit]
- Karim, M.A.A.; Ismail, W.Z.W.; Shuib, F.M.M.; Aziz, N.A.A.; Ghazali, A.K. Water Quality Monitoring and Assessment Using Machine Learning: A Review of Formulation, Modeling Approaches, and Explainable Artificial Intelligence. Environments 2026, 13, 267. [Google Scholar] [CrossRef] [Scilit]
- Mamat, N.; Razali, S.F.M.; Hamzah, F.B. Enhancement of water quality index prediction using support vector machine with sensitivity analysis. Front. Environ. Sci. 2023, 10, 1061835. [Google Scholar] [CrossRef] [Scilit]
- Diebold, F.X.; Mariano, R.S. Comparing predictive accuracy. J. Bus. Econ. Stat. 1995, 13, 253–263. [Google Scholar] [CrossRef] [Scilit]
- Dilmi, S.; Ladjal, M. A novel approach for water quality classification based on the integration of deep learning and feature extraction techniques. Chemom. Intell. Lab. Syst. 2021, 214, 104329. [Google Scholar] [CrossRef] [Scilit]
- Ooko, S.O.; Cheptegei, L.; Karume, S.M. Application of machine learning for real-time water quality monitoring in developing countries: A review. Sustain. Futures 2025, 10, 100984. [Google Scholar] [CrossRef] [Scilit]
- Dahane, A.; Benameur, R.; Souihi, S.; Naloufi, M.; Benziane, I.B.; Lucas, F.; Mellouk, A. FCL-IWQMS: Federated continual learning and IoT-based water quality monitoring system for adaptive real-time insights. In Proceedings of the ICC 2025-IEEE International Conference on Communications, Montreal, QC, Canada, 8–12 June 2025; pp. 2955–2960. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Wang, H.C.; Wang, W.; Yang, H.; Chen, J.J.; Yin, W.X.; Lv, J.Q.; Luo, X.Q.; Zhou, X.; Wang, A.J. Federated machine learning enables risk management and privacy protection in water quality. Environ. Sci. Technol. 2025, 59, 10310–10322. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Frankel, M.; De Florio, M.; Schiassi, E.; Katz, L.E.; Kinney, K.A.; Sela, L. Enhancing drinking water quality modeling: Leveraging physics-informed neural networks for learning with imperfect reaction models and partial data. Environ. Sci. Water Res. Technol. 2025, 11, 2684–2697. [Google Scholar] [CrossRef] [Scilit]
- Miller, T.; Durlik, I.; Kostecka, E.; Kozlovska, P.; Łobodzińska, A.; Sokołowska, S.; Nowy, A. Integrating artificial intelligence agents with the internet of things for enhanced environmental monitoring: Applications in water quality and climate data. Electronics 2025, 14, 696. [Google Scholar] [CrossRef] [Scilit]
- Norzeri, M.N.N.; Yatin, S.F.M.; Samsudin, A.Z.H. EcoGuard: Advancing IoT-based aquaculture with machine learning for enhanced productivity and automation. J. ISMAC 2025, 8, 112–128. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.

