Next Article in Journal
Barriers to Green Economy in the Construction Industry in Ghana
Next Article in Special Issue
Assessing WQI Using Spatial Land-Use Context Derived from Google Earth Imagery and Advanced Convolutional Neural Networks in South Korea
Previous Article in Journal
Towards Automatic Burrow Detection for Sustainable River Levees
Previous Article in Special Issue
Performance Evaluation and Model Validation of Conventional Solar Still in Harsh Summer Climate: Case Study of Basrah, Iraq
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

AI Solutions for Improving Sustainability in Water Resource Management

by
Jorge Alejandro Silva
Escuela Superior de Comercio y Administración Unidad Santo Tomás, Instituto Politécnico Nacional, Mexico City 11350, Mexico
Sustainability 2026, 18(4), 2154; https://doi.org/10.3390/su18042154
Submission received: 25 December 2025 / Revised: 10 February 2026 / Accepted: 15 February 2026 / Published: 23 February 2026

Abstract

Water systems experience increasing sustainability challenges from climate variability, aging infrastructure, and energy and chemical intensity demands, but AI has typically been assessed against prediction accuracy rather than demonstrated operational success. This PRISMA 2020 systematic review analyzed the role of AI solutions on sustainability in distribution, treatment, and basin management. The database search identified 920 records; after deduplication (n = 185), screening was conducted on n = 735 titles/abstracts and examination of the full text for n = 85, providing a total of n = 41 included peer-reviewed studies for qualitative synthesis and n = 38 for quantitative/bibliometric synthesis with the additional analysis of seven grey-literature sources. Evidence mapping reveals high growth post-2020, and distribution and wastewater operations are dominated by a few companies. The most deployable evidence is found with monitoring, anomaly/leak detection, and short-term forecasting, while optimization and reinforcement-learning control are primarily simulation validated with limited field applications. While accuracy metrics are often reported, transformation into water saved, kWh/m3, chemicals, compliance/reliability/resilience/equity measures are inconsistently and less frequently operationalized. In general, AI is most believable when it is part of analysis-ready workflows, bounded decision support, and measurement-and-verification.

1. Introduction

Water resources are transitioning to a new normal where the past is an increasingly poor guide to the future. International evaluations have documented across-the-board adverse impacts from “too much or too little water”, including: widespread departures from averages in the quantities of water flowing in rivers, stored at reservoirs and groundwater stations, covering lake and glacier areas; and a call for clear enhancements in monitoring systems/data sharing to help manage risk under climate extremes [1]. Simultaneously, global policy debate underlines that water security is not only a biophysical problem but also one of development, equity, and peace as scarcity and variability heighten competition among users and reveal structural governance shortcomings [2].
Only limited advances have been made in achieving Sustainable Development Goal 6 (SDG 6), and they are insufficient at the current rate of progress to attain the related targets for 2030. Reporting on the UN’s SDG 6 highlights incremental progress in access to safely managed drinking water and sanitation during 2015–2022 but also emphasizes a lingering service divide between households and uneven progress across regions [3,4]. In the wastewater sector, the recent global update emphasizes that people must pick up the pace to reach SDG Target 6.3 and underscores how treatment coverage, operational efficiency, and surveillance are key for public health and protection of ecosystems [5]. The range of kinds of independent mid-term reviews for several SDG 6 indicators describes progress as “alarmingly off-track”; they caution that, in some respects, progress is either stalled or regressing—reinforcing a need for decision systems capable of handling uncertainty and synthetic prioritization [6].
Operational water resource management, however, involves linked decisions about allocation, storage, conveyance, treatment, quality assurance, and ecosystem maintenance—frequently among fragmented institutions and capital-intensive infrastructure that is challenging to instrument. Utilities often also incur ongoing technical losses, largely described as non-revenue water, which reduces supply and compromises financial sustainability, while minimizing the ability to reinvest [7]. These strains are now further amplified by rapid urbanization, increasingly stringent environmental expectations, and skyrocketing energy prices, all making the sustainability performance of water system resource efficiency, emissions, affordability, resilience, and fairness an increasingly explicit focus rather than simply a secondary co-benefit.
Within this context, AI has been promoted as a “key enabling layer” for sustainable water resources management as it (i) extracts actionable signals from heterogeneous data streams; (ii) enhances forecasting and situational awareness, and (iii) optimizes sequential decisions under uncertainty. The practical AI portfolio now encompasses smart sensing and anomaly detection in water distribution networks; predictive models for demand, leakage, and asset failure; data-informed and physics-calibrated hydrologic forecasting; operational optimization in pumping, storage, and treatment planning; and decision support systems that enable the handling of multi-objective trade-offs (e.g., cost/energy/water quality/service contour. Recent synthesis around water distribution also reflects the potential for AI-based leakage analytics to be successful at minimizing losses and mitigating economic and environmental influences of undiscovered leaks [8,9]. At a broader level, applied work highlights different cases of machine learning applications: infrastructures, demand prediction [10], and water quality monitoring, and it notes the persistent challenges in data quality, transferability, as well as operational deployment at scale.
Another rapidly expanding frontier is the “smart operations” layer, where AI is linked to simulators or virtual models of infrastructure and basins. Digital twins, which correspond to both component-level and system-wide models, are being developed for scenario evaluation, failure anticipation, and operations control. A recent review in the water sector enumerated digital twin use cases, common architectures, and challenges pertaining to data integration, governance, and lifecycle of maintenance [11]. Simultaneously, reinforcement learning (RL) is also considered for sequential control of reservoirs and urban water systems by finding policies that operate under uncertain inflows, conflicting objectives, and hard constraints. A systematic review of RL in the context of water management states that numerous RL agents are trained only in simulations but not directly on data from historical operations and emphasizes explainability, since decisions about water are safety-critical [12].
A third thread relates to the scientific integrity and robustness of AI models across changing climate and hydrological conditions. For water applications, pure data-driven models can deteriorate when they are extrapolated from training conditions, and this renders hybrid models that incorporate physical structure or constraints preferable. Hydrology (and Earth system modeling) are increasingly concerned with differentiable, physics-informed approaches that advance gradient-based learning and enable us to more tightly couple physically based explanations and machine learning [13]. Sustained efforts targeting such advances are of direct relevance to sustainability, as they determine how well AI-driven decision-making remains dependable in the face of extremes, regime shifts, and regions with sparse data, all of which disproportionately affect marginal users and ecosystems.
But the case for the sustainability of AI in water management is not a slam dunk. One, AI systems can create new vulnerabilities: biased prioritization (say, for areas that have better sensors), opaque operational decisions, and “automation complacency” in mission-critical infrastructure. Second, the rollout of AI may generate governance externalities concerning data ownership, privacy, and security, notably in relation to smart metering and high-resolution consumption data that allow for sensitive inferences about household activities. For instance, existing work on privacy-protecting smart water meter databases explicitly casts the trade-off between information value and privacy risk as a function of temporal resolution and publication decisions [14]. Third, water utilities are facing higher levels of cyber threat, and the digital modernization introduces new surfaces for attack; recent guidance for drinking water and for wastewater systems emphasizes the need for a systematic evaluation of cybersecurity practices, prioritized risk mitigation actions, and planning to integrate resilience [15]. Lastly, AI has an ecological footprint (energy use, embodied impacts, and computational intensification) itself and can enhance inequities if not managed properly—an aspect emphasized in previous green AI reviews, which emphasize the need for proactive management of risks alongside opportunity narratives [16].
These opportunities and constraints drive the necessity for a systematic literature review that synthesizes evidence on AI solutions to enhance water resource sustainability in lieu of an end-state discussion of AI performance. Syntheses to date are typically partitioned by subdomains (e.g., leak detection, water quality monitoring, treatment process optimization, or hydrological forecasting) and differ considerably in how they articulate “sustainability” outcomes; assess implementation readiness; or tackle governance and risk. Therefore, the focus of this review is three-fold (1) to provide an analytical overview of how AI has been used in practice across the end-to-end water management cycle (monitor–predict–optimize–govern), (2) to type sustainability outcomes using a clear framework (e.g., resource efficiency, emitting/energy, service reliability, affordability and equity), and (3) to recognize cross-cutting barriers as well as enabling conditions for real-world usage data governance, cybersecurity, explainability and institutional capacity. It was an articulated research agenda that highlights the importance of replicable evaluation, impact measures that go beyond accuracy, and decision-based validation in operational settings.

1.1. Purpose and Scope of the Review

The purpose of this review is to consolidate and critically analyze novel approaches whereby AI coverage will also include machine learning (ML), deep learning (DL), reinforcement learning (RL), or hybrid/physics-informed methodologies—are employed with the goal of enhancing the sustainability performance of water systems for resource management applications. The scope covers the end-of-the-end water cycle and decision hierarchy, from (i) water resources planning and allocation (drought/flood risk, reservoir operations), (ii) water supply and distribution (leakage, pressure, asset management), to monitoring of compliance with water quality regulations and circularity-oriented operations. Consistent with reporting recommendations for PRISMA 2020, the review is structured as an audit-ready synthesis that is reproducible [17].
The review defines “sustainability” as a multi-dimensional outcome space that goes beyond predictive accuracy measures. More specifically, evidence is evaluated in relation to sustainability-relevant performance outcomes (for example, water losses/efficiency, energy intensity and emissions, reliability and resilience under extremes, risk reduction for compliance, and—where data are available—equity and affordability). The scope further explicitly encompasses cross-cutting deployment conditions, such as data governance, cybersecurity, and trade-offs for privacy surrounding the use of smart metering and high-resolution operational data [14,15].

1.2. Research Questions

The overarching question that this systematic review seeks to answer is the following: What are the conditions under which AI-enabled interventions in water resource management can yield real, sustainable gains, as opposed to simply increases in predictive or diagnostic accuracy?
To help to maintain the question empirically tractable and auditable, the review is organized according to two staged research questions.
RQ1 (evidence of impact).
What AI-enabled applications show reported measured operational sustainability outcomes (Non-Revenue Water Reduction, Energy intensity: kWh/m3, chemical dose reduction, compliance performance, reliability and resilience, affordability and/or distributional equity) rather than just upstream proxy metrics of model performance?
RQ2 (impact conditions).
For those studies that report measured operational outcomes, what are the common deployment conditions and governance mechanisms (data readiness, workflow integration, human-in-the-loop practices, monitoring and verification, and cybersecurity/privacy/accountability controls) that are associated with the materialization and stability of those outcomes?
To steer away from ambiguous or contradictory judgments about what constitutes credible evidence, we classify each study into one of three predefined impact-and-deployment evidence tiers (Tier 1–3) by aggregating (i) whether or not operational outcomes are measured, and (ii) the reported level of deployment maturity.
Tier 1 (quantitative impact, implementation): utility-, plant-, or basin-authority-level field implementation providing measurements of sustainability (e.g., before-after study, interrupted time series, controlled comparison). Tier 2 (limited evidence, pilot/limited deployment): pilot or pre-operational deployment with at least one measured operational outcome but limited scale, duration, or counterfactual strength. Tier 3 (proxy evidence only): simulation, laboratory, or offline (retrospective) non-field evaluation reporting model performance or proxy indicators, but not measured operational outcomes: studies reporting deployment without measured outcomes are treated as Tier 3 for sustainability-impact claims.

1.3. Originality and Contribution of This Review

More recent reviews have significantly pushed synthesis in more focused parts of the AI–water landscape—such as synthetic digital twins [11], control-oriented application reinforcement learning applications [12], and operational tasks for water distribution responding to leaks, based on AI/ML methods [8,9,10]. But, this piecemeal reporting results in two enduring barriers for sustainability-focused decisions: (i) numerous syntheses are still structured around methods and case studies rather than on measurable sustainability impacts and the decision paths that bring them about; and (ii) evidence robustness is variedly reported without a standardized assessment modal separating simulation-overloaded findings from pilot-supported, operationally deployed innovations [10,11,12]. Furthermore, these limitations are very relevant in safety- and compliance-critical water applications, where evaluation should go beyond prediction performance to also encompass decision and implementation validity [8,9,12].
This review is unique in three specific ways. Firstly, it consolidates the evidence for AI, which spans interventions across the cycle into a unified evidence map with three structuring dimensions, water-management function, and solution role (forecasting/detection/decision support/optimization and control as well as digital twin architectures), and sustainability outcome space simultaneously. This crosswalk can be used to find where the literature is thick with examples but weak in outcome reporting, and where control-focused research is still underpinned by simulation-heavy evaluation [11,12]. Second, the review operationalizes an “AI-to-impact chain” (data → model → decision → implementation → outcomes) by which to assess whether anything is “doing so” than having the reduction-in judgmental, rather than relying on improvements in predictive performance as a sufficient proxy of AI’s potential impact and acting more consistently with reporting that is transparent and audit-ready [17]. Third, it factors in governance and risk limits (privacy, cybersecurity, accountability, and operational safety) on the assessment of sustainability plausibility, highlighting that efficiency gains that elevate exposure to critical-infrastructure risk may just be unsustainable [14,15,16].
In so doing, the review adds a decision-relevant synthesis that disambiguates what has been proven to drive sustainability impact from what could drive such in AI-enabled solutions; reveals where evidence is currently concentrated in “the field”; and describes under which socio-technical conditions AI (may) constitute solutions that can credibly enhance resource efficiency, compliance, resilience based on hard data and—where measurable—equity outcomes [1,3,4,5,6].

1.4. Specific Contributions and Outputs

In this review, efforts aimed at addressing these challenges are attempted to be advanced in academia and practice along four dimensions:
(1)
Sustainability-Based Categorization of AI Applications in Water Resources Management: Taking the results of recent surveys and sector analysis as a basis, this paper develops an extensive taxonomy for AI tools specifically in the domain of water management. It categorizes these tools based on their operational goals—infrastructure assessment, demand prediction, and water quality surveillance—and the timescale of decision-making procedures they support, such as real-time operations, tactical planning, or long-term resources management [10,11].
(2)
An Evidence Map Association of AI Techniques to Sustainability Impact: Instead of exclusively emphasizing models’ accuracy or prediction performance, it relates AI techniques to material sustainability impacts. It exposes which studies are dependent on proxy measures (RMSE, accuracy) without converting these into actionable dimensions such as water savings, energy efficiency, emission reduction, or regulatory compliance. This framework enables more informed decision-making for utility managers and policy makers against the backdrop of SDG 6 and rising global water scarcity patterns [1,5].
(3)
A Systematic Review of Quality and Maturity of the Evidence Base: The current paper evaluates the methodological strength of the reviewed studies, including validation methods, reproducibility, and deployment readiness, emphasizing remaining disparities, especially in advanced control systems such as reinforcement learning models that are usually primarily confirmed in simulation settings or applied to small-scale real-world scenarios, with less consideration for the wide applicability and interpretability of these methods [12].
(4)
Agenda for the Future: Trustworthy AI in Water Systems: Finally, the review describes key enablers for AI uptake in water infrastructure, including data privacy (e.g., smart metering), cybersecurity measures, and governance structures facilitating algorithmic decision making. It also suggests the following priority areas for further investigation: production of benchmark datasets; uncertainty-aware evaluation environments; conjunction with hybrid modeling approaches and validation methodologies aimed at operational decision support [14,15].

1.5. Structure of the Article

The rest of this article is organized as follows. Section 2 Conceptual framing, such as system boundaries for water resource management and a sustainability outcomes framework to assess AI interventions. Section 3 describes the review methods, databases, search strategy, eligibility criteria, screening process, data extraction variables, and synthesis processes: presented according to PRISMA 2020 [17].
Section 4 presents results organized along the themes of water management and AI function (monitoring, prediction, optimization/control, decision support), as well as a summary of sustainable impacts reported and the maturity of evidence. Section 5 synthesizes cross-cutting enablers and constraints (governance, cybersecurity, privacy, explainability, and organizational capacity) in an integrated discussion, culminating with a prioritized research agenda. Section 6 ends with the practical implications for utilities, basin authorities, technology providers, and regulators.

2. Conceptual and Analytical Framework: AI Solutions for Sustainability in Water Resource Management

This section presents the analytical framework developed for (i) operationalizing “sustainability” as an outcomes space to inform WRM decisions, and (ii) categorizing “AI solutions” in a transparent manner that enables screening, coding, and synthesis. This is not intended to be a full, detailed technical background, but just introduce some coding-related constructs, which the following evidence map is auditable and decision-ready from a coding viewpoint.

2.1. Sustainability Problem—Space in Water Resource Management

Operational WRM is performed increasingly under interacting stressors: climate-driven variability and extremes, aging and leakage-prone infrastructure, increasing cost and energy of chemicals, tightening regulatory expectations, as well as institutional fragmentation that constrains investment and coordinated operations. In the real world, utilities and basin authorities need to make decisions pertaining to allocation, conveyance, treatment, and compliance in environments where they face: (a) incomplete instrumentation and telemetry; (b) heterogeneous and imperfect data streams; (c) limited OPEX/CAPEX and technical capacity. These limitations guide the possibilities of AI-enabled interventions, in terms of which sustainability claims can be reasonably based on evidence [1].
From a sustainability perspective, three pressure points are particularly conspicuous within the WRM value chain:
Reliability and allocation of water quantity: fluctuations in the amount of available water from drought to flood stress reservoir operations, conjunctive surface groundwater management, and demand planning.
Water quality and pollution control: diffuse and point source pollution, inadequate treatment capacity, monitoring gaps undermined both ecosystem and public health protection.
System efficiency and resource intensity—NRW, energy-intensive pumping and treatment, reactive maintenance contribute to financial impacts as well as environmental burdens; strategic NRW reduction and energy efficiency are often considered a set of primary levers for cost-profitable reductions [18].
These pressure points can pressure AI to address broader appeal than “just” giving predictions: they require assistance and control in decision-making (e.g., optimized pumping and pressure management, early warning systems, adaptive operations) and the institutional mechanisms to ensure safe, secure, accountable deployment.

2.2. Operationalizing “Sustainability” Outcomes for This Review

To maintain this focus on review relevance, “sustainability” is conceived of as a multi-dimensional territory of outcomes, rather than a single variable. During coding, each study is checked for whether it reports operational performance measures (e.g., reduced NRW, reduced kWh/m3) compared with more upstream proxy metrics of performance (For example, RMSE, classification accuracy). The results space is organized in a framework with four dimensions that are related to typical observable variables in the WRM application:
(A)
Environmental performance
Decreased conceptual burden (e.g., greater resource allocation efficiency; maintenance of environmental flows)
Better water quality and lower pollutant loads
Improved ecosystem status and resilience (where indicators exist)
(B)
Resource efficiency and circularity
Reduced NRW and avoidable losses
Reduced energy intensity of pumping/treatment: kWh/m3 and related proxies
Enhanced performance in wastewater treatment and its potential for reuse (including maintenance of monitoring and control)
That wastewater performance matters to sustainability are illustrated by the ongoing global monitoring of progress in wastewater treatment alone, alongside other associated targets [5].
(C)
Economic and operational performance
Operating costs (maintenance, energy, chemicals) lessened
Better asset management (condition monitoring, risk-based ranking)
Reduced downtime and faster time to recover from incidents
(D)
Social and governance outcomes
Reliability of supply, cost, and equity of access
Transparency and responsibility in making decisions
Enterprise risk management (including cybersecurity and privacy)
This framing has been purposely adapted to work with an applied utility and basin decision context—it says nothing about what evidence may be coded in terms of (e.g., the measurement at the wellhead; leak detection accuracy, energy savings, treatment compliance rates), nor the sustainability implications for a system (e.g., reduced withdrawals, avoidance, improved public health protection).

2.3. A Taxonomy of AI Solutions Across the WRM Value Chain

Given the heterogeneous design of studies and application scopes, AI solutions are categorized along two complementary axes: the functional role (what AI does in the decision-making process) and the WRM domain (where this is applied). This two-dimensional taxonomy helps in evidence mapping and avoids over-interpreting the algorithm families.

2.3.1. Functional Roles of AI in WRM

  • Sensing and perception augmentation.
Anomaly detection in sensor networks.
Quality control of data, treatments for missing data, and sensor fusion.
2.
Prediction and forecasting.
Demand/Inflow/WQ/Failure probability forecasting.
3.
Diagnosis and attribution.
Root cause analysis (e.g., anomalies, pollution events, or pressure transients’ drivers).
Classes of events and, in some cases, source identification.
4.
Optimization and decision support.
Regulating pump schedule, pressure control, chemical feed optimization, allocation of resources subject to constraints.
5.
Adaptive control.
Closed-loop feedback policies (reinforcement learning with believable simulation/field evidence if available).
6.
Strategic planning and prioritization.
Risk-based asset renewal planning; prioritization of interventions in a budget-constrained environment.
This functional taxonomy maps to applied research treatment of “smart water” context problem framing (infrastructure analysis, demand analysis, and water quality monitoring) and tensions in all three areas where ML/DL methods have been used [10].

2.3.2. WRM Domains (Where AI Is Applied)

Hydrology and reservoir operations: inflow prediction, water management rule modification, early warning of floods, drought planning.
Water distribution networks (WDN): leak detection, burst prediction, pressure management, and NRW analysis.
Monitoring of water quality: contamination prediction, event detection, forecasting of compliance risk.
Wastewater and Reuse: DO control, energy efficiency optimization, performance monitoring, process reliability.
Cross-cutting government: demand programs, customer service analytics, and regulatory reporting assistance (if applicable).
This domain layer is crucial because the “same” family of algorithm (e.g., LSTM or gradient boosting) can have very different sustainability implications depending on where it is run/what it is targeting (e.g., demand forecasting for peak shaving vs. treatment compliance prediction vs. basin allocation).

2.4. Modeling Paradigms: From Black-Box Prediction to Hybrid Digital Decision Systems

A pragmatic SLR in Sustainability would need to differentiate between “AI as a model” and the “AI as a social technical capability”, which is part and parcel of the operational systems. There are three isolated modeling paradigms that reappear in WRM research.

2.4.1. Data-Driven ML/DL Models

These include tree-based ensembles, support vector machines, neural networks (e.g., LSTM/GRU for time series), and graph-based ones where the network structure is important. While they are routinely chosen for prediction because of their superior performance, interpretability can be greatly reduced due to generalization issues in the face of non-stationarity (e.g., climate variability, shifts in demand) as well as high-stakes operational settings where explainability is required [10].

2.4.2. Hybrid (Physics-Informed/Physics-Guided) Approaches

The latest methodological critiques warn that tools of post-hoc explainability, for example, SHAP or LIME, which have become widely known in water application cases, are suitable to explain how a model is partial towards its entry variables rather than establish causal relationships or stable mechanisms. In a review of large-sample hydrology, Slater et al. explain that AI (XAI) alone is insufficient to replace causal inference and complementary methods—statistical inference, causal modeling, quasi-experimental validation, and mechanistic–ML hybrids—are required not only for developing faithful representations of the underlying complex data generative processes but also in answering meaningful queries of these models. These perspectives should be imbued within the model, and this is important for sustainability-driven decision-making where decisions are made within policy (regulatory) settings, requiring causal rather than simply feature attribution alone [19].

2.4.3. Digital Twins and Decision-Centric Architectures

Digital twins in the water industry are increasingly presented as a system of systems involving the interconnection between telemetry, models (physical and/or ML), and decision workflows in planning and real-time control. Recent reviews remark that “digital twin” is frequently applied inconsistently; hence, to synthesize it, it is crucial to code what is adopted (data integration, model coupling, calibration strategy, decision loop closure) rather than reading labels on packs [11].
In the context of wastewater management, digital twins are often conceived as technology-specific decision support (e.g., for surveillance, control evaluation, optimization, and scenario analysis) relying on model correctness, data processing, and operational implementation [20].

2.5. Reinforcement Learning and Adaptive Control: Promise, Constraints, and Evidence Expectations

Reinforcement learning (RL) is also seen as a direction toward adaptive WRM under uncertainty (e.g., reservoir operation, pump control, and dynamic allocation). But RL outcomes are extremely sensitive to (i) simulator realism, (ii) reward design and constraint enforcement, and (iii) the governance environment that governs whether learned policies can be deployed.
A recent systematic overview on RL in WRM highlights both the increasing interest and returning shortcomings: most studies are still confined to simulations, and translation to real-life operations implies considering safety constraints, uncertainty, and interpretability carefully [12].
As such, this SLR will consider RL-based contributions to be sustainability-relevant only if the work offers sound evidence on one or more of the following:
  • constraint satisfaction and safety (not only average reward);
  • robustness under scenario shifts, interpretation or policy validation against domain baselines, operational integration possibility (data needs, compute time latency, governance approval).

2.6. Trustworthy AI in WRM: Risk, Privacy, Cybersecurity, and Accountability

As WRM constitutes an essential service, all claims of sustainability are inextricably bound up with risk principles. An AI-powered intervention that creates efficiency but raises cyber risk, privacy exposure, or unaccountable decision-making may be “unsustainable” in practice.

2.6.1. AI Risk Management and Governance Principles

The AI RMF 1.0 is also presented as a general across-domain framework for the management of AI-related risks throughout the lifecycle (design, development, deployment, and monitoring), focusing on Safety-Security-Resilience-Privacy by design, Privacy enhancement, Transparency, Bias mitigation/management with adaptation for specific domains’ features [21].
Counterpart to OECD’s intergovernmental AI principles. The Council recommends trustworthy AI which upholds fundamental rights and democratic values—highly pertinent if water decisions affect equity, affordability, and public health [22].

2.6.2. Cybersecurity as a Sustainability Prerequisite

Water and wastewater systems are more dependent on networked instrumentation and SCADA/OT environments for smarter operations; AI implementations can widen an attack surface with new connections, cloud services, or third-party integrations. Sector-specific guidance emphasizes the importance of gap analysis and “installation of key controls by drinking water and wastewater operators” [15].

2.6.3. Privacy Risks in Smart Metering and Demand Analytics

Fine-resolution water utilization may disclose private household activities, making the sharing of data and external analysis unsuccessful. Empirical investigation, where the privacy of smart water meter databases is protected, shows that there are trade-offs between privacy and utility that depend on temporal resolution, population size, and sampling design—a useful boundary if research begins to suggest open data-driven optimization or benchmarking [14].
These considerations directly inform how the review codes “AI solutions,” as outside of algorithm type, it will also extract evidence regarding data governance, security controls, privacy protections, and human oversight; these qualities condition whether sustainability benefits are sustainable (in the ontological sense or durable), scalable, and socially acceptable.

2.7. An “AI-to-Impact” Chain for Synthesizing Sustainability Evidence

Building on top of the elements discussed above, an AI-to-impact chain is proposed that will serve as a framework for synthesis:
  • Input layer (inputs): data coverage/quality, telemetry and integration, governance entitlements, security posture.
  • Analysis layer (methods): ML/DL/hybrid/RL; uncertainty treatment; explainability; validation strategy.
  • Decision layer (outputs): alerts, predictions, suggestions for action, schedules of action to be made, policies of control.
  • Implementation layer (implementation): workflow integration, latency, operator trust and adoption, maintenance, and monitoring.
  • Outcome layer (impacts): gains of efficiency (NRW/energy), compliance improvement, resilience, cost reduction, equity aspects.
  • Risk layer (constraints): cybersecurity, privacy, accountability, failure modes, distributional harms.
This chain of reasoning creates an analytical framework for understanding why reported “high accuracy” does not guarantee sustainability impact—the latter depends on whether decisions change, if changes persist, and if risks are tamed.

3. Materials and Methods

3.1. Review Design and Reporting Standard

This work is set up as a systematic literature review (SLR) of artificial intelligence (AI) approaches for water resources management (WRM) with an explicit perspective of sustainability. Reporting: This review was reported following the PRISMA 2020 statement, with the inclusion of a PRISMA flow diagram and checklist to enhance transparency and reproducibility [17].

3.2. Conceptual Framing and Operational Definitions

For screening and coding exercises, AI solutions are models where the computer can learn from data patterns (e.g., machine learning, deep learning, reinforcement learning, hybrid physics-informed ML) and/or automate decision support (e.g., optimization, intelligent control, anomaly detection) in WRM contexts. In WRM, sustainability is understood across at least one family of outcomes: (i) environmental performance; Sustainable: (e.g., reduced extraction, leakage energy use emissions pollutant loads); (ii) economic performance (e.g., life cycle cost O&M efficiency avoided losses); and (iii) social/governance dimensions: (e.g., service reliability equity proxies transparency compliance risk management). To do that in the context of synthesizing AI, the review codes governance and risk management against AI risk guidance [21], so as to avoid leaving behind considerations for trustworthiness specific to AI.

3.3. Eligibility Criteria

Inclusion and exclusion criteria are pre-specified, uniform at all sites:
  • Inclusion criteria
Relevance: This study contributes to AI applications in WRM, ranging from (but not limited to) demand forecasting and allocation, risk of drought/flood, groundwater assessment, distribution networks such as pipe leak detection or pressure management, treatment optimization, water quality monitoring, irrigation/watershed management, up to integrated water–energy decision support.
AI as a substantive method: AI is the core of the method (not just an on-the-side mention).
Sustainable connection: There is a minimum of one sustainability-related result, metrics, or explicit discussion linking the AI solution to sustainability goals (environmental, economic, and social/governance).
Types of publication: Peer-reviewed journal articles and conference papers (if the methods/results are adequately described). Relevant targeted grey literature (Section 3.5) is included where it presents authoritative standards, governance frameworks, or sector baselines applicable to AI-in-water deployment.
  • Exclusion criteria
Non-water domains (unless one can transfer and explicitly validate them in water mediums).
Conceptual essays that do not include methods and evidence (unless they are characterized under “conceptual/governance-only” grey literature and deliberately justified).
Patents, editorials, slides, and non-archived materials.
Publications with insufficient methodology details to evaluate the AI method (e.g., data sources and implementation/validation).
A structured screening guide is used to apply these criteria, minimizing discretionary exclusions and facilitating auditable exclusion logs [17].

3.4. Information Sources

To ensure an equitable coverage of engineering technology, environmental science, and AI/CS contributions, the search strategy is adopted to cover multidisciplinary bibliographic databases (e.g., Scopus, Web of Science Core Collection) and technical indices (e.g., IEEE Xplore) where applicable. Grey literature is targeted to capture sectoral standards and authoritative reports that inform sustainability baselines and deployment guidelines in water systems (e.g., global wastewater treatment status reporting, water-cycle climate extremes reporting, cybersecurity guidance for water utilities) [1,5,15].

3.5. Search Strategy and Grey Literature Procedure

3.5.1. Academic Database Search

The search pairs will be concatenated using AND in the final query string.
AI block: “artificial intelligence” OR “machine learning” OR “deep learning” OR “neural network” OR “reinforcement learning” OR “physics-informed” OR “digital twin” or “anomaly detection” or “predictive maintenance”.
Water management block: “water resource* management” OR “water distribution” OR “water network*” OR “leak detection” OR “water treatment” OR “wastewater networks” OR “groundwater resources” OR “watershed hydrological models/monitoring systems”, or irrigation/hydrology-related information.
Sustainability block: sustainability OR “energy efficiency” OR “greenhouse gas*” OR carbon OR “life cycle” OR resilience OR equity OR “circular economy”.
A pilot search is made for the purpose of ensuring sensitivity (the recall ratio) and specificity (control of irrelevant retrieval). The ultimate search strings, database limits (document type, language), and date when conducted are thoroughly reported to allow for duplication as recommended by PRISMA 2020 reporting guidelines [17]. The database searching of literature spanned from 2000 to 2 April 2025 (date last search). This long window allowed early AI applications to be included and took into consideration the rapid post-2020 surge in digital water technologies. Studies after April 2025 were excluded from the review, and it was recognized that future advances may contribute to new evidence-based evolution.

3.5.2. Grey Literature Search (Targeted)

Grey literature is included by choice, but in a limited and auditable manner. Specifically, sources are restricted to (i) those that offer authoritative benchmarks or sectoral indicators related to WRM sustainability; (ii) operational limitations that constrain the deployment of AI (such as utility cybersecurity demands); or (iii) principles relevant to trustworthy AI and critical infrastructure as codified in AI risk governance. Potential sources are UN system reports (UN-Water, UNESCO WWDR), intergovernmental organizations (WMO), standards organizations (NIST), and national regulators (EPA) [1,5,15,21].
The extremely heterogeneous nature of grey literature makes the assessment of methodological soundness necessary for each non-peer-reviewed resource to ensure precision and scientific rigor. In practice, it was (i) determined if the document was issued by a bona fide organization with demonstrated expertise in this area (e.g., UN agencies, national regulators, standards bodies); (ii) confirmed if it presented objective data or guidance as opposed to just opinion; and (iii) corroborated key facts against peer-reviewed sources where feasible. Only the articles that met these criteria were kept. This methodical review addresses the use of grey reports that could bias findings.

3.6. Study Selection Process

Before screening, records are downloaded and duplicated in the reference manager software. The selection is performed in two phases:
Title/abstract screening against eligibility criteria.
Full-text screening with reasons for exclusion recorded.
For the sake of clarity, below is presented a brief overview of the most common reasons for exclusion at full text. Studies were excluded if (i) they had no substantial AI component (e.g., descriptive surveys/proposals speculating on AI without running and evaluating model); or (ii) no sustainability-related indicator was reported, nor discussion of links with environmental, economic or social results; or (iii) methodological details were not provided in a way that enables replicability (e.g., missing sources of data, algorithm description/validation steps); or (iv) full-text paper was not accessible despite inquiries to authors and search for alternative repositories. This information also provides an audit trail to show that the final corpus of studies (n = 41 peer-reviewed and seven grey literature sources) was chosen on clearly defined criteria rather than excluded through unsystematic means, thus reducing potential biases.
To reduce selection bias and improve reliability, two authors independently carried out the screening procedures for both of its phases. All differences were settled by discussion to consensus; unresolved discrepancies were adjudicated by a third reviewer.
A codebook was established and piloted. Data were written into a predefined spreadsheet from the included records to include the following terms:
Bibliography details: Author(s) and year, Title (separated by comma), source.
Field of Application: Upland municipal water supply, wastewater treatment, hydrology/catchment (with rich experience).
AI Model Features: Family type (e.g., tree-based, deep learning, RL), data types, if feature engineering was applied, interpretability techniques used, and whether it is hybrid/physics-informed.
Methodological quality: Train/validation/test split, use of baseline comparisons, external validation, sensitivity/uncertainty analysis.
Maturity of Deployment: Conceptual, laboratory scale/pilot-scale, operational upon utility/plant-scale.
Outcomes for sustainability: Measured indicators concerning water savings (NRW), energy use (kWh/m3), chemical management, compliance records, system reliability, and affordability/equity aspects.
Governance & Risk Controls: Reference to workflow integration, human-in-the-loop design, data governance, cybersecurity (i.e., safeguards measures), privacy protection, or accountability.
Level of Evidence: For each study, this has been classified according to its reported deployment level (tier 1 to 3) and the measurement of operational sustainability outcomes, as described in Section 1.2.
The overview of the selection process, also demonstrating reasons for exclusion at full-text screening, can be found in the PRISMA 2020 flow diagram [17].

3.7. Data Extraction and Coding Scheme

Field of application: urban water distribution networks, drinking water treatment, wastewater collection/treatment, hydrology/catchments/basins, irrigation, and cross-sector water energy support tools.
AI model: family of models (for example, tree-based, deep learning, and reinforcement learning), input data types, how to engage in feature engineering, interpretability/uncertainty technique applied—was it physics-informed or a hybrid method?
Methodological rigor: a train/validation/test setup, leakage controls, baselines/benchmarks, c external validation, d sensitivity analysis.
Level of deployment maturity: conceptual, lab, pilot, or operational (utility, plant, or basin authority scale).
Sustainability performance indicators, environmental/resource-efficacy (NRW). Type of indicator: Quantity Metric: kWh/m3 Emissions Chemicals Pollutant load Economic/Operation metrics Cost Downtime Productivity Social/Governance metrics Reliability Service Continuity Affordability Equity Transparency/Compliance.
Operational governance and risk controls: workflow integration, human-in-the-loop ingestion, monitoring, and error-correction process, data governance principles for developing a pipeline of content transformation activities. Deploying cybersecurity and privacy protection involves people in appropriate oversight to take AI risk management guidance (concerning critical infrastructure) [15,21].
Tier 1–3 evidence level of proof (impact and deployment): The included studies were categorized based on the predefined rule introduced in Section 1.2, considering both whether monitored operational sustainability outcomes were reported and the reported state of deployment maturity.
The definitions and decision rules for tiers are available in Section 1.2 and have been consistently implemented in the coding and synthesis.

3.8. Quality Appraisal and Risk-of-Bias Assessment

Study quality assessment Review component (including information search) The typical methodological diversity of AI-in-water studies (quantitative modeling, case studies, or mixed methods) renders it challenging to make holistic judgments of study quality, and so a matching tool assessment was applied:
The Mixed Methods Appraisal Tool (MMAT) 2018 [23] is used to assess hybrid/mixed methods or mixed research design.
The CASP Qualitative Studies Checklist [24] is used to assess the quality of qualitative research (e.g., deployment analyses with a governance focus).
These are also potentially included as evidence input in the review, e.g., umbrella components, and their methodological quality can be judged with AMSTAR 2 [25] or risk of bias with ROBIS where applicable [26].
In parallel, since AI papers often under-report reproducibility-relevant information, we used (and ran a mini-rubric) the AI credibility rubric, the tools code if and only if data provenance and representativeness; validation strategy and leakage controls; baseline competitiveness; interpretability/monitoring; and transferability constraints. The trustworthy AI and governance markers are interpreted according to established AI risk management guidelines [21].

3.9. Evidence Synthesis and Optional Bibliometric Mapping

Narrative synthesis is incorporated and organized according to the water system segment and the role of the AI solution. Statements on sustainability impact are linked explicitly to the evidence-tiering scheme: Tier 1 studies serve as a basis for statements about measured operational impacts; Tier 2 studies are referred to in terms of emerging but limited evidence, while Tier 3 studies are discussed as articulating technical potential and research directions, though not establishing sustainability impact.
Note that the review purposefully refrains from applying a rigid numerical scoring system to compare AI solutions against sustainability objectives. The reason is that it can be applied in a variety of contexts, including distribution networks, wastewater treatment plants, and basin-scale operations, as well as very different baseline situations. Furthermore, equivalent measurable outcome reports are relatively scarce. In this situation, any scoring scheme is going to recommend a level of accuracy that cannot be justified. Rather, the method is based on an evidence-tier ranking along with a qualitative evaluation of sustainability claims. With the advent of more standardized and comparable results, potential opportunities could emerge in future research to develop domain-specific benchmarking strategies.

3.10. Transparency and Reproducibility

To allow the audit of the project, the following were kept: (i) the final database search strings; (ii) the rules applied for de-duplication; (iii) the decisions made when screening (along with reasons for rejection); and (iv) the extraction codebook. These items are made available as supporting files where appropriate, to support replication, in line with PRISMA 2020’s focus on transparency of reporting [17].

4. Results

4.1. Study Selection

The PRISMA checklist is available in the Supplementary File (See Table S1); please refer to Figure 1 also for the presentation of the PRISMA 2020 flowchart. A search of the database provided 920 hits. After deduplication (185 removed), 735 titles/abstracts were screened, and 650 were excluded based on prespecified eligibility criteria (e.g., not water-sector relevant, no AI component, commentary only, inadequate methodological transparency). Full texts were evaluated for 85 articles, and n = 41 studies were included in qualitative synthesis and n = 38 in quantitative/bibliometric synthesis (when applicable). Grey literature (institute frameworks, operational guidance) contributed to seven more sources and was analyzed separately so as not to conflate empirical evidence with normative guidance.

4.2. Descriptive Characteristics of the Included Studies

4.2.1. Publication Trends and Venues

In the analyzed corpus n = 41, publication volume grows significantly after 2020, reflecting an increased pace of digitalization in utilities and the general spread of smart sensing, IoT telemetry, and cloud analytics within water operations. The evidence is focused on engineering, environmental informatics, and water-utility operations channels, with a smaller but expanding presence in interdisciplinary sustainability spaces.

4.2.2. Water System Segments Covered

The majority of included studies could be categorized under the following three operating segments:
Urban WDSs: detection of leakage/anomaly, monitoring and management of pressure levels, prediction of demand volumes, assessment for risk at the asset level. In application-oriented surveys of smart WDS research, these are usually divided into infrastructure analysis, demand analysis, and water-quality monitoring, which corresponds quite well with evidence mapping for the WDS subset in this study [10].
Wastewater collection and processing: process oversight, aeration control, early alerting for effluent quality delinquencies, preventive maintenance for key assets.
Decision support facing the catchment/basin and hydrology: Runoff/streamflow forecasting, drought/flood early warning, reservoir or allocation optimization.
Within the evidence base, digital twins are disproportionately prevalent in distribution networks and wastewater treatment, relatively thinly represented in the desalination and reuse/reclamation settings—which anecdotally mirrors emerging findings of recent domain reviews—as well as seeming to align with the broader spread of “applied” digital twin studies across the water sector [11].

4.3. Classification of AI Solution Types

Regarding the coding framework described in Section 3, the selected papers can be summarized as falling into five main families of AI solutions:

4.3.1. Prediction and Forecasting (Time-Series and Spatiotemporal)

Applications for forecasting include the prediction of demand at multiple temporal scales, short-term hydraulic state surrogates, rainfall–runoff/streamflow models, and early warning water quality deterioration (e.g., turbidity/nutrient spikes). Methods include tree-based ensembles, deep sequence models (e.g., LSTM/GRU), and the growing trend towards hybrid methods that incorporate physical constraints or simulation outputs.

4.3.2. Detection and Diagnosis (Anomalies, Leaks, Events)

Most of them focus on reducing non-revenue water by detecting and diagnosing leaks, bursts, and abnormal consumption detection/diagnosis. In this work, the analysis focuses on (i) event detection performance and (ii) localization accuracy under sparse sensing. More recent reviews for leak-detection stress how “smart water management” sensing and analytics have expanded the method families beyond acoustic/manual methods, but also champion deployment trade-offs (cost versus coverage versus alarm precision) [8,9].

4.3.3. Optimization and Control (Including Reinforcement Learning)

Optimization issues under these considerations cover pump scheduling, pressure control, reservoir operation policies, and wastewater process control. When the technology choice of reinforcement learning (RL) is adopted, one common empirical trend is that agents are primarily trained and tested on simulated environments rather than real historical operation data directly, where deep Q-networks are among the most widely used algorithms [12]. Operationally, the gap between simulation success and real-world reliability has been one of the most noted limits in RL-based synthesis.

4.3.4. Digital Twins and Hybrid “AI + Physics” Stacks

The typical implementation of a digital twin tends to take the form of connected systems: telemetry + hydraulic/process simulation + analytics layers (state estimation, prediction, control). In digital twin research, among the included contributions, the most mature are monitoring and predictive maintenance, whereas closed-loop optimization is not yet been systematically validated in real environments (and under safety constraints). Larger-scale sector-specific reviews also show that digital twin activity is focused on the distribution and wastewater areas [11].

4.3.5. Decision Support and Prioritization (Risk Scoring, Investment Planning)

A smaller but relevant policy set builds multi-criteria decision support- for example, ranking pipe replacement candidate locations, prioritizing sensor deployment, or allocating scarce operation and maintenance resources. These are often the closest to “sustainability” decision contexts (i.e., equity, affordability, resilience), though many of them still ostensibly operationalize sustainability indirectly (see Section 4.4).

4.4. Sustainability Outcomes Reported and How They Are Operationalized

To match technical outputs with sustainability claims, each study was coded for (i) type of outcome and (ii) level of evidence:
Resource efficiency impacts (water, energy, chemicals): Previous research converts the model output into avoided losses (e.g., reduced non-revenue water), energy savings due to better quality pumping/aeration, or improved chemical dosing efficiencies. But most are reported in papers as predictive accuracy, without the gains being translated into operational savings or environmental benefits.
Water quality and compliance performance: While a sub-section of studies attempts to model parameters that are associated with compliance, to support early intervention, external validation against regulatory monitoring regimes is frequently constrained.
Resilience measures: Resilience is often approximated by the recovery or failure detection time, and robustness under interventions.
Equity and social outcomes: More rarely (if at all) operationalized, though present as service continuity proxies or vulnerability-sensitive allocation rules, rather than measured distributional impact.
As global progress on wastewater treatment is monitored by SDG indicator 6.3,1 (safely treated flows), studies purporting SDG contribution were evaluated for whether they report outcomes that might conceivably be mapped to compliance/safe-treatment gains rather than just model performance [5].

4.5. Evidence Robustness and Evaluation Practices

4.5.1. Validation Regimes

In the corpus, there were three predominant validation regimes:
In hindsight validation (very popular in demand/quality prediction).
Synthetic/simulation validation (as in RL control and some hydraulic anomaly detection).
Few prospective or on-time pilots (minority; mainly utility-backed initiatives).
As with synthesis in the RL setting, simulation-intensive evaluation precludes moving from result to safe and regulatory operation [12].

4.5.2. Reproducibility and Transparency

Reporting quality differed: some described datasets, pre-processing, and hyperparameters; others did not. At a part level, when digital twins are asserted, characteristics of coupling architecture (e.g., data pipeline and synchronization and latency handling) frequently are inadequate to reproduce at the “system” level.

4.5.3. Risk, Governance, and Operational Constraints

Despite the direct relevance for water utilities, the number of these papers that directly consider cybersecurity and OT risk is limited. This contrast is particularly striking in the face of institutional advice for water/wastewater operators to identify and improve cybersecurity controls in their IT/OT environments [15].
Likewise, very few studies directly relate AI risks (e.g., robustness, bias) to structured risk management frameworks that are increasingly used in high-stakes infrastructure settings [21].

4.6. Evidence Map of AI Solutions by Water-Management Function

Table 1 summarizes the included studies as an evidence map, which is crossed by water system function (e.g., distribution, wastewater, basin planning) and AI solution type (forecasting, detection, optimization/RL/IPM/digital twins, decision support), and by sustainability outcome space (water savings; energy/GHG; compliance; resilience; equity). Three high-level patterns emerge:
Idle focus: Evidence is thickest in distribution and (leaks/anomalies, demand) wastewater process monitoring/optimization.
Outcome imbalance: Predictive metric is abundant, quantification of real-world impacts (water/energy/GHG reductions) is less pervasively reported.
Translation gap: There are more simulation-dominated than field evaluations in terms of control-oriented applications (e.g., RL and closed-loop optimization).

4.7. Summary of Gaps Observed Within the Results

Based on the included data and the results of the appraisals, these lacunae were identified in the evidence base (and justify why this section can be a discussion/thought piece but should not replace it):
Limited field evidence for closed control loop: Especially for RL and digital twin guided optimization [11,12].
Weak sustainability reporting: Despite being set out in documents, many papers do not represent resource savings, and the impact of these leads to compliance and service on output measures that are traceable.
Misreporting cybersecurity and governance limitations: Low-hanging fruit with a real-world impact [15].
Partial alignment with climate-forced variability: Despite growing evidence of the need for global water resource assessments, many AI models are tested under stationarity or short data windows [1].
Though based primarily on qualitative evidence, a few Tier 1 cases gave enough information to estimate the scale of sustainability gains attributed to AI. BeChained, e.g., showed that an AI-enabled pump-scheduling solution, implemented on an industrial water utility in Spain, resulted in a 18.7% reduction in pump operating energy consumption with unchanged service levels [60]. In the same way, the Xylem leak-detection solution deployed in Dallas Water Utilities led to savings of about 7.2 million gallons of water per day and a 50% decrease in the number of leaks [61]. These real-world applications serve as a reminder that AI can deliver double-digit efficiency improvements when integrated into operations with defined decision loops and measurement-and-verification processes. But these were few—only two studies had reported auditable before–after comparison. Furthermore, the reported benefits are very sensitive to the context (size of utility, initial efficiency, and regulation) and should not be transferred without hesitation. However, these quantitative examples illustrate that semi-quantitative synthesis has the potential to provide an indication of the order of magnitude of possible savings and suggest that there should be increased standardization in how operational outcomes are reported.

5. Discussion

5.1. Reframing “AI for Sustainability” in Water Management: From Model Accuracy to System Performance

A common analytical thread through the 41 studies’ contents is that AI-induced value enhancements in water management continue to be largely described as better prediction, rather than measurably better sustainability. This is important because sustainability for water systems is finally achieved in the realm of operational decisions (maintenance dispatch, pump schedules, aeration setpoints, chemical dosing policies, asset renewal plans, drought rules, emergency response) rather than only through model-based metrics. It is this issue that one reads through in the reviews of smart WDS applications: the main bottleneck to move machine learning algorithms into operational conditions is not the algorithm by itself but its application under data constraints, missingness problems, interoperability gaps, and utility-grade adoption concerns [10,55]. Accordingly, the literature must be considered evidence-based on readiness for decision support rather than a guarantee of delivered impact at scale.
A practical implication is that the pathway “AI → sustainability” should be considered a chain of evidence:
Data quality and representativeness → (1) model performance and robustness → (2) decision integration (human/automation) technology, → (3) operational change, → (4) quantified outcomes (water, energy, chemicals compliance resilience wherever possible, equity if relevant).
A large pool of existing research notes (1) and in some cases (2), while perhaps fewer are systematically recording (3)–(4); this is one reason for the continuing unevenness with which sustainability accounting continues to be practiced across domains [10,28].
A helpful way to discipline the interpretation of results is to treat “predictive performance” and “sustainability impact” as analytically separate claims. More generally, better prediction does not necessarily imply good decisions that are more useful to the outcome (prediction is not explanation, and the latter requires linking model outputs to actions, interventions, or practical operational changes) [62]. This distinction is relevant for your Tier logic: Tier 3 level studies can be useful proof-capability but should be sponsored as performance evidence unless they have been partnered with a plausible impact pathway (the decision + operation you want to influence) and an evaluation that will measure change. To aid in “closing the loop,” you can refer readers to established measurement-and-verification practice, where ex ante impacts are defined, baselined, and quantified using transparent accounting rules [63]. And third, as real-world deployment introduces risk (data drift, security, accountability), capturing an operational view of AI systems in the paper is appropriate to explicitly situate its framing around “credible evidence” (as formulated here) in a risk-management context—i.e., that validity and reliability are but necessary conditions for trustworthiness within operational contexts [21].

5.2. Where the Evidence Is Strongest: Monitoring, Detection, and Forecasting as “High Readiness” Families

5.2.1. Leakage, Bursts, and NRW Reduction: Strong Technical Momentum, Mixed Outcome Quantification

Leakage and burst analytics is one of the most developed clusters in this corpus. The domain has evolved from narrow signal-processing techniques to AI’s wider family (supervised, semi-supervised, deep learning, and hybrid methods), but the contextualization remains similar: performance should be tested under realistic sensor deployment coverage, false-alarm control, and localization needs, aiming at reducing the time interval between alarm response while minimizing water losses [8,9]. Network-based methods are symptomatic of a sector trend towards topology-conscious learning: graph neural architectures are built to tap into network structures that the traditional procedure ignores [37,43]. Burst localization techniques are geared towards operational relevance by maximizing burst detectability in a sparse monitoring network—an important consideration for utilities with limited resources to invest in dense instrumentation [54].
Yet even in this high-readiness family, sustainability claims are often implicit. A handful of studies make an excellent case that more timely detection leads to water savings, but few provide a widely applicable “NRW reduction accounting” that can be used to map improvements in detection lead time and localization accuracy into the expected volume of savings and associated energy/emissions savings resulting from reduced production and pumping. Asset analytics start to tackle this by targeting proactive repair or renewal [31,59], but the connection of predictive maintenance outputs with audited reductions in bursts/leaks are still a rare occurrence. In the context of Environmental Governance, one would interpret this gap to mean that future reporting standards should not categorize “water saved” as a secondary effect or downstream benefit.
For use in leakage/NRW-focused applications, “measured sustainability impact” is under its sector-specific accounting—no such single observation where NRW is a balance-flow dialectic embedded within metering integrity, sampling on match sequence of flights and sub flights, and disentangling apparent from real losses [64]. This has direct implications for Tiering: if the utility does not have (i) a defendable baseline water balance, (ii) a validated audit process, and (iii) an intervention protocol that converts detection into repair and/or pressure-management response, it can easily report great event detection but never demonstrate lower abstractions or energy/chemical savings. You can make the case stronger by specifically linking your “impact” language to established NRW utility and regulator performance frameworks, such as IWA performance indicators for water supply services [65] and AWWA’s audit-and-loss-control guidance, so that Tier 1 “measured impact” involves changes in auditable rather than model-only metrics [66].
The technical horizon includes detection, localization, and quantification methods, but practical performance is limited by sensor configuration, DMA concept, water hydraulics, and stochasticity of ambient conditions [67]. Recent synthesis as well as stress that the main roadblock for utilities is not the existence of the algorithms, but their operational feasibility: reliable under realistic noise, integrable with utility workflows, and able to ensure actionable prioritization [68]. This outside literature is consistent with your central finding: the most believable evidence Tier 1 occurs when (a) it is a human-in-the-loop intervention, (b) it has a short path between detection and conservation, and (c) the utility can do the repair/pressure management works with traceable work orders and their audits updated [66,67].

5.2.2. Water Demand and Supply Forecasting: Operational Relevance Is Clear; Sustainability Quantification Is Still Inconsistent

Yet another relatively mature family is the demand and supply forecasting. Deep sequence and hybrid-correction pipelines are optimized to minimize prediction error for short horizons, which can be factored into processes such as production planning or pump scheduling [27,45,46]. The sustainability significance is believable, as avoidable energy consumption, excess chemical dosing, and pressure fluctuations (which can cause leaks and stress in pipes) all result from supply–demand discrepancy. However, like the NRW literature, many forecasting studies end at error measures rather than systematically translating these improvements into resource gains (e.g., energy consumed per cubic meter treated, chemical usage per cubic meter treated, and avoided overflow/violation in downstream processes).
The analysis of demand-forecasting findings could be enriched by juxtaposing the above insights against the demand-forecasting review literature, which has repeatedly provided evidence that method selection is contingent on the nature of the decision (short-term operations versus long-term planning), lead-time horizons, and the costs of error [69]. In particular, the field has constantly highlighted the need for probabilistic/uncertainty-aware forecasting when utilities are required to make decisions under uncertainty (e.g., scheduling of pumping operations from wells and storage management or imposing drought restrictions) as point-accuracy metrics do not capture their operational impact alone [69]. Related to this, a growing literature on urban water-demand modeling suggests that demand is characterized by socio-demographic structure, pricing, climate variability, and policy—such that model transferability and governance constraints can be as important, if not more important, than model class [70]. These external results bolster the Tier framing: a lot of studies can keep being Tier 3 even with high predictive scores, unless they demonstrate that the predictions were incorporated into an actionable decision-aid and led to measurable operational insights (total kWh reduction, not just improved accuracy) rather than only better out-of-sample performance.
Furthermore, input feature selection and engineering are so important that they exert a more dominating impact on forecasting performance than architectural novelty. For instance, in a recent work on biogas production forecasting, 11 out of the top 15 most important variables were not directly observed features but were generated through feature engineering from raw SCADA data. Systematic, including domain-specific features (such as lagged hydrological drivers, control constraints, derived indicators) could lead to a more robust and generalizable approach, and documenting such feature engineering decisions would result in AI models being more transferable and interpretable [71].

5.2.3. Hydrology-Facing Early Warning: Resilience Benefits Are Plausible but Must Be Evaluated Under Non-Stationarity

In basin and catchment settings, AI is often put forward for resilience: flood forecast, flash-flood warning, and prediction of water level [32,44,57]. Prediction of draught and analysis of long-range trends take AI use into anticipation, planning, and allocation considerations [47]. The warning here is that it is in hydrologic extremes where non-stationarity and regime shifts are the most important, so for model evaluation, demonstrating robustness over climatic variability should be prioritized over short-window performance. Research that explicitly harnesses long historical baselines is better poised to inform resilience than research based on recent decades, and it continues to raise questions about generalizability [47].
The latter discussion can be sharpened at the basin/catchment scale by linking “credibility” to the fact that many hydrological relationships are nonstationary under climate and land use change, and models trained on historical regimes may fail just when decision support is most critical [72]. This suggests that a “test” of whether impact claims for (for instance) models should be conditioned on deployment assumptions: it is not sufficient to report high hindcast performance if the system needs to possess robustness under regime shift, extrapolation, or policy-driven operational changes. The deep-learning synthesis literature within water resources analogously emphasizes the need for evaluation to be connected to intended application (and not just performance on a test set) when it comes to generalization, interpretability, and uncertainty, as well as performance, which is not synonymous with decision validity alone [73]. This framing enables a more forceful conclusion in the paper: non-stationary, basin-facing end-use Tier 1 impact evidence is only likely to be rare cases when (a) studies show decision integration and (b) robustness checks that match nonstationary planning contexts are demonstrated.
Beyond the basin setting, non-stationarity and transferability represent ubiquitous sustainability issues across all water-management sectors. In the studies reviewed here, many models worked well under site-specific and temporally static conditions; however, few assessed how well they could be generalized to new locations, climate regimes, or operational statuses. Studies in hydrological extremes show that models cannot always replicate the features of hydrological components for different climate change scenarios and must thus relate to the knowledge of a physical process model to guarantee reliable predictions [74]. Incorporating cross-site validation, regime-shift testing, and explicit transferability analysis in future research will be key to sustainable implementation, as climate change changes water resources and operational regimes.

5.3. Where the Evidence Is Promising but Less Deployable: Optimization, RL Control, and the “Simulation-to-Field” Gap

Optimization and control-related studies directly aim at sustainability measures (energy, compliance), with an emphasis on pumping and wastewater processes [40,41,56]. However, the control-oriented AI literature is narrowed by a common shortcoming: evaluation contexts are often simulated for an application where the desired deployment is in real-time.
This is seen especially in the literature on reinforcement learning (RL). The systematic review of RL applications to WRM reveals a tendency that agents are mainly trained and tested in simulators, while the key remaining issues are safe transfer, constraint handling, interpretability, and monitoring under real operational risk [12]. RL reviews in WDS-related leakage management reiterate this finding: simulation outcomes cannot be taken as field-grade evidence, as real systems have safety constraints (e.g., bursting), operational policies, human override, and failure consequences that are never characterized in setting up research scenarios [49]. Even when DRL comes in the guise of “real-time control,” it typically looks at retrospective evaluation with simulation models, not extended trials on in-situ research [34].
From the viewpoint of Sustainability, it is seen that any claims on energy saving for emission reduction based on RL-driven optimization would be interpretatively best framed as conditionally credible in a modeled world but not yet strongly evidenced at scale unless tested prospectively with strong constraints and governance. This implies a research agenda that emphasizes “deployment science” for water infrastructure: not only algorithm design, but also testing pathways for safe adoption of models, monitoring model drift, and operator-centric integration.
To help the transition to treatment and network operations, the “simulation-to-field gap” can be strengthened by introducing the wastewater digital twin review literature that openly positions digital twins as an interface between mechanistic understanding, live operational data and closed-loop optimization—and at the same time flags up that real value depends on instrument quality, model governance, maintenance (with very long-run skills richest available) over time [20]. This is also consistent with Tier’s argument: optimization/control studies often remain at Tier 3, as they deliver in silico potential savings, contrasting with Tier 1 evidence.

5.4. Digital Twins as a Bridging Architecture: AI Is Becoming a System, Not a Standalone Model

This DT information in the corpus supports that AI value, also in water management, relies on efficient runtime embedding into an operational architecture that couples telemetry → models → analytics → decisions. It is also worth mentioning that the majority of the reviewed applications tend to arise from WDS and wastewater sectors; the number of applications related to some services, such as desalination and reuse, is less common, suggesting an unbalanced maturity along the water cycle but a mixed trend for DT when the research perspective is not only considered [11,20]. Crucially, DT studies change the unit of analysis: it is no longer a model itself but the full pipeline—data synchronization, latency, calibration, uncertainty resolution, and user interaction.
Recent work in the “digital water” and digital twin literature supports this re-framing by arguing that utilities’ value capture stems less from any individual algorithm than from the socio-technical system that maintains models synced up with assets, operators, and decision routines. In application, the digital twin has recently been characterized as an operational architecture that connects sensing, data governance, analytics, and work orders into a decision loop that is auditable and maintainable over time rather than a one-and-done modeling campaign [75,76]. This is important if the end-to-end pipeline is the unit of analysis; then, “effectiveness” depends on architecture-level factors (data quality controls, update cadence, exception handling, human-in-the-loop interfaces) that determine if analytics can be actioned reliably and at scale. In line with your maturity-gradient argument, a recent systematic review of smart water systems also indicates that digital solutions do not mature uniformly along the water cycle, achieving higher coverage levels in networked utility contexts than in less instrumented segments—underlining the need to approach digital twins as an infrastructure for decision and not merely a better prediction machine [55].
The case of ML-enabled digital twins suggests that system-level reproducibility and integration details (data pipeline, model update, operational use) are also at play in terms of adoption and credible claims for sustainability impact [53]. The DT-centric approach also articulates a practical distinction between near-term and long-term value:
Short term: monitoring, diagnostics, predictive maintenance (easier to validate and lower risk).
Longer term: closed-loop optimization and self-optimizing control (higher risk, more stringent governance needs).
This difference matches the larger theme of RL/control and underscores that the most compelling evidence is focused presently on decision support and monitoring as opposed to autonomy.

5.5. Governance, Privacy, and Trust as Adoption Constraints: Evidence Is Emerging but Uneven

An implicit theme that pervades the table is that the adoption of all forms of water infrastructure depends on more than technical performance: it depends on trust, legitimacy, and governance fit. Smart metering analytics exemplify this especially because the consumption data can encode sensitive household periodic behavior. Privacy-preserving “activity- and resolution-aware” that read how privacy preservation can be engineered without sacrificing analytic utility; a prerequisite to the adoption of responsible AI at scale [14]. On the other hand, meta-analyses of anomaly detection in smart metering reveal that despite advancements in the field, practical deployment concerns, including class imbalance, noise, and non-technical opportunities, prevail [35]. This has two sustainability implications:
Data governance is sustainability governance. Without appropriate privacy and acceptable use of the data, smart metering analytics cannot be widely deployed, thereby curtailing conservation results.
Seeing is not enough—processes of response are important. Anomaly detection does not mean anything in terms of savings unless utilities have validation, customer communication, and a routine to close the loop.
Recent technical work on privacy engineering in utility analytics further supports your claim that “data governance is sustainability governance” by operationalizing privacy as a design space with tangible trade-offs. E.g., activity- and resolution-aware privacy guards have emerged that can dampen inference risk from high-frequency utility data, yet retain downstream analytical utility—an approach that conforms closely to the responsible scaling of smart metering and sensing solutions [14]. And that is directly related to your adoption logic: privacy controls are regulatory, not just in the “check box” sense, but by providing the stamp of public legitimacy and organizational operation (or permit) that signals whether utilities can apply analytics at the depths of coverage needed to produce meaningful conservation or NRW results.
Another analogous theme is that of interpretability. The AI-augmented pump operation is formulated as interpretable decision support, which is in line with the requirements of operator-centric deployment, particularly when a high standard of safety and service reliability is to be maintained [33]. About interpretable and knowledge-based ML, in the treatment setting, it is straightforward to see a path for adopting such methods as controls can be encoded or/and/or decisions rendered auditable [20]. However, reviews on AI in chemical dosing also highlight that federated training is crucial: “interpretability, standardization and generalization are a constant issue even when the technical promise is high” [50].
Another governance problem revolves around the danger of algorithmic bias (and insufficient data variety) in an “AI-marketplace” that is controlled by a handful of private companies. If training sets reflect historical inequalities or are derived from a limited geographic or socio-economic group of people, then models may output discriminatory patterns unintentionally. This could result, for example, in service enhancements being concentrated in well-instrumented places and little consideration for those with no metering or from marginalized communities. Nelson suggests that water utilities to conduct fairness audits, involve a diverse set of stakeholders, and mandate more transparency from vendors on model development processes and data sources to mitigate such risks [77]. The incorporation of fair-aware design principles and mandatory open reporting of algorithmic provenance in procurement criteria would help encourage environmental as well as institutional sustainability of AI deployments.

5.6. How Sustainability Outcomes Are Operationalized: What the Literature Measures Versus What It Should Measure

A big point is a synthesis from me: what is the easiest to measure (accuracy, F1, AUC, RMSE) has diverged widely from what sustainability audiences need (water saved, energy saved, emissions reduced, compliance improved, resilience strengthened, inequities remedied). A few studies approach outcome-based reporting by connecting models to operational aims, such as energy use and effluent quality [41] or reporting measurable gains in treatment efficacy, and cost-related proxies [20]. However, the forecasting and anomaly-detection literatures are still mostly concerned with performance metrics that are not converted into resource or compliance values of common currency [27,35,46].
From a review article perspective, this means that the next generation primary research must include a minimum set of sustainability reporting elements, including at least:
Water: estimated volume savings (Avoided NRW, avoided sewer bursts, irrigation efficiency benefits).
Energy: Scheduling control vs. Energy changes (kWh/m3).
Chemicals: reduction in dosing amount per m3 treated with quality constraints.
Adherence: percentage of time in adherence or non-adherence.
Resilience: time to lead to the extreme; time to recover from disruption.
Equity (where relevant): spread of benefits in improved continuity of service, reduction in risk for the poor areas, and affordability effects.
In this regard, agricultural demand-side applications such as advanced ET0 estimates for irrigation scheduling are strategically important, wherein AI is connected to most of the water use in many basins to counter utility-centric WDS/WWTP evidence [52]. However, like utility settings, it is uncertain whether the sustainability benefit of better estimates is further reflected in altered irrigation behavior and measured efficiency outcomes.
In applied analytics, predictive gains do not directly translate into gains in decision, as operational value would require that predictions bring changes in action under real constraints (budgets, drivers’ availability, safety, and service targets) [62]. Which is exactly why outcome-based reporting should be thought of as an evidentiary standard rather than some sort of nice-to-have: in the absence of a definition on baseline, treatment description, and toward a plausible measurement-and-verification (M&V) pathway, performance measures are simply untethered from claims around resource/compliance/equity. So, “minimum reporting elements” can be framed as an alternative to a particular well-known but rarely acknowledged failure mode of evaluation—relying too heavily on predictive indicators without effectively vetting the decision path that leads from them to sustainability outcomes.
Changes in infrastructure and sustainability research (and recent publications) can also offer you specific, single-sector metrics that you are able to point to as the ‘common currency’ outcomes. For instance, using methodological reviews of WWTP energy benchmarking, it is emphasized that comparisons of energy intensity are not meaningful unless plants have been truly normalized (influent characteristics, treatment goals, and boundary definitions), which directly backs up your call for standardized reporting, as opposed to isolated proxies [78]. Similarly, new urban-water GHG studies reveal that the emissions profile is incredibly sensitive to system boundaries and “new water” options (e.g., desalinated, long-distance transfers), highlighting the need for GHG sustainability reporting to directly specify boundary definition and scenario assumptions [79]. Case-specific GHG accounting for integrated urban water systems is another issue where, especially characteristic of sanitation processes frequently dominating operational emissions, reporting of outcome metrics (tCO2e, kWh/m3, compliance time) becomes important in addition to predictive accuracy [80]. And from a global scale analysis of water treatment energy demand, energy burdens are massive at the system level and highly variable by technology class: making it clear that sustainability outcomes must be operationalized as energy and emissions linked metrics (not just model-fit statistics) is also clear already, [81].

5.7. Implications for Practice: A Staged Adoption Roadmap Grounded in the Evidence Base

Based on Table 1, utilities and agencies might read the evidence to support a phased implementation approach:
Stage 1: Readiness (monitoring and evidence for decision making)
Leak and burst detection, anomaly detection, pressure estimation/state inference, demand prediction, early warning [8,9,27,38,43,45,54].
Human-in-the-loop decision-making can be used to implement these applications, while sustainability gains are most likely when connected to response protocols.
Stage 2: Implement Core operational analytics (optimization with constraints)
Optimum pumping scheduling and aeration energy saving when the models are explicitly designed to achieve efficiency targets, with interpretable constraints and auditability [33,40,41].
But at this point, governance is about change management, operator training, and performance monitoring: see smart systems syntheses [10,55].
Deployment of a digital twin for system-level scaling
DT systems, which integrate sensing information, models, and analytics, may help reduce fragmentation and enable continuous improvement but need explicit attention for their data processing pipeline stages, including synchronization process, model updating, and operational utility [11,20,53].
Stage 3: Closed-loop AI control Risk factor: Severe (Evidence still being determined)
The RL-style scheduling/control and the DT-derived optimization shall be carefully applied, with strong safety constraints, fallback strategies, as well as in-field justification [12,34,49,56].
The reading both validates the potential of these strategies but also supports the cautious assertions for readiness for deployment.

5.8. Research Agenda from “More Models” to “Better Evidence”

According to the patterns and constraints repeatedly cited in all the included studies and reviews, five priorities were identified:
Develop a common framework for sustainability accounting and measurement-and-verification (M&V).
For example, it could be beneficial for evidence synthesis if studies consistently reported outcome conversions (water/energy/chemicals/compliance), not mere model metric ones [10,20,28].
Expand prospective and multi-site validation.
Generalizability remains a central gap. Multi-utility and multi-basin studies are necessary to assess portability, national calibration, and recalibration needs, more specifically in the case of detection and forecasting applications [8,9,55].
Control and RL: Closing the interface-to-reality gap from the simulator to the field.
RL, SLR, and domain reviews coincide in their findings of safety, constraints & transfer as the constraints [12,49]. Research should focus on constrained RL, human override, drift detection, and robust operational monitoring.
Develop hybrid physics–AI methods in data-poor or non-stationary regimes.
PINNs and PIF-based scheduling methods offer a promising route to enhance plausibility and robustness [29,40], particularly for cases with a low level of monitoring or where the hydroclimatic regimes are changing [47].
Strengthen privacy, interpretability, and governance-by-design.
Privacy-preserving smart metering schemes and transparent operation models are not optional requirements; they are underlying conditions for continued acceptance [14,20,33,50].

5.9. Limitations of the Evidence Base and of This Synthesis

Although this summary encompasses a diverse range of applications, the inferential capacity is tempered by heterogeneity in study designs and reporting quality. Indeed, several families—pages named control, driven by RL and digital twin, intuitive-heavy and domain-agnostic systems in the loop control—are currently simulation dominant or have limited case evidence [12,20,49,53]. Furthermore, the corpus is biased towards WDS and wastewater operationally focused contexts, fewer of which explicitly quantify distributional (equity) impacts or make direct SDG-aligned outcome mappings. Third, as most studies highlight the novelty of method comparison, what is feasible and sustainable might differ; this review necessarily distinguishes feasibility (model works in study) from impact that sustains operationally (models change decisions and outcomes), a distinction which may require further direct research.

6. Practical Implications

6.1. Implications for Water Utilities: Prioritizing “High Readiness” AI That Closes the NRW and Reliability Loop

Table 1 demonstrates that the most defensible near-term pathway for AI-enabled sustainability in water utilities is to focus on less ambitious applications with high deployability, including monitoring, detection, and short-horizon forecasting. The leakage family and the burst detection families are two of the most mature in distribution operations, but they need to provide practical value, i.e., effectiveness under realistic sensing constraints, false-alarm tolerances, and localization requirements to be able to quickly repair [8,9]. The trend to topology-aware learning is also indicative of utilities having an inclination to treat distribution data as connected—rather than independent—such that graph-based leakage detection and pressure estimation methods expect to leverage the dependencies present, for example, in enhancing diagnostics and state inference when instrumented monitoring is imperfect [37,43]. Burst localization studies contribute to a fundamental procurement consideration for many utilities—the lack of sensors—by focusing on detectability under sparse coverage [54]. However, the sustaining effect is condition-based unless utilities implement a full chain from detection to verification and dispatch. On the other hand, ML-enabled asset prioritization and failure prediction are pragmatically useful, as historical failure patterns can be translated into proactive measures that mitigate bursts and losses –but only if seamlessly integrated within work-order systems and capital renewal planning instead of as isolated analytics [31,59].

6.2. Implications for Demand and Supply Planning: Translating Forecasting Skills into Resource Efficiency

Demand–supply forecasting is presented in the evidence base as a high-readiness family with clear operational relevance, since a demand–supply mismatch results in later unneeded energy consumption, chemical dosing variability, and pressure instability that in turn can contribute to further leakages and pipe stress. Deep flow sequence and hybrid model correction workflows are developed to minimize short-term forecasting errors required for operational planning, such as production scheduling and downstream pump operations [27,45,46]. The upshot is that utilities should be checking the performance of forecasting systems not only on statistical accuracy but also on efficiency indicators available further downstream, which capture value in terms of sustainability (e.g., changes in kWh per cubic meter, chemical use per cubic meter treated, peak-load exposure, or reductions in instability-driven operational events). If this is not made explicit in their translation from forecast improvement to operational outcomes, then claims of sustainability are likely to remain implicit rather than evidenced.

6.3. Implications for Drinking Water and Wastewater Treatment: Adopting Constraint-Aware and Interpretable AI for Energy–Quality Coupling

Treatment settings are appealing for AI-enabled sustainability because results can be readily formulated as energy intensity, chemical usage, and compliance performance. The use of AI can be used to simultaneously optimize effluent quality and energy usage; however, this is particularly true for aeration processes where large quantities of energy are required and where regulatory boundaries are of utmost significance [41]. For actual deployment, this means that optimization should be deployed on some set of explicit constraints and auditable logic, because “energy savings” are not fungible, where you can trade one against your regulatory exposure. Interpretable as well as knowledge-embedded ML methods offer a more operationally feasible path by satisfying constraints and allowing decision transparency, which becomes even more critical in safety and compliance-sensitive domains [20]. In line with this, recent reviews on AI in chemical dosing highlight enduring challenges such as standardization, interpretability, and generalizability [50], which suggests that simple black-box dosing solutions should not be scaled without ensuring drift monitoring, conservative constraints, and operator-oriented validation. For all defense that R is imagined not cavalierly, but as a position roughly and in ideology among the plausible positions that could be adopted, it seems worth attending to quiet data points such as how few of those myriad “real-time control” claims are still in crescent in data-driven simulation domains rather than backed by continued in-situ validation [34], what the focus of safety, transfer, constraint handling, and monitoring should be for infrastructure adoption when one reads between the lines of a broader R synthesis [12].

6.4. Implications for Basin Authorities and Risk Managers: Strengthening Early Warning and Groundwater Decision Support Under Non-Stationarity

In catchment and basin scales, the most credible benefits are likely to manifest when AI is used as a means of reinforcing early warning and anticipatory planning, rather than direct automation. Flood and flashflood forecast systems have the potential for increasing warning lead times and enhancing coordination of response, but their benefit to sustainable development is contingent on their entrapment within operational protocols for issuing warning orders, deploying emergency action plans, and learning from past events [32,57]. IoT-based telemetry of water levels can improve near-real-time preparedness; however, this again requires decision-making integration to deliver tangible resilience [44]. Drought forecasting and trend identification based on long historical sequences is a more robust approach for planning purposes when compared with short-window models, but the implication is that an evaluation must explicitly consider non-stationarity and decision value in varying hydroclimatic conditions [47]. Groundwater applications demonstrate that remote sensing and ML have utility in decision-making where in situ monitoring is sparse, whereas physics-informed approaches allow the opportunity for physically consistent inference at data-sparse locations [29,30,42]. In practical terms, the priority of uptake is not just to produce maps or predictions but to connect these outputs into rules for allocation, which in turn allow decisions and the design of monitoring networks that can validate models through time.

6.5. Implications for Digital Twin Programs: Treating AI as a System-Level Capability Rather than a Standalone Model

Digital twins are not a single analytic product, but rather an integration architecture of telemetry through process/hydraulic models to analytics and decision interfaces. Reviews suggest that the maturity of DT is most prominent in distribution systems and wastewater, but relatively limited for other domains, e.g., desalination and reuse, and this indicates that adoption should be on a case-by-case basis [11,20]. The implication here is that for utilities that are interested in utilizing DTs, they will instead need to focus on system-level design requirements—pipeline reliability, synchronization and control, and latency management (e.g., timekeeping), calibration routines, uncertainty handling, and user interaction—as these will serve as a gatekeeper to whether DTs can credibly support sustainability impacts. Empirical case studies on ML-enabled DTs, moreover, stress the importance of system-level documentation and operational integration for reproducibility, governance, and ongoing performance monitoring [53].

6.6. Implications for Governance, Privacy, and Trust: Enabling Adoption Without Undermining Legitimacy

The literature base suggests governance and legitimacy dimensions are directly relevant to whether AI-enabled sustainability can scale, given the smart metering context where household behavior is inferable from patterns of consumption. Privacy-preserving approaches that secure activity and resolution level data while retaining analytic utility offer pathways for responsible adoption and broader public acceptability [14]. And systematic reviews of anomaly detection in smart metering also demonstrate that class imbalance, noise, and deployability are central constraints too, such that the sustainability value of detection systems rests upon response protocols to determine validation, customer communication, and repair or enforcement interventions [35]. In application, principles of governance and sustainability become the same: without assuredness over privacy and appropriate use of data, the most scalable conservation and loss-reduction analytics will not be deployable at scale.

6.7. Measurement, Verification, and Procurement: Aligning Contracts and KPIs with Sustainability Outcomes

A common real-world limitation across Table 1 is that the literature base frequently reports varying model performance without adjusting its capacity to translate performance gains into validated sustainability outcomes. Smart water systems and WDS applications reviews continually highlight that the translation gap is not an issue of algorithm availability but instead a problem of data constraints, interoperability, or operationalization [10,55]. As a result, it is up to practitioners to demand that all deployments explicitly identify the operational lever it seeks to modify, the sustainability outcome at which it aims, and the counterfactual strategy used to test for impact. This need is especially true in optimization and RL control settings, where simulation-to-field transfer remains a key bottleneck, and where statements of savings or reductions should come with conditions (i.e., “may save up to” or “could reduce emissions”) until tested prospectively under robust constraints and monitoring [12,34,49]. From a procurement point of view, sustainability-focused contracts should thus include deployability rather than just model metrics (including measurable interoperability), documentation of pipelines as well as failure modes, indications for human override, and outcome-driven performance commitments.

7. Conclusions

In this paper, the current state of knowledge on AI-enabled approaches to enhancing sustainability in the complete water cycle (i.e., aiming both toward decentralized systems for treating drinking water and wastewater) as well as decision support systems at basin and catchment scales is discussed. Applying a PRISMA 2020 method for selecting evidence and an evidence mapping approach, the research included 41 peer-reviewed articles as part of the review, supplemented by selected grey literature reviews, which were appraised independently.
One of the key lessons from this review is that AI’s sustainability potential is best achieved if it is part of larger operational decision-making frameworks, governance systems, and measurement-and-verification protocols. The role of AI must move from purely seeking technical improvements in model acuity to a broader embedding within contextualized, actionable institutional environments.
The results show a strong technological maturity gradient. The strongest, most deployable evidence relates to applications of monitoring, detection, and forecasting—including leak detection, pipe burst analytics, network state estimation, anomaly identification, and short-range demand and supply forecasting in networks. These applications are also becoming closer to actual deployment since they can be integrated into human-in-the-loop systems and are more in alignment with operational protocols, and they provide logical pathways between early warnings and better resource efficiency and reliability, at least if utilities have the capability to act upon the insights [8,9,27,43,54].
Yet even in these high-readiness sectors, reported sustainability benefits are typically implied rather than empirically measured. A few analyses highlight improved detection accuracy or forecasting precision without repeatedly translating these effects into measured savings in non-revenue water, energy consumption, chemical use, or service disruptions [10,55].
In treatments, the references more commonly relate AI to such sustainability-relevant targets as energy quality, economy, and compliance consistency, as well as chemical efficiency. Here, the most operationally feasible strategies fall back to constrained and interpretable decision support, acknowledging that safety, regulatory compliance, and operator responsibility place demand on adoption expected beyond predictive performance [20,41,50]. These results imply that “responsible deployability”—interpretability, auditable constraints, ongoing monitoring—should be treated as a first-order design criterion for AI use in the treatment context.
On the other hand, optimization in control applications (notably with reinforcement learning) is promising but less applicable across most of the existing evidence. Perhaps the most pervasive of these constraints is the simulation-to-field gap: existing control-centric works tend to evaluate performance chiefly in simulated or data-driven emulation-based environments, as opposed to sustained situ operation wherein safety limiters, human override, institutional practices, and failure implications are prevalent. The review thus cannot but recommend treating all claims regarding energy savings, reduced emissions, or increased system reliability resulting from RL-driven or closed-loop AI control as conditional unless prospective validation under operational governance and constraint regimes [12,34,49] are presented.
Digital twins appear as a mediating architecture transferring the unit of analysis from model to system. The established fact is that the most effective digital twin advancements are obtained when telemetry, hydraulic, and/or process models, analytics layers, and decision interfaces come together to make a cohesive operational pipeline supported by clear synchronization, updating, and validation mechanisms [11,20,53]. The result of this is the recognition that digital twin programs should not be seen simply as single-tool rollouts, but rather as multi-year integration and governance processes—adoption success being a function of data infrastructure availability, interoperability effectiveness, and organizational ability as much (or perhaps even more) than modeling.
In all areas of application, governance limitations—privacy, legality, comprehensibility, and fit in institutions—are not marginal. Smart metering analytics demonstrate that privacy-preserving design itself may operate as a precondition for scaling conservation-oriented analytics, while systematic evidence shows how practical deployment constraints, including imbalanced data, noise, and response protocols, determine if anomaly detection will lead to real savings or unintended harm [14,35]. These results provide further evidence that sustainability governance in digital water can incorporate data governance.
Collectively, the review provides four general conclusions. To start, AI-enabled sustainable water management should be evaluated based on system performance and the validated outputs—water saved, reduced energy and chemical intensity of water treatment operations, improved compliance and resilience—rather than prediction accuracy alone. Secondly, the most defensible near-term roll-out plan is gradual: scale up from monitoring and decision support first, move into resource-constrained optimization where auditability is possible, deploy digital twins where system integration can be shown, and experiment very carefully in pilots using AI closed loops with substantial safety back stops. Thirdly, a more standardized sustainability reporting set would be desirable, which bridges outputs from the model to measurement and verification metrics, to ensure better cross-comparability and robust evidence synthesis. Fourth, future research should move away from developing more models and focus on creating better evidence, including through multi-site, prospective validation, explicit handling of non-stationarity, and governance-by-design, especially for high-stakes control applications.
Finally, although the review shows that AI may in principle support sustainability ambitions in water resources management, the overarching message of the evidence base is that any actual impact at scale rests upon operational integration, governance capacity, and rigorous outcome accounting. Accordingly, moving the field forward will involve more closely integrating AI innovation, deployment science, utility-grade engineering, and sustainability-oriented measurement to ensure AI advances not only lead to better predictions but also result in demonstrable improvement to water system performance.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/su18042154/s1, Table S1: PRISMA 2020 checklist [17].

Funding

The research receives no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study.

Conflicts of Interest

The author declares no conflicts of interest.

References

  1. WMO. State of Global Water Resources Report 2024; World Meteorological Organization (WMO): Geneva, Switzerland, 2025. [Google Scholar] [CrossRef]
  2. UNESCO World Water Assessment Programme; Koncagül, E.; Connor, R.; Abete, V. The United Nations World Water Development Report 2024: Water for Prosperity and Peace; UNESCO: Paris, France, 2024. [Google Scholar]
  3. Division, U.S. The Sustainable Development Goals Report 2024; United Nations: New York, NY, USA, 2024; pp. 1–48. [Google Scholar]
  4. United Nations. Goal 6: Ensure Access to Water and Sanitation for All. Available online: https://www.un.org/sustainabledevelopment/water-and-sanitation/ (accessed on 23 December 2025).
  5. UN-Water. Progress on the Proportion of Domestic and Industrial Wastewater Flows Safely Treated; World Health Organization: Geneva, Switzerland, 2024; pp. 1–102. [Google Scholar]
  6. UN Environment Programme. Mid-Term Status on SDG 6 Indicators: 6.3.2, 6.5.1, & 6.6.1. Available online: https://www.unep.org/resources/report/mid-term-status-sdg-6-indicators-632-651-661-2024 (accessed on 23 December 2025).
  7. Kingdom, B.; Roland, L.; Philippe, M. The Challenge of Reducing Non-Revenue Water (NRW) in Developing Countries; World Bank: Washington, DC, USA, 2006; pp. 1–40. [Google Scholar]
  8. Farah, E.; Shahrour, I. Water Leak Detection: A Comprehensive Review of Methods, Challenges, and Future Directions. Water 2024, 16, 2975. [Google Scholar] [CrossRef]
  9. Farah, E.; Shahrour, I. Use of Data-Driven Methods for Water Leak Detection and Consumption Analysis at Microscale and Macroscale. Water 2024, 16, 2530. [Google Scholar] [CrossRef]
  10. Taloma, R.J.L.; Cuomo, F.; Comminiello, D.; Pisani, P. Machine learning for smart water distribution systems: Exploring applications, challenges and future perspectives. Artif. Intell. Rev. 2025, 58, 120. [Google Scholar] [CrossRef]
  11. Bam, P.G.; Rezaei, N.; Roubanis, A.; Austin, D.; Austin, E.; Tarroja, B.; Takacs, I.; Villez, K.; Rosso, D. Digital Twin Applications in the Water Sector: A Review. Water 2025, 17, 2957. [Google Scholar] [CrossRef]
  12. Kåge, L.; Milić, V.; Andersson, M.; Wallén, M. Reinforcement learning applications in water resource management: A systematic literature review. Front. Water 2025, 7, 1537868. [Google Scholar] [CrossRef]
  13. Song, Y.; Knoben, W.J.M.; Clark, M.P.; Feng, D.; Lawson, K.; Sawadekar, K.; Shen, C. When ancient numerical demons meet physics-informed machine learning: Adjoint-based gradients for implicit differentiable modeling. Hydrol. Earth Syst. Sci. 2024, 28, 3051–3077. [Google Scholar] [CrossRef]
  14. Cardell-Oliver, R.; Cominola, A.; Hong, J. Activity and resolution aware privacy protection for smart water meter databases. Internet Things 2024, 25, 101130. [Google Scholar] [CrossRef]
  15. U.S. Environmental Protection Agency. EPA Guidance on Improving Cybersecurity at Drinking Water and Wastewater Systems; U.S. Environmental Protection Agency: Washington, DC, USA, 2024; pp. 1–12. [Google Scholar]
  16. Toderas, M. Artificial Intelligence for Sustainability: A Systematic Review and Critical Analysis of AI Applications, Challenges, and Future Directions. Sustainability 2025, 17, 8049. [Google Scholar] [CrossRef]
  17. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, 21. [Google Scholar] [CrossRef]
  18. The World Bank. Reducing Nonrevenue Water and Improving Energy Efficiency. Available online: https://documents1.worldbank.org/curated/en/539681593009432732/pdf/Guidance-Note.pdf (accessed on 23 December 2025).
  19. Slater, L.; Blougouras, G.; Deng, L.; Deng, Q.; Ford, E.; van Dijke, A.H.; Huang, F.; Jiang, S.; Liu, Y.; Moulds, S.; et al. Challenges and opportunities of ML and explainable AI in large-sample hydrology. Philos. Trans. R. Soc. A Math. Phys. Eng. Sci. 2025, 383, 20240287. [Google Scholar] [CrossRef]
  20. Wang, A.J.; Li, H.; He, Z.; Tao, Y.; Wang, H.; Yang, M.; Savic, D.; Daigger, G.T.; Ren, N. Digital Twins for Wastewater Treatment: A Technical Review. Engineering 2024, 36, 21–35. [Google Scholar] [CrossRef]
  21. National Institute of Standards and Technology. Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1); National Institute of Standards and Technology: Gaithersburg, MD, USA, 2023. [Google Scholar]
  22. OECD. Recommendation of the Council on OECD Legal Instruments Artificial Intelligence. Available online: https://legalinstruments.oecd.org/api/print?ids=648&lang=en (accessed on 23 December 2025).
  23. Hong, Q.N.; Fàbregues, S.; Bartlett, G.; Boardman, F.; Cargo, M.; Dagenais, P.; Gagnon, M.P.; Griffiths, F.; Nicolau, B.; O’Cathain, A.; et al. The Mixed Methods Appraisal Tool (MMAT) version 2018 for information professionals and researchers. Sage J. 2018, 34, 285–291. [Google Scholar] [CrossRef]
  24. CASP. C.A.S.P. CASP Checklist: CASP Qualitative Studies Checklist. Available online: https://casp-uk.net/casp-tools-checklists/qualitative-studies-checklist/ (accessed on 23 December 2025).
  25. Shea, B.J.; Reeves, B.C.; Wells, G.; Thuku, M.; Hamel, C.; Moran, J.; Moher, D.; Tugwell, P.; Welch, V.; Kristjansson, E.; et al. AMSTAR 2: A critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both. BMJ 2017, 358, 4008. [Google Scholar] [CrossRef]
  26. Whiting, P.; Savović, J.; Higgins, J.P.T.; Caldwell, D.M.; Reeves, B.C.; Shea, B.; Davies, P.; Kleijnen, J.; Churchill, R. ROBIS: A new tool to assess risk of bias in systematic reviews was developed. J. Clin. Epidemiol. 2016, 69, 225–234. [Google Scholar] [CrossRef] [PubMed]
  27. Shan, S.; Ni, H.; Chen, G.; Lin, X.; Li, J. A Machine Learning Framework for Enhancing Short-Term Water Demand Forecasting Using Attention-BiLSTM Networks Integrated with XGBoost Residual Correction. Water 2023, 15, 3605. [Google Scholar] [CrossRef]
  28. Abunama, T.; Dellieu, A.; Nonet, S. Advancements in machine learning modelling for energy and emissions optimization in wastewater treatment plants: A systematic review. Water Environ. J. Promot. Sustain. Solut. 2024, 38, 554–572. [Google Scholar] [CrossRef]
  29. Ali, A.S.A.; Jazaei, F.; Clement, T.P.; Waldron, B. Physics-informed neural networks in groundwater flow modeling: Advantages and future directions. Groundw. Sustain. Dev. 2024, 25, 101172. [Google Scholar] [CrossRef]
  30. Elmotawakkil, A.; Sadiki, A.; Enneya, N. Predicting groundwater level based on remote sensing and machine learning: A case study in the Rabat-Kénitra region. J. Hydroinformatics 2024, 26, 2639–2667. [Google Scholar] [CrossRef]
  31. Farajzadeh, N.; Sadeghzadeh, N.; Jokar, N. Water distribution pipe lifespans: Predicting when to repair the pipes in municipal water distribution networks using machine learning techniques. PLoS Water 2024, 3, e0000164. [Google Scholar] [CrossRef]
  32. Hahn, Y.; Kienitz, P.; Wönkhaus, M.; Meyes, R.; Meisen, T. Towards Accurate Flood Predictions: A Deep Learning Approach Using Wupper River Data. Water 2024, 16, 3368. [Google Scholar] [CrossRef]
  33. Marzouny, N.H.; Dziedzic, R. AI-Assisted Pump Operation for Energy-Efficient Water Distribution Systems. Eng. Proc. 2024, 69, 3. [Google Scholar] [CrossRef]
  34. Hu, F.; Zhang, X.; Lu, B.; Lin, Y. Real-time control of A2O process in wastewater treatment through fast deep reinforcement learning based on data-driven simulation model. Water 2024, 16, 3710. [Google Scholar] [CrossRef]
  35. Kanyama, M.N.; Bhunu Shava, F.; Gamundani, A.M.; Hartmann, A. Machine learning applications for anomaly detection in Smart Water Metering Networks: A systematic review. Phys. Chem. Earth Parts A/B/C 2024, 134, 103558. [Google Scholar] [CrossRef]
  36. Li, Z.; Ma, W.; Zhong, D.; Ma, J.; Zhang, Q.; Yuan, Y.; Liu, X.; Wang, X.; Zou, K. Applications of machine learning in drinking water quality management: A critical review on water distribution system. J. Clean. Prod. 2024, 481, 144171. [Google Scholar] [CrossRef]
  37. Li, X.; Wu, Y. A Convolutional Graph Neural Network Model for Water Distribution Network Leakage Detection Based on Segment Feature Fusion Strategy. Water 2024, 16, 3555. [Google Scholar] [CrossRef]
  38. Liu, J.; Wu, D.; Mohammed, H.; Seidu, R.; Liu, J.; Wu, D.; Mohammed, H.; Seidu, R. A Novel Method for Anomaly Detection and Signal Calibration in Water Quality Monitoring of an Urban Water Supply System. Water 2024, 16, 1238. [Google Scholar] [CrossRef]
  39. Luo, W.; Huang, L. Prediction-based Real-time Anomaly Detection for Water Quality Time Sequences. In Proceedings of the 4th International Conference on Computational Modeling, Simulation and Data Analysis (CMSDA ’24), Hangzhou, China, 6–8 December 2024; Association for Computing Machinery: New York, NY, USA, 2025; pp. 364–371. [Google Scholar]
  40. Ma, H.; Wang, X.; Wang, D. Pump Scheduling Optimization in Urban Water Supply Stations: A Physics-Informed Multiagent Deep Reinforcement Learning Approach. Int. J. Energy Res. 2024, 2024, 9557596. [Google Scholar] [CrossRef]
  41. Mao, Z.; Li, X.; Zhang, X.; Li, D.; Lu, J.; Li, J.; Zheng, F. Optimization of effluent quality and energy consumption of aeration process in wastewater treatment plants using artificial intelligence. J. Water Process Eng. 2024, 63, 105384. [Google Scholar] [CrossRef]
  42. Sarkar, S.K.; Rudra, R.R.; Talukdar, S.; Das, P.C.; Nur, M.S.; Alam, E.; Islam, M.K.; Islam, A.R.M.T. Future groundwater potential mapping using machine learning algorithms and climate change scenarios in Bangladesh. Sci. Rep. 2024, 14, 10328. [Google Scholar] [CrossRef]
  43. Truong, H.; Tello, A.; Lazovik, A.; Degeler, V. Graph Neural Networks for Pressure Estimation in Water Distribution Systems. Water Resour. Res. 2024, 60, e2023WR036741. [Google Scholar] [CrossRef]
  44. Widiasari, I.R.; Efendi, R. Utilizing LSTM-GRU for IOT-Based Water Level Prediction Using Multi-Variable Rainfall Time Series Data. Informatics 2024, 11, 73. [Google Scholar] [CrossRef]
  45. Zhang, Y.; Li, J.; Sun, S.; Li, G.; Yang, Q.; Sun, Y.; Wang, X.; Xu, C. Short-Term Water Supply Forecasting for Water Treatment Plant Using Temporal Multi-Scale Features. Water 2024, 16, 3573. [Google Scholar] [CrossRef]
  46. Zubaidi, S.L.; Al-Bugharbee, H.; Alattabi, A.W.; Ridha, H.M.; Hashim, K.; Al-Ansari, N.; Yaseen, Z.M. Forecasting urban water demand using different hybrid-based metaheuristic algorithms’ inspire for extracting artificial neural network hyperparameters. Sci. Rep. 2024, 14, 24042. [Google Scholar] [CrossRef]
  47. Bouaziz, M.; Abid, M.A.; Medhioub, E.; John, A. A Century of Data: Machine Learning Approaches to Drought Prediction and Trend Analysis in Arid Regions. Water 2025, 17, 3567. [Google Scholar] [CrossRef]
  48. Ding, X.; Chen, Y.; Zeng, H.; Du, Y. Time Series Prediction of Water Quality Based on NGO-CNN-GRU Model—A Case Study of Xijiang River, China. Water 2025, 17, 2413. [Google Scholar] [CrossRef]
  49. Javed, A.; Wu, W.; Sun, Q.; Dai, Z. Leak Management in Water Distribution Networks Through Deep Reinforcement Learning: A Review. Water 2025, 17, 1928. [Google Scholar] [CrossRef]
  50. Jin, J.; Liu, M.; Chen, B.; Wu, X.; Yao, L.; Wang, Y.; Xiong, X.; Wei, L.; Li, J.; Tan, Q.; et al. Artificial Intelligence in Chemical Dosing for Wastewater Purification and Treatment: Current Trends and Future Perspectives. Separations 2025, 12, 237. [Google Scholar] [CrossRef]
  51. Kanyama, M.N.; Bhunu Shava, F.; Gamundani, A.M.; Hartmann, A. AI-Driven Anomaly Detection in Smart Water Metering Systems Using Ensemble Learning. Water 2025, 17, 1933. [Google Scholar] [CrossRef]
  52. Khandappa, P.K.; Haladappa, M.S. Improving Irrigation Scheduling through Deep Learning-Based Reference Evapotranspiration Estimation. Eng. Technol. Appl. Sci. Res. 2025, 15, 30185–30190. [Google Scholar] [CrossRef]
  53. Ma, Z.; Zhu, Y.; Chen, C.; Li, T.; Li, Y.; Li, X.; Wang, Y.; Waite, T.D.; Guan, J. Towards the digitalization of water treatment facilities: A case study on machine learning-enabled digital twins. J. Water Process Eng. 2025, 77, 108316. [Google Scholar] [CrossRef]
  54. Min, K.; Kim, J.H.; Jung, D.; Lee, S.; Kang, D. Pipe Burst Detection and Localization in Water Distribution Networks Using Faster Region-Based Convolutional Neural Network. Water 2025, 17, 3380. [Google Scholar] [CrossRef]
  55. Quintana, D.; Felix-Herran, L.C.; Tudon-Martinez, J.C.; Lozoya-Santos, J.d.J. On Smart Water System Developments: A Systematic Review. Water 2025, 17, 2571. [Google Scholar] [CrossRef]
  56. Pei, S.; Hoang, L.; Fu, G.; Butler, D. Real-Time Pump Scheduling in Water Distribution Networks Using Deep Reinforcement Learning. J. Water Resour. Plan. Manag. 2025, 151, 04025012. [Google Scholar] [CrossRef]
  57. Soares, J.A.; Ozelim, L.C.; Bacelar, L.; Ribeiro, D.B.; Stephany, S.; Santos, L.B. ML4FF: A machine-learning framework for flash flood forecasting applied to a Brazilian watershed. J. Hydrol. 2025, 652, 132674. [Google Scholar] [CrossRef]
  58. Sseguya, F.; Jun, K.S. Deep Reinforcement Learning for Optimized Reservoir Operation and Flood Risk Mitigation. Water 2025, 17, 3226. [Google Scholar] [CrossRef]
  59. Yılmaz, S. Failure Analysis and Machine Learning-Based Prediction in Urban Drinking Water Systems. Appl. Sci. 2025, 15, 12887. [Google Scholar] [CrossRef]
  60. BeChained. Optimizing Water Pump Operations Using BeChained AI. Available online: https://bechained.ai/use-case-water-pump-optimition (accessed on 8 February 2026).
  61. Inc., X. Utility Saves an Average of 7 Million Gallons of Water Per Day by Utilizing Innovative Technologies to Support Their Leak Detection Program. Available online: https://www.xylem.com/en-in/resources/case-studies/utility-saves-an-average-of-7-million-gallons-of-water-per-day-by-utilizing-innovative-technologies-to-support-their-leak-detection-program/ (accessed on 8 February 2026).
  62. Shmueli, G. To explain or to predict? Stat. Sci. 2010, 25, 289–310. [Google Scholar] [CrossRef]
  63. Efficiency Valuation Organization (EVO). International Performance Measurement and Verification Protocol (IPMVP): Core Concepts (EVO 10000-1:2016). Available online: https://evo-world.org/images/corporate_documents/Evo-Guides_Family-v05-12mars2019-page-low-res.pdf?utm_source= (accessed on 23 December 2025).
  64. Lambert, A.; Hirner, W. Losses from Water Supply Systems: Standard Terminology and Recommended Performance Measures; IWA Publishing: London, UK, 2000. [Google Scholar]
  65. Alegre, H.; Baptista, J.M.; Cabrera, E., Jr.; Cubillo, F.; Duarte, P.; Hirner, W.; Merkel, W.; Parena, R. Performance Indicators for Water Supply Services, 2nd ed.; IWA Publishing: London, UK, 2006. [Google Scholar]
  66. American Water Works Association. Water Loss Control. American Water Works Association. Available online: https://www.awwa.org/resource/water-loss-control/ (accessed on 23 December 2025).
  67. Puust, R.; Kapelan, Z.; Savic, D.A.; Koppel, T. A review of methods for leakage management in pipe networks. Urban Water J. 2010, 7, 25–45. [Google Scholar] [CrossRef]
  68. Romero-Ben, L.; Alves, D.; Blesa, J.; Cembrano, G.; Puig, V.; Duviella, E. Leak detection and localization in water distribution networks: Review and perspective. Annu. Rev. Control 2023, 55, 392–419. [Google Scholar] [CrossRef]
  69. Donkor, E.A.; Mazzuchi, T.A.; Soyer, R.; Roberson, J.A. Urban water demand forecasting: Review of methods and models. J. Water Resour. Plan. Manag. 2014, 140, 146–159. [Google Scholar] [CrossRef]
  70. House-Peters, L.A.; Chang, H. Urban water demand modeling: Review of concepts, methods, and organizing principles. Water Resour. Res. 2011, 47, W05401. [Google Scholar] [CrossRef]
  71. Schroer, H.W.; Just, C.L. Feature Engineering and Supervised Machine Learning to Forecast Biogas Production during Municipal Anaerobic Co-Digestion. ACS EST Eng. 2023, 4, 660–672. [Google Scholar] [CrossRef]
  72. Milly, P.C.D.; Betancourt, J.; Falkenmark, M.; Hirsch, R.M.; Kundzewicz, Z.W.; Lettenmaier, D.P.; Stouffer, R.J. Stationarity is dead: Whither water management? Science 2008, 319, 573–574. [Google Scholar] [CrossRef] [PubMed]
  73. Shen, C. A transdisciplinary review of deep learning research and its relevance for water resources scientists. Water Resour. Res. 2018, 54, 8558–8593. [Google Scholar] [CrossRef]
  74. Kumar, N.; Patel, P.; Singh, S.; Goyal, M.K.; Kumar, N.; Patel, P.; Singh, S.; Goyal, M.K. Understanding non-stationarity of hydroclimatic extremes and resilience in Peninsular catchments, India. Sci. Rep. 2023, 13, 12524. [Google Scholar] [CrossRef]
  75. Expósito, A.; Cebollero, E.D. Digital revolution reshaping water management: Policy recommendations. Util. Policy 2025, 86, 101653. [Google Scholar] [CrossRef]
  76. Grigg, N.S. Digital transformation in water utilities: Status, challenges, and prospects. Smart Cities 2025, 8, 99. [Google Scholar] [CrossRef]
  77. Nelson, J. Governance and Ethics in AI Adoption for Water Utilities. Available online: https://www.trinnex.io/insights/governance-and-ethics-in-ai-adoption-for-water-utilities (accessed on 8 February 2026).
  78. Gallo, M.; Malluta, D.; Del Borghi, A.; Gagliano, E. A critical review on methodologies for the energy benchmarking of wastewater treatment plants. Sustainability 2024, 16, 1922. [Google Scholar] [CrossRef]
  79. Yan, G.; Kenway, S.J.; Lam, K.L.; Lant, P.A. Greenhouse gas emission dynamics and trajectories in urban water supply and wastewater systems. Water Res. 2025, 275, 123153. [Google Scholar] [CrossRef]
  80. Shim, I. Assessing greenhouse gas emissions in urban water management scenarios: Analysis for mitigation. Sustainability 2025, 17, 1959. [Google Scholar] [CrossRef]
  81. Magni, M.; Jones, E.R.; Bierkens, M.F.; van Vliet, M.T. Global energy consumption of water treatment technologies. Water Res. 2025, 277, 123245. [Google Scholar] [CrossRef] [PubMed]
Figure 1. PRISMA flow chart diagram.
Figure 1. PRISMA flow chart diagram.
Sustainability 18 02154 g001
Table 1. Selected literature.
Table 1. Selected literature.
IDTitleType of DocumentAuthors and DateKey Findings Relevant to AI-Enabled Sustainability in Water Management
1A Machine Learning Framework for Enhancing Short-Term Water Demand Forecasting Using Attention-BiLSTM Networks Integrated with XGBoost Residual CorrectionJournal articleShan et al. (2023) [27]Suggests a deep learning add demand-forecasting pipeline (attention and sequence modeling) with residual correction for short-term improvements—aiding in operational planning, a smoother pressure/production management, and further downstream energy and water-efficiency gains through scheduling.
2Advancements in machine learning modeling for energy and emissions optimization in wastewater treatment plants: A systematic reviewJournal article (review)Abunama et al. (2024) [28]A scoping review specifically concentrating on ML algorithms for energy/consumption optimization in real world operation data levels, organizing approaches and pointing to evidence/research-to-practice gaps as being of interest to utilities that are interested in transferring models into actual efficiency savings.
3Physics-Informed Neural Networks in Groundwater Flow Modeling: Advantages and Future DirectionsJournal article (review/tutorial)Ali (2024) [29]Scrutinizes the PINN methodology for groundwater problems highlighting advantages of meshless formulations and exposing potential for integrating governing physics into learning (applicable to data-poor aquifer management and more physically consistent inference).
4Activity- and Resolution-Aware Privacy Protection for Smart Water Meter DatabasesJournal articleCardell-Oliver et al. (2024) [14]Suggests privacy protection mechanisms for smart metering databases that retain analytic utility—directly applicable to responsible AI enablement and governance and data-sharing terms for sustainability analytics.
5Predicting groundwater level based on remote sensing and machine learning: a case study in the Rabat-Kénitra regionJournal articleElmotawakkil et al. (2024) [30]Leverages those REMS products (i.e., GRACE/MODIS-like variables) with ML to estimate groundwater-levels as a tool for anticipatory planning and climate-sensitive monitoring, which will be particularly useful in regions where the availability of in situ well networks is limited.
6Water distribution pipe lifespans: Predicting when to repair the pipes in municipal water distribution networks using machine learning techniquesJournal articleFarajzadeh et al. (2024) [31]Leverages one-class ML methods (e.g., OC-SVM, Isolation Forest) for the detection of pipes in need of repair employing municipality supplied data sets; introduces a low instrumentation approach to proactive maintenance with reductions in bursts/leaks and water saving potential.
7Use of Data-Driven Methods for Water Leak Detection and Consumption Analysis at Microscale and MacroscaleJournal articleFarah & Shahrour (2024a) [9]Presents a data-driven approach to leak detection + consumption profiling; This relied on AMR-type telemetry and pattern-based logic to find leaks; this aspect enables for the identification of leak events occurring on various temporal scales; Supports NRW reduction and operational targeting.
8Water Leak Detection: A Comprehensive Review of Methods and Future DirectionsJournal article (review)Farah & Shahrour (2024b) [8]Synthesizes AI-empowered leak identification (signal-based, signal–data-driven and hybrid methods) illuminating performance measures, sensor-coverage tradeoffs, and deployment limitations that impact reduction in non-revenue water.
9Towards Accurate Flood Predictions: A Deep Learning Approach Using Wupper River DataJournal articleHahn et al. (2024) [32]Uses deep learning for forecasting floods in the river-catchment, toward climate-resilience and disaster-risk reduction, to help enhance planning of preparedness/response by water-resource managers.
10AI-Assisted Pump Operation for Energy-Efficient Water Distribution SystemsConference proceeding (MDPI)Hedaiaty Marzouny et al. (2024) [33]Introduces an AI-enabled interpretable framework for pump operation/scheduling, defining the problem of decision support for operators aiming to increased energy efficiency and system performance (paying special attention to interpretability as an enabler of use).
11Real-Time Control of A2O Process in Wastewater Treatment Through Fast Deep Reinforcement Learning Based on Data-Driven Simulation ModelJournal articleHu et al. (2024) [34]Presents a DRL-based model-free real-time control strategy for a wastewater biological process, with a data-driven simulator Allows the settlement of a multi-objective problem (stability under regulatory compliance vs. operational efficiency) in controlled assessment settings.
12Machine learning applications for anomaly detection in Smart Water Metering Networks: A systematic reviewJournal review articleKanyama et al. (2024) [35]Surveys ML techniques for anomaly detection in SWM, structuring the evidence around families of algorithms and practical constraints (like data imbalance, noise and deployment feasibility). Confirms your Results statement that supervision/diagnosis is more developed than closed-loop control in application.
13Applications of Machine Learning in Drinking Water Quality Management within Water Distribution Systems: A Critical ReviewJournal article (review)Li et al. (2024) [36]Reviews the roles of ML in drinking water quality control, focusing on WDS environments Shows that the evolution of WDS monitoring is being transformed from expert knowledge-based to data-driven Identifies main application areas (quality prediction and event/anomaly detection) and ongoing deployment barriers (data quality, generalizability and operationalization).
14A Convolutional Graph Neural Network Model for Water Distribution Network Leakage Detection Based on Segment Feature Fusion StrategyJournal articleLi & Wu (2024) [37]Introduces a graph (deep) learning leakage detection mechanism based on network topology and the fusion of segment-level features; enhances event detectability in presence of network-structured dependencies (which relates to faster response for repair and smaller losses).
15A Novel Method for Anomaly Detection and Signal Calibration in Water Quality Monitoring of an Urban Water Supply SystemJournal articleLiu et al. (2024) [38]Recommends a data-driven solution to adaptive water quality monitoring toward more timely information verses time-based sampling, by viewing anomaly detection/signal processing as a route to earlier intervention and lower public health/service risks.
16Prediction-Based Real-Time Anomaly Detection for Water Quality Time SequencesConference/journal proceeding (ACM)Luo et al. (2024) [39]Presents a predictive, deep learning-based online anomalous time series detector for water quality data with a focus on early warning approach using flowing sensor readings and stresses the operational perspective as an implementation for chemotactic monitoring of pipelines.
17Pump Scheduling Optimization in Urban Water Supply Stations: A Physics-Informed Multiagent Deep Reinforcement Learning ApproachJournal articleMa et al. (2024) [40]Showcases an AI-based scheduling methodology that combines physics-informed deep learning and reinforcement leaning to minimize energy usage in pump scheduling—this is related to sustainability through the reduced energy waste and more flexible control under constraints.
18Optimization of effluent quality and energy consumption of aeration process in wastewater treatment plants using artificial intelligenceJournal articleMao et al. (2024) [41]Introduces an AI model to predict effluent quality and support aeration energy optimization, thereby tying analytics outcomes directly to compliance-facing performance and energy metrics—a critical sustainability linkage in WWTP operations.
19Future groundwater potential mapping using machine learning algorithms and climate change scenarios in BangladeshJournal articleSarkar et al. (2024) [42]Integrates ML-based hydrogeological risk from climate-change analysis, with multi-parameter geospatial predictors into groundwater potential zoning to inform long-term resource planning and risk-informed allocation decisions.
20Graph Neural Networks for Pressure Estimation in Water Distribution SystemsJournal articleTruong et al. (2024) [43]Shows GNN-based pressure estimation in district metering areas; offers the possibility to decrease reliance on dense instrumentation and facilitate more efficient monitoring/diagnostics—supporting leak analytics, operational control and digital twin-ready state estimation.
21Utilizing LSTM-GRU for IOT-Based Water Level Prediction Using Multi-Variable Rainfall Time Series DataJournal articleWidiasari et al. (2024) [44]Applies LSTM–GRU models to IoT-based river water level prediction; enables near real-time early warnings, operational preparedness and resilience.
22Digital Twins for Wastewater Treatment: A Technical ReviewJournal article (review)Wang (2024) [20]Aggregates digital twin (DT) theory and practice for both treatment plants and sewer networks for water-based systems with a focus on DT system components data–model coupling, telemetry integration), including challenges faced when scaling up from the World of in-silico.
23Short-Term Water Supply Forecasting for Water Treatment Plant Using Temporal Multi-Scale FeaturesJournal articleZhang et al. (2024) [45]Proposes a prediction method that is specially developed for the water treatment plant supply operation, considering both multi-scale temporal characteristics; facilitates production scheduling, mitigates under/over-treatment risk and potentially improves energy/chemical efficiency by better understanding short-term water demand.
24Forecasting urban water demand using different hybrid-based metaheuristic algorithms’ inspire for extracting artificial neural network hyperparametersJournal articleZubaidi et al. (2024) [46]Compares hybrid/metaheuristic generated models for urban demand forecasting; improved predictive capability; better capacity planning thus more reliable supply operation and hence indirectly econ-efficiency (energy/chemicals) via robust demand–supply fit.
25A Century of Data: Machine Learning Approaches to Drought Prediction and Trend Analysis in Arid RegionsJournal articleBouaziz et al. (2025) [47]Leverages super-long time series to assess the efficacy of ML techniques in drought prediction and trend analysis; underpins forward looking management and allocation decisions with impacts for ecological integrity and water security.
26Time Series Prediction of Water Quality Based on NGO-CNN-GRU Model—A Case Study of Xijiang River, ChinaJournal articleDing et al. (2025) [48]Proposes a CNN–GRU model for river water quality prediction to enhance early warning and management responses; provides operational support for sustainability through better compliance preparation, risk reduction and possible ecosystem protection results.
27Digital Twin Applications in the Water Sector: A ReviewJournal article (review)Ghorbani Bam et al. (2025) [11]Surveys DT applications in WDS and wastewater (dominant) vs. reclamation/desalination (thinner coverage); highlights the coupling telemetry + models + analytics still needs stronger validation to substantiate sustainability claims.
28Leak Management in Water Distribution Networks through Deep Reinforcement Learning: A ReviewJournal article (review)Javed et al. (2025) [49]Puts DRL for leakage management (policy learning for control/response) in perspective, emphasizes that many works are still simulation-based and safe transfer is challenging due to strong monitoring/constraints/validation requirements—optimization fundamentals needed for trustful NRW and energy savings.
29Artificial Intelligence in Chemical Dosing for Wastewater Purification and Treatment: Current Trends and Future PerspectivesJournal review articleJin et al. (2025) [50]Chemical dosing in wastewater treatment with AI based methods: a review on data-driven performance optimization, monitoring/control integration and conceptual barriers (standardization, interpretability, generalization).
30AI-Driven Anomaly Detection in Smart Water Metering Systems Using Ensemble LearningJournal articleKanyama et al. (2025) [51]Introduces an AI anomaly detection framework on smart water metering networks with ensemble learning + resampling to prevent class-imbalance problem; frames anomaly detection as a tool for water conservation and loss reduction.
31Reinforcement Learning Applications in Water Resources Management: A Systematic Literature ReviewJournal article (systematic review)Kåge et al. (2025) [12]Aggregates RL applications to allocation, control, and operations (where most agents are trained and evaluated on simulators) with transfer, safety and constraints being the primary missing pieces for deployment in real world infrastructure.
32Improving Irrigation Scheduling through Deep Learning-Based Reference Evapotranspiration EstimationJournal articleKhandappa & Haladappa (2025) [52]Utilizes deep learning to determine reference evapotranspiration (ET0) which will enhance irrigation scheduling decisions; encourages water-efficient agriculture.
33Towards the digitalization of water treatment facilities: A case study on machine learning-enabled digital twinsJournal article (case study)Ma et al. (2025) [53]A case study of data-driven digital twins enabled by ML for water treatment plants: enables discussions on system-level generalization and operation coupling.
34Pipe Burst Detection and Localization in Water Distribution Networks Using Faster Region-Based Convolutional Neural NetworkJournal articleMin et al. (2025) [54]Object-detection (Faster R-CNN) based detection/localization of bursts in case of partial sensor coverage; to help in immediate response, water-loss abatement and service resilience.
35On Smart Water System Developments: A Systematic ReviewJournal article (systematic review)Quintana et al. (2025) [55]Analyzes smart water advances and barriers to adoption; contributes framing for efficiency/resilience co-benefits.
36Real-Time Pump Scheduling in Water Distribution Networks Using Deep Reinforcement LearningJournal articlePei et al. (2025) [56]Utilizes deep RL (PPO) for real-time pump scheduling; discusses energy optimization opportunity and transfer/safety concerns.
37ML4FF: A Machine Learning Framework for Flash Flood ForecastingJournal articleSoares et al. (2025) [57]A machine learning framework for flash-flood prediction: From risk management to climate-resilience applications.
38Deep Reinforcement Learning for Optimized Reservoir Operation and Flood Risk MitigationJournal articleSseguya & Jun (2025) [58]DRL for reservoir operations with multi-objectives against each other; related to climate-driven extremes and robustness.
39Machine Learning for Smart Water Distribution Systems: Exploring Applications, Challenges and Future PerspectivesJournal article (review)Taloma et al. (2025) [10]It examines ML in both smart metering and WDS operations of references may result from limitations on operationalization and data.
40Knowledge embedding and interpretable machine learning optimize comprehensive benefits for water treatmentJournal article (open access)Wang et al. (2025) [20]Explainable ML for dosing control subject to operational constraints; presents quantified improvements (turbidity and dosing cost reduction), which directly serve sustainability accounting beyond accuracy.
41Failure Analysis and Machine Learning-Based Prediction in Urban Drinking Water SystemsJournal articleYılmaz et al. (2025) [59]ML for failure prediction to plan maintenance; contributes to the decreasing of bursts, improved reliability and decreased water losses.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Silva, J.A. AI Solutions for Improving Sustainability in Water Resource Management. Sustainability 2026, 18, 2154. https://doi.org/10.3390/su18042154

AMA Style

Silva JA. AI Solutions for Improving Sustainability in Water Resource Management. Sustainability. 2026; 18(4):2154. https://doi.org/10.3390/su18042154

Chicago/Turabian Style

Silva, Jorge Alejandro. 2026. "AI Solutions for Improving Sustainability in Water Resource Management" Sustainability 18, no. 4: 2154. https://doi.org/10.3390/su18042154

APA Style

Silva, J. A. (2026). AI Solutions for Improving Sustainability in Water Resource Management. Sustainability, 18(4), 2154. https://doi.org/10.3390/su18042154

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop