Next Article in Journal
Conjunctive-Use Frameworks Driven by Surface Water Operations: Integrating Concentrated and Distributed Strategies for Groundwater Recharge and Extraction
Previous Article in Journal
Water Inrush in Roof Bed Separation Due to Extra-Thick Seam Mining and Its Control
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Proposed Application of a Tree-Based Model for a Priority Scenario Restoration Plan for a Water Distribution Network †

by
Samantha Louise N. Jarder
1,2,* and
Lessandro Estelito O. Garciano
2
1
Department of Innovation and Sustainability, School of Innovation and Sustainability, De La Salle University, Manila 1004, Philippines
2
Department of Civil Engineering, Gokongwei College of Engineering, De La Salle University, Manila 1004, Philippines
*
Author to whom correspondence should be addressed.
This article is the expanded version of Application of Regression Decision Trees for Scenario-Priority WDN Restoration Strategy, which was presented at the JSCE 4th AI/Data Science Symposium Information, Kanazawa University Kakuma Campus, Kanazawa, Japan, 16 November 2023.
Water 2026, 18(1), 131; https://doi.org/10.3390/w18010131
Submission received: 14 November 2025 / Revised: 19 December 2025 / Accepted: 30 December 2025 / Published: 5 January 2026
(This article belongs to the Topic Geospatial AI: Systems, Model, Methods, and Applications)

Abstract

Hazard impacts are increasing in complexity as the world population grows. No universal strategies are available to minimize or eliminate the impacts of all scenarios. In this paper, a priority scenario-based strategy methodology is proposed using a Decision Tree (DT) machine learning tool. This approach identifies the parameters and combinations that contribute to high impact and loss from a hazard event conditioned on a priority scenario. The method is applied to a local water distribution network under seismic hazards. The priority scenarios in this study are vulnerability (VPS), damage (DPS), and cost (CPS). Each priority scenario identifies different affected areas. Some areas were repeatedly affected in different priority scenarios, showing an overlap of effects and making them a high crucial priority. Based on the analysis, a priority-based map was generated, highlighting areas that should be given priority for restoration or protection. The DTs were compared with other ML tools and Tree-based models to ascertain the best tool that determines the affected parameters. Competition tests compared the results from the ML tools and showed acceptable predictions; however, the DT was demonstrated to be the most ideal tool for this proposed method, showing an r2 of 0.6745, 0.9259, and 0.7343 for VPS, DPS, and CPS, respectively.

1. Introduction

There have been countless destructive impacts of natural hazards on the built environment that have been recorded due to their profound impact. The parameters of the hazard vary from event to event, e.g., magnitude, location, and type. With the scale and diversity of the parameters and scenarios, it is challenging to identify a specific intervention or strategy that is universally applicable to each scenario [1,2]. Water distribution networks, or WDNs, are especially susceptible to earthquakes primarily because the pipes are spatially distributed, and pipe connectivity can be an issue after a large seismic event. The effect of a large-magnitude earthquake on a WDN can be functional and/or physical loss. These losses result in water disruptions to the service area, which can severely impact health, industries, businesses, agriculture, education, and other sectors [3,4,5]. With these disruptions, economic activities can stop, while firefighting and disaster response capabilities are also affected. Several factors influence the impact of hazards on a WDN, e.g., pipe material, pipe diameter, exposure, and liquefiable areas [6].
When damage and losses occur post-event, recovery is a necessary goal to return a community utility to the same operational level as before the disaster event. However, there is no universal procedure to determine the priority for restoration and recovery [7]. Poor decisions may affect the recovery of a community [8], thus leading to a delay in recovery and improper distribution of limited resources [9]. According to Mushtaha et al., decision-makers face many factors in prioritization, thus adding to the difficulty of recovery for a community [7].
With various data on parameters and circumstances, it is challenging to identify an optimal solution that is collectively applicable to each scenario [1,2]. Analyzing every factor and the associated impact, loss, and damage is tedious and complicated in a deterministic scenario. It is more difficult if uncertainties of the hazard are considered in a probabilistic sense (with multiple probabilities). With this, a methodology that compartmentalizes the parameters that result in high loss or damage is proposed. The goal of this paper is to identify the sets of parameters and combinations that contribute to high loss, damage, vulnerability, cost, etc., so that interventions can minimize impacts and a prepared strategy for recovery can be ready.
Another issue is that lifeline systems are complex and continually evolving to meet the community’s demands [5,10,11]. Thus, the risk of a specified location changes over time. The uncertainty of seismic impact increases the difficulty of identifying a universal restoration strategy. Numerous factors influence the effects of hazards on pipeline systems [6]. The purpose of this study is to determine the sequence of priorities and to identify areas prone to physical seismic impacts on the pipes.
As human societies continue to expand and develop, complexities related to hazards, disaster risk, and recovery become increasingly complex and complicated. However, technology has also significantly progressed and mainstreamed; thus, Artificial Intelligence and machine learning (ML) in hazard modeling and risk reduction should be explored and maximized [12,13].
In this paper, the authors propose utilizing the output from a Tree-based Model (TBM) to determine a recovery plan, prioritizing which parameter to improve and minimizing the loss incurred during a hazard event as well as identifying key areas to prioritize recovery. Tree-based Models include Decision Tree (DT), Random Forest (RF), and Gradient Boosting (GB). The results show that the DT is the most desirable tool to achieve this study’s objective. The TBMs were also compared to other ML tools to check if these can be applied. This idea was proposed to address numerous combinations of variables of the dataset and the distinctiveness of impacts for different areas and hazards. Recent studies have utilized DTs as a predictive tool to either categorize or determine an output based on a set of data [13,14]. However, in this study, a DT was used to identify parameters that could produce a specific output. Here, a model will be used to determine the factors that contribute most to the highest outcome, based on data obtained from a previous study.
When the DT identifies the variables and conditions that contribute to the peak vulnerability of pipes in the area, damage of pipes, or cost of damage for replacement, areas with these combined parameters can be visualized using ArcGIS Pro 3.0. Importance weights were also applied to check the effects of the different combinations of the priority scenarios.
The investigation solely focuses on the physical impacts, damages, and losses resulting from damaged WDN due to an earthquake event. Results show that various parameters influence the prioritization of vulnerability, damage, and cost, and can visually aid decision-makers in identifying crucial areas.

2. Theory

The methodology was derived from the general equation of Hazard Risk:
R i s k = H a z a r d   ×   V u l n e r a b i l i t y   ×   E x p o s u r e C a p a c i t y
The concept of this paper is to utilize the Risk Equation to aid in identifying crucial areas. The proposed method was designed to be used in any type of hazard, at any location, at any type of structure. Each event has unique characteristics. There are also different types of hazards, vulnerabilities, exposures, and capacity. The combinations of these parameters can be measured in numerous ways. For example, in the case of hazards, earthquakes can be measured in Peak Ground Acceleration (PGA), flooding in meters deep, and typhoons in kilometers/hour.
The equation shows the output is risk, when the input is hazard, vulnerability, exposure, and capacity. Since risk can be interpreted and estimated in numerous ways such as life, loss, or damages, higher risk has a higher value [15,16]. In this methodology, risks are ranked depending on their value; the higher the value, the higher the rank. The ranking system shows one level of difference between one area to another, to show that the highest value has the highest risk, thus it should be the most prioritized followed by the rank after. It also shows uniformity of units.
The proposed methodology was designed to be applied to different cities, hazards, and lifeline systems.

3. Case Example

In this paper, a WDN located south of the Philippines was used as a case study to apply the concepts stated in the previous sections. In this regard, the following data, obtained from a previous study [17], were used: WDN and city layout, peak ground acceleration (PGA), liquefaction hazard, soil conditions, pipe materials, pipe diameters, probability of failure, average damage rate, average damage spot, and cost of replacement. It is worthwhile noting that most of the data were obtained from local government agencies.
The city layout was divided into a 500 m by 500 m grid in the ArcGIS Pro 3.0 software, thereby obtaining the city′s boundary. The shapefile of the city was obtained from igismap.com. Each grid in the city layout has a unique set of data when analyzed, thus yielding distinct results.
The methodology for generating a DT for each priority scenario is illustrated in Figure 1. The input parameters were divided into testing and training data for each priority scenario. The latter were utilized to produce the DT model, and the former to validate it. This was conducted for each of the priority-based scenarios.

3.1. Probabilistic Seismic Hazards Analysis (PSHA)

PSHA was performed in a previous study [17] and data such as seismic sources, magnitude, depth, source-to-site distance, etc., were obtained from the Philippine Institute of Volcanology and Seismology. The important results of the analysis [18] are the peak ground acceleration (PGA) that was used as an input parameter to the DT, and the peak ground velocity (V) that was utilized to calculate the estimated damages. This covers the hazard aspect of the methodology.

3.2. Estimated Damages

Equations developed by Isoyama et al. (2000) [19] were used to estimate the seismic damage. Each mesh has data for the parameters below. Equation (1) is the standard rate of damage.
R(V) = 3.11 × 10−3 × (V − 15)1.3
nij = Pp × Pd × Pg × Pl × R(V)
y i j = n i j × L i j
The correction coefficients for types of materials of pipes, diameter size of pipes, classification of ground surface, and liquefaction hazard level are variables Pp, Pd, Pg, and Pl, respectively. Table 1 presents the average values for the correction coefficients applied to the Standard Damage Rate. These values will affect the vulnerability of the WDN to the applied PGV, and this will result in the modified damage rate (Equation (2)).
The modified damage rate per mesh is then multiplied to the number of pipes present in that specific mesh to estimate the Average Damage Spot (Equation (3)). Where y i j is the ADS per mesh i in category j , km; and L i j is the length of pipes per mesh i in category j , km.
The probability of failure (Pf) of the pipes per mesh area was then estimated using the ADS. Here, it was considered that at least one break is already considered failure in the mesh area. It was assumed that the distribution of damage behaves in a Poisson Distribution, thus Pf = 1 − Ppoisson. Equation (4) shows the general formula for the probability of failure of pipes at any value of x, where x is the number of breakages. To simplify the calculations, the Ppoisson at no chance of breakage was estimated. Pf was then obtained knowing the Ppoisson is where there is no breakage. Thus, the equation of the probability of failure of at least 1 break of pipe is presented in Equation (6).
P p o i s s o n = y i x e y i x !
P p o i s s o n ( x = 0 ) = e y i
P f = 1 e y i

3.3. Estimation of Damage Cost

The estimated cost of the impact was calculated. Since the ADS refers to the average damage that may occur in the seismic event, it was assumed that it is the total pipe required for replacement in the mesh. This would result in obtaining the average loss in the mesh area (Equation (7)). Where W ¯ i is the amount of loss of pipes in each mesh area per variable; in this paper, m, the variable, considered is the diameter size, and a is the unit loss for each variable.
W ¯ i = i = 1 m a · y i j
Here, loss refers to the cost of replacements estimated during a seismic event. The price of pipe replacement is shown in Table 2. The data provided used the local cost in 2018 of pipe retailed in the study area.

3.4. Priority Scenarios Strategy

In real-life cases, numerous factors can be used to prioritize restoration. This proposed methodology was designed to modify or add priority scenarios depending on the available data. The level of risk differs depending on the hazard, vulnerability, and exposure of the structure. Numerous models were developed to identify the risk of a structure in a given location to every hazard. Risk can be quantified in different ways, such as life, monetary units, damage, etc.; however, for the method to be applicable to other cases, risks were designed to be ranked from 1 being the lowest risk to j the highest risk depending on the number of meshes. This was performed for the uniformity of units. Theoretically, it can identify the combination of parameters that contribute to the level of risk, no matter the hazard, structure, and location.
In this paper, only the parameters of the physical loss of a WDN were considered due to limited and accessible data. Three scenario priorities are Vulnerability Priority Scenario (VPS), Damaged Priority Scenario (DPS), and Cost Priority Scenario (CPS), which are the three scenario-based priorities. Various combinations of the three priority scenarios were also considered, where the weight of importance varies for every priority scenario. This is the Combined Priority Scenario (Combined PS). VPS focuses on structures and areas that are very exposed and susceptible to loss or damage from the hazard. Areas with high probability of breakage or failure were highlighted and identified. In this paper, DPS refers to the projected number of damaged pipes in the area, despite the impacts of vulnerability. And for CPS, this priority scenario identifies parameters and regions that contribute to high cost. This priority strategy focuses on recovering or avoiding expensive damage. In this paper, replacement cost was used.

3.4.1. Vulnerability Priority Scenario

In this paper, VPS aims to detect areas with high susceptibility and exposure, regardless of the number of structures present in the area. The output variable utilized in the VPS DT is the modified damage rate, nij. Input parameters involve PGA, material correction coefficient, and diameter size.
This identifies the area that has the highest combination of high exposure to hazard and low resistance against the hazard. The results from this priority scenario can help locate which structures need priority for reinforcement to strengthen against hazards.

3.4.2. Damaged Priority Scenario

Here, DPS was evaluated as areas having a higher likelihood of failure. In this study, the input parameters under DPS include PGA, pipe diameters, pipe lengths, pipe materials, and liquefaction coefficient. Meanwhile, the ranked index of the probability of at least 1 break, P f , is the output parameter. This priority scenario locates areas with densely damaged pipes during an event.

3.4.3. Cost Priority Scenario

Here, CPS identifies areas with high cost when a hazard occurs. The general idea for this type of priority scenario model output is the cost, while the branches from the DT would identify the conditions and combinations for a higher cost. CPS input variables are identical to DPS, while the output variable is the ranked index of the average loss per mesh, W ¯ . Due to the availability of data, the replacement cost of pipes was used as the basis for CPS. Other types of cost can be used in future studies. Using this priority scenario, decision makers can identify and plan how to minimize costs incurred in a hazard event and identify which areas, and which materials were likely damaged during the event.

3.4.4. Combined Priority Scenario

The methodology above showed that each priority scenario was assumed independently and did not impact on one another. However, there might be instances where the decision-makers would want to incorporate the effects of each priority scenario to one another. Thus, the concept of combining the priority scenarios was applied in this paper. Depending on which priority scenario, i , and the output data parameter, P r j i , the meshes were arranged in an increasingly ordered manner. Each mesh, j, was given an arbitrary number based on its ascending order or ranking. Here, the arbitrary number 1 represents the least prioritized output parameter, and k represents the peak priority to give immediate action. The Combined Priority Index, I j T , for each mesh is calculated using Equation (8). The equation is derived from the concept of weighted average [20].
I j T = α 1 I j 1 + + α n I j n i = 1 n α n
i = 1 n α n = 1.0
I j i = r a n k ( P r j i ) k  
To establish a uniform unit, the ranking and ranking index were obtained as shown in Equation (10), where Iji represents the Scenario Priority Index per mesh. Combining independent priorities was considered. Here, the priorities were combined, varying in the assigned weight of importance, α n . Here, n is equivalent to 1, 2, and 3, representing VPS, DPS, and CPS, respectively. The summation of all weights of importance is assumed to be equal to 1. The decision-maker can assign a weight of importance to each priority scenario, provided that the sum of all weights equals 1.

3.5. Decision Tree

The goal of this study is to identify the parameters that would lead to a high priority; in this case, it would be preferable to use a Decision Tree. In this paper, Classification and Regression Tree (CART) analysis is employed to identify the parameters that contribute most to the highest cost, damage, or loss in the network. Using CART, it can identify the input variables that contribute most to the output [20,21,22,23,24]. This tool is preferred over other ML tools because of its simplicity, and it is visually easy for interpreters to comprehend since the objective of this methodology is for different types of decision-makers at different expertise to comprehend and implement the plans.
Python 3.12.2 was used to generate DTs. Most of the random data (80%) was used to produce the DTs, while the remaining 20% was used to validate the model.

Restoration Strategy

To test and analyze the models, the values of the coefficient of determination, R2, and root mean squared error, RMSE, are compared and analyzed. To avoid overfitting, the dataset is divided into training data and test data. Training data are used to create the model, while test data are used to validate and check the consistency of the model.
Using the same data split and parameters from the Tree-based Model, other machine learning tools were used to test whether these tools are applicable to the goals of this study. All three Tree-based models, Artificial Neural Network, Support Vector Machine, and Multiple Linear Regression, were compared to one another using the scatter test for equality and comparisons of residual errors. The hyperparameters for ANN, SVM, and MLR were set in the most basic setting to see if they can cluster the dataset in a discrete manner or show a combination of parameters of grouped data.
  • Decision Trees
Decision Trees (DTs) are supervised learning techniques, which have a non-parametric aspect [7,23,24]. The most common type is the Classification and Regression Tree (CART). A DT trains the data through partitioning. This process starts from the root node and splits the data for each node based on a given condition until the data can no longer be separated [25].
Here, Classification and Regression tree analysis (CART) is used to determine which factors contribute the most to cost, damage, or loss demand in the system. Using CART, it can show which factor contributes more to the output. Due to its simplicity and visually easy to understand for users and interpreters, this method is preferred over other ML tools. The goal of this is for different decision-makers to understand and apply the methodology.
In this study, the CART DT was used. Numerous trees were created, and the search for the highest R2 and the lowest RMSE for both the test and train datasets was conducted. The splitting for each node was performed through the concept of variance reduction.
To create the Decision Trees, Python was used. DTs were produced from 80% of the data; the rest was used to validate the model. Due to the endless possible ways to model a Decision Tree, the following conditions were set for iterations of the DT to determine the optimized hyperparameters for this study:
  • The random state was set to 0.
  • The minimum samples in the leaf of 10.
  • The maximum number of leaf nodes ranged from 2 to 20.
  • The maximum depth of None and ranged from 2 to 20.
Using the GridSearch, numerous combinations of hyperparameters were tested. The scoring includes the highest r2 and lowest RMSERMSE for both the test and train datasets. Only the minimum sample leaf, maximum number of leaf nodes, and maximum depth were iterated to avoid complicating the tree and avoiding a deterministic model. The minimum number of samples was set to avoid overfitting the tree and to identify a group of data having similar conditions that have high priority.
2.
Random Forest
A Random Forest is an ensemble technique of numerous trained Decision Trees. This technique randomly selects samples from the original data and uses new, randomly generated data in the model [23,25].
Based on the best hyperparameters of the Decision Tree, the RF model was generated with 1000 estimators.
3.
Gradient Boosting
Gradient Boosting is a tool with a technique that combines simple Decision Trees to obtain an effective prediction. Unlike DTs and RF, a loss function is used to test the accuracy of the GB model [25].
Like the RF method, the GB used the same hyperparameters that were iterated from GridSearch from DT. Here, 100 unique and shallow Decision Trees were generated to create the GB model.
4.
Artificial Neural Network
An Artificial Neural Network is an ML tool designed to mimic the human brain’s functionality. This tool identifies patterns and learns from previous iterations [26]. This technique can also be used for nonlinear and incomplete datasets [27].
There were two hidden layers in the ANN. The activation was set to a sigmoid function. The optimizer for the model was set to Adam.
5.
Support Vector Machine
The Support Vector Machine is another tool within the supervised learning category. The algorithm technique of this tool evolved from the vector theory. The approach is that data are treated as vectors when plotted [28].
The data for both input and output were transformed to scalar. The model was set to the most basic setting. The kernel used was linear.
6.
Multiple Linear Regression
Multiple Linear Regression is a technique that models the influence of the independent variables (input) on the dependent variable (output). This technique extends the classic linear regressions by using numerous independent variables [29].
The generation of the MLR model was the most direct, with no additional settings applied. The input and output data were fed to create the model.

4. Results and Discussion

4.1. Probabilistic Seismic Hazards Analysis

Vulnerability Priority Scenario

  • Tree-Based Model
The scatter plot for the best hyperparameter for DT is shown in Figure 2, where Figure 2a is for the train data, and Figure 2b is for the test data. Figure 3 and Figure 4 are the scatter plots for the RF and GB, respectively. The x-axis for these plots represents the actual data, while the y-axis shows the predicted value from the generated models. Table 3 presents the R2 and RMSE values for the Tree-based Models. Although both RF and GB perform better in predictions than DT, it can be observed in the DT scatter plot (Figure 2) that the predicted data is grouped, which is more suitable for the objective of this study. This is because both RF and GB produce numerous tree models in one run and do not display a single tree or a summary of one tree.
Comparing Figure 2, Figure 3 and Figure 4 shows that only Figure 2 groups the data samples in discrete rule-based clusters, where each cluster has a single predicted value, while Figure 3 and Figure 4 show each actual value has a unique predicted value. The RF and GB produced numerous trees and assembled them to create one tree. They do not produce a single optimized tree that shows the combination of parameters contributing to high priority.
  • Competition Test
The method was tested in other machine learning tools. Here, Artificial Neural Networks (ANNs), Support Vector Machine (SVM), and Multiple Linear Regression were the ML tests used and compared to the results from the Tree-based Models.
Different models were developed, and their predicted outputs were compared to the actual values. Figure 5 shows the competition analysis of each model vs. the actual values. The dashed line shown in the figure is the Equality Line, which represents a predicted value equal to the exact value. This shows that each model is produced, consisting of predicted values. This indicates that regardless of the tool used, the predicted values are relatively the same. However, the DT’s results show the relationship between the parameters and are grouped into different groups depending on conditions to reach a specific interval.
Similar to RF and GB, the models for SVM, ANN, and MLR show a deterministic result rather than a clustering of data contributing to high priority. These models show good predictability but could not identify the combination of parameters contributing to high priority. The clustering of output data for DT can be observed in Figure 5 and Figure 7, while the other models show a deterministic predicted value for each actual value.
Figure 5 and Figure 6 show the distribution of residual errors for all the machine learning tools used. It shows that each ML tool can be used for prediction, but reiterates the point that DT is the most suitable tool for this method. Figure 7 shows the residuals of all the ML tools.
2.
Final Results
  • Vulnerability Priority Scenario
The optimal DT model produced for VPS shows 13 splits as shown in Figure 8. Numerous DTs were created as the number of splits increased, starting from three. Although subsequently adding the number of splits for model-predicted results versus training data can still increase the R2 and lower the RMSE, the r2 and RMSE for model-predicted results versus training data will reverse after exceeding the optimum number of splits.
The optimal DT indicates which combination of parameters contributes to areas with the highest vulnerability to earthquakes. The split of the data depends on the condition as stated in the node. If the split dataset satisfies the condition, it will go to the “True” node as directed by the arrow; if not, it will direct itself to the “False” node. The coefficient of determination, r2, and RMSE are 0.675 and 0.1669, respectively. Results show that regions with pipes having diameters of 50 to 150 mm and exposed to 0.301 g show high vulnerability, thus they require immediate priority for intervention under the VPS.
Figure 9 shows the map produced for the VPS. It shows which regions have parameters that, when combined, result in the highest vulnerability rankings obtained from the DT.
  • Damaged Priority
Similar to the VPS, scatterplots and a map identifying the location of the critical regions were produced. However, the DPS results indicate that the pipe length present in the mesh is the crucial parameter contributing to the high probability of failure. DPS has an r2 and RMSE of 0.926 and 0.078, respectively. The produced DT for DPS is shown in Figure 10, while Figure 11 shows the critical areas for the damaged priority. The results show that no matter how high the resistance or strength of the pipes is, the number of pipes present in the area is more crucial.
  • Cost Priority
Here, the results for the Cost Priority are shown. Each mesh has its own unique cost depending on the damage incurred to the pipes.
For the Cost Priority scenario, a combination of longer pipes and larger diameters contributes to a costly replacement of pipes in this sample area and hazard. Pipes with larger diameter sizes tend to be more expensive. The optimal hyperparameters have an r2 and RMSE of 0.7343 and 0.1577, respectively. Figure 12 shows the DT for CPS. Highlighted regions in the map are clustered in different parts of the city compared to the Vulnerability and Damage Priorities. Figure 13 illustrates the key areas to prioritize for Cost Priority.
  • Combined Priority
Each location can have decision-makers focusing on or prioritizing different aspects during a WDN interruption. To address the varying importance set by decision-makers, different distributions of weights (19 possible combinations) were tested and simulated.
Table 4 presents the combinations of weights used to determine the importance of each priority scenario. These weight factors are applied to each ranking of each mesh. IDs 1, 7, and 13 are the Vulnerability, Damage, and Cost Scenario Priorities, respectively.
Each combination of weights to VPS, DPS, and CPS produced unique DTs and identified the crucial areas. It was observed that though the DTs and maps were unique from each other, some areas were repeatedly identified. Due to this, all the maps were added to distinguish which area is highly prioritized, as it is often highlighted.
Table 5 shows the r2, Pearson R, and RMSE for the training data and testing data for each optimal DT of each combination of priority scenarios. The r2 ranges from 0.6311 to 0.9259, and Pearson’s r ranges from 0.7944 to 0.9622. According to Pearson′s r interpretation [30], the values show a strong to very strong correlation, showing that the optimal DT models for each combination are acceptable.
Figure 14 shows the combined priority map. After all the combinations were tested and the locations were identified, each map for each combination was overlaid on top of one another.
Figure 15 shows the zoning map of the city. The combined priority map indicates that areas repeatedly highlighted are located in the residential zone. It can also be inferred that residential areas and the community in it are most likely the ones directly affected by the loss and damage of pipes; however, more study is needed to confirm this since the calculation assumed that the pipes in each mesh are independent of the other pipes in the other mesh.

5. Conclusions

This paper showed a proposed methodology using DTs to identify the combination of parameters that contribute to high vulnerability, probability of failure, and cost. Tree-based methods and other ML tools were compared and tested to verify which is best for the proposed methodology. Although all the models presented showed good predictive power, only the DT, RF, and GB can identify the conditions of the variables that affect the priority. Although DT has the lowest r2 and highest RMSE, it is the most ideal tool, as it uses only one tree, whereas RF and GB produce numerous trees and are ensemble models, making it difficult to achieve the set goals. DT also groups the samples into a clear, discrete boundary that aligns with a hierarchical split, making the sample groupings simple and presenting a visible clustering predicted value.
This paper showed a methodology to identify the combination of conditions that can lead to high vulnerability, damage, and cost in a seismic-induced WDN using DTs. Based on the DTs, the combination of parameters can be obtained and pinpointed in GIS software. Combining the scenario priorities was also considered to examine possible interactions of each priority scenario in this study. Since the maps with crucial areas were identified, they can be compared to a land use map to determine which type of zone and population is directly affected. The methodology can aid decision-makers in identifying which areas need both immediate and subsequent attention and preparation.
It is recommended to include the impact of the functionality of pipes for comparison. Since it was assumed that each mesh is independent, the dependency and correlation of pipes need to be examined since they are connected.

Author Contributions

Conceptualization, S.L.N.J. and L.E.O.G.; methodology, S.L.N.J.; software, S.L.N.J.; validation, S.L.N.J. and L.E.O.G.; formal analysis, S.L.N.J.; investigation, S.L.N.J.; resources, S.L.N.J.; data curation, S.L.N.J.; writing—original draft preparation, S.L.N.J.; writing—review and editing, L.E.O.G.; visualization, S.L.N.J.; supervision, L.E.O.G.; project administration, S.L.N.J.; funding acquisition, S.L.N.J. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

During the preparation of this manuscript/study, the author(s) used Python 3.12, ArcGIS Pro 3.0 for the purposes of coding. The authors have reviewed and edited the output and take full responsibility for the content of this publication. This article is the expanded version of Application of Regression Decision Trees for Scenario-Priority WDN Restoration Strategy, which was presented at the JSCE 4th AI/Data Science Symposium Information, Kanazawa University Kakuma Campus, Kanazawa, Japan [31].

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
DTDecision Trees
MLMachine Learning
RFRandom Forest
GBGradient Boosting
SVMSupport Vector Machine
ANNArtificial Neural Network
MLRMulti Linear Regression
PGAPeak Ground Velocity

References

  1. Mysiak, J. Integrating Disaster Risk Reduction and Climate Change Adaptation for Risk-informed and Climate-smart Development. In Issue-Based Coalition on Environment and Climate Change Task Team on Disaster Risk Reduction and Climate Change Adaptation; United Nations Economic Commission for Europe: Geneva, Switzerland, 2021. [Google Scholar]
  2. A Global Programme on Risk Assessment and Management for Adaptation to Climate Change (Loss and Damage): Climate Risk Management—Promising Pathways to Avert, Minimise, and Address Losses and Damages. 2021. Available online: https://www.giz.org/en/downloads/giz2021-en-promising-pathways-to-avert-minimise-and-address-losses-and-damages.pdf (accessed on 12 September 2023).
  3. Makhoul, N.; Navarro, C.; Lee, J.S.; Gueguen, P. A comparative study of buried pipeline fragilities using the seismic damage to the Byblos Wastewater Network. Int. J. Disaster Risk Reduct. 2020, 51, 101775. [Google Scholar] [CrossRef]
  4. Long, L.; Yang, H.; Zheng, S.; Cai, Y. Seismic Resilience Evaluation of urban multi-age water distribution systems considering soil corrosive environments. Sustainability 2024, 16, 5126. [Google Scholar] [CrossRef]
  5. Parvaze, S.; Kumar, R.; Khan, J.N.; Al-Ansari, N.; Parvaze, S.; Vishwakarma, D.K.; Elbeltagi, A.; Kuriqi, A. Optimization of water distribution systems using genetic algorithms: A Review. Arch. Comput. Methods Eng. 2023, 30, 4209–4244. [Google Scholar] [CrossRef]
  6. Akram, M.R.; Can ZÜLFIKAR, A. Identification of factors influencing sustainability of buried continuous pipelines. Sustainability 2020, 12, 960. [Google Scholar] [CrossRef]
  7. Mushtaha, A.W.; Alaloul, W.S.; Baarimah, A.O.; Musarat, M.A.; Alzubi, K.M.; Khan, A.M. A decision-making framework for prioritizing reconstruction projects in post-disaster recovery. Results Eng. 2025, 25, 103693. [Google Scholar] [CrossRef]
  8. Mohammadnazari, Z.; Mousapour Mamoudan, M.; Alipour-Vaezi, M.; Aghsami, A.; Jolai, F.; Yazdani, M. Prioritizing post-disaster reconstruction projects using an integrated multi-criteria decision-making approach: A case study. Buildings 2022, 12, 136. [Google Scholar] [CrossRef]
  9. Mendis, K.; Thayaparan, M.; Kaluarachchi, Y.; Pathirage, C. Challenges faced by marginalized communities in a post-disaster context: A systematic review of the literature. Sustainability 2023, 15, 10754. [Google Scholar] [CrossRef]
  10. Antonowicz, A.; Brodziak, R.; Bylka, J.; Mazurkiewicz, J.; Wojtecki, S.; Zakrzewski, P. Use of EPANET solver to manage water distribution in Smart City. E3S Web Conf. 2018, 30, 01016. [Google Scholar] [CrossRef]
  11. Zhao, D.; Chen, Q.; Zhao, X.; Tong, Y.; Chen, C.; Xia, S. A resilience evolution model of urban lifeline systems during Operation Based on Performance State Transitions. J. Saf. Sci. Resil. 2025, 6, 100231. [Google Scholar] [CrossRef]
  12. Linardos, V.; Drakaki, M.; Tzionas, P.; Karnavas, Y. Machine learning in disaster management: Recent developments in methods and applications. Mach. Learn. Knowl. Extr. 2022, 4, 446–473. [Google Scholar] [CrossRef]
  13. Kuglitsch, M.M.; Pelivan, I.; Ceola, S.; Menon, M.; Xoplaki, E. Facilitating adoption of AI in Natural Disaster Management Through Collaboration. Nat. Commun. 2022, 13, 1579. [Google Scholar] [CrossRef] [PubMed]
  14. Irimia-Dieguez, A.I.; Blanco-Oliver, A.; Vazquez-Cueto, M.J. A comparison of classification/regression trees and logistic regression in failure models. Procedia Econ. Financ. 2015, 23, 9–14. [Google Scholar] [CrossRef]
  15. Panwar, V.; Noy, I.; Wilkinson, E. Calculating loss and damage from extreme weather events in Small Island Developing States. Reg. Environ. Change 2025, 25, 73. [Google Scholar] [CrossRef]
  16. Song, Y.S.; Park, M.J. A study on estimation equation for damage and recovery costs considering human losses focused on natural disasters in the Republic of Korea. Sustainability 2018, 10, 3103. [Google Scholar] [CrossRef]
  17. Jarder, S.L.N.; Garciano, L.E.O.; Maruyama, O. Probable maximum loss of a pipe network due to earthquakes: A case study in Iloilo city, Philippines. Int. J. Disaster Resil. Built Environ. 2021, 12, 223–237. [Google Scholar]
  18. Kramer, S.L. Lateral Spreading. In Encyclopedia of Natural Hazards; Bobrowsky, P.T., Ed.; Encyclopedia of Earth Sciences Series; Springer: Berlin/Heidelberg, Germany, 2013. [Google Scholar]
  19. Isoyama, R.; Ishida, E.; Yune, K.; Shirozu, T. Seismic Damage Estimation Procedure for Water Supply Pipelines. In Proceedings of the 12th World Conference on Earthquake Engineering, Auckland, New Zeland, 30 January–4 February 2000. [Google Scholar]
  20. Huynh-Cam, T.T.; Chen, L.S.; Le, H. Using Decision Trees and Random Forest Algorithms to Predict and Determine Factors Contributing to First-Year University Students’ Learning Performance. Algorithms 2021, 14, 318. [Google Scholar] [CrossRef]
  21. Chen, M.Y.; Chang, R.C.; Chen, L.S.; Shen, E.L. The key successful factors of video and mobile game crowdfunding projects using a lexicon-based feature selection approach. J. Ambient. Intell. Humaniz. Comput. 2022, 13, 3083–3101. [Google Scholar] [CrossRef]
  22. Siegel, A.F.; Wagner, M.R. Landmark summaries. In Practical Business Statistics; Academic Press: Cambridge, MA, USA, 2022; pp. 75–104. [Google Scholar] [CrossRef]
  23. Choi, R.Y.; Coyner, A.S.; Kalpathy-Cramer, J.; Chiang, M.F.; Campbell, J.P. Introduction to Machine Learning, Neural Networks, and Deep Learning. Transl. Vis. Sci. Technol. 2020, 9, 14. [Google Scholar] [CrossRef]
  24. Lejano, B.A.; Elevado, K.J.T.; Jarder, S.L.N.; Arion, J.L.M.; Capistrano, N.S.S.; Ongpeng, Z.O.; Perez, A.D. Enhancing compressive strength in concrete with waste ceramic tiles: Effects of selected aggregate modification treatments, water-cement ratio and curing periods for decision tree regression analysis. J. Eng. Sci. Technol. 2024, 19, 744–761. [Google Scholar]
  25. Ghazwani, M.; Begum, M.Y. Computational intelligence modeling of hyoscine drug solubility and solvent density in supercritical processing: Gradient boosting, extra trees, and random forest models. Sci. Rep. 2023, 13, 10046. [Google Scholar] [CrossRef]
  26. Kariri, E.; Louati, H.; Louati, A.; Masmoudi, F. Exploring the advancements and future research directions of Artificial Neural Networks: A text mining approach. Appl. Sci. 2023, 13, 3186. [Google Scholar] [CrossRef]
  27. Elevado, K.J.; Galupino, J.G.; Gallardo, R.S. Artificial Neural Network (ANN) modelling of concrete mixed with waste ceramic tiles and Fly Ash. Int. J. GEOMATE 2018, 15, 154–159. [Google Scholar] [CrossRef]
  28. Ahsaan, S.U.; Kaur, H.; Mourya, A.K.; Naaz, S. A hybrid support vector machine algorithm for big data heterogeneity using machine learning. Symmetry 2022, 14, 2344. [Google Scholar] [CrossRef]
  29. Trunfio, T.A.; Scala, A.; Giglio, C.; Rossi, G.; Borrelli, A.; Romano, M.; Improta, G. Multiple regression model to analyze the total LOS for patients undergoing laparoscopic appendectomy. BMC Med. Inform. Decis. Mak. 2022, 22, 141. [Google Scholar] [CrossRef]
  30. Akoglu, H. User’s Guide to Correlation Coefficients. Turk. J. Emerg. Med. 2018, 18, 91–93. [Google Scholar] [CrossRef]
  31. Jarder, S.N.L.; Maruyama, O. Application of Regression Decision Trees for Scenario-Priority WDN Restoration Strategy. In Proceedings of the JSCE 4th AI/Data Science Symposium Information, Kanazawa University Kakuma Campus, Kanazawa, Japan, 16 November 2023. [Google Scholar]
Figure 1. Methodology of the decision tree and mapping.
Figure 1. Methodology of the decision tree and mapping.
Water 18 00131 g001
Figure 2. Scatter plot for predicted values for test data (a) and train data (b) for the vulnerability priority.
Figure 2. Scatter plot for predicted values for test data (a) and train data (b) for the vulnerability priority.
Water 18 00131 g002
Figure 3. Scatter plot for predicted values for test data (a) and train data (b) for the vulnerability priority random forest model.
Figure 3. Scatter plot for predicted values for test data (a) and train data (b) for the vulnerability priority random forest model.
Water 18 00131 g003
Figure 4. Scatter plot for predicted values for test data (a) and train data (b) for the vulnerability priority gradient boosting model.
Figure 4. Scatter plot for predicted values for test data (a) and train data (b) for the vulnerability priority gradient boosting model.
Water 18 00131 g004
Figure 5. Competition test for the ML tools.
Figure 5. Competition test for the ML tools.
Water 18 00131 g005
Figure 6. Histogram for the residual errors of the machine learning tools.
Figure 6. Histogram for the residual errors of the machine learning tools.
Water 18 00131 g006
Figure 7. Residual box plot for each machine learning tool.
Figure 7. Residual box plot for each machine learning tool.
Water 18 00131 g007
Figure 8. Vulnerability priority decision tree: (a) whole tree model; (b) close-up of the most crucial parameter.
Figure 8. Vulnerability priority decision tree: (a) whole tree model; (b) close-up of the most crucial parameter.
Water 18 00131 g008
Figure 9. Crucial areas for vulnerability priority.
Figure 9. Crucial areas for vulnerability priority.
Water 18 00131 g009
Figure 10. Decision tree for damaged priority: (a) whole tree model; (b) close-up of the most crucial parameter.
Figure 10. Decision tree for damaged priority: (a) whole tree model; (b) close-up of the most crucial parameter.
Water 18 00131 g010
Figure 11. Crucial areas for damage priority.
Figure 11. Crucial areas for damage priority.
Water 18 00131 g011
Figure 12. Decision tree for cost priority: (a) whole tree model; (b) close-up of most crucial parameter.
Figure 12. Decision tree for cost priority: (a) whole tree model; (b) close-up of most crucial parameter.
Water 18 00131 g012aWater 18 00131 g012b
Figure 13. Crucial areas for cost priority.
Figure 13. Crucial areas for cost priority.
Water 18 00131 g013
Figure 14. The combined priority map.
Figure 14. The combined priority map.
Water 18 00131 g014
Figure 15. Iloilo city zoning map.
Figure 15. Iloilo city zoning map.
Water 18 00131 g015
Table 1. Correction coefficients.
Table 1. Correction coefficients.
Materials of PipesPpDiameter Size of Pipes, mmPd
ACP (Asbestos Cement)3.00 Φ 75 1.60
IC (Cast Iron)1.10 Φ 100 Φ 150 1.00
DCI (Ductile Cast Iron)0.30 Φ 200 Φ 450 0.80
S (Steel)0.5 Φ 600 0.5
Classification of Ground SurfacePgHazard LevelPl
Mountainous region (modified)1.10No or Low Level of Severity of liquefaction1.00
Hilly Areas (modified)1.50Moderate Level of Severity of Liquefaction2.00
Valleys with old water routes3.20High Level Severity of Liquefaction2.40
Alluvial plain1.0
Table 2. Replacement cost for pipes for every 6 m in Philippine pesos.
Table 2. Replacement cost for pipes for every 6 m in Philippine pesos.
Diameter Size, mReplacement Cost of Pipes per 6 m, PHP
0.050934.00
0.0751326.00
0.1002854.00
0.1506060.00
0.20017,001.27
0.25026,600.20
Table 3. Scatter plot summary.
Table 3. Scatter plot summary.
Decision TreeRandom ForestGradient Boosting
Train Datar20.719680.703500.95313
RMSE0.152240.156570.06225
Test Datar20.675450.704980.80878
RMSE0.166900.159120.12811
Table 4. Combination of weights of factors.
Table 4. Combination of weights of factors.
IDVulnerabilityDamageCostIDVulnerabilityDamageCost
1100110.40.20.4
20.80.10.1120.500.5
30.60.20.213001
40.40.30.3140.10.10.8
50.20.40.4150.20.20.6
600.50.5160.30.30.4
7010170.40.40.2
80.10.80.1180.50.50
90.20.60.2191/31/31/3
100.30.40.3
Table 5. Coefficient of determination, r2, and RMSE of each combination of priority scenarios.
Table 5. Coefficient of determination, r2, and RMSE of each combination of priority scenarios.
IDCriterionTraining DataTest DataIDCriterionTraining DataTest Data
1r20.71970.674511r20.8220.7824
R0.84840.8213R0.90660.8845
RMSE0.15220.1669RMSE0.12020.1402
2r20.64740.631112r20.77140.6838
R0.80460.7944R0.87830.8269
RMSE0.17060.1784RMSE0.13640.1686
3r20.73380.703513r20.85030.7343
R0.85660.8387R0.92210.8569
RMSE0.14790.1606RMSE0.10950.1577
4r20.77750.771114r20.87970.8362
R0.88180.8781R0.93790.9144
RMSE0.13490.1418RMSE0.09810.1242
5r20.88930.797915r20.85030.7531
R0.94300.8933R0.92210.8678
RMSE0.09470.1354RMSE0.10960.1517
6r20.9170.886616r20.86060.8016
R0.95760.9416R0.92770.8953
RMSERMSE0.08170.1027RMSE0.10660.1327
7r20.94830.925917r20.76140.7128
R0.97380.9622R0.87260.8443
RMSE0.06570.0778RMSE0.140.1577
8r20.93060.889018r20.72310.6538
R0.96470.9429R0.85040.8086
RMSE0.07580.0964RMSE0.15190.1691
9r20.90860.844219r20.85290.7944
R0.95320.9188R0.92350.8913
RMSE0.08660.1159RMSE0.10960.1348
10r20.86680.7945
R0.93100.8913
RMSE0.10440.1341
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Jarder, S.L.N.; Garciano, L.E.O. Proposed Application of a Tree-Based Model for a Priority Scenario Restoration Plan for a Water Distribution Network. Water 2026, 18, 131. https://doi.org/10.3390/w18010131

AMA Style

Jarder SLN, Garciano LEO. Proposed Application of a Tree-Based Model for a Priority Scenario Restoration Plan for a Water Distribution Network. Water. 2026; 18(1):131. https://doi.org/10.3390/w18010131

Chicago/Turabian Style

Jarder, Samantha Louise N., and Lessandro Estelito O. Garciano. 2026. "Proposed Application of a Tree-Based Model for a Priority Scenario Restoration Plan for a Water Distribution Network" Water 18, no. 1: 131. https://doi.org/10.3390/w18010131

APA Style

Jarder, S. L. N., & Garciano, L. E. O. (2026). Proposed Application of a Tree-Based Model for a Priority Scenario Restoration Plan for a Water Distribution Network. Water, 18(1), 131. https://doi.org/10.3390/w18010131

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop