1. Introduction
The World Economic Forum’s Global Risks Report has consistently ranked extreme weather events as one of the top global risks over the past two decades, with flooding representing one of the most hazardous climate events [
1]. Rainfall-induced floods account for nearly 40% of all natural disasters worldwide, which have affected more than 3 billion people between 1990 and 2022 [
2,
3]. The economic losses resulting from such floods exceed USD 1.3 trillion [
3,
4], and future climate projections indicate substantial increases in both the frequency and intensity of extreme precipitation events [
5]. As such, rainfall-driven floods should be one of the main priorities for climate adaptation, sustainable urban planning, and disaster risk reduction at national and international scales to reduce the global impact on humans, the environment, and the economy.
Arid and semi-arid regions experience rainfall-driven flooding in ways that differ fundamentally from humid climates [
6]. Flash floods in these regions tend to be of short duration, are highly localized, and are exceptionally intense due to sparse vegetation cover, low infiltration capacity, crusted soils, and poorly developed natural drainage systems [
7,
8]. Flash floods are very hazardous, as the time for warning, evacuation, and emergency response is minimal. Recently, arid countries have experienced rapidly expanding urbanization in locations near the wadi systems based on a false sense of security due to long periods without flooding, which have amplified vulnerability in those areas [
8,
9]. Moreover, a single storm can trigger destructive flash floods in areas with no historical flooding due to climate change [
3]. These hydrological characteristics make arid regions sensitive to extreme rainfall events projected under climate change [
4,
5].
The Middle East and North Africa have been experiencing increased flash flood-intensified events since the beginning of the 21st century [
10,
11,
12]. Jeddah in Saudi Arabia experienced flash floods in 2009 and 2011; these events resulted in hundreds of fatalities and more than USD 1 billion in total [
13,
14]. Cyclone Gonu (2007) and Cyclone Shaheen (2021) generated catastrophic flooding, damaging infrastructure, displacing communities, and disrupting essential services in Oman [
15,
16,
17]. Egypt has also faced repeated extreme events, specifically in the Ras Gharib area in the Red Sea governorates in 2016 and 2019, which have caused catastrophic damages to highways, oil facilities, and residential areas [
8,
18]. In Upper Egypt, Aswan experienced flash floods in 2021 that led to casualties, widespread property damage, and mass displacement from flooded burrows [
19]. Those recurring events in arid regions urge the need for rapid high-resolution flood mapping tools to aid early-warning systems and improve land-use planning.
Flood inundation mapping has traditionally relied on empirical and statistical methods, physics-based hydrodynamic models, and, recently, machine-learning approaches [
20,
21]. Empirical or statistical models (i.e., regression-type analysis, stage-damage curves) have been widely used due to their simplicity and low data requirements. These approaches can only be reliable when long-term hydrologic records exist. As such, in regions with limited or no flood history, empirical and statistical methods tend to struggle with extrapolation, making them less suitable for data-scarce arid regions [
21].
Hydrodynamic models, such as HEC-RAS, MIKE, LISFLOOD-FP, TUFLOW, and InfoWorks ICM, represent the physics-based approach for floodplain mapping. Such models simulate complex hydraulic behavior using 1D, 2D, or coupled 1D–2D formulations [
22,
23,
24,
25]. These models integrate the effect of inline infrastructure effects (i.e., bridges, culverts, storm-drain networks), variability in flow velocities, and spatiotemporal variations in water depth with high accuracy when properly calibrated. However, those models require extensive input data, such as high-resolution DEMs, channel geometry, roughness coefficient, hydrographs, and historical flood maps, for calibration and validation [
26]. In arid and semi-arid regions where gauged flow data and observed flood extents are largely absent, calibration becomes difficult [
27]. Additionally, hydrodynamic simulations can be computationally expensive, often requiring hours to days to simulate a single storm event, especially in wadis due to the huge extent of the watershed, which limits their utility for near-real-time flood forecasting or early-warning applications [
28,
29].
Machine-learning (ML) approaches such as Artificial Neural Networks, Convolutional Neural Networks, Random Forests, and hybrid deep-learning models have gained significant attention in recent years for flood susceptibility mapping and rapid hazard classification [
30,
31]. These models can capture nonlinear hydrologic behaviors and process large spatial datasets with minimal computational cost once trained. In data-rich environments, ML models have demonstrated strong predictive skills, outperforming traditional statistical methods in terms of accuracy and hydrodynamic models in terms of efficiency [
28,
29,
32]. However, their performance is heavily dependent on the quantity and quality of training datasets, which include rainfall–runoff records with the associated flood boundaries. This requirement limits their applicability in arid regions, where flood events are infrequent, observational data are sparse, and the use of transfer-learning between basins is highly uncertain due to the unique hydrologic characteristics of wadi catchments.
Hydrodynamic models remain widely used in arid and semi-arid regions, even with limited meteorological and hydrological data; thus, the results are not reliable [
33]. The produced inundation maps are highly uncertain and often rely on subjective parameter estimation due to lack of calibration. If enough data are available, hydrodynamic models are highly computationally demanding, which makes them inefficient for near real-time flood forecasting in flash flood-prone environments where response time is critical. GIS-based flood mapping techniques provide a promising and operationally feasible alternative [
27,
34]. These methods require minimal input data, execute within seconds to minutes, and produce first-order floodplain inundation estimates suitable for early-warning systems and preliminary risk assessments.
The Height Above Nearest Drainage (HAND) is one of the most widely adopted GIS tool models used for flood susceptibility analysis, national hazard mapping, and near-real-time inundation estimation [
35,
36,
37]. Although HAND is computationally efficient and requires minimal hydrologic input, its performance is strongly dependent on the accuracy of the underlying drainage network and DEM quality. HAND is very sensitive to drainage lines, DEM artifacts, and channel extraction, particularly in flat terrains, braided systems, or urbanized watersheds. GeoNet was developed as an advanced terrain-processing tool to enhance the HAND tool by introducing nonlinear DEM filtering, curvature metrics, and flow-accumulation thresholds [
38,
39]. GeoNet substantially improves drainage-network extraction, which in turn enhances the reliability of HAND flood mapping and related terrain-driven modeling frameworks. GeoFlood emerged as a next-generation GIS-based inundation approach that synthesizes GeoNet drainage networks with HAND elevation metrics to generate synthetic rating curves for each stream reach [
40,
41]. A key advantage of GeoFlood is its ability to leverage GPU acceleration, allowing for the rapid processing of high-resolution DEMs across large basins with speeds far exceeding traditional hydrodynamic models. GeoFlood converts estimated discharge into water-surface elevations and inundation extents using simplified conceptual hydraulics, achieving reported match rates of 60–90% with the US Federal Emergency Management Agency (FEMA) reference flood products [
37,
40]. This balance of accuracy and computational efficiency positions GeoFlood as a valuable tool for emergency-response operations, large-scale regional assessments, and rapid scenario analysis.
GeoFlood is inherently more suitable for data-scarce regions, as it relies primarily on terrain characteristics rather than continuous hydrological observations. However, GeoFlood performance has not been comprehensively evaluated under arid-region conditions, and several of its core assumptions related to channel slope estimation, drainage segmentation, and DEM-derived hydraulic behavior require further investigation for flat terrain areas. The GeoFlood tool sensitivity has to be checked for uncertainties that might arise in arid regions due to limited data availability.
Accordingly, the objective of this study is to advance the understanding of GeoFlood performance in arid, data-scarce environments through two detailed case studies that (1) evaluates the accuracy of GeoFlood-generated inundation maps by comparison with a benchmark HEC-RAS 2D hydrodynamic model with different topographies; (2) investigates the sensitivity of inundation outputs to Manning coefficient, different return periods, segmentation lengths, different wadi structures, and slope-calculation methods during stream–network extraction; and finally (3) proposes two enhancements to the GeoFlood tool by incorporating a penalized cost approach (via the CPOP algorithm) to automatically detect slope breakpoints and improve the accuracy and stability of slope calculation via an alternative Theil–Sen slope trend estimation. By improving the GeoFlood tool and quantifying the resulting changes in inundation-mapping accuracy, this study provides new insights into optimizing terrain-based flood mapping for arid-region flood hazards and improving the robustness and reliability of the rapid GIS-based inundation tool GeoFlood.
2. Methods
As illustrated in
Figure 1, the conventional approach for generating flood inundation maps is based on coupling hydrological and hydraulic modeling. The hydrological model begins with the acquisition of a digital elevation model (DEM), land use/land cover (LULC), and rainfall data. The DEM is preprocessed to delineate the watershed at the location of interest and to extract key characteristics such as catchment area, mainstream length, and average slope. LULC is then intersected with the watershed boundary to estimate the runoff coefficient. Rainfall data are analyzed using frequency analysis and rainfall distribution curves to represent different return periods. These watershed characteristics and rainfall data are subsequently processed within a hydrological model (HEC-HMS was used in this study due to its widespread use in the literature [
42,
43]) to generate discharge hydrographs corresponding to each return period.
The hydraulic model uses the generated flow hydrographs together with the DEM and LULC data for the river area to simulate flood depth, velocity, and the inundation extent. Different hydraulic models can be employed, where each model uses different numerical formulations. In this case study, HEC-RAS was used as the benchmark model, which mainly relies on solving the shallow water equations. LULC in the HEC-RAS model is translated into Manning coefficients. Model performance in HEC-RAS is strongly influenced by DEM quality, Manning coefficient estimation, and computational grid resolution and direction. HEC-RAS is widely reported in the literature to produce high-accuracy results; however, it typically requires substantial data preparation and computational resources.
Both the hydrological model (HEC-HMS) and the hydraulic model (HEC-RAS) require calibration and validation; however, in arid regions where there is a lack of data measurements, the models are built based on common engineering practice, where the model’s parameters are calculated based on the code of practice and standards, then sensitivity analysis is performed [
44,
45,
46].
In this study, the hydraulic modeling is replaced by the GIS-based tool GeoFlood to evaluate its reliability as a computationally efficient alternative for flood inundation mapping. The following subsection describes the core principles and workflow of the GeoFlood model, while the subsequent subsection presents the proposed enhancements aimed at improving GeoFlood performance and robustness.
2.1. GeoFlood Tool
The GeoFlood tool begins with data preprocessing, which involves several critical steps: the acquisition of a high-resolution Digital Elevation Model (DEM), identification of the stream network and catchment of interest, derivation of the rating curve based on the assigned Manning coefficient, and determination of the design “hydrologic” discharge value to ensure accurate simulation [
40]. The DEM provides essential information on terrain morphology, flow paths, and potential water accumulation areas [
35,
36]. The DEM is subsequently used for watershed delineation using the Watershed Modeling System (WMS) to define the stream network and catchment boundaries, ensuring consistency with real hydrological features. The Manning coefficient is assigned based on channel soil type/land cover to develop stage–discharge relationships. Finally, river discharge values are obtained from HEC-HMS simulations to produce the required corresponding inundation map. Following preprocessing, the GeoFlood tool is executed, encompassing four main processes: (1) DEM processing, (2) network extraction, (3) application of the HAND module, and (4) hydraulic network calculations, as shown in
Figure 2.
DEM processing consists of a series of sequential steps, which begin with the application of a nonlinear diffusion filter to reduce random noise [
40], and then Grass GIS is used to calculate slope, curvature, flow direction, and flow accumulation to define the path that surface water follows [
47,
48].
The main purpose of network extraction is to obtain a hydrologically consistent drainage network by applying a skeleton extraction to focus on the main streams, node reading for an accurate definition of start and end points to the main stream of interest, Negative Height Above Nearest Drainage (NHND) computation to identify areas located below the elevation of the nearest drainage channel and subsequently to calculate the difference between its actual elevation and the elevation of the nearest mainstream per cell, and, finally detailed hydraulic calculations such as slope, contributing area, and flow properties, which are extracted [
37,
40,
41].
The Height Above Nearest Drainage (HAND) module is a topographic index that quantifies the vertical distance between each DEM cell and its nearest drainage cell, providing a spatial measure of relative elevation with respect to the extracted stream network [
36,
37]. HAND module relies on TauDEM tools for DEM processing to fill small depressions, generate flow direction and flow accumulation maps, and define drainage pathways, which is different from the original DEM processing that happened in the previous step [
37,
43,
49]. HAND is the foundation tool for flood inundation mapping. The HAND module compares every cell’s HAND value to the simulated water surface elevation derived from rating curves and hydrologic model discharge outputs. Each grid cell is classified as inundated when its HAND value is less than or equal to the modeled water surface elevation (stage) corresponding to the simulated discharge; cells with HAND values exceeding the predicted stage are classified as non-inundated.
Hydraulic network calculations encompass several integrated steps to represent river hydraulics and flood dynamics, as shown in
Figure 2. The stream of interest is divided into smaller segments to enable spatial variability in flow representation. Afterwards, catchment delineation is performed for each segment to define the contributing drainage areas. Then, river attributes such as channel length, slope, width, depth, and cross-sectional shape are calculated for each segment, where each segment is linked with its catchment by network mapping to maintain the connectivity from upstream to downstream. Finally, a hydraulic property base table and a full table detailing hydraulic relationships across multiple stage levels, such as flow area, wetted perimeter, hydraulic radius, velocity, fixed Manning coefficient, and discharge, are generated for each segment. To simulate the flood inundation map, discharge from HEC-HMS is integrated with rating curves to compute the water depth and find the flood extent. More details about GeoFlood and its operation can be found in the archived paper by Kyanjo et al. [
50].
2.2. GeoFlood Slope Calculation Enhancement
GeoFlood slope estimation affects the tool performance as it controls hydraulic network calculation and river attribute estimation, which defines the critical understanding of flow dynamics and energy distribution along the channel. Slope calculation is heavily influenced by segmentation, where shorter segments yield high details but can capture local noise, while longer segments smooth out the variability, as established by Zheng et al. [
37,
40]. To study slope calculation sensitivity, equal segmentation was applied to the stream length, with segment lengths varying from 100 m up to the total length of the stream.
In the original GeoFlood approach, the slope of each segment was calculated using two fixed points along the segment, located at 10% and 85% of its total length. The slope was then computed as the difference in elevation between these two points divided by the cumulative distance along the segment, as shown in Equation (1) [
50]. However, if the resulting slope is either zero or negative, a small positive slope (0.00001) was assigned to maintain the flow direction. All segments that were assigned this hypothetical value were rechecked with the previous segments and merged iteratively until a positive slope was obtained. Alternatively, the geometric mean of previous positive slopes could be used. While this method ensures hydrological continuity and reduces the influence of DEM noise at segment ends, it often overlooks local variations within the segment, particularly in meandering or steep areas, potentially producing unrealistic slope representations. Furthermore, it is very prone to being trapped in local erroneous elevation values.
where
represents the elevation at location
i and
mli is the cumulative distance from the beginning of the segment to that point
Li.
In this study, alternative slope estimation methods were tested to enhance the original slope estimation and not to solely rely on just two points that could catch the DEM noise. Each segment is subdivided into small uniform intervals (10 m) so that the slope for each subsegment will be based on those subsegments points. Then, after computing the subsegment slopes, slope aggregation method is applied to estimate the segment slope. In this study, the average, the median, and the Theil–Sen trend method [
51,
52,
53] were investigated. Similar assumptions were made for subsegment slopes that were adverse or zero to be replaced by (0.00001). The aggregation method using both the average and median was applied once by taking all values and another time by only considering the positive slopes while excluding assumed slope values of 0.00001. The Theil–Sen method was employed by calculating the median slope between all possible pairs of points within a segment, as shown in Equation (2). After computing the slope for each segment, the slope had to be checked with upper and lower limits based on hydrological knowledge and expert judgment to ensure a realistic tool representation for the area of interest.
where
and
are elevations of points
i and
j, and
and
are their corresponding cumulative distances along the segment.
Another enhancement was introduced to the GeoFlood tool related to automatic reach segmentation based on the Continuous Piecewise Optimal Partitioning (CPOP) approach, which is a dynamic programming algorithm for detecting multiple slope breakpoints in a reach elevation profile, finding the “best” piecewise-linear fit by minimizing residual error while penalizing segmentation complexity using a penalty on changes in slope, proving a more reliable number of changepoint locations. CPOP is used in this research for both automatic segmentation and slope calculation. CPOP is different than the original GeoFlood approach as it does not rely on a prior decision on the number of segmentations or the length of segments, but the method itself identifies the optimal segmentation and the corresponding optimal slope without user/expert judgement. Nevertheless, the CPOP method divides the streams into unequal reach segments.
The CPOP approach works by extracting the longitudinal profile of the stream of interest from the DEM by sampling elevation values at regular intervals with a fixed spacing
Δd (10 m is used in this study similar to other sub segmentation). Then, the CPOP approach is based on a total cost function (presented in Equation (3)) that identifies breakpoints by separating segments with different longitudinal behaviors to better represent the longitudinal profile. Those segments are found by minimizing the total cost function, which is composed of a data-fitting term and a penalty term. The penalty term is essential, as without it, the tool will treat each
Δd as a new segment.
where
k is the number of longitudinal segments,
is the breakpoint location,
are the slope and intercept of the linear representation for each segment, and
is a penalty parameter that controls segmentation sensitivity and tool complexity. The parameter
can be assumed based on the literature, then adjusted based on the output. More details are available in the
Supplementary Material Appendix SA.
After identifying the breakpoints, each segment length (
is computed directly from its geometry, while the segment slope is calculated as the average rate of elevation change with respect to distance along the segment, according to Equation (4).
is the horizontal distance between consecutive points within the segment.
The CPOP approach is more representative of the river morphology than the fixed length approach because it reflects the natural downstream direction of the stream from higher to lower elevations and does not fix a priori the segmentation lengths and breakpoints locations. For more details about how the CPOP approach works, refer to Fernhead et al. [
54].
All the updated codes for GeoFlood in this section are available in the data section, and the detailed explanation of the steps is in the
Supplementary Material Appendix SA.
2.3. Performance Evaluation
The main focus of the inundation maps is to find the flood extent without a detailed representation of the flood depth. As such, the tool evaluation is conducted by comparing GeoFlood-generated inundation maps with either observed flood extents or benchmark hydraulic model outputs (which is what is adopted in this study) based on pixel-based comparison. This comparison relies on four quantities: True Positive (TP), False Positive (FP), False Negative (FN), and True Negative (TN), which generate a confusion matrix. Other composite metrics are derived based on these quantities. For instance, TP represents the number of pixels correctly predicted as flooded, FN corresponds to flooded pixels missed by GeoFlood, FP indicates pixels incorrectly classified as flooded, and TN describes pixels correctly identified as non-flooded. To be able to quantify the tool performance based on those four elements, composite performance metrics were used.
Most performance metrics used in the literature include precision, recall, F1-score, the Fowlkes–Mallows Index (FM), the Matthews Correlation Coefficient (MCC), accuracy, the Critical Success Index (CSI), and Error Bias (EB). Precision is a measurement of the proportion of correctly predicted flooded pixels among all predicted flooded pixels, while recall quantifies the proportion of observed flooded pixels that were correctly captured by the tool, as shown in Equations (5) and (6), respectively. F1-score represents a balance between precision and recall into a single metric, as presented in Equation (7). FM is another representation of precision and recall by evaluating the geometric mean, as shown in Equation (8). MCC is a measuring index that combines all four confusion matrix components into a balanced measure, as shown in Equation (9). Accuracy expresses the proportion of the correctly classified pixels, while CSI focuses on flooded areas by excluding TN, as shown in Equations (10) and (11), respectively. EB quantifies the tool tendency to either overestimate (EB < 1) or underestimate (EB > 1), where an EB that equals one presents no-bias perfect results, as shown in Equation (12).
It is important to note that the extent of the evaluation boundary has a strong influence on the number of TN pixels, which can significantly inflate accuracy and affect MCC values. As such, the boundary of interest was created by applying a buffer of 1 m to the combined flood extent maps generated for each corresponding discharge related to a certain flood return period. While this approach restricts evaluation to the hydraulically relevant zones, the buffer value might still introduce a certain bias in non-flood areas. To avoid biased results, all performance metrics that include the TN value were omitted from further analysis. Precision, recall, F1-score, FM, CSI, and EB were the main performance metrics adopted in this study, as their calculation relies mainly on TP, FP, and FN, which are not sensitive to boundary buffer selection as TN is. The choice of those parameters ensures the meaningful evaluation of GeoFlood sensitivity to key parameters and its ability to reproduce benchmark flood extents.
3. Study Area
To assess GeoFlood performance and robustness in arid regions and its sensitivity to key parameters (terrain, Manning coefficient, return period), two contrasting case studies were chosen in the Kingdom of Saudi Arabia (KSA), as illustrated in
Figure 3. Each case study represents distinct geomorphological and hydraulic conditions. The upstream catchments shown in
Figure 3 are used to calculate the flow hydrograph using HEC-HMS. Further details are available in the
Supplementary Material Appendix SB. Case Study (A) represents a section of Wadi Hanifa in Riyadh, which is represented by the red square
Figure 3a. The section modeled for Wadi Hanifi is characterized by a relatively steep slope of 0.4% and a well-defined dry stream (wadi) corridor, as shown in
Figure 4. Wadi Hanifa usually experiences moderate flooding, with a peak flow of 154.87 m
3/s, equivalent to the 100-year return period. Case Study (B) represents the most downstream reach of Wadi Bayer located in Al-Jawf on the boundary between KSA and Jordan, as shown in
Figure 3b. The reach section used in this study has slopes ranging from 0.01% to 0.3%, located within a steep mountainous region transitioning between two lowland areas, as shown in
Figure 5. This reach is more complex in hydraulic conditions, as it consists of variable slopes. Wadi Bayer also experiences a higher flood discharge magnitude of 880.24 m
3/s, equivalent to the 100-year storm event. GeoFlood main parameters are summarized in
Table 1. Case Study (B) provides a rigorous testbed example for GeoFlood to examine its performance better under both highly dynamic flow conditions and complex stream networks.
GeoFlood performance in this manuscript was compared against the flood inundation maps generated by the hydrodynamic benchmark model (HEC-RAS 2D) according to the boundary extent shown in
Figure 4b and
Figure 5b, developed by Hamouda et al. [
27], as there is no available flood extent or flood depth data available in this region. The HEC-RAS model is a widely accepted benchmark for flood modeling when no available measured flood data or extents are available for calibration. The HEC-RAS model had a mesh size of 3 × 3 m for Case Study (A) and 10 × 10 m for Case Study (B). The GeoFlood mesh was matched with those mesh sizes rather than the DEM resolution for comparison and consistency. More details on the HEC-HMS and HEC-RAS models, data requirements, and their sources are available in the
Supplementary Material Appendix SB [
55,
56,
57].
5. Discussion
This study systematically evaluated the performance of the GeoFlood tool for flood inundation mapping in arid-region environments, with a particular focus on how terrain characteristics, discharge magnitudes, Manning coefficients, stream segmentations, and slope calculation methods influence tool accuracy and stability.
Terrain characteristics were the most dominant control on GeoFlood tool performance. For Case Study (A), where the stream is well-defined wadis with clear longitudinal gradients, the tool consistently achieved excellent agreement with the benchmark HEC-RAS results, with an F1-score of more than 0.95. In contrast, flatter terrains with weak morphological confinement, as presented in Case Study (B), introduced substantial uncertainty, which led to overestimation and a substantial reduction in the F1-score to approximately 0.89. These results are in agreement with Zhen et al. (2018) [
40] study that showed that HAND-based tools are less reliable in flat terrains. It is worth mentioning that the DEM resolution will affect the terrain characteristics, and as such, it is beneficial to use a high-resolution DEM, such as the one used in this study, rather than relying on low-resolution ones.
Changing return periods only affected Case Study (B), where the terrain was not well defined and where the F1-score improved from approximately 0.87 at low return periods to over 0.90 at higher return periods. This improvement reflects the fact that larger floods produce broader and more continuous inundation extents, reducing sensitivity to local elevation errors and improving boundary detectability. This bias of F1-score and CSI towards better performance in higher discharge cases is in line with previous findings from Stephens et al. (2014) [
59] and Cohen et al. (2025) [
60]. However, the error bias also showed that at low flow, there is a tendency of underestimation, while at high flows, GeoFlood tends to overestimate. This highlights the importance of jointly interpreting accuracy and bias rather than relying solely on a single performance metric.
The Manning coefficient was shown to be a key parameter controlling GeoFlood behavior through its influence on rating curves and HAND thresholds. For both case studies, low values tend to make the GeoFlood tool underestimate the flood extent, whereas higher values make the tool tend to overestimate, where the flatter terrain regions are mostly affected. This behavior is consistent with the expected hydraulic logic that lower Manning values result in lower water stages and higher flow velocities, producing a smaller flood extent, and vice versa. Accordingly, the sensitivity analysis aligns with theoretical explanation as well as prior studies emphasizing the need for Manning calibration and even its calibration with respect to the flood return period, as presented in the studies by Zhen et al. (2018) [
37,
40]. Applying spatial variable Manning coefficients along the reach would be beneficial when the land cover changes significantly; also, incorporating transverse spatial variability would be beneficial for higher flow return periods, as roughness typically differs between wadi/river and overland flow.
Stream segmentation was identified as the most critical factor affecting GeoFlood tool stability. Short segmentation lengths representing less than 20% of the total stream length consistently produced unstable results. These instabilities were primarily attributed to DEM-induced noise affecting the segment slope calculation method. When segmentation lengths exceeded 20% of the total stream length, performance improved markedly and became more consistent, with F1-scores exceeding 0.96 and CSI values surpassing 0.93. These results suggest 500 m and 2000 m as the minimum segment lengths for Case Studies (A) and (B), respectively; meanwhile, the study by Zheng et al. (2018) [
40] recommended 1.5 km as the minimum segment length. Later, the study by D’Angelo et al. (2022) [
58] tested lengths of 2 km, 5 km, and 10 km, and found that 5 km of a total length of 168 km indicates the best balance between inundation extent and water level accuracy. With the improvement introduced in this study by using the CPOP algorithm, there is no need to test multiple segmentation lengths in search of the optimum segmentation, as the algorithm identifies the optimum breakpoints and changes in stream slope correctly.
The Theil–Sen trend estimator and CPOP both enhanced robustness. The Theil–Sen method enhanced the original slope method, particularly for shorter segments, while the CPOP method provided the most robust and reliable performance by automatically identifying optimal segmentation locations directly from the longitudinal elevation signal. CPOP effectively eliminated the noise associated with arbitrary or overly fine segmentation. Across both study areas, CPOP achieved balanced and stable results with EB values close to 1.00 for both case studies, unlike the original and Theil–Sen methods, which poorly behaved in Case Study (B) when the segmentation was not properly selected.
This study has several limitations, as the model relied on the HEC-RAS as a benchmark instead of actual measurements due to the absence of such data. The Manning coefficient was assumed and was not calibrated, when in reality, when flow depth measurements are present, Manning calibration would give more reliable results, as both HEC-RAS and GeoFlood are sensitive to the Manning coefficient.
Future work should explore hybrid approaches that combine CPOP-based segmentation with Theil–Sen slope estimation. Additionally, embedding spatially variable Manning coefficients to better distinguish between channelized flow and overbank inundation could enhance the tool’s predictability. Moreover, CPOP–GeoFlood needs to be evaluated on longer river systems and more complex systems with multiple reaches, tributaries, and branching wadis. Such proposed research will further clarify the scalability and transferability of the proposed enhancements.
6. Conclusions
This study shows that GeoFlood can provide reliable flood inundation maps in arid environments, particularly in well-defined wadis, and maintain acceptable performance across a wide range of flood magnitudes when key parameters such as Manning coefficient and stream segmentation are properly selected. Terrain characteristics remain the dominant factor controlling GeoFlood tool accuracy where flatter areas exhibit higher sensitivity and flood extent maps are uncertain and unreliable.
Embedding the CPOP approach significantly improves the GeoFlood framework by removing the need for arbitrary segmentation and reducing errors in slope estimation. By automatically identifying terrain-driven breakpoints and representative slopes, CPOP enhances GeoFlood stability and keeps bias close to one, making the results more reliable in data-scarce regions. This makes the CPOP–GeoFlood framework well-suited for early-warning applications where detailed hydraulic models are not available. However, GeoFlood currently cannot account for hydraulic structures (such as bridges, culverts, and dams) and uses simplified conceptual hydraulics. Thus, GeoFlood cannot replace detailed hydrodynamic models like HEC-RAS in complex urban and engineering environments.
Future research should focus on the further testing and validation of the enhanced CPOP–GeoFlood framework in larger and more complex wadi systems, including areas with mixed land cover, urban environments, and microtopographic variations, and assess the model sensitivity against DEM resolution and different return periods. Comparative studies against other modeling approaches, including machine learning techniques, would also be valuable when sufficient observational data are available. Additionally, exploring the spatial variation of the Manning coefficient and the hybrid segmentation methods in more diverse hydrological and geographic contexts would be necessary.