1. Introduction
Urban pluvial flooding has emerged as one of the most critical water-related hazards in rapidly urbanizing regions. Intensified short-duration rainfall under climate change, combined with the expansion of impervious surfaces and aging drainage infrastructure, has increased both the frequency and severity of urban flood events worldwide [
1,
2,
3,
4]. The 8–9 August 2022 storm over the central region of Korea, for example, produced daily rainfall of up to 381.5 mm and an hourly maximum of 141.5 mm in the Seoul metropolitan area, values that rank among the highest in the national record, and caused widespread damage to underground spaces and transport systems [
5]. Such events highlight the growing vulnerability of dense Asian megacities, in which storm sewer systems were often designed for historical conditions and are incapable of handling present-day extremes [
6].
In South Korea, more than 60% of urban inundation studies have relied on physically based models such as SWMM coupled with one- or two-dimensional overland flow models and are typically driven by design storms derived from intensity–duration–frequency curves and temporal distribution models (e.g., Huff) [
7]. These approaches have been instrumental for infrastructure planning and regulatory flood mapping, but they exhibit several limitations when applied to real-time risk management. First, dynamic one- and two-dimensional (1D–2D) models at resolutions compatible with dense urban fabrics require substantial effort for model construction and calibration, including detailed representation of storm sewer networks, building footprints, and land surface elevations. Second, high-resolution hydrodynamic simulations are computationally expensive; when run over large-scenario ensembles or long forecast horizons, they are often too slow to support time-critical decision-making during rapidly evolving events [
7,
8]. Finally, design-storm-based studies do not directly exploit the growing availability of real-time hydro-meteorological observations and forecasts that could underpin anticipatory flood warnings at street scale.
In parallel with advances in physics-based models, the hydrology and water resource community has rapidly adopted machine learning (ML) and deep learning (DL) approaches for forecasting hydro-meteorological variables. Recent reviews document the proliferation of recurrent neural networks (RNNs), long short-term memory (LSTM) networks, gated recurrent units (GRUs), convolutional neural networks (CNNs), and hybrid architectures in rainfall–runoff modeling, streamflow and water-level prediction, and other hydrological forecasting tasks [
9,
10].
Building on this foundation, several recent studies have proposed DL architectures specifically tailored for flood forecasting. Hybrid LSTM- or GRU-based models have been used to improve streamflow prediction under non-stationary climate and land-use conditions and to integrate multiple data sources such as radar rainfall, satellite precipitation, and numerical weather prediction outputs [
9,
10].
At the same time, CNN-based models have been explored as surrogates for hydrodynamic simulators, learning mappings from rainfall fields or hydrograph descriptors to spatially distributed inundation depths. For example, Frame et al. [
11] used ML models trained on outputs of the US National Water Model to rapidly map flood extent over continental scales, while Fraehr et al. [
8] conducted a systematic comparison of inundation surrogate models and demonstrated that appropriately designed DL surrogates can approximate high-resolution hydraulic models with orders-of-magnitude speed-ups. Dense U-Net-type architectures have also been employed to super-resolve coarse hydraulic outputs into detailed inundation maps suitable for urban applications [
12].
More recently, DL surrogates have been embedded within end-to-end flood prediction frameworks that emulate the full rainfall-to-inundation process. Farfan-Duran et al. [
13] proposed a DL-based surrogate that integrates net rainfall estimation via the SCS-CN method with a 2D hydrodynamic model (Iber-SWMM), enabling flexible analysis of antecedent moisture conditions. Other studies have explored CNN-based inundation models trained on long archives of LISFLOOD-FP simulations or regional hydraulic models, demonstrating that once trained, such surrogates can produce inundation depth fields within seconds per event [
14,
15,
16]. These developments suggest a pathway toward real-time or near-real-time urban flood nowcasting, especially when coupled with rapid forecasts of boundary conditions from DL-based hydrological models.
Despite this progress, important gaps remain, particularly for small, highly urbanized basins in East Asia. The review by Lee et al. [
7] showed that in South Korea, data-driven approaches to urban inundation modeling remain at an early stage relative to the extensive body of physics-based work. From an operational standpoint, urban catchments can respond on very short time scales, such that the timing and shape of rainfall within events can strongly affect peak flow and inundation dynamics [
17]. In South Korea, urban inundation studies have largely relied on physics-based modeling (with SWMM being dominant), while data-driven approaches are increasingly recognized but remain constrained by the scarcity of extreme-event data [
7,
18]. Recent studies further indicate that RNN-family models can exhibit peak lag and peak underestimation, and that these errors become more problematic as lead time increases, underscoring the need for careful multi-horizon evaluation for operational flood forecasting [
18].
Most Korean studies focus either on optimizing physically based models or on assessing the sensitivity of inundation to design-storm selection and rainfall temporal distribution models, rather than on building operational surrogates tightly coupled to real-time observations and forecasts [
19]. Moreover, few documented examples integrate DL-based stream-stage forecasting with DL-based inundation mapping within a unified framework evaluated using recent extreme events, such as the 2022 flood that affected the Anyang and Seoul metropolitan areas [
5].
This study addresses these gaps by developing and testing a two-stage AI framework for urban flood nowcasting in a densely developed reach of the Anyang and Hagui Streams in Bisan-dong, South Korea. In the first stage, a GRU-based recurrent neural network is trained on multi-year rainfall and stream-stage records from multiple gauges to produce short-lead (10–60 min) nowcasts of the water level at key locations. In the second stage, an ANN–CNN surrogate model is trained on an ensemble of 1D–2D XP-SWMM simulations forced by synthetic design storms and a range of downstream water-level conditions. The surrogate converts compact descriptors of rainfall and boundary conditions into high-resolution (256 × 256) maps of maximum inundation depth over the urban basin. Once trained, the two components can be combined such that GRU-predicted stages and scenario-based rainfall information drive the surrogate to generate near-instantaneous inundation maps.
The specific contributions of this work are threefold. First, we present one of the few documented applications of GRU-based water-level nowcasting for a small urban stream network in Korea, using operational gauge and meteorological data and rigorously evaluating performance across multiple lead times. Second, we develop an ANN–CNN inundation surrogate that emulates a detailed XP-SWMM dual-drainage model for the Bisan-dong study area and quantify its accuracy in reproducing both inundation area and grid-scale depth metrics. Third, we demonstrate how the coupled GRU–CNN system can support urban flood nowcasting by generating high-resolution inundation maps for the August 2022 flood and other scenarios, thereby linking point-scale forecasts with spatially explicit flood information. The framework is intended as a step toward operational AI-enabled urban flood early-warning systems that complement, rather than replace, established physically based modeling tools.
This study aims to establish and verify a two-stage AI framework for urban flood nowcasting that combines GRU-based short-term stream-stage prediction with an ANN–CNN inundation surrogate trained on XP-SWMM simulations. The key evaluation focuses on multi-horizon stage accuracy (10–60 min) at multiple gauges and on the surrogate’s ability to reproduce inundation extent and depth fields under diverse rainfall and downstream boundary conditions.
3. Results
3.1. Water-Level Forecasting
The GRU stream-stage forecasting model was evaluated at four gauging sites along the Anyang–Hagui stream network: Anil Bridge (ANL), Hoan Bridge (HOA), Daehan Bridge (DAH), and Indeogwon Bridge (IDW). For ANL, HOA, and DAH, the period 2011–2018 was used for training and 2019–2022 for independent validation, whereas for IDW—where observations began later—the training and validation periods were 2015–2020 and 2021–2022, respectively. Forecasts were generated at lead times of 10, 20, 30, and 60 min. Model performance was assessed using the mean absolute percentage error (MAPE), the coefficient of determination (
), and the Nash–Sutcliffe efficiency (NSE). The results are summarized in
Table 2,
Table 3,
Table 4 and
Table 5 and illustrated by the hydrographs in
Figure A1,
Figure A2,
Figure A3 and
Figure A4.
At Anil Bridge, the GRU model exhibited very high performance for lead times up to 30 min. During both the training and validation periods, R2 and NSE exceeded 0.96 for 10–30 min forecasts, and MAPE remained below 5% in both periods. For short lead times, the model reproduced the timing and magnitude of water-level fluctuations with high fidelity, including the rapid rise and recession during extreme events. At a 60 min lead time, however, R2 and NSE decreased to approximately 0.88–0.89, and MAPE increased to roughly 8–11%, indicating a clear degradation in predictive performance. The corresponding hydrographs show that 60 min forecasts tend to respond later than the observations and that peak stages are sometimes overestimated relative to the gauge records.
The results at Hoan Bridge were similar regarding the overall performance but exhibited several distinct features. For 10 and 20 min forecasts, R2 and NSE again exceeded 0.95 in both periods, while 30 min forecasts maintained values above 0.90. MAPE at HOA was generally low (approximately 1.4–3.9% for 10–30 min), although the increase from 10 to 20 min was more pronounced than at ANL, reflecting a greater sensitivity to lead time. At 60 min, R2 and NSE dropped below 0.80, and MAPE increased to around 7–8%. Unlike ANL, the HOA forecasts showed relatively little systematic overestimation; instead, the main deficiencies at the 60 min horizon were delayed response and underestimation of observed peak levels during intense events.
At Daehan Bridge, the GRU model also showed excellent performance at short lead times. During the training period, R2 and NSE exceeded 0.96 at 10–20 min and remained above 0.90 at 30 min; in the validation period, all three short-lead horizons (10, 20, and 30 min) achieved R2 and NSE values greater than 0.95. MAPE values mainly ranged between 1.6% and 6.2%, comparable with or slightly higher than those at ANL and HOA. As at the other gauges, 60 min forecasts displayed reduced performance, with R2 and NSE below 0.90. An interesting feature at DAH is that, during the training period, the 60 min MAPE (4.7%) was slightly lower than the 30 min MAPE (5.6%), even though the correlation-based metrics decreased with lead time. This pattern suggests that error magnitudes remained moderate while timing errors increased. Visual inspection of the hydrographs confirms that 10–30 min forecasts track the observed patterns very closely, whereas 60 min forecasts show delayed peaks and a tendency to underestimate maximum stages.
The results at Indeogwon Bridge differed from those at the other sites because of the relatively short training record. During 2015–2020, R2 and NSE remained below 0.80 for all lead times, and MAPE was relatively high (approximately 6–8%), indicating limited learning from the sparse data. In contrast, the validation period (2021–2022) showed markedly improved performance: R2 and NSE exceeded 0.94 at all lead times, and MAPE decreased to approximately 4–7%. The forecasts at IDW generally followed the observed stages well, with little systematic overestimation; instead, they tended to slightly underestimate water levels, particularly at peak times and at the 60 min horizon. This pattern suggests that once several significant events became available in the record, the GRU model effectively captured the local hydrological response despite the shorter observation history.
Taken together, the results from the four gauges demonstrate that the GRU architecture is well suited for short-lead (10–30 min) stream-stage nowcasting in this small, rapidly responding urban basin. Across all sites, R2 and NSE values above approximately 0.95 and MAPE values typically below 5% at these lead times indicate that the model reliably reproduces both baseflow stages and sharp rises associated with convective storms. Forecast performance decreases systematically at the 60 min lead time, reflecting the intrinsic difficulty of longer-horizon prediction in such flashy catchments. Nevertheless, even the 60 min forecasts exhibit moderate performance and provide useful qualitative guidance on forthcoming high-water conditions. Accordingly, subsequent sections emphasize the 10–30 min forecasts when coupling the GRU outputs with the inundation surrogate model for urban flood nowcasting.
3.2. Inundation Forecasting
The performance of the ANN–CNN surrogate model for urban inundation prediction was evaluated by comparing its outputs with the XP-SWMM 1D–2D simulations used as training labels. As described in
Section 2.3, a total of 864 scenarios were generated by combining design storms of different durations and depths with 16 downstream boundary water-level conditions at the storm sewer outfalls. For each scenario, the surrogate generated a 256 × 256 map of maximum inundation depth over the study area. Model performance was assessed using the mean absolute percentage error (MAPE) for total inundation area and grid-based water depth (
Table 6).
Overall, the ANN–CNN model reproduced the XP-SWMM inundation patterns with good agreement in flood extent and moderate errors in local water depth. The MAPE for inundation area was 8.89%, indicating that the surrogate captured the areal extent of flooding with less than 10% relative error on average across all scenarios. In contrast, the MAPE for grid-based depth was 19.49%, reflecting larger discrepancies at the cell scale.
Qualitative comparisons between XP-SWMM simulations (
Appendix A.2,
Figure A5,
Figure A6,
Figure A7,
Figure A8,
Figure A9 and
Figure A10) and AI-predicted inundation maps (
Appendix A.3,
Figure A11,
Figure A12,
Figure A13,
Figure A14,
Figure A15 and
Figure A16) further support these findings. For low and high downstream boundary water levels (
Table 1) and storm durations of 1–3 h with total rainfall ranging from 50 to 200 mm, the surrogate correctly identified the main flood-prone streets, intersections, and low-lying blocks. As rainfall depth and downstream water level increased, both models showed a consistent expansion of inundated areas toward upstream sections and adjacent areas. Differences were more evident in narrow flow paths and small depressions, where the ANN–CNN output tended to smooth sharp gradients and slightly underestimate peak water depths, particularly under the most intense storm and backwater conditions.
Despite these local discrepancies, the surrogate model strikes a useful balance between accuracy and computational efficiency. Once trained, the ANN–CNN model generates high-resolution inundation maps in near real time for new combinations of rainfall and boundary conditions, whereas running the full XP-SWMM model for the same scenario ensemble would be computationally expensive. For many practical applications—such as identifying flood hotspots, screening structural and non-structural measures, or supporting nowcasting-based risk communication—the achieved accuracy in inundation area and depth is sufficient to inform decision-making. However, for detailed hydraulic design at specific locations (e.g., sizing critical drainage structures), use of the complete hydrodynamic model remains advisable.
In summary, the results in
Table 6 and
Appendix A.2 and
Appendix A.3 demonstrate that the proposed ANN–CNN surrogate is capable of emulating a complex dual-drainage model over a wide range of rainfall and downstream boundary scenarios. When coupled with the GRU-based stream-stage forecasts, this capability forms the basis for rapid, scenario-based urban flood nowcasting in the Bisan-dong study area.
4. Discussion
The results of this study demonstrate that a GRU-based short-term stream-stage forecasting model and an ANN–CNN inundation surrogate trained on XP-SWMM simulations can achieve practical performance for urban flood prediction. For stream-stage forecasting, the GRU model achieved high coefficients of determination and Nash–Sutcliffe efficiencies, along with low error rates, at very short lead times of 10–30 min, suggesting that the GRU effectively learns nonlinear temporal dependencies in rainfall–stage time series. In contrast, performance degraded at 60 min lead times, and the predicted peak stages tended to be delayed or insufficiently reproduced. This behavior can be attributed to the combined effects of flashy urban catchment characteristics (e.g., a time of concentration of approximately one hour) and increasing uncertainty in future rainfall.
In highly urbanized basins such as the study site, runoff generation and stage rise occur rapidly during short-duration storms with abrupt changes in rainfall intensity, and water levels frequently surge and recede within tens of minutes. Accordingly, at longer lead times (e.g., 60 min), post-forecast rainfall variability and drainage–channel interactions exert a more decisive influence, and uncertainty in recurrent neural network-based models driven solely by rainfall and stage time-series increases. Peak values, in particular, occur infrequently and are therefore underrepresented in the training data; moreover, training procedures that minimize average errors often result in smoothed peaks and lagged peak timing. Differences in predictive performance among gauging stations may also reflect variations in observation quality (e.g., measurement noise) and drainage conditions, including backwater effects. These factors should be examined more rigorously in future work using longer observation records, improved data quality, and the inclusion of additional explanatory variables.
For inundation prediction, the surrogate model yielded a relatively low MAPE for inundation area (8.89%) but a higher MAPE for grid-based inundation depth (19.49%), reflecting the inherent difference in difficulty between predicting flood extent and predicting localized depth. Inundation area has a quasi-classification character when defined by a threshold (e.g., depth > 0) and is therefore more stable, whereas inundation depth is a continuous variable for which small spatial displacements or subtle boundary shifts can translate into substantial cell-wise relative errors. Errors can be amplified near flood boundaries, where depths are close to zero, and in locations with strong micro-topographic controls (e.g., roads, intersections, and local depressions), where steep depth gradients exist and where the surrogate may represent boundaries and peak depths more smoothly.
In addition, because the inundation surrogate uses scenario descriptors that summarize rainfall and downstream boundary conditions as inputs, it has limited ability to represent differences in runoff and inundation responses arising from finer-scale variations in rainfall temporal patterns or spatial distributions, even for identical total rainfall. Notably, the “ground truth” for the surrogate is not observed inundation but rather the XP-SWMM 1D–2D simulation outputs. Therefore, surrogate accuracy ultimately depends on the assumptions of the physics-based model, the quality of sewer-network and topographic datasets, parameter settings, and boundary-condition specifications. While the ability of the surrogate to closely emulate XP-SWMM results is a significant operational advantage, the possibility that uncertainties in the physics-based model may be inherited by the surrogate should also be acknowledged.
Despite these limitations, the proposed approach offers clear advantages aligned with practical needs in urban flood prediction. Detailed 1D–2D hydrodynamic models can reproduce scenario-specific inundation processes with high fidelity, but their computational cost often restricts real-time operation. In contrast, once trained, the ANN–CNN surrogate can generate high-resolution inundation depth maps very rapidly for the same input conditions, enabling timely situational awareness and decision support. The GRU-based stage nowcasting also showed strong performance at very short lead times, indicating its potential use for early warning and rapid response within the 10–30 min window (e.g., road closures, proactive drainage operation at vulnerable locations, and deployment of emergency resources). Nonetheless, the surrogate is not intended to fully replace the physics-based model; rather, its primary value lies in complementing hydrodynamic simulations by overcoming computational constraints and enabling rapid flood-map production during operations. For applications requiring high-precision hydraulic analysis, such as detailed infrastructure design or site-specific depth estimation, physics-based modeling remains essential, whereas the surrogate is better suited for rapid hotspot identification, scenario screening, and operational decision support.
Several limitations of this study warrant further investigation. First, the analysis is based on a single case study focused on the Bisan-dong confluence area; therefore, additional validation is required to assess generalization to other urban basins. Second, the stage forecasting model relies primarily on rainfall and stage time series and does not incorporate additional predictors that may improve longer-lead performance, such as radar rainfall fields, numerical weather prediction outputs, upstream discharge, or soil moisture states; this limitation may partly explain reduced peak fidelity at lead times of 60 min or longer. Third, because the inundation surrogate was trained on XP-SWMM simulations, direct validation against observation-based flood marks or inundation maps is limited, and uncertainties in XP-SWMM inputs (e.g., sewer-network data, terrain representation, and roughness parameters) can affect the results. Future work should therefore focus on improving longer-lead forecasting by integrating radar- and forecast-based rainfall and downstream-stage predictions, enhancing peak reproduction through loss-function design (e.g., peak-weighted or quantile-based losses) and event-based training strategies, improving depth accuracy using multi-scale architectures such as U-Net and boundary-aware loss formulations, and strengthening credibility through hybrid validation that combines simulation outputs with observation-based flood evidence (e.g., flood-mark surveys and inundation maps). With these advancements, the coupled framework of GRU-based stage nowcasting and ANN–CNN inundation surrogates proposed in this study is expected to contribute to real-time flood prediction in fast-responding urban basins and to support the enhancement of urban flood emergency-response systems.