3.1. Analysis of Observations and Comparison with Model Precipitation Forecasts
An extreme heavy-rainfall event affected the Pearl River Delta (PRD) on 8 September 2022. Based on the hourly CMPAS-V2.1 precipitation product, the 24 h accumulated precipitation maxima were mainly distributed over central-southern Guangzhou, with the maximum accumulated-precipitation grid point marked by the five-pointed star in
Figure 2a. Accumulated precipitation at this location exceeded 100 mm, indicating the pronounced local character of this event. It should be noted that CMPAS-V2.1 is a gridded multi-source merged precipitation product, and its grid-cell values may differ from observations at an individual rain gauge. Thus,
Figure 2 is used to characterize the regional precipitation distribution and its hourly evolution, whereas station-based extremes represent rainfall intensity at a smaller spatial scale.
Figure 2c presents the hourly rainfall evolution at the maximum accumulated-precipitation grid point. Rainfall intensified rapidly around 10:00 BT and peaked during 11:00–12:00 BT, with hourly precipitation of approximately 56 and 62 mm/h, respectively, followed by rapid weakening after 13:00 BT. This evolution indicates that the event was not characterized by persistent rainfall throughout the day; instead, a short-lived and highly concentrated convective rainfall burst contributed most of the accumulated precipitation. As shown in
Figure 2b, the rainfall center was located in the relatively low-elevation area near the PRD river network and estuary, while higher terrain is found to the north and northwest. This topographic setting may modulate low-level moisture transport and local convergence.
The event was characterized by localized and short-duration heavy rainfall in the absence of an obvious frontal passage, tropical cyclone, strong low-level vortex, or other dominant large-scale lifting system over the Pearl River Delta. The ERA5 fields instead indicate a favorable convective environment with instability, moisture, multilayer flow, and localized ascent. Therefore, the term describes the relative absence of a clearly dominant synoptic-scale forcing mechanism in this case; it does not represent a universal operational classification criterion.
Figure 3 provides a combined view of the convective environment and circulation conditions over the study area and its surroundings at 08:00 BT on 8 September 2022. In
Figure 3a, relatively high CAPE values are distributed over the study area and adjacent regions, indicating that substantial convective instability had accumulated and that the thermodynamic environment was favorable for the development of deep convection.
Figure 3b further shows relatively high relative humidity at 500, 700, and 850 hPa. The middle- and lower-tropospheric moist layers are vertically connected near the study area, indicating sufficient moisture for the initiation and maintenance of deep convective clouds.
Figure 3d–f present the wind speed and wind direction at 500, 700, and 850 hPa, respectively. The differences among the wind fields at these levels indicate a vertically varying environmental circulation. The lower-level winds describe the background moisture-transport conditions, while the middle-level flow characterizes the environmental airflow in which the convective system developed. The spatial co-occurrence of enhanced CAPE, high relative humidity, and the multilayer wind fields therefore indicates a favorable environment for intense convective rainfall.
Figure 3c presents the pressure vertical velocity integrated from 1000 to 200 hPa, providing a tropospheric-column perspective of the vertical motion. Negative values indicate enhanced column-integrated ascent, whereas positive values indicate relatively stronger subsidence or weaker ascent. The spatial correspondence between the column-integrated ascent and the regions of high CAPE and relative humidity provides additional dynamical support for the development and maintenance of deep convection. However, CAPE and moisture primarily describe favorable thermodynamic preconditioning rather than a sufficient trigger for convection. For this case, the lower-tropospheric wind configuration, together with the spatial correspondence between moist air and column-integrated ascent, suggests that low-level moisture transport and localized convergence may have contributed to convective initiation and subsequent maintenance. Because the present analysis does not include a full moisture-budget or convergence diagnosis, this interpretation is limited to a physically plausible environmental association rather than a uniquely demonstrated triggering mechanism.
The 3-km ensemble members exhibit substantial differences in the spatial distribution, location of heavy-rainfall centers, and areal coverage of precipitation. Members with relatively higher ETS values for precipitation ≥25 mm/day are outlined in red for visual reference only (The highlighting criterion is not used in predictor identification, clustering, or final member selection). Members 3, 10, 17, and 26 exceed the unified ETS threshold of 0.10. Since ETS accounts for hits expected by random chance, these members show genuinely higher skill than the remaining members in predicting heavy rainfall (>25 mm/d). Nevertheless, their absolute ETS values remain modest, indicating considerable uncertainty in accurately forecasting the location of this weakly forced extreme-rainfall event. Although these members reproduce localized heavy-rainfall centers to varying degrees, the position, structure, and extent of the heavy-rainfall band differ markedly among members.
The FBI values further show that members with relatively high ETS do not necessarily have a realistic event frequency (
Figure 4). Member 3 has an FBI of 0.84, which is comparatively close to the unbiased value of 1.0, suggesting that its ETS is associated with a relatively reasonable heavy-rainfall coverage. In contrast, members 10, 17, and 26 have FBI values of only 0.45, 0.38, and 0.43, respectively, indicating pronounced underforecasting. Their limited skill mainly results from capturing parts of the heavy-rainfall cores rather than accurately reproducing the full heavy-rainfall area. Therefore, a member with a relatively high ETS but an FBI far below 1 still substantially underestimates the spatial coverage of heavy rainfall.
In the 9-km configuration, as shown in the figure, members with an ETS exceeding the unified threshold of 0.10 are highlighted with red boxes for consistent visual comparison. By applying the identical highlighting threshold (ETS > 0.10) across both resolutions, a consistent visual reference is maintained. These members retain more effective heavy-rainfall prediction skill than the other members after random hits are removed. Member 25 has the highest ETS (0.26), indicating a relatively good ability to capture the heavy-rainfall region. Compared with the 3-km members, the best 9-km members attain higher ETS values.
FBI provides further insight into the forecast bias associated with these higher ETS values(
Figure 5). The highest ETS of member 25 is accompanied by an FBI close to unity, indicating that its performance was not obtained primarily by enlarging the forecast rain area and increasing false alarms; it therefore provides more credible overall forecast skill. In contrast, members 28 and 30 have FBI values of 0.72 and 0.41, respectively, indicating underforecasting of the heavy-rainfall area. Although they capture parts of the rainfall core and achieve relatively high ETS, they do not adequately reproduce the full spatial coverage. The joint use of ETS and FBI thus prevents a one-sided assessment based only on TS or ETS.
Both model configurations show pronounced member-to-member differences in precipitation forecast performance. These differences are evident not only in ETS, which measures forecast skill after accounting for random hits, but also in FBI, which diagnoses systematic overforecasting or underforecasting. For this weakly forced extreme-rainfall event, some members can capture localized heavy-rainfall cores but may still fail to represent the overall coverage realistically; members with higher ETS and FBI close to 1 provide more robust operational guidance. However, because precipitation forecast accuracy differs substantially among ensemble members for such extreme events, real-time manual selection of members is difficult to implement in operational practice. We therefore argue that, if the key circulation features controlling such weakly forced events can be identified, ensemble members can be objectively classified and fitted according to these features. This approach can exclude members with poor precipitation forecast skill and thereby improve the overall forecast accuracy of the model ensemble. The following section further investigates this classification and optimization framework.
To reduce the double-penalty effect caused by small spatial displacement errors in strict gridpoint verification, a multi-scale Fraction Skill Score (FSS) analysis was conducted for heavy rainfall exceeding 25 mm/day (
Figure 6). The FSS values of both R3 and R9 increase consistently as the neighborhood scale expands from 5 to 125 km, indicating that both configurations contain some displacement and scale-mismatch errors in the predicted heavy-rainfall structures. Their spatial skill improves when increasing positional tolerance is allowed. The ensemble-mean FSS of R9 increases from approximately 0.38 at 5 km to 0.81 at 125 km, while its median increases from approximately 0.40 to 0.83. In comparison, the ensemble-mean and median FSS values of R3 increase from approximately 0.22 and 0.23 to approximately 0.51 and 0.53, respectively.
At all neighborhood scales, the ensemble-mean and median FSS values of R9 are higher than those of R3, suggesting that R9 retains more stable overall spatial skill for this weakly forced extreme-rainfall case even after spatial tolerance is introduced. Thus, the better performance of R9 in the strict gridpoint ETS evaluation cannot be attributed entirely to the double-penalty effect. Nevertheless, the member ranges of the two configurations overlap substantially at most scales, and some R3 members attain relatively high FSS values at larger neighborhood scales, demonstrating considerable member-dependent uncertainty. Therefore, the results indicate higher ensemble-level spatial consistency of R9 for this particular event, rather than supporting a general conclusion that the 9-km configuration is universally superior to the 3-km configuration.
The strict gridpoint and neighborhood verification results provide complementary information. The ETS and FBI values quantify categorical skill and frequency bias at the verification-grid scale, whereas FSS quantifies spatial agreement after progressively increasing positional tolerance. For the present case, R9 has higher ensemble-mean and median FSS values than R3 at all examined neighborhood scales. These quantitative differences indicate higher spatial consistency of the R9 ensemble for this event; however, the overlapping member ranges indicate substantial member-dependent uncertainty and do not support a general resolution-ranking conclusion.
The clustering and member-selection procedures were implemented separately but identically for R3 and R9. Therefore, differences between the two circulation-conditioned subset means may reflect not only resolution-dependent circulation representation but also differences in the size and composition of the selected subsets. The R3-R9 comparison is consequently interpreted as event-dependent and conditional on the respective ensemble structures.
3.2. Extraction of Key Circulation Factors Influencing Precipitation
Figure 7 presents the spatial distributions of the Spearman rank correlations between area-mean hourly precipitation and the environmental variables over the study region on 8 September 2022. Each gridpoint correlation was calculated from N = 24 paired hourly samples. Statistical significance was evaluated at the field level, rather than by interpreting pointwise confidence intervals independently: the 99th percentile of the maximum absolute correlations from 2000 precipitation-series permutations was used as the domain-wide threshold (α = 0.01). The correlation patterns exhibit pronounced spatial and vertical variations, indicating that the rainfall event was not controlled by a single thermodynamic or dynamical factor. Instead, the precipitation evolution was associated with the combined variations in convective instability, lower-tropospheric moisture availability, and horizontal circulation at different levels.
CAPE is positively correlated with precipitation over most of the northern and western portions of the domain, with relatively extensive areas passing the field-significance test. This indicates that increases in convective instability generally occurred simultaneously with enhanced precipitation. However, weak negative or nonsignificant correlations are present over parts of the southeastern domain, demonstrating that the relationship between CAPE and precipitation was spatially heterogeneous. For heavy rainfall over Guangdong, high CAPE is generally a favorable thermodynamic background condition, but it does not necessarily provide a unique indication of the rainfall location, intensity, or hourly evolution. The column-integrated vertical velocity exhibits generally weak correlations, with alternating positive and negative areas and only limited significant regions. This suggests that the vertically integrated vertical-motion index alone cannot fully characterize this short-duration and highly localized convective rainfall event.
Relative humidity is predominantly positively correlated with precipitation, particularly at 700 and 850 hPa, where the positive correlation areas are more coherent and include extensive significant regions. This result highlights the importance of middle- and lower-tropospheric moisture in regulating precipitation variability. The 500 hPa relative humidity pattern is comparatively weaker, suggesting that the rainfall event was more directly associated with moisture variations in the middle and lower troposphere.
The wind-related correlation patterns show a distinct vertical structure. At 500 and 700 hPa, both the U- and V-winds are predominantly negatively correlated with precipitation across broad portions of the domain. The negative correlation associated with the 700 hPa U-wind is particularly extensive and spatially coherent. In contrast, the 850 and 925 hPa V-winds exhibit more pronounced positive correlations over the southern part of the domain and its surrounding area, with extensive significant regions. This vertical contrast suggests that variations in the mid-level zonal and meridional flow, together with changes in the low-level southerly flow, may jointly regulate moisture transport, horizontal convergence, and the organization of the convective system.
Relative vorticity displays strong spatial heterogeneity and dipole-like positive and negative structures at several levels. Relatively extensive significant areas appear at 500, 700, and 925 hPa, indicating that local rotational flow and dynamical disturbances in the middle and lower troposphere were closely associated with precipitation evolution. However, the coexistence of positive and negative correlation regions implies that the role of relative vorticity was strongly dependent on location. Its contribution to the regional precipitation signal therefore cannot be inferred solely from an isolated local maximum.
Figure 8 summarizes the spatial correlation fields in
Figure 7 by calculating regional-mean correlation coefficients over 22–24° N and 112–115° E. The resulting bar chart demonstrates substantial differences among the environmental variables. The 700 hPa U-wind has the largest absolute correlation coefficient and therefore exhibits the strongest regional association with precipitation among all the examined variables. The negative sign indicates that, during the analyzed period, changes toward weaker or more negative 700 hPa U-wind anomalies were generally accompanied by enhanced regional precipitation. This result is consistent with
Figure 7, where the 700 hPa U-wind shows broad negative correlations and extensive significant areas. It identifies the 700 hPa zonal flow as the most prominent middle-tropospheric dynamical indicator of this weakly forced heavy-rainfall event.
The 850 and 925 hPa V-winds both have correlation coefficients of +0.58, representing the second-largest absolute correlations in
Figure 8. Their positive correlations indicate that enhanced low-level southerly flow was closely associated with increased precipitation. Together with the spatial patterns in
Figure 7, these results suggest that the two low-level V-wind variables may reflect the variability of warm-moist airflow from the south, as well as its contribution to moisture transport and low-level convergence. The 500 hPa V-wind has a relatively strong negative correlation, indicating that meridional circulation changes in the middle troposphere also contributed to the organization or maintenance of the rainfall system.
CAPE has a regional-mean correlation coefficient of +0.33, which is lower than those of the most strongly correlated circulation factors but still represents a moderate positive relationship. The correlation coefficients of relative humidity at 500 hPa, 700 hPa, and 850 hPa confirm the supportive role of moisture conditions, although their regional-mean correlations are weaker than those of the key wind factors. The correlation coefficients of vertical velocity are −0.33 at 500 hPa, +0.06 at 700 hPa, and +0.17 at 850 hPa, while the column-integrated vertical velocity has a value of −0.14. This inconsistency indicates that vertical-motion correlations depend strongly on pressure level and are not sufficiently stable when represented by a single regional-mean index. The regional-mean correlations of relative vorticity are +0.07, +0.16, +0.23, and +0.07 at 500, 700, 850, and 925 hPa, respectively. Although relative vorticity shows extensive significant regions in parts of
Figure 7, the cancellation between positive and negative correlation areas leads to relatively small regional-mean coefficients in
Figure 8.
Considering both the spatial significance patterns in
Figure 7 and the regional-mean correlation coefficients in
Figure 8, the 700-hPa U-wind, 500-hPa V-wind, 850-hPa V-wind, and 925-hPa V-wind were selected as the primary circulation predictors for the subsequent clustering analysis. The 700-hPa U-wind exhibited the largest absolute regional correlation coefficient, while the 850-hPa and 925-hPa V-winds highlighted the importance of low-level warm-moist airflow and moisture transport. The 500-hPa V-wind provided an additional indicator of middle-tropospheric meridional circulation.
CAPE was not included in the clustering analysis because the purpose of the clustering procedure was to identify circulation-related predictors that could distinguish circulation-dependent ensemble differences. Although CAPE is an important thermodynamic prerequisite for deep convection, its positive correlation does not necessarily distinguish the rainfall center or the precipitation differences among ensemble members. Accordingly, the clustering analysis focused on the four circulation predictors listed above.
3.3. Clustering Analysis of Key Circulation Factors in the Model
Figure 9a presents the bootstrap distributions of the ARI and target-cluster Jaccard similarity for the eight resolution–predictor combinations. Both metrics show considerable variation across resamples and combinations, indicating that the reproducibility of the complete cluster partition and the target-cluster membership is not uniform. Nevertheless, the distributions cover a broad range of similarity values rather than being concentrated exclusively near zero, suggesting that the clustering results retain a detectable degree of robustness under perturbation of the 30-member sample. The differences among combinations also indicate that clustering stability depends on both model resolution and circulation predictor.
Figure 9b shows that the mean silhouette coefficients are generally highest at, or close to, k = 2, while they tend to decrease or fluctuate at relatively low levels as the number of clusters increases to k = 12. This indicates that increasing the number of clusters does not produce a clearer overall separation of the circulation patterns and may instead lead to weaker or less distinct groupings. Accordingly, k = 2, marked in the figure, provides the most appropriate and relatively parsimonious clustering solution among the candidate values examined.
Agglomerative hierarchical clustering using Ward’s minimum-variance linkage and Euclidean distance was applied independently to the four key circulation predictors for both the R3 and R9 configurations. As shown in
Table 2, Silhouette analysis selected two clusters for all eight resolution–predictor combinations, although the degree of separation varied among predictors and resolutions, with mean silhouette coefficients ranging from 0.1868 to 0.4059. The CCC values ranged from 0.4987 to 0.6836. In particular, the CCC values for the 925-hPa V-wind were relatively low for both R3 and R9 (0.5053 and 0.4987, respectively), approaching or falling below the empirical reference threshold of 0.5. This indicates that the hierarchical clustering structure of the 925-hPa V-wind was relatively weak; therefore, this predictor was not used independently as a decisive criterion for member selection.
Instead, the target cluster was defined using a three-predictor consensus criterion. Members were selected only when they belonged to the target clusters identified from all three more reliable predictors: 850-hPa V-wind, 700-hPa U-wind, and 500-hPa V-wind. This procedure retained 13 members for the R3 configuration and 7 members for the R9 configuration. The selected subsets were therefore based on cross-predictor agreement rather than on any single clustering result, providing a more conservative and physically consistent basis for subsequent analysis.
Figure 10 compares the composite spatial anomaly fields of the target-cluster members and the remaining members for the 500-hPa V-wind, 700-hPa U-wind, and 850-hPa V-wind under the R3 and R9 configurations. The differences between the target-cluster members and the remaining members are mainly reflected in the spatial extent, gradient strength, and location of the anomaly centers rather than in a complete reversal of the circulation signs across the entire domain. The 500-hPa V-wind exhibits relatively coherent middle-tropospheric anomaly structures in both configurations. The 700-hPa U-wind shows a southwest-negative and northeast-positive pattern, whereas the 850-hPa V-wind shows relatively positive anomalies over the western and central parts of the domain and weaker or negative anomalies toward the east. These differences indicate that the spatial organization of the lower- and middle-tropospheric anomalies provides information for distinguishing different ensemble forecast scenarios.
The broad spatial configurations are similar between R3 and R9, but the anomaly centers and gradient strengths are not identical, indicating that model resolution affects the representation of local lower- and middle-tropospheric circulation among ensemble members. Because the target clusters were independently determined for the three selected predictors,
Figure 10 represents case-specific relative circulation characteristics rather than universally applicable weather regimes. Under the strict consensus criterion requiring a member to belong to the target cluster for all three predictors, 13 R3 members and 7 R9 members were retained. The precipitation composites in
Figure 11 were therefore calculated as equal-weight means of these two circulation-conditioned subsets.
Figure 11 presents the 24 h accumulated precipitation composites for the circulation-conditioned subsets selected using the three-predictor consensus criterion and for the remaining members. In the R3 configuration, the subset mean concentrates the heavier precipitation over the central and southern parts of the study domain and provides better spatial coverage of the observed heavy-rainfall region near Guangzhou-Foshan. In contrast, the remaining-member mean is more spatially diffuse and exhibits a less concentrated heavy-rainfall center. In the R9 configuration, the subset mean also shows a more organized heavy-rainfall band than the mean of the remaining members. Overall, the members that belonged to the target clusters for all three predictors produced a precipitation pattern with spatial organization closer to the observed rainfall distribution.
The circulation-conditioned subset mean is defined as the equal-weight arithmetic mean of the 24 h accumulated precipitation forecasts of the selected members. No distance-based weighting, additional spatial smoothing, or posterior fitting of the precipitation field was applied. The R3 and R9 subsets contain 13 and 7 members, respectively. It is worth reiterating that the subset mean presented in
Figure 11 is evaluated primarily as a visual and spatial diagnostic scenario. The intention is not to claim an optimized deterministic post-processing product that unconditionally maximizes quantitative scores, but rather to demonstrate how circulation-consistency selection extracts physically meaningful spatial indications from ensemble spread. Consequently, the comparison in
Figure 11 is therefore intended to assess differences in the spatial organization of the subset mean relative to the remaining-member mean, rather than to claim statistically significant forecast improvement from a single case. The subset means show a more concentrated heavy-rainfall pattern and a precipitation region closer to the observed Guangzhou-Foshan area. However, because only one event is available, these differences should be interpreted as case-specific conditional guidance. A robust quantitative assessment of improvement relative to the full ensemble mean, remaining-member mean, and alternative selection strategies requires additional independent events.
3.4. Discussion and Limitations
This study focuses on the localized extreme heavy-rainfall event that occurred on 8 September 2022, for which the circulation predictors were identified from the statistical associations between hourly precipitation and ERA5 environmental fields. These predictors were subsequently used to classify the 30 CMA-TRAMS ensemble members. Accordingly, the circulation predictors and target-member subsets identified in this study characterize the ensemble uncertainty associated with this particular event. They should not be regarded as universal predictors or circulation regimes applicable to all weakly forced heavy-rainfall events, nor as evidence that one model resolution is generally superior to another. Nevertheless, localized and abrupt heavy-rainfall events, although less frequent than ordinary precipitation systems, can pose substantial risks because of their rapid development, strong spatial intermittency, and potential for simultaneous misses by both subjective human forecasting and objective numerical-model prediction. The present case study therefore provides an initial assessment of the circulation associations of such an event and examines the feasibility of a circulation-conditioned ensemble-member selection approach.
The CCC diagnostics further indicate that the clustering structure associated with the 925-hPa V-wind was relatively weak. This predictor was therefore excluded from the final member-selection procedure. The final selection was based on the 700-hPa U-wind, 500-hPa V-wind, and 850-hPa V-wind, with a strict criterion requiring each retained member to belong to the target cluster for all three predictors. This criterion helps limit the influence of the less reliable predictor and ensures greater consistency among the selected members; however, it also reduces the size of the conditional subset. This reduction is particularly notable for R9, for which only seven members were retained. The mean of the R9 conditional subset should therefore be interpreted with appropriate caution, as it may be more sensitive to sampling variability.
The circulation-conditioned subset mean is calculated as the equal-weight arithmetic mean of the 24 h accumulated precipitation forecasts from the selected members. It represents a conditional deterministic scenario constrained by circulation consistency, rather than a replacement for the full 30-member ensemble. The full ensemble remains necessary for representing forecast spread and probabilistic uncertainty.
Figure 11 suggests that the circulation-conditioned subset provides a more organized spatial representation of heavy precipitation for this particular event. However, this result should be viewed as event-specific evidence and should not yet be interpreted as demonstrating stable correction skill or general superiority across different cases or model resolutions.
Because this study relies on predictors identified retrospectively using observed precipitation and reanalysis data from the same event, it serves primarily as a proof-of-concept diagnostic study. For potential operational application in future real-time forecasting, these key circulation predictors would need to be pre-established beforehand by constructing climatological predictor libraries or weather-pattern look-up tables from extensive historical case archives. Consequently, a larger and independent sample of events is needed to evaluate the transferability of the identified predictors and the stability of the clustering-based precipitation guidance. Future work will use the numerical-model archive to identify additional representative localized extreme-rainfall events and construct a broader event database. Multi-event ensemble verification and comparative analyses will then be conducted to examine forecast biases, circulation associations, and possible formation mechanisms more systematically.