Review Reports
- Violet Ishak *,
- Danieli Mara Ferreira and
- José Eduardo Gonçalves
- et al.
Reviewer 1: Anonymous Reviewer 2: Anonymous Reviewer 3: Anonymous
Round 1
Reviewer 1 Report (Previous Reviewer 2)
Comments and Suggestions for AuthorsThe authors have substantially revised the manuscript and adequately addressed the comments raised in my previous review. The revised version more clearly defines the research objectives and principal contributions, improves the description of the datasets, common evaluation period, and chronological training–testing split, and provides additional implementation details for QDM and the dry/wet occurrence-correction models. The statistical evaluation has also been strengthened through the exact McNemar test, paired permutation tests, bootstrap confidence intervals, and Holm adjustment. Moreover, the results are now interpreted more cautiously and appropriately with respect to the evaluation metric, basin, forecast lead time, and forecast–observation pairing. The methods, results, and conclusions are now sufficiently consistent, and I have no further substantive concerns. I therefore recommend acceptance of the manuscript in its present form.
Author Response
We sincerely thank the reviewer for their time, their highly constructive feedback throughout the review process, and for recognizing the substantive improvements made to the manuscript. We are thrilled that the revisions have fully addressed your previous concerns and that the manuscript is now recommended for acceptance. Your insights have significantly strengthened the rigor and clarity of this research.
Reviewer 2 Report (Previous Reviewer 1)
Comments and Suggestions for AuthorsThe authors have satisfactorily addressed the major concerns raised in the previous review. The revised manuscript has been substantially improved, particularly through the extension of the occurrence-correction analysis across all three basins, clarification of the training/testing and dry/wet verification procedures, improved statistical evaluation, more cautious interpretation of the QDM results, and strengthened discussion and literature positioning. The manuscript is now scientifically sound, clearly presented, and suitable for publication in its present form. I therefore recommend acceptance.
Author Response
We sincerely thank the reviewer for their constructive and detailed feedback during the review process, which was instrumental in improving the quality of this manuscript. We are very pleased that the extension of the occurrence-correction analysis, the refined statistical evaluations, and the expanded discussion have satisfactorily addressed your concerns. We deeply appreciate your time and your recommendation to accept the manuscript for publication.
Reviewer 3 Report (Previous Reviewer 3)
Comments and Suggestions for AuthorsThe manuscript has been substantially revised. I appreciate the authors' thorough effort in addressing my previous comments, particularly regarding the evaluation of the ensemble predictors and extending the analysis to three catchments. These changes have significantly improved the quality of the paper.
However, there are still some critical methodological uncertainties that need to be addressed. The main comments are provided below:
- However, a significant issue related to the use of future information still remains regarding the QDM method. In Equation (1), the empirical cumulative distribution function (CDF) covering the entire test period is used to determine the quantile position for a given forecasted value (lines 308–310). This means that correcting a given forecast utilizes information that would not yet be available at the time the forecast is issued. This directly contradicts the claimed "strictly out-of-sample" approach (lines 328–329) as well as the stated capability of utilizing the method operationally in real-time conditions (lines 213–214). Therefore, the results should be recalculated while preserving the full chronology of data availability, ensuring that the correction of each forecast does not utilize information from the future. If this approach is not feasible, it must be clearly stated in the manuscript that the applied procedure utilizes information from the entire test period and, consequently, does not reflect real-time operational conditions.
- Lines 400–405: in the case of the AR(1) predictor, it is not clearly specified what the lagged variable represents. It should be clarified whether, for the individual lead times (L0–L7), the predictor uses the actual precipitation observation from the previous day or a forecast value. If the observation is used, there is an additional issue concerning the availability of such observations for longer lead times, e.g. L3 or L5. This requires clarification in the context of a real-time operational system, since it is necessary to exclude the possibility of using information that would not yet be available at the time the forecast is issued.
- Lines 325–326: the multiplicative QDM correction factor was limited to a value of 2.0 (ratio.max = 2). Such a limitation may artificially constrain the ability to correct the mean and variability of the adjusted precipitation. It is therefore worth considering whether the deterioration of the individual KGE components, described in lines 691–698, actually results from a weakness of the QDM method. It cannot be ruled out that the technical limitation imposed here plays an important role. A sensitivity analysis using higher values of ratio.max (e.g. 5 or 10) would be justified. If such a test was not performed, this issue should at least be addressed in the Discussion as a possible technical explanation for the lack of improvement in KGE.
- Lines 201–203: calculating the spatial mean precipitation using a simple arithmetic mean based on only two measurement points for a catchment is questionable. Given the convective nature of precipitation in the region, such an approach may lead to an inaccurate estimation of precipitation. This may also have a significant effect on the underestimation of the performance of the spatial forecast models. The authors should consider using weighted methods (e.g. Thiessen polygons) or clearly state in the Discussion that some of the errors attributed to the forecasts may in fact result from an inadequate spatial representation of the observations themselves based on the available rain-gauge data.
Author Response
We thank the reviewer for his time and constructive comments on our manuscript. We have carefully considered all the points raised and provide our responses and clarifications below. The manuscript has been revised where appropriate to further improve its clarity and presentation.
Comment 1: However, a significant issue related to the use of future information still remains regarding the QDM method. In Equation (1), the empirical cumulative distribution function (CDF) covering the entire test period is used to determine the quantile position for a given forecasted value (lines 308–310). This means that correcting a given forecast utilizes information that would not yet be available at the time the forecast is issued. This directly contradicts the claimed "strictly out-of-sample" approach (lines 328–329) as well as the stated capability of utilizing the method operationally in real-time conditions (lines 213–214). Therefore, the results should be recalculated while preserving the full chronology of data availability, ensuring that the correction of each forecast does not utilize information from the future. If this approach is not feasible, it must be clearly stated in the manuscript that the applied procedure utilizes information from the entire test period and, consequently, does not reflect real-time operational conditions.
Response 1:
We fully agree that operational forecasting evaluation must be strictly out-of-sample and preserve exact chronology to avoid look-ahead bias. However, the data leakage concern stems from an ambiguity in our manuscript regarding how the cumulative distribution function (CDF) is constructed dynamically, rather than statically, for each specific lead time. Our approach preserves full chronology and does not utilize future information.
Independent Lead-Time Series The core of our methodology treats each forecast lead time (e.g., L0, L1 ... L6) as a completely independent, parallel time series. The QDM correction is applied to each lead time separately using a dynamic rolling window, not by extracting a single CDF from an entire static future block.
The Chronology of the Rolling Window To demonstrate that no future information is utilized in Equation (1), consider a 6-day lead time (L6) forecast with a target date of January 7th.
- The Issue Date: This forecast is generated and issued on January 1st.
- The Test Window: For this specific lead time, we construct a 360-day rolling window of past forecasts ending on the target date (January 7th). This window contains the L6 forecast targeting Jan 7th, the L6 forecast targeting Jan 6th, Jan 5th, and so on, rolling backward.
- The Out-of-Sample Proof: Because this is an independent L6 series, the forecast targeting Jan 7th was issued on Jan 1st. The forecast targeting Jan 6th was issued on Dec 31st. Therefore, at the exact moment of issuance (January 1st), every single forecast value within that 360-day test window has already been generated. No forecast in this calculation was issued after January 1st.
- The Training Window: The remaining historical period used for the baseline CDFs (both historical observations and forecasts) entirely precedes this 360-day window. All observations used in the correction occurred well before the January 1st issue date.
By anchoring the rolling window to the target date of an independent lead-time series, the latest issue date inside the calculation window is always exactly the current issue date. The mathematical chronology remains intact.
We recognize that this dynamic adaptation of QDM was not explained clearly enough. We have revised the manuscript [339-355] to explicitly detail this diagonal data preparation and rolling window structure, clarifying how the methodology remains strictly out-of-sample and viable for real-time operational conditions.
Comment 2: Lines 400–405: in the case of the AR(1) predictor, it is not clearly specified what the lagged variable represents. It should be clarified whether, for the individual lead times (L0–L7), the predictor uses the actual precipitation observation from the previous day or a forecast value. If the observation is used, there is an additional issue concerning the availability of such observations for longer lead times, e.g. L3 or L5. This requires clarification in the context of a real-time operational system, since it is necessary to exclude the possibility of using information that would not yet be available at the time the forecast is issued.
Response 2:
In the implementation used in our analysis, the AR(1) predictor utilizes the actual precipitation observation from the day immediately preceding the target day, rather than a forecasted value.
We fully acknowledge the operational issue identified by the reviewer. Under the retrospective evaluation framework adopted in this study, the observed precipitation from the previous day was available for constructing the AR(1) predictor for all lead times. However, in a strict real-time operational application, this information would not necessarily be available at the time the forecast is issued for longer lead times. For example, a forecast issued on Day 1 targeting Day 4 (L3) would require the precipitation observation from Day 3, which would not yet be available on Day 1. Therefore, directly using the observation from the day immediately preceding the target day for longer lead times could introduce information that would not be available at the forecast issuance time.
To preserve strict chronology in a real-time operational system, the lagged predictor must be constructed sequentially. For the shortest lead times where the previous day's observation is physically available at issuance, the actual observation serves as the AR(1) predictor. For subsequent lead times, the predictor dynamically substitutes the actual observation with the corrected forecast from the preceding lead time (e.g., the predictor for L2 utilizes the corrected forecast for L1). This chained approach ensures that only information physically available at the time of forecast issuance is utilized.
We have revised the manuscript to explicitly clarify what the lagged variable represents, distinguish the retrospective implementation adopted in this study from a strictly real-time operational implementation, and discuss this limitation and the possible approach for overcoming it [427-442].
Comment 3: Lines 325–326: the multiplicative QDM correction factor was limited to a value of 2.0 (ratio.max = 2). Such a limitation may artificially constrain the ability to correct the mean and variability of the adjusted precipitation. It is therefore worth considering whether the deterioration of the individual KGE components, described in lines 691–698, actually results from a weakness of the QDM method. It cannot be ruled out that the technical limitation imposed here plays an important role. A sensitivity analysis using higher values of ratio.max (e.g. 5 or 10) would be justified. If such a test was not performed, this issue should at least be addressed in the Discussion as a possible technical explanation for the lack of improvement in KGE.
Response 3: The imposed value of ratio.max = 2 may constrain the magnitude of the multiplicative correction, particularly for precipitation events requiring larger adjustments. However, we did not perform a sensitivity analysis using higher values of ratio.max (e.g., 5 or 10), and therefore we cannot determine from the current results whether this constraint contributed to the observed deterioration in the KGE components.
We have consequently addressed this point in the Discussion as a possible technical explanation and identified sensitivity testing of the ratio.max parameter as an important topic for further investigation. In particular, our results showed that improvements in RMSE after QDM correction were not consistently accompanied by improvements in KGE. The decomposition of KGE indicated that the largest deterioration was associated with the variability component. Although this behavior may be partly related to the imposed upper limit on the multiplicative correction factor, this explanation remains a hypothesis because the corresponding sensitivity analysis was not performed.
The Discussion has therefore been revised to state that the fixed ratio.max = 2 may have limited the magnitude of the precipitation correction and that future analyses should evaluate higher values of this parameter to determine its influence on the mean, variability, and resulting KGE [848-864].
Comment 4: Lines 201–203: calculating the spatial mean precipitation using a simple arithmetic mean based on only two measurement points for a catchment is questionable. Given the convective nature of precipitation in the region, such an approach may lead to an inaccurate estimation of precipitation. This may also have a significant effect on the underestimation of the performance of the spatial forecast models. The authors should consider using weighted methods (e.g. Thiessen polygons) or clearly state in the Discussion that some of the errors attributed to the forecasts may in fact result from an inadequate spatial representation of the observations themselves based on the available rain-gauge data.
Response 4: The basin-average precipitation was calculated using the arithmetic mean of the available rain gauges. We recognize that, particularly for basins represented by only two stations and under convective precipitation conditions, this approach may not fully capture the spatial variability of precipitation across the basin.
A spatial weighting method, such as Thiessen polygons, was not applied in the present analysis. Therefore, we have clarified in the Discussion that the station-based basin precipitation represents an estimate based on the available gauge network and that uncertainties in its spatial representation may also contribute to differences between the observations and the spatial forecast products.
This limitation has been added to the Discussion to provide appropriate context for the interpretation of the forecast evaluation results [875-887].
Round 2
Reviewer 3 Report (Previous Reviewer 3)
Comments and Suggestions for AuthorsThank you to the Authors for the responses to all questions / comments.
The authors clarified the method of creating the CDF in the QDM method and the application of independent time series for individual lead times. They explained the operation of the rolling window and the preservation of data chronology. Furthermore, they addressed the application of the AR(1) predictor, indicating both the use of observations from the day preceding the forecast and the limitations of this approach for longer lead times. The proposed chained solution helps maintain data chronology in operational forecasting. Regarding the ratio.max = 2 limitation, the authors acknowledged the lack of sensitivity analysis for higher values, but included this limitation in the discussion as a possible factor influencing the KGE results and a topic for further research. They also addressed the representativeness of precipitation data, pointing out the limitations of the arithmetic mean with two stations.
This manuscript is a resubmission of an earlier submission. The following is a list of the peer review reports and author responses from that submission.
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe manuscript entitled “Evaluating rainfall forecast skill in numerical weather prediction models and the effects of bias correction on the PCJ (Piracicaba-Capivari-Jundiaí) river basins in Brazil” addresses a relevant topic for operational hydrology, particularly in relation to precipitation forecasting, hydrological modeling and water resources management.
The manuscript may be reconsidered after major revision. The main issues concern the limited spatial generalization of the occurrence-correction analysis, the need for a clearer explanation of the categorical verification framework, the inconsistency between the abstract and the results, the metric-dependent interpretation of the QDM results, and the need for more detailed methodological and reproducibility information, particularly for the Bayesian models.
Detailed reviewer comments and suggestions are provided in the attached Word file. The authors should address these points carefully in order to improve the scientific clarity, methodological robustness and presentation quality of the manuscript.
Comments for author File:
Comments.pdf
The English language is generally understandable; however, it requires revision to improve clarity, grammar and scientific expression. Several sentences are overly long, some terms are used inconsistently, and minor grammatical errors are present. A careful language edit is recommended before publication.
Author Response
Response for Reviewer 1
Comment 1: The occurrence-correction analysis should not be overgeneralized. The manuscript evaluates three basins, but the extended dry/wet occurrence-correction analysis, especially the hierarchical Bayesian part, is restricted to the Atibaia-Atibaia basin. The authors should either extend at least a simplified occurrence-correction analysis to the remaining basins or clearly state throughout the abstract, discussion and conclusions that these findings are exploratory and basin-specific.
Response 1: We thank the reviewer for this important comment. We agree that restricting the hierarchical Bayesian occurrence-correction analysis to a single basin limited the generalizability of the original conclusions. In response, we have extended the occurrence-correction analysis to all three basins, including both the frequentist and hierarchical Bayesian approaches. The revised analysis therefore evaluates the two occurrence-correction frameworks across the three basins, four forecast–observation pairings, and seven forecast lead times, providing a more comprehensive assessment of their consistency and applicability across different basin characteristics and forecast configurations.
Accordingly, we have revised the Abstract, Results, Discussion, and Conclusions to avoid basin-specific overgeneralization and to frame the findings in terms of the evidence obtained across the three study basins. This expansion also strengthens the assessment of whether the observed performance of the occurrence-correction approaches is consistent beyond the Atibaia basin.
Comment 2: The dry/wet categorical verification framework must be clarified. The manuscript appears to treat dry or zero-rainfall days as the target event, which is acceptable but may confuse readers because many precipitation verification studies define wet/rain occurrence as the target event. The definitions of hits, false alarms, misses and correct negatives should be explained more explicitly, preferably with a short example.
Response 2: We thank the reviewer for pointing out the potential ambiguity in the categorical verification framework. We have clarified in the revised manuscript that dry days are defined as the target event, rather than wet days, because accurately identifying dry conditions is of particular interest in the study area. We recognize that precipitation verification studies commonly define wet/rain occurrence as the target event, and therefore we have made this convention explicit to avoid confusion.
Specifically, a day is classified as dry when precipitation is ≤0.1 mm and as wet when precipitation is >0.1 mm. With dry conditions defined as the target event, the contingency table is:
|
FCST/OBS |
OBS dry (0) <= 0.1 mm |
OBS wet (1) > 0.1 mm |
|
FCST dry (0) <= 0.1 mm |
Hits (H) |
False alarms (FA) |
|
FCST wet (1) > 0.1 mm |
misses (M) |
correct negative (CN) |
The corresponding definitions and table have been added/clarified in Section 2.3 (Assessment Metrics).
Comment 3: The abstract should be revised for consistency with the results. The abstract states that FCST-SIM consistently outperformed FCST-ENS, whereas the results indicate that FCST-ENS shows stronger categorical performance for several metrics, including higher ACC/CSI and lower FAR. The abstract should reflect the metric-dependent nature of the findings.
Response 3: We thank the reviewer for identifying this inconsistency. We have substantially revised the Abstract to ensure that its conclusions are consistent with the results. In particular, the previous statement that FCST-SIM consistently outperformed FCST-ENS has been removed. The revised Abstract now emphasizes that the effects of occurrence correction depend on the forecast–observation pairing, basin, and forecast lead time, rather than indicating a consistent superiority of one forecasting system.
Comment 4: The interpretation of QDM performance should be more cautious. The results suggest that QDM can reduce RMSE in some cases, especially at longer lead times, but this improvement doesn't consistently translate into improved KGE. The authors should clearly explain that bias correction performance depends on the selected metric, lead time, basin and forecast-observation pair.
Response 4: We thank the reviewer for this valuable comment and agree that the interpretation of the QDM results should be more cautious. The revised manuscript has been expanded to emphasize that the performance of the QDM bias-correction method is metric-dependent and varies according to the verification metric, lead time, basin, and forecast–observation pair. In addition, we have included a decomposition of the Kling–Gupta Efficiency (KGE) to provide further insight into the aspects of model performance that are improved or degraded by the QDM correction. These additions provide a more balanced interpretation of the results and clarify the strengths and limitations of the method.
The corresponding definitions and table have been added/clarified in Section 3.3 (Bias-Correction Performance Using the Complete Chronological Series) [570, 606].
Comment 5: The training/testing design and threshold selection procedure need clearer reporting. The exact training and testing periods should be provided after the common overlapping period and missing-value removal. The authors should also confirm that QDM parameters, monthly wet-day probabilities, threshold optimization and all model-selection decisions were derived strictly from the training data to avoid temporal leakage.
Response 5: We thank the reviewer for this valuable comment. We agree that the description of the training and testing strategy and the procedures used for calibration and threshold selection should be presented more clearly. In the revised manuscript, we have expanded the Methods section to explicitly describe the chronological training and testing periods after defining the common overlapping period and removing missing values. Specifically, 1,318 daily observations from 7 September 2019 to 21 April 2023 were used for training, while the remaining 330 observations from 22 April 2023 to 20 March 2024 were reserved for out-of-sample testing. [190, 193]
We have also clarified that all parameters and decisions requiring calibration or optimization—including the QDM parameters, monthly wet-day probabilities, occurrence-correction model parameters, classification thresholds—were derived exclusively from the training dataset. [205,207]; [215, 216].
Comment 5: The use of the ECMWF ensemble median as the main predictor should be justified. Using only the ensemble median may discard useful probabilistic information such as ensemble spread, wet member frequency and ensemble quantiles. The authors should discuss why this simplification was preferred and whether probabilistic ensemble predictors could improve the correction of dry/wet occurrence.
Response 6: We thank the reviewer for highlighting the need to justify the use of the ECMWF ensemble median as the main predictor. In response, we have added a new subsection (Section 2.2.1, Selecting the Ensemble Predictor) that evaluates four ensemble summary statistics—mean, median, 25th percentile, and 75th percentile—as candidate deterministic predictors. Their performance was compared using Bias, RMSE, and the Kolmogorov–Smirnov (KS) statistic across the two observational references, the three basins, and the seven forecast lead times. The ensemble median provided the best overall balance across the evaluated criteria and therefore was selected as the deterministic ensemble predictor for the subsequent analyses.
Comment 7: The added value of the Bayesian framework should be explained more clearly. Since the frequentist and Bayesian approaches show similar performance, the manuscript should clarify what the Bayesian model contributes beyond the frequentist model. If uncertainty quantification is the main advantage, posterior intervals, convergence diagnostics and basic calibration checks should be reported.
Response 7: We thank the reviewer for this comment. The revised manuscript clarifies that the Bayesian and frequentist frameworks were evaluated as competing approaches to determine which could provide more accurate and consistent dry/wet occurrence forecasts. The comparison was performed under the same experimental conditions across all three basins, four forecast–observation pairings, and seven forecast lead times. Although the Bayesian models did not consistently outperform the frequentist models across all configurations, HBLR-AR1 showed the most balanced overall performance among the four occurrence-correction approaches. The Discussion has been revised to reflect this result and to avoid implying a systematic predictive advantage of the Bayesian framework.
Comment 8: The computational-cost argument should be strengthened. Stating that the full Bayesian analysis required approximately three hours per basin does not by itself fully justify restricting the analysis to only one basin. The authors should either provide a stronger operational rationale or include a lighter comparison across the other basins.
Response 8: We thank the reviewer for this comment. We agree that computational cost alone did not provide sufficient justification for restricting the Bayesian analysis to a single basin. Following the reviewer’s suggestion, we therefore extended the Bayesian occurrence-correction analysis to all three basins in the revised manuscript. The revised analysis comprises 168 Bayesian model fits (3 basins × 4 forecast–observation pairings × 7 forecast lead times × 2 model configurations), allowing the Bayesian and frequentist approaches to be compared under the same experimental design across the complete study domain. Consequently, the original computational-cost limitation is no longer used as a justification for restricting the Bayesian analysis to one basin.
Comment 9: The reproducibility of the Bayesian models should be improved. The manuscript should report key MCMC and software details, including the number of chains, burn-in, thinning, posterior sample size, convergence diagnostics such as R-hat/effective sample size, prior sensitivity checks and package versions.
Response 9: We thank the reviewer for highlighting the need for greater transparency in the implementation of the Bayesian models. We have revised the Methods section to explicitly report the MCMC and software settings used in the analysis. The revised Bayesian analysis comprises 168 model fits (3 basins × 4 forecast–observation pairings × 7 forecast lead times × 2 model configurations), all implemented using JAGS with three MCMC chains. For each fit, 5,000 iterations were discarded as burn-in, followed by 20,000 sampling iterations per chain, with no thinning, resulting in 60,000 post-burn-in draws per parameter. The software and package versions used for the Bayesian analysis have also been added to the revised manuscript. [302,304]
Given the substantial expansion of the Bayesian analysis to all three basins, and the primary focus of the study on comparing the predictive performance and generalizability of the frequentist and Bayesian occurrence-correction approaches, we have kept the presentation focused on the resulting forecast-performance evaluation rather than providing an extensive model-by-model presentation of MCMC diagnostics and prior-sensitivity analyses. The revised manuscript now provides the MCMC specifications and implementation details necessary to reproduce the Bayesian sampling procedure.
Comment 10: The conclusions should distinguish robust findings from exploratory findings. Results supported across basins should be separated from findings derived only from the single-basin Bayesian occurrence-correction experiment. This will prevent overinterpretation and improve the manuscript's scientific balance.
Response 10: We thank the reviewer for this important comment. In the revised manuscript, the occurrence-correction analysis, including the hierarchical Bayesian models, was extended to all three basins. Therefore, the conclusions are no longer based on a single-basin Bayesian experiment. The Conclusions section has also been revised to distinguish the findings that were consistently observed across basins and forecast configurations from those that were more dependent on the specific basin, forecast–observation pairing, or lead time, thereby avoiding overgeneralization of the results.
Comment 11: The literature review is relevant but should be strengthened with a few recent and directly related studies. In particular, the claimed research gap regarding the limited use of QDM in short-term operational NWP post-processing should be supported with recent references. The authors may also add recent work on operational precipitation post-processing, NWP-based rainfall forecast bias correction, dry/wet occurrence correction, Bayesian or frequentist occurrence modeling, and rainfall-runoff forecasting driven by corrected precipitation forecasts. The authors should select the most relevant studies themselves; the purpose should be to clarify novelty and positioning, not merely to increase citation numbers.
Response 11: We thank the reviewer for this valuable suggestion. We have strengthened the literature review by adding recent and directly relevant studies on NWP precipitation post-processing, QDM and other bias-correction methods, dry/wet occurrence correction, and frequentist and Bayesian approaches. The revised literature review now better establishes the three main gaps addressed by this study: (1) the comparatively limited application of QDM to short- and medium-range operational NWP precipitation forecasts; (2) the limited direct comparison between frequentist and Bayesian logistic-regression approaches for dry/wet occurrence correction; and (3) the need to evaluate these post-processing approaches across hydrologically and meteorologically complex basins, where local precipitation characteristics may affect forecast representativeness and correction effectiveness. These additions clarify the novelty and positioning of the present study within the existing literature.
Comment 12:The figures and tables need minor improvements. Figure 2 and Figure 3 are useful but crowded; font sizes, legends and captions should be improved. Table 3 should use decimal points rather than decimal commas in an English manuscript.
Response 12: We thank the reviewer for these suggestions. Figure 2 has been revised to improve readability, including adjustments to the font size, legend, and caption. Figure 3 and Table 3 have been removed from the revised manuscript as part of the restructuring of the presentation. In addition, all numerical values in the revised manuscript now use decimal points, consistent with English-language conventions.
Comment 13: The manuscript would benefit from careful language editing. Examples include replacing 'sensibility' with 'sensitivity', using 'which explicitly preserves' when referring to QDM, changing 'highlight' to 'highlights' where the subject is singular, using 'was applied to both' rather than 'was implemented using the QDM method to both', and writing 'et al.' instead of 'et. al.'. These language issues are secondary to the scientific comments but should be corrected before publication.
Response 13: We thank the reviewer for these helpful suggestions. The manuscript has been carefully revised for English language, grammar, terminology, and consistency. The specific issues identified by the reviewer have been corrected throughout the manuscript, including the use of “sensitivity” instead of “sensibility,” “which explicitly preserves” when referring to QDM, appropriate subject–verb agreement, “was applied to both,” and the correct citation format “et al.” We have also conducted a broader language review to improve the clarity and readability of the revised manuscript.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThis manuscript evaluates rainfall forecast skill and bias-correction methods over the PCJ River Basin in Brazil by comparing the SIMEPAR operational forecast and the ECMWF ensemble forecast, and by testing QDM and dry/wet occurrence correction methods. The topic is operationally relevant, the study period is reasonably long, and the attempt to combine intensity correction with occurrence correction is valuable; however, substantial issues remain in the consistency of conclusions, reproducibility of methods, generalizability of results, and support for hydrological application claims.
1. There is a clear inconsistency between the Abstract and the main results. Lines 17–19 state that FCST-SIM consistently outperformed FCST-ENS, whereas Section 3.1, lines 386–392, and the Discussion, lines 509–512, indicate that FCST-ENS generally achieved higher ACC and CSI and lower FAR with more stable lead-time behavior. The authors should recheck the calculations and clearly state which forecasting system performs better under which metric, lead time, basin, and observational reference.
2. The manuscript repeatedly links the results to hydrological-model inputs and operational hydrological forecasting, but no streamflow simulation or discharge-forecast verification is actually conducted. If improving operational hydrological applications is a central claim, at least one experiment showing how corrected precipitation propagates into streamflow forecasts should be added; otherwise, the conclusions should be restricted to precipitation forecast post-processing.
3. The data-processing workflow is not sufficiently transparent, especially the conversion of datasets with different spatial resolutions into basin-averaged precipitation for the three sub-basins. Lines 146–189 describe OBS-STN, OBS-GRD, FCST-SIM, and FCST-ENS, but the manuscript does not adequately explain area weighting, grid resampling, accumulation windows, time-zone consistency, missing-data handling, or rain-gauge quality control. Given that some basins contain only two to four gauges, these details can materially affect the verification results.
4. The definition of dry/wet events should be made consistent throughout the manuscript. Lines 232–237 define wet days using precipitation > 0.1 mm, whereas lines 336–350 formulate the contingency table using forecast = 0 and observed = 0 as the dry target event. In addition, the Discussion describes FCST-SIM as having a wet bias, while lines 517–518 state that it tends to overpredict dry conditions; these interpretations are physically inconsistent and need to be corrected.
5. The description of the QDM bias-correction procedure is too brief to allow reproducibility. The authors should provide the exact QDM equations, empirical distribution estimation procedure, extrapolation strategy for extreme quantiles, treatment of zero precipitation, whether calibration was performed separately by season and lead time, and whether the 80%/20% split represents a strict out-of-sample chronological validation. Since Figure 4 shows RMSE improvements but frequent KGE degradation, the KGE components should also be decomposed to clarify what QDM improves and what it degrades.
6. The dry/wet occurrence-correction analysis has limited generalizability. Lines 325–334 state that model comparison was performed only for the Atibaia–Atibaia basin, lead time L0, and the ENS-STN pair, and the extended analysis still remains limited to one basin. This design does not support broad conclusions for all three PCJ basins or both forecast systems; at minimum, the computationally cheaper frequentist logistic regression should be extended to all basins and lead times, or this part should be explicitly framed as an exploratory case study.
7. The ensemble information in FCST-ENS is underused, and uncertainty assessment is lacking. Although ECMWF provides 51 ensemble members, the analysis mainly relies on the ensemble median after QDM and also uses the median as the predictor in the occurrence models. The authors should consider adding ensemble mean, spread, and probability-based predictors, and should report bootstrap confidence intervals or significance tests for the improvements shown in Figures 2–4 and Table 3; otherwise, it is difficult to judge whether some small improvements are statistically meaningful.
Author Response
Comment 1: There is a clear inconsistency between the Abstract and the main results. Lines 17–19 state that FCST-SIM consistently outperformed FCST-ENS, whereas Section 3.1, lines 386–392, and the Discussion, lines 509–512, indicate that FCST-ENS generally achieved higher ACC and CSI and lower FAR with more stable lead-time behavior. The authors should recheck the calculations and clearly state which forecasting system performs better under which metric, lead time, basin, and observational reference.
Response 1: We thank the reviewer for identifying this inconsistency. We agree that the statement in the Abstract was incorrect due to an editing error and does not accurately reflect the results presented in the manuscript. We have carefully rechecked the reported results and confirmed that the analyses and conclusions in the Results and Discussion sections are correct. The revised Abstract now emphasizes that the effects of occurrence correction depend on the forecast–observation pairing, basin, and forecast lead time, rather than indicating a consistent superiority of one forecasting system.
Comment 2: The manuscript repeatedly links the results to hydrological-model inputs and operational hydrological forecasting, but no streamflow simulation or discharge-forecast verification is actually conducted. If improving operational hydrological applications is a central claim, at least one experiment showing how corrected precipitation propagates into streamflow forecasts should be added; otherwise, the conclusions should be restricted to precipitation forecast post-processing.
Response 2: We appreciate the reviewer's comment and will make it more explicit in the revised manuscript that the aim of the present study is to evaluate precipitation bias-correction methods and assess their ability to improve the raw forecast data. This represents an important step toward the broader objective of improving hydrological forecasting. The impact of the corrected precipitation forecasts on hydrological simulations and forecasting will be analyzed and evaluated in a subsequent study. [701,706]
Comment 3:The data-processing workflow is not sufficiently transparent, especially the conversion of datasets with different spatial resolutions into basin-averaged precipitation for the three sub-basins. Lines 146–189 describe OBS-STN, OBS-GRD, FCST-SIM, and FCST-ENS, but the manuscript does not adequately explain area weighting, grid resampling, accumulation windows, time-zone consistency, missing-data handling, or rain-gauge quality control. Given that some basins contain only two to four gauges, these details can materially affect the verification results.
Response 3: We thank the reviewer for this valuable comment and agree that the data-processing workflow should be described more clearly. For the OBS-GRD dataset, no additional spatial resampling was performed because the product is already provided on a regular grid covering the entire study region. Basin-averaged precipitation was obtained by extracting the grid cells located within each basin and calculating their spatial average. [167,169]
The FCST-SIM and FCST-ENS the forecast data were referenced to the same local time (UTC−3), ensuring temporal consistency across all datasets and aggregated to a common daily temporal resolution [154, 156], after which basin-average precipitation was calculated using the same procedure adopted for the observational gridded dataset.
For the OBS-STN dataset, the rain-gauge observations were first subjected to quality-control procedures, including the removal of negative values and physically unrealistic precipitation values. The quality-controlled observations were then averaged within each basin to obtain basin-mean precipitation. [157,163].
The revised manuscript has been expanded to describe these processing steps more clearly and to improve the transparency and reproducibility of the methodology.
Comment 4: The definition of dry/wet events should be made consistent throughout the manuscript. Lines 232–237 define wet days using precipitation > 0.1 mm, whereas lines 336–350 formulate the contingency table using forecast = 0 and observed = 0 as the dry target event. In addition, the Discussion describes FCST-SIM as having a wet bias, while lines 517–518 state that it tends to overpredict dry conditions; these interpretations are physically inconsistent and need to be corrected.
Response 4: We thank the reviewer for this careful observation. We agree that the description of the event definition and the contingency table can be made clearer.
In this study, wet days are defined as days with precipitation greater than 0.1 mm, while days with precipitation less than or equal to 0.1 mm are classified as dry days. The contingency table was constructed using this threshold, with dry days considered as the target event. To avoid ambiguity, we have revised the manuscript so that the contingency table is described using the precipitation threshold (≤ 0.1 mm and > 0.1 mm) rather than the notation "0" and "≠ 0". The corresponding definitions and table have been added/clarified in Section 2.3 (Assessment Metrics).
Regarding the interpretation of FCST-SIM, the discussion of the raw forecasts and the Bayesian-corrected forecasts referred to different stages of the analysis, but we agree that this distinction was not sufficiently clear. The revised manuscript has been edited to explicitly distinguish between these two analyses and to clarify the interpretation of the verification metrics, thereby avoiding potential misunderstandings.
Comment 5: The description of the QDM bias-correction procedure is too brief to allow reproducibility. The authors should provide the exact QDM equations, empirical distribution estimation procedure, extrapolation strategy for extreme quantiles, treatment of zero precipitation, whether calibration was performed separately by season and lead time, and whether the 80%/20% split represents a strict out-of-sample chronological validation. Since Figure 4 shows RMSE improvements but frequent KGE degradation, the KGE components should also be decomposed to clarify what QDM improves and what it degrades.
Response 5: We thank the reviewer for this valuable comment. The revised manuscript now explains the QDM formulation and reports the parameters used in its implementation with the QDM() function from the MBC R package. The training and testing datasets were separated chronologically, with the first 80% of the time series used for calibration and the remaining 20% reserved for an independent out-of-sample evaluation. This has also been clarified in the revised manuscript. [248, 270]
Regarding the reviewer's suggestion to decompose the Kling–Gupta Efficiency (KGE) into its individual components, we agree that this is a valuable analysis, as it would help explain which aspects of model performance are improved and which are degraded by the QDM correction. Accordingly, this analysis has been added to the revised manuscript, and the corresponding results and discussion have been expanded. [572, 608]
Comment 6: The dry/wet occurrence-correction analysis has limited generalizability. Lines 325–334 state that model comparison was performed only for the Atibaia–Atibaia basin, lead time L0, and the ENS-STN pair, and the extended analysis still remains limited to one basin. This design does not support broad conclusions for all three PCJ basins or both forecast systems; at minimum, the computationally cheaper frequentist logistic regression should be extended to all basins and lead times, or this part should be explicitly framed as an exploratory case study.
Response 6: We thank the reviewer for this valuable comment. We agree that limiting the occurrence-correction analysis to a single basin reduced the generalizability of the results. In the revised manuscript, this analysis has been extended to include all three study basins, providing a more comprehensive evaluation of the proposed methodology. This revision strengthens the generality of the results and supports the conclusions drawn for the different hydrological settings considered in this study. The corresponding results have been added/clarified in Section 3.2 (Dry/Wet occurrence correction ).
Comment 7: The ensemble information in FCST-ENS is underused, and uncertainty assessment is lacking. Although ECMWF provides 51 ensemble members, the analysis mainly relies on the ensemble median after QDM and also uses the median as the predictor in the occurrence models. The authors should consider adding ensemble mean, spread, and probability-based predictors, and should report bootstrap confidence intervals or significance tests for the improvements shown in Figures 2–4 and Table 3; otherwise, it is difficult to judge whether some small improvements are statistically meaningful.
Response 7: We thank the reviewer for these valuable suggestions. We agree that incorporating additional ensemble characteristics, such as the ensemble mean, spread, and probability-based predictors, together with formal statistical significance analyses, could provide further insight into the performance of the proposed methodology. However, these additions would substantially expand the scope of the present study, whose primary objective is to evaluate and compare precipitation occurrence correction methods using a representative predictor derived from the ECMWF ensemble. In response, we have added a new subsection (Section 2.2.1, Selecting the Ensemble Predictor) that evaluates four ensemble summary statistics—mean, median, 25th percentile, and 75th percentile—as candidate deterministic predictors. Their performance was compared using Bias, RMSE, and the Kolmogorov–Smirnov (KS) statistic across the two observational references, the three basins, and the seven forecast lead times. The ensemble median provided the best overall balance across the evaluated criteria and therefore was selected as the deterministic ensemble predictor for the subsequent analyses. We also acknowledge that exploiting the full ensemble information and quantifying the statistical significance of the reported improvements are important topics for future research and represent a natural extension of the present work.
To support the significance of these subtle improvements, the updated manuscript details the methodological framework used for assessing the occurrence-correction techniques [370, 410], as summarized below:
|
Assessment |
Purpose |
Unit of analysis |
Statistical / descriptive approach |
Interpretation |
|
Paired classification analysis |
Assess whether occurrence correction changes individual dry/wet classifications relative to the raw forecast |
Daily classifications, separately for each lead time |
Exact McNemar test; OR = b/c; exact 95% CI; Holm adjustment across seven lead times |
OR > 1 indicates more corrected errors than newly introduced errors |
|
Metric-based analysis |
Assess whether occurrence correction systematically changes categorical verification metrics |
Seven lead-time differences for each basin × forecast–observation pairing × method |
One-sided exact paired permutation test; bootstrap 95% CI; Holm adjustment across the four methods |
Positive differences indicate improvement; p<0.05 indicates evidence of systematic improvement after adjustment |
|
Multi-metric method ranking |
Identify the method providing the most balanced overall performance across the four metrics |
Seven-lead-time mean for each basin × forecast–observation pairing × method |
Rank POD, CSI, ACC, and FAR from 1 (best) to 4 (worst); overall rank = mean of four metric ranks |
Lower overall rank indicates better and more balanced performance |
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript has practical value. The authors describe an operational hydrometeorological forecasting system for the economically important PCJ River Basin. The Introduction discusses issues related to reservoir operation and streamflow simulation. Figure 1 presents the locations of the hydrological stations. This suggests that the manuscript addresses broader hydrometeorological problems. In reality, however, the study is limited solely to the statistical correction of precipitation forecasts. It does not examine whether this correction translates into improved streamflow forecasts. The practical benefit of the proposed methodology is also questionable. In the best-case scenario (Figure 3, lead time L0), the maximum improvement corresponded to only 21 additional days with a correct forecast over the entire 330-day test dataset. Moreover, the average results presented in Table 3 for the gridded datasets (ENS/GRD) show that the advanced Bayesian models increased the overall accuracy (ACC) by only 0.37% (HBLR-Seasonal) and 0.5% (HBLR-AR1) compared with the raw forecasts. Implementing computationally demanding methods to achieve such small improvements in precipitation forecast skill, while not evaluating their impact on streamflow modelling, raises doubts about the practical value of the proposed approach.
Specific comments are given below.
Figure 1. Without the river network, the hydrological stations appear to be floating on the map background. A similar comment applies to the reservoirs discussed in the text. The authors emphasize the importance of the retention reservoirs, but their locations are not shown. In addition, the figure caption refers to "control stations" (three of which are mentioned in the text), whereas the map displays more than a dozen points labelled only as "Hydrological stations" in the legend. Separate symbols should be used for the main SPHM-PCJ control stations and for the remaining reference stations and rain gauges.
Lines 147–189. The authors describe different temporal coverage for the OBS-STN dataset (until January 2026) and the OBS-GRD dataset (until March 2024), and then state that all analyses were performed for the common period from 7 September 2019 to 20 March 2024. This should be stated more clearly in the data description. The current presentation suggests that some analyses use data extending to 2026, making the validation procedure difficult to follow.
Lines 211–214. The authors state that the models with the "best performance according to the verification metrics" were selected. However, they do not specify which metric determined the final selection. It is also unclear how cases were handled when different models performed best according to different verification metrics.
Lines 207–217. Reducing the 51 ECMWF ensemble members to a single median solely because of "computational costs", and subsequently developing a probabilistic Bayesian model based on this median, raises serious methodological concerns. Likewise, justifying the analysis of only one catchment solely by the computation time (approximately three hours) is not scientifically convincing. A three-hour runtime is not a meaningful limitation for modern numerical analyses. This approach removes almost all information about forecast uncertainty contained in the ensemble prediction. No analysis is provided to demonstrate that the ensemble median is an adequate representation of the full ensemble. The impact of this decision on the final results has also not been evaluated. This is the most important methodological issue in the manuscript.
The test period is too short. Although the authors have access to a valuable reference dataset extending back to 1961, the need to match it with the available FCST-ENS forecasts (available only since September 2019) reduced the analysis period to just 1,648 days. As a result, the 20% test dataset contains approximately 330 days, which is less than one complete hydrological cycle. This period is too short for a reliable validation of the proposed models.
Figure 2. The use of different y-axis ranges across the panels is misleading. Since all categorical metrics (ACC, CSI, FAR and POD) are bounded between 0 and 1, the same y-axis limits should be used for all panels (preferably 0–1). The current scaling hampers visual comparison among basins and may exaggerate apparent differences in forecast performance.
Technical issues. Decimal separators are used inconsistently. Table 3 uses commas as decimal separators (e.g. "34,1", "−7,9"), whereas the rest of the manuscript uses decimal points, as required in English (e.g. "0.637").
Author Response
Comment 1: In reality, however, the study is limited solely to the statistical correction of precipitation forecasts. It does not examine whether this correction translates into improved streamflow forecasts. The practical benefit of the proposed methodology is also questionable. In the best-case scenario (Figure 3, lead time L0), the maximum improvement corresponded to only 21 additional days with a correct forecast over the entire 330-day test dataset. Moreover, the average results presented in Table 3 for the gridded datasets (ENS/GRD) show that the advanced Bayesian models increased the overall accuracy (ACC) by only 0.37% (HBLR-Seasonal) and 0.5% (HBLR-AR1) compared with the raw forecasts. Implementing computationally demanding methods to achieve such small improvements in precipitation forecast skill, while not evaluating their impact on streamflow modelling, raises doubts about the practical value of the proposed approach.
Response 1: We thank the reviewer for raising this important point regarding the practical implications of the precipitation forecast corrections. We agree that evaluating whether improved precipitation forecasts translate into improved streamflow forecasts is an important subsequent step. However, assessing the impact of the corrected precipitation forecasts on hydrological model performance and streamflow forecasting is beyond the scope of the present study. The objective of this work is to evaluate whether statistical post-processing methods can improve the quality and categorical representation of precipitation forecasts. The effect of these corrections on streamflow forecasts will be investigated in a subsequent stage of the research.
Regarding the magnitude and practical relevance of the reported improvements, the revised manuscript has substantially expanded the analysis to include all three basins and has revised the assessment framework. Rather than relying on individual percentage improvements or isolated lead-time results, the revised analysis evaluates changes using the appropriate statistical tests, confidence intervals, and descriptive multi-metric ranking across basins, forecast–observation pairings, and lead times. The revised Results and Discussion therefore place greater emphasis on the consistency and statistical evidence of the improvements, rather than interpreting small changes in individual metrics as evidence of broad practical superiority. The conclusions have also been revised accordingly to avoid overstating the practical benefits of the correction methods.
Comment 2: Figure 1. Without the river network, the hydrological stations appear to be floating on the map background. A similar comment applies to the reservoirs discussed in the text. The authors emphasize the importance of the retention reservoirs, but their locations are not shown. In addition, the figure caption refers to "control stations" (three of which are mentioned in the text), whereas the map displays more than a dozen points labelled only as "Hydrological stations" in the legend. Separate symbols should be used for the main SPHM-PCJ control stations and for the remaining reference stations and rain gauges.
Response 2: We appreciate the reviewer's valuable suggestion and agree that Figure 1 can be improved to better represent the study area and the monitoring network. In the revised manuscript, we will update the figure by including the river network and the locations of the main reservoirs discussed in the text. We will also distinguish the SPHM-PCJ control stations from the remaining hydrological stations and rain gauges using different symbols and an updated legend, thereby improving the clarity and consistency between the figure, its caption, and the manuscript text.
Comment 3: Lines 147–189. The authors describe different temporal coverage for the OBS-STN dataset (until January 2026) and the OBS-GRD dataset (until March 2024), and then state that all analyses were performed for the common period from 7 September 2019 to 20 March 2024. This should be stated more clearly in the data description. The current presentation suggests that some analyses use data extending to 2026, making the validation procedure difficult to follow.
Response 3: We appreciate the reviewer's comment. We agree that the preceding description of the individual datasets may give the impression that different temporal periods were used in the analyses. In the revised manuscript, we have clarified the data description by explicitly stating at the beginning of the section that, although the observational and forecast datasets have different temporal coverages, all analyses presented in this study were performed exclusively over their common overlapping period (7 September 2019 to 20 March 2024). The descriptions of the individual datasets have also been revised to emphasize that the extended records beyond this period are provided only to describe the full availability of each dataset and were not used in the validation. [152,156]
Comment 4: Lines 211–214. The authors state that the models with the "best performance according to the verification metrics" were selected. However, they do not specify which metric determines the final selection. It is also unclear how cases were handled when different models performed best according to different verification metrics.
Response 4: We appreciate the reviewer's valuable comment and agree that the model selection criterion should be described more explicitly. The updated manuscript details the methodological framework used for assessing the occurrence-correction techniques [370, 410], as summarized below:
|
Assessment |
Purpose |
Unit of analysis |
Statistical / descriptive approach |
Interpretation |
|
Paired classification analysis |
Assess whether occurrence correction changes individual dry/wet classifications relative to the raw forecast |
Daily classifications, separately for each lead time |
Exact McNemar test; OR = b/c; exact 95% CI; Holm adjustment across seven lead times |
OR > 1 indicates more corrected errors than newly introduced errors |
|
Metric-based analysis |
Assess whether occurrence correction systematically changes categorical verification metrics |
Seven lead-time differences for each basin × forecast–observation pairing × method |
One-sided exact paired permutation test; bootstrap 95% CI; Holm adjustment across the four methods |
Positive differences indicate improvement; p<0.05 indicates evidence of systematic improvement after adjustment |
|
Multi-metric method ranking |
Identify the method providing the most balanced overall performance across the four metrics |
Seven-lead-time mean for each basin × forecast–observation pairing × method |
Rank POD, CSI, ACC, and FAR from 1 (best) to 4 (worst); overall rank = mean of four metric ranks |
Lower overall rank indicates better and more balanced performance |
Comment 5: Lines 207–217. Reducing the 51 ECMWF ensemble members to a single median solely because of "computational costs", and subsequently developing a probabilistic Bayesian model based on this median, raises serious methodological concerns. Likewise, justifying the analysis of only one catchment solely by the computation time (approximately three hours) is not scientifically convincing. A three-hour runtime is not a meaningful limitation for modern numerical analyses. This approach removes almost all information about forecast uncertainty contained in the ensemble prediction. No analysis is provided to demonstrate that the ensemble median is an adequate representation of the full ensemble. The impact of this decision on the final results has also not been evaluated. This is the most important methodological issue in the manuscript.
Response 5: We sincerely thank the reviewer for this important comment and agree that the original manuscript did not provide an adequate justification for using only the ECMWF ensemble median. We also agree that attributing this decision primarily to computational cost was not scientifically sufficient. In response, we have added a new subsection (Section 2.2.1, Selecting the Ensemble Predictor) that evaluates four ensemble summary statistics—mean, median, 25th percentile, and 75th percentile—as candidate deterministic predictors. Their performance was compared using Bias, RMSE, and the Kolmogorov–Smirnov (KS) statistic across the two observational references, the three basins, and the seven forecast lead times. The ensemble median provided the best overall balance across the evaluated criteria and therefore was selected as the deterministic ensemble predictor for the subsequent analyses.
It should be emphasized that QDM was applied to all 51 ensemble members. However, the evaluation compared observed values against the median, which was also the metric utilized in the regression-based approaches.
Finally, we agree that restricting the evaluation to a single catchment weakened the generality of the original study. During the revision process, we expanded the analyses to include all three study catchments. The revised manuscript now presents results for all three basins.
Comment 6: The test period is too short. Although the authors have access to a valuable reference dataset extending back to 1961, the need to match it with the available FCST-ENS forecasts (available only since September 2019) reduced the analysis period to just 1,648 days. As a result, the 20% test dataset contains approximately 330 days, which is less than one complete hydrological cycle. This period is too short for a reliable validation of the proposed models.
Response 6: We thank the reviewer for this valuable comment. We acknowledge that the available common analysis period is relatively short and that the resulting 330-day test period covers less than one complete hydrological cycle. However, the ECMWF ensemble forecasts used in this study are only available from September 2019, which limits the common period with the observational datasets. Therefore, the analysis was conducted over the longest common period available across all forecast and observational datasets, resulting in 1,648 daily observations, of which 330 were reserved for independent out-of-sample testing. We recognize that the limited length of the available test period constrains the extent to which the results can be generalized across multiple hydrological cycles; nevertheless, this period represents the longest common period available for the complete set of datasets evaluated in this study.
Comment 7: Figure 2. The use of different y-axis ranges across the panels is misleading. Since all categorical metrics (ACC, CSI, FAR and POD) are bounded between 0 and 1, the same y-axis limits should be used for all panels (preferably 0–1). The current scaling hampers visual comparison among basins and may exaggerate apparent differences in forecast performance.
Response 7: We thank the reviewer for this helpful observation and agree that using different y-axis ranges may hinder visual comparison among the basins and exaggerate apparent differences in model performance. In the revised manuscript, Figure 2 has been updated so that all categorical performance metrics (ACC, CSI, FAR, and POD) are displayed using a common y-axis range of 0 to 1. This modification improves the consistency of the figure and facilitates direct comparison of model performance across the study basins.
Comment 8: Technical issues. Decimal separators are used inconsistently. Table 3 uses commas as decimal separators (e.g. "34,1", "−7,9"), whereas the rest of the manuscript uses decimal points, as required in English (e.g. "0.637").
Response 8: We thank the reviewer for identifying this inconsistency. Table 3 has been removed from the revised manuscript as part of the restructuring of the presentation. In addition, all numerical values in the revised manuscript now use decimal points, consistent with English-language conventions.
Author Response File:
Author Response.pdf