Next Article in Journal
Soil CO2 Efflux in Scots Pine Forests in Central Siberia After Wildfire and Logging: Diurnal and Seasonal Patterns
Previous Article in Journal
Linking Rainfall Intensity Variability to Local Adaptation Responses and Traditional Knowledge: A Mixed-Methods Case Study for Food Security Resilience in Boja, Indonesia
 
 
Article
Peer-Review Record

Evaluation and Post-Processing of Precipitation Forecast Skills at Short Lead Times for Hydrological Applications over the Ouémé Basin

Climate 2026, 14(7), 146; https://doi.org/10.3390/cli14070146
by Yaovi Aymar Bossa 1,2,* and Jean Hounkpè 1,2
Reviewer 1:
Reviewer 2:
Climate 2026, 14(7), 146; https://doi.org/10.3390/cli14070146
Submission received: 23 April 2026 / Revised: 20 June 2026 / Accepted: 21 June 2026 / Published: 10 July 2026
(This article belongs to the Topic Numerical Models and Weather Extreme Events (2nd Edition))

Round 1

Reviewer 1 Report (Previous Reviewer 1)

Comments and Suggestions for Authors

The revised manuscript has made some adjustments; however, many substantive reviewer comments from the previous round were either not addressed or insufficiently incorporated. As a result, the manuscript still requires significant reworking before it can be considered for publication.

Major Issues

  1. Abstract
    • The abstract remains descriptive rather than informative.
    • Key findings should be reported quantitatively (e.g., ranges or median improvements in KGE after bias correction).
    • Study period and forecast lead times (1–7 days) should be explicitly stated.
    • The metric list should be condensed to a general phrase such as “using complementary continuous and event-based verification metrics”.
  2. Introduction
    • This section requires a substantial overhaul.
    • Discussion of prior related studies is limited and not sufficiently up to date.
    • The research gap is not clearly articulated.
    • The novelty of the study (e.g., combined multi-metric evaluation and bias correction for West African precipitation forecasts) should be explicitly stated.
    • Objectives should be clearly listed at the end of the introduction rather than implied.
  3. Materials and Methods
    • Several sections are overly long yet lack clarity where most needed.
    • The number of rain gauge stations used should be clearly stated in the manuscript, with justification, particularly since some stations lie outside the basin.
    • The explanation of Thiessen polygons is unnecessarily detailed; a concise description with a standard reference is sufficient.
    • The source link for hindcast precipitation data should be explicitly provided.
    • Clarifications are required regarding:
      • Resampling from 1° to 0.25° (to avoid confusion with statistical downscaling),
      • Alignment of hindcast lead times with observations,
      • Whether ensemble means or individual members were analysed,
      • Separation (or not) of calibration and validation periods in bias correction.
  4. Bias-Correction Methods
    • The section is mathematically dense and would benefit from simplification.
    • The rationale for polynomial regression requires stronger justification.
    • Stationarity assumptions of quantile mapping should be acknowledged.
  5. Results and Discussion
    • Some repetition exists across subsections (e.g., KGE and APBias interpretation).
    • Basin-scale analysis should clarify statistical methods and emphasise that findings are indicative.
    • The discussion would benefit from clearer separation between interpretation, limitations, and future work.
    • Recent literature should be more fully integrated into the discussion.
  6. Conclusions
    • The conclusion should be condensed into a single coherent paragraph.
    • Core methodological contributions should be restated explicitly.
    • Practical hydrological implications (e.g., for flood forecasting and lead-time utility) should be briefly highlighted.

Minor Issues

  • Standardise mathematical notation throughout.
  • Improve clarity in descriptions of resampling and hindcast usage.
  • Reorder and standardise keywords (alphabetical order; consistent capitalisation).

In general, while the study addresses a relevant topic and has scientific potential, substantial revisions are still required to improve structure, clarity, and methodological transparency.

Comments for author File: Comments.pdf

Comments on the Quality of English Language

The manuscript demonstrates scientific relevance, but the authors did not adequately address many substantive reviewer comments from the previous round, particularly regarding the introduction, abstract, and methods. The issues are not cosmetic; they affect interpretability and transparency. In its current form, the paper should not be accepted without major revision. A carefully revised version could be publishable if the authors fully engage with the reviewer feedback.

Author Response

We thank the reviewer for the thorough and critical assessment of our manuscript. We appreciate the recognition of the relevance of the topic, particularly in the context of precipitation forecasting and hydrological applications in West Africa. We also acknowledge the concerns raised, and have undertaken a substantial revision of the manuscript. We believe these revisions significantly improve the clarity, robustness, and interpretability of the study. We respectfully hope that the revised version addresses the main concerns raised and demonstrates the scientific value of the work for assessing precipitation forecast skill and bias correction in data-sparse regions such as the Ouémé basin.

Author Response File: Author Response.pdf

Reviewer 2 Report (New Reviewer)

Comments and Suggestions for Authors

This manuscript presents a systematic evaluation of precipitation forecast skill from six numerical weather prediction (NWP) models over the Ouémé River basin in the West African monsoon region—a topic that carries both scientific value and operational promise. The authors adopt a verification framework spanning continuous metrics (Kling–Gupta Efficiency, KGE; Absolute Percentage Bias, APbias) and categorical measures (Likelihood Ratio, LHR). The construction of this dual-index system is generally sound and appropriate for the stated purpose.
However, the manuscript must be addressed before it can be considered for publication.
1. The manuscript is conspicuously silent on the provenance of the reference precipitation data. Are these observations drawn from national rain-gauge networks, satellite retrievals, or a merged product? If gauge data are used, how do the authors reconcile sparse station coverage—an endemic problem in West Africa—with the coarse 1° model grid? A station distribution map should be provided to demonstrate the spatial representativeness of the observational network across the study domain. If satellite-derived precipitation is employed, the product version must be stated, together with a brief account of its known biases in this region.
2.The authors do not clarify whether the historical record was partitioned into independent training and validation periods for the bias-correction experiments. For distribution-based methods such as quantile mapping, the risk of overfitting is non-negligible if the same data are used for both calibration and verification. I strongly recommend that the authors implement a rigorous cross-validation or hold-out validation scheme and report the performance of each correction method on data not seen during training.
3. In Figure 7, the y-axis is labeled “k-mean order,” yet the text refers to ten precipitation intensity classes. The mapping between these two descriptors is confusing. The axis label should be changed to something intuitive such as “Precipitation intensity class” or simply “Intensity level.”
4.Section 4 reads more like a recapitulation of results than a genuine scholarly discussion. The authors should contextualize their findings against existing NWP evaluation studies in the West African monsoon region and probe the physical reasons behind the inter-model performance gaps. For example, what specific advantages in convective parameterization, data assimilation strategy, or ensemble configuration explain ECMWF’s superior skill relative to the other models? A mechanistic discussion along these lines would materially enhance the academic value of the manuscript.
5.The abstract and introduction explicitly frame the ultimate goal as improving precipitation inputs for flood forecasting. Yet the body of the paper stops at precipitation verification and never engages with the hydrological response. While a fully coupled hydrological model may lie beyond the scope of this study, the authors should at least acknowledge this limitation candidly in the Discussion and sketch a credible technical roadmap from precipitation evaluation to operational streamflow forecasting. As currently structured, the narrative arc is incomplete; the study stops halfway.

The manuscript is not acceptable in its present form. Substantive improvements are required, above all in the validation design for bias correction and in the depth of the Discussion. I look forward to evaluating a revised version.

Author Response

We thank the reviewer for the thorough and critical assessment of our manuscript. We appreciate the recognition of the relevance of the topic, particularly in the context of precipitation forecasting and hydrological applications in West Africa. We also acknowledge the concerns raised, and have undertaken a substantial revision of the manuscript. We believe these revisions significantly improve the clarity, robustness, and interpretability of the study. We respectfully hope that the revised version addresses the main concerns raised and demonstrates the scientific value of the work for assessing precipitation forecast skill and bias correction in data-sparse regions such as the Ouémé basin.

Author Response File: Author Response.pdf

Round 2

Reviewer 1 Report (Previous Reviewer 1)

Comments and Suggestions for Authors

 

The manuscript presents a valuable and methodologically relevant contribution to multi-model precipitation forecast evaluation and bias correction for hydrological applications in West Africa. The authors have made noticeable improvements; however, several areas still require further strengthening to reach publication-quality standards.

Major revisions required:

  1. Abstract

    • Include quantitative performance improvements (e.g., percentage improvement in KGE and PBIAS).
    • Clearly state the methodological novelty.
  2. Introduction

    • Sharpen the research gap and avoid absolute claims; rephrase to reflect limited existing studies.
    • Improve coherence by merging overlapping paragraphs.
    • Explicitly justify:
      • The novelty of combining LHR, KGE, and PBIAS.
      • The scientific relevance of 1–7 day hindcasts.
  3. Study Area

    • Add hydrological characteristics (e.g., runoff response, lag behaviour).
    • Improve clarity and precision of wording.
  4. Methods

    • Ensure consistency in the study period versus the hindcast period.
    • Correct the KGE formulation.
    • Replace APBias with standard Percentage Bias (PBIAS) and correct its formulation.
    • Improve transparency and reproducibility of methodological steps.
  5. Results

    • Reduce descriptive repetition across lead times.
    • Provide stronger synthesis and comparative interpretation.
    • Include statistical support (e.g., confidence intervals or significance testing).
    • Justify the selection of the “Top 3” ranking approach.
  6. Discussion

    • Reduce redundancy with results.
    • Compare findings with African-region studies.
    • Add a clear operational implications subsection.
  7. Limitations

    • Explicitly discuss:
      • Coarse spatial resolution (~1°)
      • Gauge sparsity
      • Absence of hydrological model coupling
  8. Conclusions

    • Reduce repetition from the abstract.
    • Emphasise key quantitative findings.

Minor revisions:

  • Improve grammatical consistency.
  • Standardise terminology throughout.
  • Enhance clarity and self-containment of figures and captions.
  • Replace subjective statements with properly referenced arguments.

In general, the manuscript is promising but requires substantial refinement in clarity, methodological rigour, and synthesis before it can be accepted.

Find the attached comprehensive report for your reference.

 

Comments for author File: Comments.pdf

Comments on the Quality of English Language

 

The manuscript is generally understandable; however, the English requires moderate editing for clarity, grammatical consistency, and flow. Issues include subject–verb agreement, inconsistent terminology, and occasional awkward phrasing. Careful proofreading or professional language editing is recommended to improve readability and ensure the scientific message is communicated precisely.

 

Author Response

Please, see the attached file

Author Response File: Author Response.pdf

Reviewer 2 Report (New Reviewer)

Comments and Suggestions for Authors

The manuscript has been revised, and its scientific content and methodological design are now largely in line with the standards required for publication in Climate. The overall scientific framework is essentially sound. However, the quality of language and presentation remains a significant obstacle. In its current form, the quality of the English, the standardization of the figures and tables, and the completeness of the reference list are not sufficient for the manuscript to proceed to the copyediting and production phase. I recommend that the authors address the following points before resubmitting. The manuscript requires professional English language editing. The current writing contains numerous basic grammatical errors and inconsistencies in terminology, which seriously affect readability and the overall professional tone. Please carefully edit the entire manuscript for language. Upon resubmission, please provide either a certificate from a professional editing service or a tracked-changes version highlighting the language modifications. Regarding the figures and tables, there are multiple obvious labeling errors and incomplete captions. Please check and correct each of the following: Figures 2 & 3 (pp. 9-10): The captions are incomplete, ending with "based on the" and "over the Ouémé", respectively. Please complete the sentences. Figure 4 (p. 11): The color bar is labeled "Cor". To avoid ambiguity, I suggest changing this to "Pearson r" or providing the full name ("Correlation"). Table 4 (p. 12): The model name abbreviations are inconsistent (e.g., "Meteo" and "UK" vs. "Météo-France" and "UK Met Office" used elsewhere in the text). Please unify them to match the main text. Figures 5 & 6: The legend uses the term "Prediction" to refer to the raw, uncorrected forecasts. I recommend changing this to "Raw forecast" or "Uncorrected" to distinguish them more clearly from the bias-corrected results. Figure 7 (p. 15): The color bar on the right is labeled "ROC". However, the main text and the figure caption discuss the Likelihood Ratio (LHR). This appears to be a clear copy-paste error and must be corrected. Section 3.4 places strong emphasis on the positive effects of Empirical Quantile Mapping (EQM), but the discussion of cases where the correction effect is limited or even negative (e.g., for the ECMWF forecasts at lead times TL6–TL7 in sub-basins such as Bonou and Zangnanado, where the raw forecast occasionally outperforms the corrected one) is insufficient. Please expand the discussion to include possible reasons for these exceptions (e.g., non-stationarity at longer lead times, over-correction) to make the conclusions more robust. Problems with the references: Reference 1 and Reference 6 are completely identical (both Collischonn et al., 2005). Please remove the duplicate and renumber the reference list accordingly. Reference 23 (Willett et al.) lists the publication year as 2026. Please verify whether this paper has been formally published and, if so, provide the complete volume, issue, and page information. If it is a preprint, please update the citation format accordingly. Reference 31 (Knoben et al.) has a truncated title in the PDF. Please provide the full title.

Comments on the Quality of English Language

The English prose remains below the standard expected for an international journal

Author Response

Please, see the attached file

Author Response File: Author Response.pdf

This manuscript is a resubmission of an earlier submission. The following is a list of the peer review reports and author responses from that submission.


Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

Thank you for the opportunity to review this manuscript titled “Evaluation and optimisation of precipitation predictive skills of six numerical weather prediction models over the Ouémé basin.” The study is timely, well-motivated, and addresses an operationally significant challenge in monsoon-influenced West African basins, using a combination of continuous and categorical verification metrics and bias-correction techniques. The manuscript is scientifically valuable; however, several revisions are required to improve clarity, structure, coherence, and completeness.

  1. Abstract
  • The abstract is overly extended and should be reduced to 150–300 words.
  • Summarise all key elements concisely: aim, dataset, methods, key quantitative results, and implications.
  • Reduce repeated statements on model performance.
  • Present key findings with numerical values.
  • Add explicit practical implications (e.g., relevance for flood early‑warning systems).
  • Improve readability by breaking long, multi-clause sentences.
  1. Keywords
  • Restrict to 3–6 keywords.
  • Arrange alphabetically.
  • Remove conjunctions.
  1. Introduction
  • Improve overall flow and coherence.
  • Correct opening sentence: “several days ahead”.
  • Introduce structure using a top‑down approach: global → regional → basin‑scale forecasting.
  • Avoid repeated use of “However” in consecutive sentences.
  • Incorporate more recent studies, including those relevant to West Africa.
  • Provide a clearer statement of novelty at the end of the section.
  • Integrate relevant past studies and emphasise identified research gaps.
  1. Materials and Methods

Study Area

  • Ensure consistent spelling of place names (e.g., Zagnanado).
  • Streamline climatic zone descriptions.
  • Justify climate periods used (1950–1969 and 1970–2004).
  • State clearly the number of stations used in this study.
  • Remap Figure 1 to include country-level and Africa inserts.

Data

  • Justify the data period (1985–2015). Clarify apparent conflict with references to 1985–2024.
  • Explain briefly the rationale for Thiessen polygons over alternatives like kriging.
  • Provide the data source links for all hindcast datasets.
  • Discuss implications of resampling from 1° to 0.25°, especially uncertainty.
  • Use “nearest‑neighbour interpolation” instead of “nearest method”.
  • Use “sub‑basin” consistently.

Hindcast Products

  • Reduce repetitive descriptions of NWP systems.
  • Include a comparative table summarising main model characteristics—resolution, assimilation, ensemble size, strengths, and dataset links.
  • Add a concluding sentence on how model differences may affect precipitation predictability.

Performance Evaluation

  • Present evaluation metrics in a summary table, followed by concise descriptions.
  • Ensure formulas are consistently formatted.
  • Simplify explanation of the likelihood ratio (LHR).
  • Introduce Table 1 earlier to improve flow.

Bias Correction Methods

  • Justify the choice of quantile mapping (QM) compared with other methods.
  • Specify the software or packages used (e.g., R‑package qmap).
  • Justify inclusion of polynomial regression, given known limitations for rainfall.
  1. Results

Raw Model Performance

  • Provide more scientific, quantitative explanations of findings.
  • Include a sentence summarising the general decline in skill with lead time.
  • Ensure figure captions fully describe symbols, colours, and station names.
  • Resolve inconsistencies in basin names.
  • Where appropriate, connect findings to known monsoon predictability limits.

APBias Analysis

  • Add a supporting summary table.
  • Soften categorical phrases, using objective language such as “persistent high bias suggests systematic error”.
  • Include quantitative statements (e.g., “bias exceeding 40% at TL6–TL7”).

Basin Area Effects

  • Specify the significance test used (Pearson or Spearman, α‑level).
  • Provide brief reasoning for observed negative correlation in CMCC.
  • Increase figure font sizes.

Model Ranking

  • Explain clearly how percentage contributions were calculated.
  • Briefly justify why ECMWF outperforms others at most lead times except TL1.

Bias-Correction Results

  • Consolidate common observations across basins to reduce repetition.
  • Explain why over-adjustment occurs at longer lead times.
  • Use labelled figure panels (a–f) for clarity.
  1. Discussion
  • Streamline to avoid repeating earlier results.
  • Add comparisons with other West African basins where available.
  • Discuss operational implications for climate‑service providers in Benin.
  • Incorporate more recent literature.
  • Mention hydrological challenges such as data scarcity and convective rainfall dominance.
  1. Conclusions
  • Rewrite as a concise single paragraph.
  • Include clear statements of key findings and their implications.
  • Add explicit limitations (e.g., sparse observations, grid‑basin mismatch).
  • Strengthen recommendations (e.g., “ECMWF + EQM performs best for TL1–4”).
  • Provide suggestions for hydrological forecasting improvements and uncertainty estimation.
  1. Minor and Editorial Issues
  • Standardise units (e.g., km², °C).
  • Ensure consistent naming of rainfall stations.
  • Maintain consistent English style (American vs British).
  • Correct punctuation and ensure figures appear in numerical order.
  • Limit excessive parenthetical citations.

In general, the manuscript has strong scientific merit and clear potential for publication after the authors make major but achievable revisions to improve clarity, structure, methodological justification, and presentation quality.

 

NOTE: The authors should check the attached comprehensive report.

Comments for author File: Comments.pdf

Comments on the Quality of English Language

The English can be improved, and the authors should be consistent with the use of English (American or British).

Author Response

"Please see the attachment."

Author Response File: Author Response.pdf

Reviewer 2 Report

Comments and Suggestions for Authors

This study focuses on West African river basins, evaluating and optimizing the precipitation forecasting capabilities of six numerical weather prediction (NWP) models for the Ouémé River basin in Benin. Using both continuous and event-based validation metrics—including the Kling–Gupta Efficiency (KGE), absolute percentage bias, and likelihood ratio—this study aims to reveal lead-time dependence, basin-scale effects, and the added value of statistical bias correction. Although this research holds significant regional value, to enhance the overall quality of the paper, improvements are recommended in the following areas:

 

Major comments

  1. Please specify the details for the six datasets:

(1) Product name, version number, spatial resolution, temporal resolution, reporting frequency;

(2) Whether ensemble forecasting is used (ensemble size), and whether ensemble mean/median/member is employed;

(3) Definition of lead time (TL1: cumulative +24h? Or cumulative on Day 1?).

  1. Observation data: Number of rainfall stations, station distribution, missing data handling, and quality control methods must be specified. Thiessen polygon averaging for sub-basin totals is acceptable, but the number of stations and representativeness must be reported.
  2. The manuscript states “ROC,” but the actual method used is LHR = POD/FAR. Please standardize terminology and methodology. If reporting ROC, include AUC or the ROC curve; if insisting on LHR, provide definitions and formulas for POD and FAR.
  3. Likelihood Ratio (LHR): “10 times greater is good” is a subjective rule of thumb. It is recommended to provide an appropriate explanation; otherwise, the results may be unreliable.

Minor comments

  1. Lines 43: “Flood forecasting at the basin scale requires precipitation forecasts several days”. Syntax error, should be revised to “Flood forecasting at the basin scale requires precipitation forecasts with lead times of several days”.
  2. Lines 44: “are mostly provided” should be revised to “is commonly provided”.
  3. Lines 97-98: References must be formatted consistently (using a numbered system).
  4. Lines 105-107: The description is overly long and confusing; it is recommended to condense it and present the key information.
  5. Lines 165: The manuscript actually uses APBias and LHR, not Pbias and ROC; consistency is required.
  6. Lines 198: Modify the table content format to [0, 10), [10, 15), [15, 25), [25, ∞).
  7. Lines 259: It should be Figure 2.

Author Response

"Please see the attachment."

Author Response File: Author Response.pdf

Reviewer 3 Report

Comments and Suggestions for Authors

This manuscript evaluates the performance of six numerical weather prediction (NWP) precipitation products over the Ouémé basin in Benin and explores the potential of several statistical bias-correction techniques to improve rainfall forecasts for hydrological applications. The topic is relevant, as reliable precipitation forecasts are important for flood forecasting in West Africa, and relatively few studies have examined forecast skill in this region.

However, in its current form the manuscript suffers from several significant methodological and conceptual issues that limit the robustness and interpretability of the results. In particular, the description and selection of the forecast dataset, the spatial processing of precipitation fields, and the verification framework raise concerns about the validity of the conclusions.

The manuscript currently reads more as a collection of methodological tools than as a study guided by a clearly articulated scientific question. The motivation for selecting specific verification metrics and bias-correction approaches is not sufficiently explained, and several methodological choices are insufficiently justified.

For these reasons, I do not believe the manuscript is suitable for publication in its present form and therefore recommend rejection. However, the general research question—evaluating precipitation forecast skill and post-processing approaches for hydrological applications in the Ouémé basin—remains potentially valuable. If the authors substantially revise the methodological framework, the study could form the basis of a stronger manuscript in the future.

Additional Comments:

Line 43: Do you mean “several days in advance”?

Lines 49-50: Reliable precipitation forecasts are necessary, but not sufficient, for reliable flood forecasts.

Lines 53-57: Please provide a reference for each model.

Line 61: “assessment performance evaluation” -> this sentence does not make sense.

Lines 74-76: Please provide references to support this point.

Line 92: Please explain what “unimodal” and “bimodal” mean.

Lines 94-98: The sentence appears to mix spatial scales. It first describes climatic conditions at the basin scale, but then reports rainfall values measured at a single climatic station (near 9°N). This may create ambiguity regarding whether the values represent basin-wide averages or station observations. The authors may wish to clarify this distinction and explicitly state that the rainfall data refer to station measurements.

Lines 103-113:

The study evaluates precipitation skill, yet gauge observations are first aggregated to sub-basin means using Thiessen polygons. Given the highly heterogeneous and convective rainfall regime of the region, as well as the uneven and sparse gauge distribution shown in Fig. 1, this approach introduces substantial and spatially variable representativeness error.

For a study focused on precipitation (not discharge), verification should be performed directly at the gauge scale (e.g., model value interpolated to station locations). Aggregating sparse gauge data into basin averages using Thiessen polygons effectively imposes an artificial spatial structure on the observations and may strongly influence model ranking.

As currently implemented, the evaluation framework does not convincingly isolate model skill from artefacts of the spatialisation procedure. This issue substantially limits the robustness of the conclusions.

Lines 114-120:

The description of the precipitation dataset and processing workflow lacks clarity and raises several methodological concerns.

First, the manuscript refers to “seasonal forecast” data downloaded from the Copernicus Climate Data Store, including hindcasts (1993–2016) and forecasts (2017–present). However, the specific forecasting system used (e.g., ECMWF SEAS5 or another product) is not clearly identified. Since the Copernicus platform hosts multiple datasets (e.g., ERA5 reanalysis, seasonal forecast systems, climate projections), the authors should explicitly state which product was used and provide its main characteristics.

Second, the distinction between hindcasts and forecasts requires clarification. Hindcasts are retrospective simulations used for skill assessment and are typically produced using a fixed model version over a defined period, whereas forecasts are operational real-time runs that may include system updates. It is unclear whether the hindcasts and forecasts were analyzed jointly, whether they were treated separately, or whether hindcasts were used for calibration and forecasts for validation. Mixing these datasets without explicitly accounting for potential model version differences could affect the robustness of the evaluation.

Third, the spatial processing strategy raises significant concerns. The manuscript states that 1° resolution data were resampled to 0.25° using a nearest-neighbor method prior to basin extraction. Nearest-neighbor resampling does not increase the effective spatial resolution of the data; it merely subdivides each 1° grid cell into smaller pixels with identical values. This artificial refinement does not introduce additional physical information, does not improve representation of subgrid variability, and may create a misleading impression of finer-scale data. In other words, this procedure constitutes a form of artificial downscaling rather than true resolution enhancement.

Furthermore, the spatial representativeness of the precipitation forcing should be carefully examined. A 1° grid cell corresponds to roughly 10,000 km² at the equator. Given that some of the analyzed sub-basins are as small as 7,035 km², certain basins may be smaller than a single native grid cell. This creates a clear scale mismatch between the forcing resolution and the hydrological units. Simply subdividing the coarse grid into smaller pixels does not resolve this mismatch and does not improve basin-scale representativeness. The rationale for resampling prior to extraction is therefore unclear.

Finally, there appears to be a conceptual inconsistency in using a seasonal forecast product while evaluating short lead times (1–7 days). Seasonal forecast systems are primarily designed for longer lead times (monthly to seasonal scales), and their performance characteristics differ from those of medium-range forecast systems optimized for short-range prediction. The suitability of a seasonal forecast product for daily lead times of 1–7 days should therefore be clearly justified.

Overall, the data source, spatial processing methodology, and evaluation framework require clearer explanation and stronger methodological justification to ensure reproducibility and scientific robustness.

Lines 114-120: Why isn’t this paragraph in Section 2.1.3?

Section 2.1.3: please provide references for each of the NWP models.

Section 2.1.3: This section only aims at presenting the NWP models. And it takes almost a entire page, which is disproportionately long compared to the length of the manuscript. Many of the details on these models provided here are irrelevant to the main discussion of the paper, and could in fact much better be summarized on a table for example.

Line 177: “Mean flow benchmark” -> Isn’t the manuscript dealing with precipitation?

Lines 184-191: The description of the LHR computation is somewhat unclear. It is not entirely evident whether “detection” refers to (i) exceedance of each threshold (i.e., binary event defined as precipitation > T), or (ii) correct classification within the k-means-derived precipitation intervals.

The procedure used to define the ten precipitation intervals for the computation of the likelihood ratio (LHR) is not sufficiently described. The manuscript states that a one-dimensional k-means clustering (k = 10) was applied to precipitation observations, but it is unclear which dataset was used (all stations combined, each basin separately, or basin-averaged values) and how the resulting clusters were converted into precipitation thresholds. Since k-means returns cluster centers rather than interval boundaries, additional clarification is required on how the class limits were defined. As currently written, the methodology is not reproducible.

Line 195: “mean flow” -> isn’t the paper on precipitation?  

Table 1: The formatting of the precipitation intervals should be revised. The lower bound should precede the upper bound, and interval notation should use a comma (e.g., ) rather than a dash.

Lines 257-264: The verification framework described in Section 2.1.2 appears inconsistent with the presentation of the results. The methods state that gauge observations were spatialised to sub-basin averages using Thiessen polygons, suggesting that the evaluation is performed at the basin scale. However, the results section refers to model performance at individual “stations” (e.g., Zagnanado, Kaboua, Bétérou). It is therefore unclear whether the verification was conducted at the station scale or at the basin scale. This distinction is important because the use of Thiessen polygons implies basin-averaged precipitation, whereas station-based verification would require interpolation of model fields to gauge locations. The authors should clarify the verification framework and ensure consistency between the methods and results sections.

Lines 265-268: The interpretation of the results in this paragraph appears overstated. The manuscript attributes the higher KGE values of the ECMWF forecasts to the “robustness of ECMWF’s data assimilation system and model physics.” However, the analysis presented in the study does not examine model physics, data assimilation systems, or forecast system design. The results simply indicate that ECMWF exhibits higher KGE values in this particular comparison. Attributing the observed performance differences to specific components of the forecast system is therefore speculative and should either be supported by additional evidence or toned down.

Figures 2, 3, 4, 5, 6, 7 & Table 2: It would make things much clearer to explicitly specify the lead times rather than writing TL1, TL2. Etc.

Section 3.2: The analysis presented in Section 3.2 is not statistically meaningful. The authors attempt to infer a relationship between basin area and forecast skill using only six sub-basins. Performing correlation analysis and claiming statistical significance with n = 6 observations is fundamentally unreliable. With such an extremely small sample size, correlation estimates are highly unstable and significance tests have essentially no inferential value. A single data point could completely change the result.

Despite this limitation, the manuscript reports statistically significant relationships (α = 10%) and proceeds to interpret them physically. These claims are not supported by the data. With only six basins, the analysis cannot provide any robust evidence of a scale dependency in model performance.

Table 2: It is unclear what this table represents. The table header is too brief and does not allow the reader to understand the meaning of the values presented. In addition, the caption does not explain why some cells are highlighted in yellow.

Author Response

"Please see the attachment."

Author Response File: Author Response.pdf

Reviewer 4 Report

Comments and Suggestions for Authors

see attachment

Comments for author File: Comments.pdf

Author Response

"Please see the attachment."

Author Response File: Author Response.pdf

Back to TopTop