A Case Study on the Stability of Neural Network Climate Prediction Models with Different Training Stop Criteria
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsPlease read the attached file.
Comments for author File:
Comments.pdf
The English is generally clear and readable.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThe manuscript addresses an interesting and relevant topic and is potentially suitable for publication. The investigation of randomness and its impact on neural network stability is clearly presented and methodologically sound. However, in its current form, the paper appears to be strongly focused on algorithmic and methodological aspects, with limited connection to the underlying physical and climatic context. For a journal like Atmosphere, a deeper discussion of the physical implications and relevance of the results would be necessary. In addition, the manuscript would benefit from a more comprehensive engagement with the existing literature in this topic, particularly studies that address similar questions within the considered climate and geographical area.
Overall, the work currently reads more as a computational exercise than as an analysis of the potential and implications of these approaches in a climate science context. Strengthening the link to physical processes and situating the results within the broader climate literature would significantly improve the manuscript. In particular, pay attention to Introduction and Conclusions.
In addition, there are several specific comments to be addressed:
L 35-36 Clarify better the concept and improve the English.
L 52-53 The authors do not specify how the randomness has impact (performance? feature selection? convergence?).
L 69 Please specify the period, in terms of the years considered.
L 72-74 Please check the English style.
L 85-86 Please check the English style.
L 88-90 If I well understand, the model cannot accept an input number larger than 30. As the PC are 30, why did the author not consider all, but only 6?
L 96 Are you sure that the gradient is computed using the backpropagation algorithm?
L 107-108 This procedure is quite common, it is not necessary to specify it.
L 141 rephrase as: “the default settings are used”.
L 158-159 Please check the English style.
L 179-180 The concept of standard deviation is well known, this sentence can be removed.
L 184 rephrase as: "a high Sta indicates poor stability".
L 198 change “not being particularly small” with “which is not particularly small”.
L 202 rephrase as: “the model performances do not improve in a relevant manner when…”
L 213 add “either” before “positive”.
L 216 rephrase as “of a factor 0.1 with respect to the original value”.
L 218 rephrase as “will be” with “is”
L 252 change “runs” with “was run”
L 398-400 rephrase as: “Compared to a single-model approach, an ensemble-model strategy is advisable, as it yields greater stability and superior predictive skill”.
L 403-405 There is something wrong in the English style.
L 405-407 rephrase as: among the three stopping criteria, Validation must be avoided because….
L 417 change “it introduces” with “introduced”.
L 429 Which section are you talking about?
L 431-443 this paragraph basically repeats concepts already explained. There is no need for this repetition, the paragraph can be removed.
Comments on the Quality of English Languagethere are several sentences that need improvement. An English revision is needed.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for AuthorsThe manuscript under review is devoted to the study of climate models, developed by neural networks and trained with different stop criteria. The authors investigate the prognosis of July precipitations in Yangtze River (middle and lower reaches) on the data of early winter sea surface temperature as a case study. They useed neural network model with three sources of randomness. The authors got the data for 1951 to 2021 from the National Climate Center of the China Meteorological Administration. They describe their model in rather detailed way. The results of realisation of different model configuration are presented as graphics and are described and discussed.
The results prove rather efficiently the main conclusion of the authors: “improvements in stability are achieved while predictive skill is also enhanced. When using multi-model ensemble techniques, the weaker stability of individual models is probably beneficial to achieving better predictive skill in the ensemble.”
This paper, of course, will be of interest for the specialists in this field.
The only remark, the reviewer can do is - it is better to remove the abbreviation from the abstract (l. 12). There is no need in it at all, and the abstract must be free of abbreviations.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsDear Authors,
Thank you for your thorough and careful response to the initial review. The revised manuscript represents a genuine and substantial improvement over the original submission, and the effort invested in addressing each comment is evident. Below I provide a final assessment of the revisions.
The large majority of comments have been fully and satisfactorily addressed. The addition of the ENSO index experiment meaningfully strengthens the generalizability claims, and the revised title ("A Case Study on...") correctly reflects the study's scope. The removal of the Dropout discussion eliminates an unsupported claim, and the revised framing of the ensemble instability observation — now explicitly linked to ensemble diversity theory via Hansen and Salamon (1990) and Krogh and Vedelsby (1995) — is appropriately cautious. The restructuring of the original Figure 2 into two separate figures is a clear improvement in readability. The language revisions, the explicit Weight Decay distinction, the added sentence on the Loss stopping criterion insight, and the trade-off qualifier in the Conclusions are all well-executed.
Two points require attention before the manuscript can be considered complete.
First, the sensitivity analysis justifying the 0.1 scaling factor (Table R1) and the validation of the Sta metric over 30 runs (Table R2) are presented only in the response letter and do not appear in the manuscript. Since the Initialization0.1 technique is the paper's principal methodological contribution, readers of the published article will not have access to the reasoning that justifies its key parameter. Please incorporate a concise summary of these findings into the manuscript body. For the scaling factor, this could be a brief paragraph in Section 2.2 noting the sensitivity sweep and explaining why values below 0.1 produce diminishing returns. For the Sta validation, a sentence or two in Section 2.3 would suffice.
Second, the LOO computational cost comment was addressed implicitly through the revised Introduction, but the actual scale of the experiments was not stated. Please add a sentence in Section 2.3 noting the number of training runs entailed by the LOO strategy (e.g., 71 × 30 = 2,130 per experiment), so that readers working with larger models are aware of the scaling implications.
These are minor additions that do not require new experiments or structural changes. Once incorporated, I would consider the manuscript ready for publication.
I thank the authors for their engagement with the review process and congratulate them on a well-revised manuscript.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have properly addressed my previous comments, and the manuscript is almost ready for publication. I appreciate they strengthened the link to climatic context in the Introduction. However, two sentences are somewhat convoluted and require refinement. If I have correctly understood the authors’ intentions, these sentences can be reformulated as follows:
L 42-48:
Addressing the challenge of limited number of climate event samples, two main approaches are generally adopted. The first approach relies on numerical climate models: if these models can generate simulations that closely match observations, simulated data can be used to expand the training set, making it feasible to employ relatively complex machine learning models [11–14]. The other approach uses only observational samples; in this case, simpler machine learning models, typically containing only a few dozen parameters, are preferred, such as the widely applied artificial neural networks.
L 56-72:
In recent years, a few studies have highlighted the non-negligible impact of randomness on model performance in such contexts [21,22]. Beyond the challenge of limited sample size, climate-event datasets are also characterized by low information quality. Specifically, the predictors used as model inputs influence the target variables only to a limited extent, implying a weak statistical relationship between them. As an example, consider the prediction of July precipitation over the middle and lower reaches of the Yangtze River (MLYR) using early-winter sea surface temperature (SST) anomalies as predictors. Although El Niño–like SST anomalies in the preceding winter increases the likelihood of above-normal July precipitation, there are also years in which precipitation is below normal (e.g., 1966 and 2019). This behavior differs from that observed in many other machine learning applications. In handwritten digit recognition, for instance, images of the digit “1” are consistently associated with the target label “1”. Climate-event prediction, by contrast, is intrinsically probabilistic and affected by substantial internal variability. Motivated by the combined challenges of small sample sizes and low information quality, this study systematically investigates the sources of instability in NN models, quantifies their relative contributions, and proposes targeted mitigation methods, thereby supporting the construction of more robust NN-based climate prediction models.
Please check carefully!
Comments on the Quality of English LanguageMinor corrections are needed.
Author Response
Please see the attachment.
Author Response File:
Author Response.pdf

