Review Reports
- Yui Watanabe *,
- Takuya Tomoda and
- Takeshi Nagata
- et al.
Reviewer 1: Anonymous Reviewer 2: Anonymous Reviewer 3: Anonymous
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsMy main comment on this overall interesting study is related to the considerably larger proportion of lung cancers (42.8% vs 15.2% for the breast cancers as the second most frequent cancer) and of high risk cases compared with other cancers and low risk cases, respectively. Can you reasonably assume that such a skewed data distribution did not result into a biased assessment of the overall performance of the DeepSurv model under investigation?
Further comments:
1) Abstract
1a) Its various subsections should be clearly identified with appropriate headers (e.g. Introduction / Background, Materials and Methods, Results, Conclusion).
1b) The findings mentioned at lines 36-42 should be reported in quantitative terms along with the results of a formal statistical analysis, including p-values for statistical significance. The same applies to the corresponding parts of the Results section.
2) Introduction. OK.
3) Materials and methods
3a) Lines 99-100. Please briefly summarize the CT image acquisition protocol used to generate CT simulation images.
3b) Lines 118-119. Please provide a brief explanation of the main features DeepSurv model, including any relevant literature references or supplementary material if needed.
4) Results. See comment 1b above.
5) Discussion. Generally ok. At lines 313-315, please briefly comment on whether and to what extent the first two limitations (retrospective and single center study design, no external validation) could actually limit the validity and generalizability of the study findings.
Author Response
Please see an attached file (Point_by_point_response.pdf).
Author Response File:
Author Response.pdf
Reviewer 2 Report
Comments and Suggestions for AuthorsDear authors:
This is a timely and clinically relevant study. Survival estimation can significantly inform decisions about the intensity and duration of palliative radiotherapy, and the use of SHAP and SurvLIME is an interesting and positive feature. However, several methodological aspects require clarification before the model, and in particular the apparent importance of radiation dose, can be interpreted with confidence.
Please reconsider excluding the 96 patients who were still alive but had less than 1 year of follow-up. A survival model can retain these patients as censored observations; their exclusion could introduce selection bias. A sensitivity analysis that includes all eligible patients would strengthen the study.
The radiation dose variable requires careful review. For patients who did not complete treatment, the dose was adjusted based on the number of fractions received. Because survival was measured from the start of treatment, this introduces information not available at baseline and may introduce time-related bias. The primary model should preferably use the planned dose; using the administered dose would require a landmark or time-dependent analysis.
The association between higher dose and better survival should not be taken as evidence of a treatment benefit. It is more likely that patients with a better expected prognosis receive longer or higher-dose regimens. SHAP and SurvLIME explain how the model uses a variable but do not establish causality. Therefore, claims suggesting a survival benefit from increasing the dose should be qualified.
Further methodological details are needed to ensure reproducibility. Please describe all predictor variables, how missing data are handled, variable coding, preprocessing, neural network architecture, hyperparameter selection, and the calculation of absolute survival probabilities. The data availability statement should also explain how to access the data and code. Given a cohort of 376 patients, it is important to demonstrate whether deep learning adds value compared with simpler approaches. Please compare DeepSurv with a standard or penalized Cox model and, if possible, with an established prognostic tool such as METSSS, TEACHH, Katagiri, or BMETS.
The study uses repeated internal cross-validation rather than external validation. Because the same patients appear in multiple partitions, the 50 folds are not independent, and the reported confidence intervals might be overly narrow. Please explain how these intervals were calculated and consider performing patient-level bootstrapping of the entire modeling process. Additionally, report the number of deaths, median follow-up, median overall survival, and survival estimates at relevant time points. Predictions at 3 and 6 months could be particularly useful for choosing between single-fraction and multi-fraction treatment.
Figure 1B shows Brier scores over time, not calibration. Please add a calibration plot of observed versus predicted values at one year, ideally including the intercept and calibration slope. The estimators used for the C-index, AUC, and IBS must also be specified, along with the method for handling censoring.
The explainability analysis was conducted on the internal validation sets used for early stopping, rather than on the held-out test sets. Please justify this choice and explain how the 25 patients per fold were selected. Repeated inclusion of the same patients could affect the apparent stability of the SHAP and SurvLIME results.
Please clarify whether the unit of analysis was the patient, the lesion, or the radiotherapy cycle. The number of treatment sites exceeds the cohort size, indicating that some patients received irradiation at multiple sites. Please explain how the dose and treatment site variables were assigned in these cases.
The Discussion section would benefit from more cautious wording. Dose was among the most relevant predictors, but its SHAP-based importance was much lower than that of performance status; therefore, the results do not support the claim that its importance...
Author Response
Please see an attached file (Point_by_point_response.pdf).
Author Response File:
Author Response.pdf
Reviewer 3 Report
Comments and Suggestions for AuthorsThe optimal radiotherapy schedule for bone metastases is a matter of debate. In clinical practice the determination of RT regimen requires whole person assessment including prognosis, risks to normal tissues, quality of life and patient goals. Often, a clinician's survival estimation is too optimistic. Prognostic scores can help clinicians tailor RT indications to avoid over- or undertreatment. In this context, there is interest in AI-based models capable of predicting OS in patients receiving palliative RT for bone metastases. Several randomized trials produced similar outcomes between long-course radiation therapy (30 Gy in 10 fractions) and
short-course radiation therapy (8 Gy in 1 fraction or 20 Gy in 5 fractions) (van der Liden et al Radiother Oncol 2006;78:245-253). However, the analysis reported in this article indicated that besides clinical factor, prescribed dose significantly contributed to OS in this palliative context.
Since this analysis in based on retrospective data, could there be a selection bias related to the PS and/or other clinical factors of the patient due to which patients with better PS received a higher dose? If not, how does this model avoid the bias?
I suggest to add more details on this topic in the text.
Furthermore, the data presented here refer to 3D CRT treatments (lines 106-109). Have you considered analyzing data relating to stereotactic treatments, which are becoming increasingly widespread in the context of patients with a limited number of secondary bone localizations?
Author Response
Please see an attached file (Point_by_point_response.pdf).
Author Response File:
Author Response.pdf
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThank you for your reply.
Reviewer 3 Report
Comments and Suggestions for AuthorsThank the authors for the changes made to the statistical analysis and the text. I have no further comments.