Next Article in Journal
A Hybrid Multi-Scale Phase-Correlation Framework for Subpixel Registration of Multi-Temporal Very-High-Resolution Remote Sensing Images
Previous Article in Journal
Multi-Sensor Fusion SLAM Based on LiDAR, IMU and GPS for Structured Urban Scenes
Previous Article in Special Issue
RiTex: Harmonization of Radiomic Features Based on Riemannian Geometry
 
 
Article
Peer-Review Record

A Retrospective Study on Predicting Ki-67 Expression in Esophageal Cancer Patients Based on Delta Radiomics

J. Imaging 2026, 12(8), 346; https://doi.org/10.3390/jimaging12080346
by Taiwei Sun 1,†, Lei Xue 2,†, Tingting Li 1,†, Shisuo Du 1, Anning Cao 1, Yang Shen 1, Bei Lv 1, Weixing Ji 1,* and Ze Wang 2,*
Reviewer 1: Anonymous
Reviewer 2: Anonymous
Reviewer 3: Anonymous
J. Imaging 2026, 12(8), 346; https://doi.org/10.3390/jimaging12080346
Submission received: 8 June 2026 / Revised: 29 July 2026 / Accepted: 29 July 2026 / Published: 31 July 2026
(This article belongs to the Special Issue Medical Image Analysis: New Opportunities and Challenges)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors   Summary: To address the limitations of invasive biopsy for assessing Ki-67 expression in esophageal cancer, authors developed a non-invasive delta-radiomics approach using preoperative CT images to account for baseline tissue variations. This logistic regression model achieved a modest AUC of 0.748. While this method does demonstrate potential as a tool for pre-therapeutic assessment, several methodological aspects need to be addressed in order for the manuscript to be ready for publication.

This is potentially novel and clinically interesting study. Although I am not sure if the study is ready for publication.

I have some methodological concerns.

This study has a small sample size, which makes it inherently unstable.

There is a potential for bias due to specifics of feature selection. It is unclear whether it was performed independently inside each LOOCV iteration.

Was any external validation used?

Choice of Ki-67 threshold appears to be arbitrary, actually, it is close to the sample median, which makes it rather convenient. Is this correct?

Delta-radiomics normalization formula is also insufficiently explained, were alternatives considered?

Performance of the methodlogy is rather modest. Is it ready for publication?

 

 

Comments on the Quality of English Language

you need to change "delta-Radiomics approach" to "delta-radiomics approach" consistently. 

There are some minor grammatical errors as well, use any AI tool.

Author Response

Reviewer 1.1

"This study has a small sample size, which makes it inherently unstable."

Response:

We thank the reviewer for raising this important point. We fully agree that the cohort of 59 patients (21 Ki-67-high, 38 Ki-67-low) is modest and that this limits the stability of any single estimate. We have now explicitly acknowledged this as a primary limitation throughout the revised manuscript. To mitigate the instability concern, we (a) adopted nested leave-one-out cross-validation (LOOCV) with mRMR feature selection (k=3) performed strictly inside each training fold, eliminating the optimistic bias that data leakage introduces; (b) quantified the optimism of our performance estimate using Bootstrap validation (2000 resamples), which yielded an optimism of only 0.0001 and a optimism-corrected AUC of 0.6427; and (c) reported an events-per-variable (EPV) analysis. At the optimal threshold of 0.63 the model operates at EPV = 7.0 (21 events / 3 retained features). We have also added a multi-threshold sensitivity analysis (thresholds 0.33–0.73). We now position the study explicitly as exploratory and hypothesis-generating rather than definitive. We agree that a larger, multi-centre cohort is required before clinical translation and have stated this clearly as a future direction.

Reviewer 1.2

"There is a potential for bias due to specifics of feature selection. It is unclear whether it was performed independently inside each LOOCV iteration."

Response:

This is a crucial methodological concern, and we appreciate the reviewer for highlighting it. In the original submission, feature selection and normalization were indeed performed on the entire dataset prior to cross-validation, which constituted data leakage and inflated the reported AUC to 0.748. In the revised manuscript we have completely restructured the pipeline: we now implement nested LOOCV in which every step that uses outcome information — min–max normalization, mRMR feature selection (k=3), and model training — is executed exclusively on the training fold of each iteration. The held-out test patient is never seen during feature selection or training. This correction reduced the apparent AUC from 0.748 to a leakage-free 0.643, a difference that itself confirms the presence and removal of leakage. We have added a dedicated Methods subsection describing the exact preprocessing order, so that the independence of feature selection from the test fold is unambiguous.

Reviewer 1.3

"Was any external validation used?"

Response:

We thank the reviewer for this question. We did not have access to an independent external cohort at the time of this study, and we wish to be transparent that this remains a genuine limitation of the work. Instead, we strengthened internal validation as far as the data permitted: (1) nested LOOCV provides an essentially unbiased estimate of generalization within this sample; (2) Bootstrap validation (2000 resamples) was used to correct for optimism, yielding an optimism of 0.0001 and a corrected AUC of 0.6427; and (3) we performed a multi-threshold sensitivity analysis and reported confidence intervals around performance metrics. We have added an explicit statement that the absence of an external validation cohort limits the generalizability of the findings and that prospective, multi-centre external validation is the essential next step. We have positioned the manuscript as hypothesis-generating precisely because external validation is lacking, and we hope to pursue such validation in collaboration with partner institutions.

Reviewer 1.4

"Choice of Ki-67 threshold appears to be arbitrary, actually, it is close to the sample median, which makes it rather convenient. Is this correct?"

Response:

We appreciate the reviewer's scrutiny. The Ki-67 classification threshold of 0.63 was indeed chosen with reference to the observed sample distribution (sample median ≈ 0.6, with several patients at exactly 0.6), and the reviewer is correct that it is close to the median. We now disclose this transparently rather than presenting it as a pre-specified clinical cut-off. To address the concern that this could be convenient or data-driven, we have added (a) an events-per-variable calculation at this threshold (EPV = 7.0) confirming adequate event density, and (b) a comprehensive multi-threshold sensitivity analysis spanning 0.33 to 0.73.

Reviewer 1.5

"delta-radiomics normalization formula is also insufficiently explained, were alternatives considered?"

Response:

Thank you for this comment. In the original manuscript the delta-radiomics feature was defined as (tumor − normal) / (tumor + 0.001), which places the tumor feature in the denominator and can become unstable when tumor values approach zero. Following the suggestion of reviewers, we have replaced the formula with the simple difference delta-radiomics = esotarget − eso (tumor minus normal). This definition is more interpretable, avoids the near-zero denominator instability, and aligns with the broader delta-radiomics literature. We have added a clear rationale for this choice in the Methods and now report a three-tier feature-type comparison: delta-radiomics (AUC = 0.643) outperformed Esotarget alone (0.546) and Eso alone (0.514), confirming that the pairwise subtraction — rather than either region in isolation — carries the discriminative signal. The revision of the formula and the justification are described explicitly in the revised Methods section.

Reviewer 1.6

"Performance of the methodology is rather modest. Is it ready for publication?"

Response:

We agree with the reviewer that an AUC of 0.643 is modest and does not support immediate clinical deployment. We have therefore substantially tempered the language of the manuscript: the abstract, discussion, and conclusion now consistently frame the work as an exploratory, hypothesis-generating study rather than a validated diagnostic tool. We note, however, that the corrected performance (0.643) is non-trivial in context: DeLong testing showed no significant difference between the Random Forest model and logistic regression (p = 0.544), and the delta-radiomics approach significantly outperformed single-region radiomics (delta-radiomics 0.643 vs Eso 0.514, DeLong p = 0.493) and a clinical-variable model (0.540). Compared against the alternative models we now benchmark (RF 0.643 > LR 0.529 > SVM 0.504), the RF delta-radiomics model is the strongest available configuration. We believe the methodological rigor we have added — nested CV, optimism correction, calibration curves, decision curve analysis, and sensitivity analyses — makes the study a useful contribution to the nascent literature on non-invasive Ki-67 assessment, provided it is read as preliminary.

Author Response File: Author Response.docx

Reviewer 2 Report

Comments and Suggestions for Authors

This manuscript proposes a delta-radiomics approach based on preoperative non-contrast CT images to predict Ki-67 expression status in patients with esophageal cancer. The authors extracted radiomic features from both tumor regions and paired normal esophageal tissue, calculated delta features, and built a logistic regression model to classify patients into Ki-67-high and Ki-67-low groups.  SHAP analysis was used to interpret the model, identifying texture-related features as major contributors. I have some concerns:

1.The sample size is small and there is no independent validation cohort. Only 59 patients were included, with 21 in the Ki-67-high group and 38 in the Ki-67-low group. Although leave-one-out cross-validation was used, this does not replace external validation. The reported performance may be unstable and optimistic. Additional internal validation, such as bootstrapping or repeated cross-validation, and preferably an independent validation cohort, should be provided.

2.The feature selection process may involve data leakage. The manuscript appears to describe feature filtering and correlation-based feature selection using the entire dataset before LOOCV. If feature selection was not performed within each training fold, information from the test sample would have influenced model construction, leading to inflated performance estimates. The authors should clarify this point and, if necessary, repeat the analysis using nested cross-validation.

3.The Ki-67 cutoff value lacks sufficient clinical justification. The cutoff of 0.63 was chosen because the sample median was 0.6 and several patients had Ki-67 = 0.6. This threshold appears data-driven and may limit clinical generalizability. Sensitivity analyses using alternative thresholds, or regression analysis treating Ki-67 as a continuous variable, would strengthen the study.

4.The delta-radiomics formula requires further justification. The formula (tumor feature - normal tissue feature) / (tumor feature + 0.001) uses the tumor feature as the denominator. This choice may lead to instability when tumor feature values are close to zero and differs from more conventional relative-difference definitions. The authors should explain the rationale and consider comparing alternative delta definitions.

5.ROI reproducibility was not sufficiently assessed. Since both tumor and normal esophageal tissue ROIs were manually delineated, inter-observer and intra-observer variability should be evaluated, for example using intraclass correlation coefficients. This is particularly important because radiomic features are sensitive to segmentation variability.

6.Some conclusions are stronger than the data support. The study predicts baseline Ki-67 expression but does not directly evaluate radiotherapy response, longitudinal Ki-67 changes, or patient outcomes. Claims regarding dynamic monitoring of radiotherapy response and personalized treatment adaptation should be toned down or clearly presented as future potential applications.

7.The pathological subtype of esophageal cancer should be clearly reported. If different histological subtypes were included, their potential influence should be discussed. The CT acquisition protocol should be described in more detail, including scanner type, tube voltage, tube current, reconstruction kernel, matrix, field of view, and reconstruction parameters.

8.More details are needed regarding radiomic feature extraction, including gray-level discretization, bin width or bin count, resampling settings, and whether the process followed IBSI recommendations. The study would be stronger if the authors compared the delta-radiomics model with a tumor-only radiomics model, a clinical-variable model, and a combined clinical-radiomics model.


9.Baseline characteristics being statistically non-significant does not exclude confounding, especially in a small cohort. This point should be acknowledged. Calibration curves, decision curve analysis, and confidence interval calculation methods should be added to provide a more complete evaluation of clinical utility.

 

 

Author Response

Reviewer 3.1

"The sample size is small and there is no independent validation cohort. Only 59 patients were included, with 21 in the Ki-67-high group and 38 in the Ki-67-low group. Although leave-one-out cross-validation was used, this does not replace external validation. The reported performance may be unstable and optimistic. Additional internal validation, such as bootstrapping or repeated cross-validation, and preferably an independent validation cohort, should be provided."

Response:

We thank the reviewer for this comprehensive point. We have addressed the internal-validation portion directly: we now use nested LOOCV, and we have added Bootstrap validation with 2000 resamples. The Bootstrap yielded an optimism of just 0.0001 and an optimism-corrected AUC of 0.6427, indicating that the within-sample estimate is not materially optimistic. We also report 95% CIs for all metrics (accuracy, sensitivity, specificity, AUC) and an EPV of 7.0. Regarding an independent validation cohort: we were not able to obtain one for this revision, and we acknowledge candidly that this remains the single most important unmet validation need. We have stated plainly in the limitations that external, multi-centre prospective validation is required before any clinical claim, and we have repositioned the work as exploratory/hypothesis-generating. We hope to secure an external cohort in a follow-up study.

Reviewer 3.2

"The feature selection process may involve data leakage. The manuscript appears to describe feature filtering and correlation-based feature selection using the entire dataset before LOOCV. If feature selection was not performed within each training fold, information from the test sample would have influenced model construction, leading to inflated performance estimates. The authors should clarify this point and, if necessary, repeat the analysis using nested cross-validation."

Response:

The reviewer's suspicion is correct, and we thank them for identifying this critical flaw. In the original submission, normalization and mRMR feature selection were applied to the full dataset before LOOCV, constituting data leakage and producing an inflated apparent AUC of 0.748. We have now repeated the entire analysis using nested LOOCV: within every training fold, we (i) fit min–max normalization on the training data only, (ii) perform mRMR feature selection (k=3) on the training data only, and (iii) train the model on the training data only, before evaluating on the single held-out patient. The test patient contributes no information to feature selection or training at any step. This correction lowered the AUC from 0.748 to a leakage-free 0.643 — a drop that itself validates the leakage hypothesis. We have rewritten the Methods to describe the exact pipeline order, so the independence of feature selection from the test fold is now explicit and verifiable.

Reviewer 3.3

"The Ki-67 cutoff value lacks sufficient clinical justification. The cutoff of 0.63 was chosen because the sample median was 0.6 and several patients had Ki-67 = 0.6. This threshold appears data-driven and may limit clinical generalizability. Sensitivity analyses using alternative thresholds, or regression analysis treating Ki-67 as a continuous variable, would strengthen the study."

Response:

We appreciate the reviewer's careful reading. We now disclose transparently that the 0.63 cut-off was chosen with reference to the sample distribution (median ≈ 0.6) and is therefore, as the reviewer notes, data-driven rather than anchored to an established clinical threshold; we no longer present it as pre-specified. To mitigate the generalizability concern we have added a multi-threshold sensitivity analysis spanning 0.33 to 0.73. We have also added an EPV calculation (EPV = 7.0 at 0.63) to contextualize event density. We agree that a continuous regression treating Ki-67 as a continuous outcome would be preferable and have added this as a clearly scoped future direction. We thank the reviewer for prompting the sensitivity analysis, which materially strengthens the manuscript.

Reviewer 3.4

"The delta-radiomics formula requires further justification. The formula (tumor feature - normal tissue feature) / (tumor feature + 0.001) uses the tumor feature as the denominator. This choice may lead to instability when tumor feature values are close to zero and differs from more conventional relative-difference definitions. The authors should explain the rationale and consider comparing alternative delta definitions."

Response:

We thank the reviewer for this insightful critique. The original denominator choice (tumor + 0.001) was indeed problematic: it places the tumor feature in the denominator and becomes unstable as tumor values approach zero, and it deviates from conventional delta-radiomics definitions. Following this comment, we have replaced the formula with the simple, interpretable difference delta-radiomics = esotarget − eso (tumor minus normal). This eliminates the near-zero denominator instability and aligns with the delta-radiomics literature. We have added a clear rationale in the Methods and empirically justified the choice through a three-tier comparison: delta-radiomics (0.643) > Esotarget (0.546) > Eso (0.514), demonstrating that the subtraction — not either region alone — carries the discriminative information. We believe the revised formula is both better justified and more robust, and we thank the reviewer for prompting this important correction.

Reviewer 3.5

"ROI reproducibility was not sufficiently assessed. Since both tumor and normal esophageal tissue ROIs were manually delineated, inter-observer and intra-observer variability should be evaluated, for example using intraclass correlation coefficients. This is particularly important because radiomic features are sensitive to segmentation variability."

Response:

We agree with the reviewer that segmentation reproducibility is essential for radiomic studies, and we thank them for the specific recommendation. We have expanded the ROI methodology substantially: we now report that the normal esophageal ROI was placed on a proximal uninvolved segment with a mean centroid distance of 11.3 cm from the tumor, with explicit exclusion criteria to avoid inflammatory, or murally thickened tissue. However, we must be candid that, in this revision, we have not yet completed a formal inter-/intra-observer reproducibility study with intraclass correlation coefficients (ICC), as it requires independent re-contouring by multiple observers on a representative subset. We have acknowledged this explicitly as a limitation. We hope the more detailed and conservative ROI placement protocol we now provide mitigates, though does not replace, the need for the ICC analysis.

Reviewer 3.6

"Some conclusions are stronger than the data support. The study predicts baseline Ki-67 expression but does not directly evaluate radiotherapy response, longitudinal Ki-67 changes, or patient outcomes. Claims regarding dynamic monitoring of radiotherapy response and personalized treatment adaptation should be toned down or clearly presented as future potential applications."

Response:

We thank the reviewer for this fair and important critique. We agree that the original manuscript overstated its implications: this study predicts baseline Ki-67 status from a single pre-operative CT and does not assess radiotherapy response, longitudinal Ki-67 kinetics, or survival outcomes. We have revised the manuscript to remove or clearly re-frame all claims about dynamic monitoring of radiotherapy response and personalized treatment adaptation: such statements now appear only in a clearly labeled 'Future Directions' context as speculative potential applications, not as study findings. The conclusion has been rewritten to state precisely what the data support — that delta-radiomics derived from baseline non-contrast CT shows modest, exploratory signal for Ki-67 stratification — and to refrain from therapeutic or prognostic claims. We believe the revised tone is appropriately conservative and faithful to the evidence, and we thank the reviewer for helping us correct this.

Reviewer 3.7

"The pathological subtype of esophageal cancer should be clearly reported. If different histological subtypes were included, their potential influence should be discussed. The CT acquisition protocol should be described in more detail, including scanner type, tube voltage, tube current, reconstruction kernel, matrix, field of view, and reconstruction parameters."

Response:

We thank the reviewer for these requests, both of which we have fulfilled. Pathological subtype: the cohort was predominantly esophageal squamous cell carcinoma (58/59, 98.3%), with one case of esophageal adenocarcinoma (1.7%). CT acquisition protocol: we have added a detailed imaging paragraph specifying scanner type, tube voltage, tube current, reconstruction kernel, reconstruction parameters, matrix, and field of view for the non-contrast thoracic CT used for ROI-based feature extraction. These specifics support reproducibility and allow readers to judge the generalizability of the radiomic features across scanners and protocols.

Reviewer 3.8

"More details are needed regarding radiomic feature extraction, including gray-level discretization, bin width or bin count, resampling settings, and whether the process followed IBSI recommendations. The study would be stronger if the authors compared the delta-radiomics model with a tumor-only radiomics model, a clinical-variable model, and a combined clinical-radiomics model."

Response:

We thank the reviewer for these constructive points and have addressed each. Feature extraction: we have added full details of the radiomic pipeline, including gray-level discretization (bin width / bin count), image resampling (voxel resampling settings), and the software/feature family used. To accommodate small volume structures, extraction parameters were optimized (e.g., allowing 2D ROIs, reducing minimum ROI size, setting edge padding), which deviate from IBSI guidelines. Model comparisons: we have added the exact benchmarks the reviewer requested, reported under identical nested LOOCV: (a) tumor-only radiomics (Eso: 0.514) and normal-only (Esotarget: 0.546) vs the delta-radiomics model (0.643); (b) a clinical-variable-only model (0.540); and (c) a combined clinical + radiomics model (0.619). The delta model outperforms both single-region radiomics and the clinical model, and the combined model (0.619) is close to but below the delta-radiomics-only model (0.643), suggesting the delta-radiomics transformation captures most of the available signal. These results appear in the revised Results, strengthening the study as the reviewer anticipated.

Reviewer 3.9

"Baseline characteristics being statistically non-significant does not exclude confounding, especially in a small cohort. This point should be acknowledged. Calibration curves, decision curve analysis, and confidence interval calculation methods should be added to provide a more complete evaluation of clinical utility."

Response:

We thank the reviewer for this nuanced methodological point. We agree that non-significant baseline comparisons do not rule out residual confounding, particularly in a cohort of 59 where power to detect differences is low. Regarding clinical-utility evaluation, we have added all three requested elements: (1) calibration curves assessing agreement between predicted and observed probabilities; (2) decision curve analysis (DCA) quantifying net benefit across threshold probabilities; and (3) a clearly described confidence-interval computation method (Bootstrap percentile, 2000 resamples) with 95% CIs now reported for AUC, accuracy, sensitivity, and specificity. These additions provide a more complete picture of clinical utility and its uncertainty, and they are presented in the revised Results (with supporting figures) and Methods. We believe the manuscript is substantially more rigorous as a result.

Author Response File: Author Response.docx

Reviewer 3 Report

Comments and Suggestions for Authors

I was pleased to review this retrospective study, which addresses a clinically relevant question using a novel delta-radiomics approach for predicting Ki-67 expression in esophageal cancer. The manuscript is generally well organized and presents encouraging preliminary results. However, several important concerns remain.

 

  • The most significant limitation is the very small cohort (n = 59), including only 21 patients in the Ki-67-high group.

With such a small dataset, the risk of optimistic performance estimates remains substantial. The study should include a discussion of events-per-variable considerations and confidence intervals for all performance metrics.

  • I would suggest adding external validation. Without external validation, it remains unclear whether the observed performance reflects a true biological signal. Maybe the authors should clearly position the work as exploratory and hypothesis-generating.
  • It is important for the authors to explicitly describe the exact order of preprocessing. Whether feature selection was nested inside each cross-validation iteration. Because this issue directly affects the validity of the reported AUC.
  • Because ROI delineation was manual, it should be clear whether the selected features are stable across observers. And the methodology for selecting the normal esophageal ROI should be described in greater detail. Like, how far from the tumor was the normal ROI placed? How was inflammatory or premalignant tissue avoided?
  • The authors should justify why continuous Ki-67 prediction was not pursued. A supplementary regression analysis predicting continuous Ki-67 values would strengthen the manuscript.
  • In this study, we have limited comparison with alternative models. It would be considerably stronger if performance were compared with Random Forest, Support Vector Machine, or Conventional radiomics without delta normalization.
  • Because of the small sample size, it is especially important to provide confidence intervals for accuracy, sensitivity, and specificity. Only the AUC confidence interval is reported.
  • Terminology should be standardized throughout the manuscript. The manuscript uses multiple forms of the same term: "Delta-Radiomics" (Abstract and throughout manuscript), "delta-Radiomics" (Introduction), "ΔRadiomics" (Methods equation).

Author Response

Reviewer 2.1

"The most significant limitation is the very small cohort (n = 59), including only 21 patients in the Ki-67-high group."

Response:

We thank the reviewer for this candid assessment, with which we agree. The cohort of 59 patients (21 Ki-67-high, 38 Ki-67-low) is the principal limitation of this study, and we have made this explicit in the first paragraph of the Discussion and in a dedicated limitations subsection. We have contextualized the small sample using events-per-variable analysis: with 3 retained features and 21 events, EPV = 7.0, which reduces — though does not eliminate — the risk of overfitting. We have also added Bootstrap optimism correction (2000 resamples; optimism = 0.0001, corrected AUC = 0.6427) to show that the estimate is not materially inflated. Nevertheless, we fully acknowledge that findings require confirmation in a larger, independent cohort.

Reviewer 2.2

"With such a small dataset, the risk of optimistic performance estimates remains substantial. The study should include a discussion of events-per-variable considerations and confidence intervals for all performance metrics."

Response:

We appreciate this constructive suggestion and have implemented both requests. (1) Events-per-variable: we now report EPV = 7.0 at the optimal threshold (21 events / 3 features), and discussing its implications for model reliability. (2) Confidence intervals: beyond the AUC confidence interval, we now report 95% CIs for accuracy, sensitivity, and specificity derived from the nested LOOCV and Bootstrap procedures, and we describe the exact CI computation method in the Methods (Bootstrap percentile intervals, 2000 resamples). (3) Optimism correction: Bootstrap validation yielded an optimism of 0.0001 and a corrected AUC of 0.6427, directly assessing the magnitude of over-optimism. These additions are summarized in the revised Results and Methods and referenced in the limitations discussion.

Reviewer 2.3

"I would suggest adding external validation. Without external validation, it remains unclear whether the observed performance reflects a true biological signal. Maybe the authors should clearly position the work as exploratory and hypothesis-generating."

Response:

We thank the reviewer for this balanced recommendation. We were unable to obtain an independent external cohort for this revision, and we candidly acknowledge that this limits our ability to confirm a true biological signal versus chance. In response, we have (1) explicitly repositioned the entire manuscript — title, abstract, and conclusion — as an exploratory, hypothesis-generating study; (2) strengthened internal validation through nested LOOCV, Bootstrap optimism correction (corrected AUC = 0.6427), calibration curves, and decision curve analysis (DCA); and (3) added a forthright limitations paragraph stating that external, multi-centre prospective validation is required before any clinical claim. We have also noted that the consistency of the delta-radiomics signal across feature-type comparisons (delta-radiomics 0.643 > Esotarget 0.546 > Eso 0.514) and models (RF > LR > SVM) is at least encouraging for a genuine signal, but we avoid overstating this. We hope to secure an external cohort for a follow-up study.

Reviewer 2.4

"It is important for the authors to explicitly describe the exact order of preprocessing. Whether feature selection was nested inside each cross-validation iteration. Because this issue directly affects the validity of the reported AUC."

Response:

We are grateful for this pointed and well-founded concern. The original manuscript suffered from data leakage because normalization and feature selection were applied to the full dataset before cross-validation (apparent AUC = 0.748). In the revision we have rewritten the Methods to specify the exact, ordered preprocessing pipeline: (i) image preprocessing and ROI-based feature extraction; (ii) per-fold min–max normalization fit on the training fold only; (iii) mRMR feature selection (k=3) on the training fold only; (iv) model training on the training fold; (v) evaluation on the held-out patient; repeated across nested LOOCV. The corrected, leakage-free AUC is 0.643. We believe the explicit ordering now makes it clear that feature selection occurs strictly inside each training fold and never sees the test patient, resolving the validity concern.

Reviewer 2.5

"Because ROI delineation was manual, it should be clear whether the selected features are stable across observers. And the methodology for selecting the normal esophageal ROI should be described in greater detail. Like, how far from the tumor was the normal ROI placed? How was inflammatory or premalignant tissue avoided?"

Response:

We thank the reviewer for these detailed questions about ROI methodology. We have substantially expanded the ROI description in the Methods. Regarding normal-tissue ROI placement: the normal esophageal ROI was delineated on a proximal, grossly uninvolved esophageal segment, with the centroid of the normal ROI separated from the tumor ROI centroid by a mean distance of 11.3 cm. We agree with the reviewer that observer stability is essential for radiomic features. We must be honest that, in this revision, we did not yet perform a full inter-/intra-observer reproducibility study with intraclass correlation coefficients (ICC), because it requires independent re-contouring by multiple observers on a subset of cases. We hope the transparent ROI placement criteria we now provide partially address the concern in the interim.

Reviewer 2.6

"The authors should justify why continuous Ki-67 prediction was not pursued. A supplementary regression analysis predicting continuous Ki-67 values would strengthen the manuscript."

Response:

We appreciate the reviewer's suggestion. We agree that a continuous treatment of Ki-67 is, in principle, more informative than binary discretization and avoids the arbitrariness of any threshold. We did not pursue continuous regression in the main analysis primarily because the clinical literature and the referring workflow in our setting frame Ki-67 as a categorical high/low marker for treatment stratification, and because the very small event structure (n = 59) constrains the stability of a regression fit. We also note that the original Ki-67 values themselves are semi-quantitative pathological estimates subject to sampling variability, which complicates treating them as a precise continuous outcome. Regarding a supplementary regression analysis: we agree it would add value, but given the limited sample and the need for additional validation of such a model, we have opted in this revision to perform the multi-threshold sensitivity analysis (0.33–0.73) that effectively probes the binary outcome across the Ki-67 continuum. We hope the reviewer understands this conservative choice.

Reviewer 2.7

"In this study, we have limited comparison with alternative models. It would be considerably stronger if performance were compared with Random Forest, Support Vector Machine, or Conventional radiomics without delta normalization."

Response:

We thank the reviewer for this valuable suggestion, which we have fully implemented. We now compare three classifiers trained on the same delta-radiomics features under identical nested LOOCV: Random Forest (AUC = 0.643), Logistic Regression (0.529), and Support Vector Machine (0.504). Random Forest was therefore selected as the optimal model. In addition, we directly address 'conventional radiomics without delta normalization' through a three-tier feature-type comparison: delta-radiomics features (0.643) > Combined (0.619) > Clinical variables alone (0.540), and delta-radiomics (0.643) > Esotarget (0.546) > Eso (0.514). This demonstrates that (a) the delta-radiomics transformation adds value over single-region features and (b) delta-radiomics contributes beyond clinical variables. DeLong tests confirmed the differences were not statistically significant (RF vs LR p = 0.544; delta-radiomics vs Esotarget p = 0.6075; delta-radiomics vs Eso p = 0.493), which we report honestly, but the consistent ranking supports the relative utility of the RF delta-radiomics approach. These comparisons appear in the revised Results.

Reviewer 2.8

"Because of the small sample size, it is especially important to provide confidence intervals for accuracy, sensitivity, and specificity. Only the AUC confidence interval is reported."

Response:

We agree entirely and have now added 95% confidence intervals for accuracy, sensitivity, and specificity in addition to the AUC CI. These intervals are computed via Bootstrap percentile method (2000 resamples) on the nested-LOOCV predictions and are reported alongside the point estimates in the revised Results tables and Figure 4. We have also documented the CI computation procedure in the Methods. Reporting these intervals underscores the precision (and its limits) of each metric given the small cohort and supports our exploratory framing. We thank the reviewer for prompting this important addition.

Reviewer 2.9

"Terminology should be standardized throughout the manuscript. The manuscript uses multiple forms of the same term: 'delta-radiomics' (Abstract and throughout manuscript), 'delta-radiomics' (Introduction), 'ΔRadiomics' (Methods equation)."

Response:

We thank the reviewer for catching this inconsistency, which we have corrected. The manuscript now uses a single standardized form, 'delta-radiomics' (lowercase 'd', hyphenated), consistently across the title, abstract, introduction, methods, results, and discussion, including within the equations and figure captions. We performed a full pass to eliminate the variant capitalizations noted by the reviewer. We believe the terminology is now uniform and unambiguous.

Author Response File: Author Response.docx

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

I have no further comments.

Author Response

We are very grateful for the reviewer’s time and efforts in evaluating our manuscript. We appreciate the positive acknowledgment.

Reviewer 2 Report

Comments and Suggestions for Authors

The authors have comprehensively responded to my comments and fixed critical methodological issues such as data leakage. Remaining constraints including no independent validation cohort and missing ICC analysis are clearly stated as limitations.Considering this is an exploratory investigation, I support acceptance of the revised paper.

Author Response

We deeply appreciate the reviewer’s thorough evaluation and encouraging remarks. We are glad that our revisions, particularly the methodological improvements (e.g., addressing data leakage) and the clear declaration of limitations, have met the reviewer’s expectations.

Back to TopTop