Next Article in Journal
CHaRT: An Autoregressive Transformer for Joint Forecasting of Clinical Events and Continuous Values
Previous Article in Journal
Trust, Emotion, and Skepticism in AI-Enabled Academic Marketing: Psychometric Validation and Cross-Validated Machine Learning Evidence from Higher Education
 
 
Article
Peer-Review Record

Hybrid Quantum-Classical Neural Networks for Healthcare Prediction Powered by Automated Scientific Discovery

Informatics 2026, 13(6), 98; https://doi.org/10.3390/informatics13060098
by Karthik Meduri 1,*, Ruthvik Yedla 2, Santosh Reddy Addula 1, Guna Sekhar Sajja 1, Shaila Rana 3,†, Elyson De La Cruz 3,†, Mohan Harish Maturi 1 and Hari Gonaygunta 1,†
Reviewer 1: Anonymous
Reviewer 2:
Reviewer 3: Anonymous
Informatics 2026, 13(6), 98; https://doi.org/10.3390/informatics13060098
Submission received: 31 March 2026 / Revised: 2 June 2026 / Accepted: 12 June 2026 / Published: 22 June 2026
(This article belongs to the Section Machine Learning)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors

The article presents a hybrid quantum-classical neural network (HQCNN) for clinical prediction, with a focus on parameter efficiency rather than accuracy superiority. The study is well-executed, clearly written, well organized, and methodologically sound. The authors make a compelling case for evaluating hybrid models under fair and reproducible conditions.

The paper is easy to read and presents appropriate experimental results and comparisons to classical benchmarks.

Some general remarks on presentation:

  • Rather than a full design methodology (arguable), you may want to reframe the paper as an evaluation and validation framework (more closely matching the content). I would like to remark that the HQCNN architecture is not fundamentally new. Similar hybrid architectures have been widely studied. The manuscript’s claim of introducing a "generalizable framework" is not convincingly supported.
  • The ablation study is informative but limited to a single train/test split. Can you extend the study?

Author Response

Thank you for your positive assessment of the study as "well-executed, clearly written, well organised, and methodologically sound," and for your two constructive remarks. Both points identified genuine weaknesses in the original manuscript and we have addressed each fully below.

Comment 1: Rather than a full design methodology (arguable), you may want to reframe the paper as an evaluation and validation framework (more closely matching the content). The HQCNN architecture is not fundamentally new. Similar hybrid architectures have been widely studied. The manuscript's claim of introducing a "generalizable framework" is not convincingly supported.

Response: We agree completely, and this reframing has materially strengthened the paper's honesty and precision. We have repositioned the contribution throughout the manuscript from a "design methodology" and "generalizable framework" to a reproducible evaluation and validation framework. The architecture itself is no longer presented as novel — we now explicitly state in Section 2.1 that the classical-encoder/PQC/classical-readout pattern "follows established hybrid designs" and cite the prior work that introduced it. The novelty we claim is narrower and fully defensible: a documented, shared-fold evaluation protocol with a parameter-matched classical control and a quantified epistemic-informativeness analysis.

Specific changes made:

  • Abstract rewritten to open with: "This study presents a reproducible evaluation framework for hybrid quantum-classical neural networks (HQCNNs) in healthcare classification, rather than a new architecture."
  • Section 1.3 (Problem Statement) now frames the contribution as addressing the evaluation problem, not an architectural one.
  • Section 1.5 (Contributions) reduced from five items to four, none of which claim architectural novelty.
  • The phrase "generalizable framework" has been replaced with "reusable evaluation framework, demonstrated here as a single-dataset proof-of-concept" throughout the manuscript.
  • Section 6 (Conclusion) now reads: "The contribution is the evaluation protocol — shared folds, matched control, significance testing — not a demonstrated quantum advantage."

Comment A2: The ablation study is informative but limited to a single train/test split. Can you extend the study?

Response: Yes, and we have done so. We re-ran the depth ablation across all five stratified cross-validation folds (the same shared folds used throughout the paper) for L ∈ {1, 2, 3} variational layers, reporting mean ± SD per depth with paired statistical tests between depths. The L=2 vs. L=3 difference is not statistically significant (Wilcoxon p = 1.00; paired t: p = 0.587). All three depths fall within one standard deviation of each other. We have accordingly withdrawn the claim that L = 2 is "optimal" and now state that depth does not materially affect performance in this shallow regime, and that L = 2 is adopted as a practical default rather than an identified optimum. Section 3.4.4 describes the multi-fold protocol and Section 4.2 reports the verified results. The ablation figure has been fully regenerated to show the 5-fold mean ± SD bars with the old "Optimal" and "Barren plateau effect begins" annotations removed.

Reviewer 2 Report

Comments and Suggestions for Authors

Reviewer Report

Manuscript Title: Hybrid Quantum-Classical Neural Networks for Healthcare Prediction Powered by Automated Scientific Discovery

The manuscript proposes a hybrid quantum-classical neural network (HQCNN) framework for clinical prediction, complemented by a Bayesian-surprise-guided design methodology and post hoc validation using the AutoDiscovery platform. The model is evaluated on the Wisconsin Diagnostic Breast Cancer (WDBC) dataset and compared against several classical baselines. The authors position their primary contribution as improved parameter efficiency rather than superior predictive accuracy.

Overall, the manuscript is well-written, clearly structured, and presents a thoughtful attempt to address reproducibility and design transparency in quantum machine learning. The emphasis on fair benchmarking and cautious interpretation of results is appreciated. However, several methodological and interpretational issues limit the strength of the conclusions and should be addressed before the work can be considered for publication.

 

Major Comments

Statistical Validation is Missing

While performance metrics are reported across cross-validation folds, the manuscript does not include formal statistical testing. Statements referring to “non-significant differences” or improved stability are therefore not substantiated.

Suggestion:
Include appropriate statistical tests (e.g., paired t-test or Wilcoxon signed-rank test) across folds, and report p-values and/or confidence intervals. This will strengthen the credibility of comparative claims.

 

Limited Evaluation on a Single Dataset

The study is conducted exclusively on the WDBC dataset, which is relatively small and well-studied. While suitable as a benchmark, it does not fully represent the complexity of real-world clinical data.

Suggestion:
If possible, include additional datasets to demonstrate robustness. Otherwise, the claims of generalizability should be moderated and framed as a proof-of-concept.

 

Parameter Efficiency Claim Needs Stronger Support

The manuscript highlights improved performance over a classical MLP with fewer parameters. However, there is no comparison with a classical model of similar parameter size.

Suggestion:
Include a parameter-matched classical neural network as a control, or discuss more explicitly whether the observed improvement can be attributed specifically to the quantum component rather than architectural differences.

 

Circuit Depth Ablation Requires More Robust Evaluation

The ablation study is conducted on a single train-test split and is described as exploratory. However, the conclusions drawn from it—particularly regarding optimal circuit depth—are relatively strong.

Suggestion:
Consider repeating the ablation across multiple splits or folds to provide more reliable evidence and variability estimates.

 

Clarification Needed for Bayesian Surprise Framework

The use of Bayesian surprise as an epistemic measure is interesting and potentially valuable. However, the current presentation raises some questions regarding prior selection and the use of Beta distributions.

Suggestion:
Clarify whether this framework is intended as a heuristic or a formal Bayesian model. Providing additional justification for prior choices, or briefly discussing sensitivity to these assumptions, would improve transparency.

 

Simulation-Only Evaluation

All experiments are conducted using a noiseless quantum simulator. While this is common practice, it limits the discussion around real-world applicability, especially in the context of NISQ devices.

Suggestion:
Acknowledge this limitation more explicitly and, if feasible, include results from a noisy simulator or discuss expected performance under realistic conditions.

 

Clinical Interpretation Should Be Moderated

The manuscript includes discussion on clinical deployment and governance, which is valuable. However, given that the study is limited to a benchmark dataset, these implications may be somewhat premature.

Suggestion:
Tone down claims related to clinical readiness and include a brief discussion on challenges such as data heterogeneity, external validation, and regulatory requirements.

 

Minor Comments

  • Some statements could be softened to maintain a neutral scientific tone (e.g., “provides the framework the field has long lacked”).
  • The KL divergence equation would benefit from clearer formatting and definition of variables.
  • Figures could be improved for readability (axis labels, font size).
  • Ensure consistent reference formatting and completeness of DOIs.
  • Clarify computational details such as hardware used, runtime, and reproducibility resources (e.g., code availability).

 

 

 

Comments for author File: Comments.pdf

Author Response

Thank you for your thorough and constructive major review. Your comments identified concrete methodological gaps, and we have addressed every one of them with new experiments rather than discussion alone. A note on integrity: two of the new results changed our conclusions, and we have updated the manuscript accordingly rather than defending the original phrasing. We detail each response below.

Major Comment 1 - Statistical Validation is Missing: The manuscript does not include formal statistical testing. Statements referring to "non-significant differences" or improved stability are therefore not substantiated. Include paired t-test or Wilcoxon signed-rank across folds and report p-values and/or confidence intervals.

Response: This is the most important revision and we have implemented it in full. Following the recommendation of Demšar (2006) for comparing classifiers across resamples, we ran paired Wilcoxon signed-rank tests as the primary test (n = 5 folds does not satisfy the t-test normality assumption) with the paired t-test as a secondary check. Multiple comparisons were corrected using the Holm–Bonferroni procedure. Results are reported in the new Table 2 (Section 4.1)

 

Major Comment 2 - Limited Evaluation on a Single Dataset: The study is conducted exclusively on WDBC. Moderate generalizability claims or frame as proof-of-concept if additional datasets are not added.

Response: We have chosen to moderate the claims rather than add a dataset, as recommended by your guidance. The paper is now explicitly framed throughout as a single-dataset proof-of-concept. All instances of "generalizable framework" have been replaced with "reusable evaluation framework, demonstrated as a proof-of-concept on a single benchmark dataset." The first limitation in Section 5.5 now states plainly: "External validity is untested: results are from a single dataset and constitute a proof-of-concept, not a generalization claim." Multi-dataset external validation is listed as the first and most pressing future direction in Section 5.6, framed as a prerequisite for any generalization claim rather than a suggestion.

Major Comment 3 - Parameter Efficiency Claim Needs Stronger Support: There is no comparison with a classical model of similar parameter size. Include a parameter-matched classical neural network as a control.

Response: We built and ran this control. We added a parameter-matched classical MLP with hidden layers [28, 10] totaling exactly 441 trainable parameters - identical to the HQCNN - evaluated on the same five shared folds. Results: the matched MLP reached 95.08 ± 1.81% accuracy and 99.04 ± 0.52% AUC. The HQCNN's edge at equal parameter count is +1.41 pp accuracy and +0.40 pp AUC, but this difference is not statistically significant (paired t: p = 0.056; Wilcoxon: p = 0.125). We have stated this explicitly in Sections 4.1 and 5.1: the evidence does not establish that the quantum component is responsible for any performance difference. The parameter-matched MLP is included in Table 1 and shown separately in Figure 4 (the CI plot).

Major Comment 4 - Circuit Depth Ablation Requires More Robust Evaluation: The ablation is conducted on a single train-test split. Repeat across multiple splits or folds.

Response: Addressed jointly with Reviewer A's Comment 2. We ran the full 5-fold multi-split ablation (see Table 3 in Section 4.2). Depth in L ∈ {1, 2, 3} has no statistically detectable effect on accuracy (Wilcoxon p = 1.00 for L=2 vs. L=3). The "optimal depth" conclusion has been withdrawn and replaced with the more honest finding that depth is not a sensitive hyperparameter in this shallow regime.

Major Comment 5 - Clarification Needed for Bayesian Surprise Framework: Clarify whether this framework is a heuristic or a formal Bayesian model. Justify prior choices and discuss sensitivity to Beta-distribution assumptions.

Response: The manuscript was ambiguous on this point and we have corrected it. Section 2.5 now states unambiguously: "We use Bayesian surprise as an epistemic-informativeness heuristic, not a formal generative Bayesian model of accuracy." The KL divergence equation has been added with all variables defined immediately beneath it (θ, p(θ), p(θ|D), integration domain, units in nats). Each prior in Table 4 is documented with its literature or domain-knowledge source. A prior-sensitivity note in Section 4.4 confirms that the ordinal ranking of findings (H1 > H3 > H2 > H5 > H4) is stable under reasonable prior perturbations, though the absolute nat values should be read only ordinally. The statement that this analysis "contributes nothing to prediction, it never enters model selection, training, or hyperparameter choice" appears explicitly in Sections 2.5 and 4.4.

Major Comment 6 - Simulation-Only Evaluation: All experiments use a noiseless simulator. Acknowledge this limitation more explicitly and, if feasible, include noisy simulator results.

Response: We ran the noisy-simulator experiment. Using PennyLane's default.mixed density-matrix backend, the L=2 HQCNN was re-evaluated on a held-out fold under a combined noise model: depolarising channel (p = 0.01) after each parameterised rotation and amplitude damping (γ = 0.01) after each entangling gate. Under this low-noise regime, the model was robust: accuracy 96.49%, AUC 99.71%, F1 97.10% - essentially unchanged from the noiseless result on the same fold. We report this honestly with the appropriate caveat in Section 4.5: this demonstrates robustness only at modest, uniform noise rates; real NISQ hardware has higher, device-specific gate error and decoherence, and the result should not be read as evidence of hardware readiness. The simulation-only limitation is now the second item in Section 5.5, and real hardware validation is the second item in Section 5.6.

Major Comment 7 - Clinical Interpretation Should Be Moderated: Tone down claims related to clinical readiness and discuss data heterogeneity, external validation, and regulatory requirements.

Response: All clinical-readiness language has been systematically moderated. Claims about governance and auditability are now framed as potential advantages conditional on external validation, not as demonstrated properties. Section 5.5 now includes an explicit paragraph on barriers to clinical translation: data heterogeneity across sites and equipment, the need for prospective external validation, and regulatory pathways such as software-as-a-medical-device considerations. Section 5.1 now reads: "these are potential advantages conditional on broader validation, not demonstrated properties of this single-dataset study."

 

Reviewer 3 Report

Comments and Suggestions for Authors

The manuscript titled “Hybrid Quantum-Classical Neural Networks for Healthcare Prediction Powered by Automated Scientific Discovery” presents an interesting and timely exploration of hybrid quantum-classical learning frameworks for healthcare prediction, particularly in the context of parameter-efficient clinical models. The study demonstrates strong motivation, relevant literature coverage, and a technically rich experimental setup combining HQCNN architecture with Bayesian-surprise-guided analysis. The authors have made a commendable effort to maintain reproducibility and comparative fairness through cross-validation and tuned baselines. The discussion on parameter efficiency and shallow circuit optimization is particularly valuable for researchers working in quantum machine learning applications in healthcare. However, the manuscript still requires major revision before it can be considered for publication. Several sections are excessively descriptive and repetitive, especially in the Introduction and Discussion (Pages 2–4 and Pages 12–14), which affects the overall readability and scientific conciseness of the paper. The methodological explanations related to Bayesian surprise analysis and AutoDiscovery integration need clearer mathematical justification and stronger differentiation between confirmatory analysis and predictive contribution. In Page 8, the circuit depth ablation is performed only on a single 80:20 split, which weakens the reliability of the conclusions regarding optimal circuit depth and barren plateau behavior. Similarly, the claim regarding clinical governance and interpretability advantages requires more concrete evidence and supporting evaluation metrics. The manuscript would also benefit from additional statistical significance testing between HQCNN and classical baselines, as overlapping confidence intervals currently limit the strength of the comparative claims. Figures are informative, but some captions and visual explanations can be improved for better clarity, particularly Figures 5 and 6. The limitations section is appreciated for its honesty; however, practical deployment challenges on real noisy quantum hardware remain largely theoretical and should be discussed more critically. Minor language polishing and formatting consistency are also required throughout the manuscript. Overall, the work has good research potential and addresses an emerging interdisciplinary area with practical relevance, but substantial refinement in technical justification, experimental validation, and presentation quality is necessary. Therefore, I recommend Major Revision.

Comments on the Quality of English Language

The manuscript titled “Hybrid Quantum-Classical Neural Networks for Healthcare Prediction Powered by Automated Scientific Discovery” presents an interesting and timely exploration of hybrid quantum-classical learning frameworks for healthcare prediction, particularly in the context of parameter-efficient clinical models. The study demonstrates strong motivation, relevant literature coverage, and a technically rich experimental setup combining HQCNN architecture with Bayesian-surprise-guided analysis. The authors have made a commendable effort to maintain reproducibility and comparative fairness through cross-validation and tuned baselines. The discussion on parameter efficiency and shallow circuit optimization is particularly valuable for researchers working in quantum machine learning applications in healthcare. However, the manuscript still requires major revision before it can be considered for publication. Several sections are excessively descriptive and repetitive, especially in the Introduction and Discussion (Pages 2–4 and Pages 12–14), which affects the overall readability and scientific conciseness of the paper. The methodological explanations related to Bayesian surprise analysis and AutoDiscovery integration need clearer mathematical justification and stronger differentiation between confirmatory analysis and predictive contribution. In Page 8, the circuit depth ablation is performed only on a single 80:20 split, which weakens the reliability of the conclusions regarding optimal circuit depth and barren plateau behavior. Similarly, the claim regarding clinical governance and interpretability advantages requires more concrete evidence and supporting evaluation metrics. The manuscript would also benefit from additional statistical significance testing between HQCNN and classical baselines, as overlapping confidence intervals currently limit the strength of the comparative claims. Figures are informative, but some captions and visual explanations can be improved for better clarity, particularly Figures 5 and 6. The limitations section is appreciated for its honesty; however, practical deployment challenges on real noisy quantum hardware remain largely theoretical and should be discussed more critically. Minor language polishing and formatting consistency are also required throughout the manuscript. Overall, the work has good research potential and addresses an emerging interdisciplinary area with practical relevance, but substantial refinement in technical justification, experimental validation, and presentation quality is necessary. Therefore, I recommend Major Revision.

Author Response

Thank you for your detailed and constructive assessment. We are encouraged by your recognition of the "technically rich experimental setup," "commendable effort to maintain reproducibility and comparative fairness," and the value of the parameter-efficiency and shallow-circuit discussion. We have addressed every major and minor comment below with substantive revisions.

Major Comment 1 - Sections are excessively descriptive and repetitive, especially Introduction (pp. 2–4) and Discussion (pp. 12–14). Affects readability and scientific conciseness.

Response: Agreed, and we have condensed both sections substantially. The Introduction's four background subsections were tightened, removing rhetorical phrasing such as "The models work. The question is whether they can work better" and collapsing the repeated statements of the parameter-efficiency thesis. The overlapping material between Sections 5.1 and 5.4 was merged, and the repeated barren-plateau exposition that appeared in both the Related Work and the Discussion has been consolidated into a single location. The Introduction's five-item contributions list has been reduced to four, and the discussion of each contribution in the body is no longer restated across multiple subsections.

Major Comment 2 - Bayesian surprise analysis and AutoDiscovery integration need clearer mathematical justification and stronger differentiation between confirmatory analysis and predictive contribution.

Response: We have sharpened this distinction throughout the manuscript. Section 2.5 now states explicitly and unambiguously: "We use Bayesian surprise as an epistemic-informativeness heuristic, not a formal generative Bayesian model of accuracy." The KL divergence equation is displayed with all variables defined (θ, prior, posterior, domain, units). Section 2.6 clarifies that AutoDiscovery is used as a post-hoc confirmatory instrument only. Sections 4.4 and 5.3 both contain the explicit statement: "This quantity is used purely confirmatorily and post hoc to rank how informative each finding is. It contributes nothing to prediction: it never enters model selection, training, or hyperparameter choice." The Bayesian surprise figure has also been updated so that H3's label now correctly reads "No Detectable Depth Effect (L∈{1,2,3})" rather than "Optimal Shallow Depth," eliminating the previous contradiction between the figure and the text.

Major Comment 3 - Circuit depth ablation was performed only on a single 80:20 split, which weakens conclusions regarding optimal circuit depth and barren plateau behavior.

Response: Addressed jointly with Reviewer A's Comment 2 and Reviewer B's Comment 4. We ran the full 5-fold ablation. With variability estimates now in hand, the depths fall within each other's standard deviations and the L=2 vs. L=3 contrast is non-significant (Wilcoxon p = 1.00). The "optimal depth" conclusion has been withdrawn. Section 4.2 now states: "Depth in this shallow regime does not materially affect performance; we adopt L=2 as a default, not an optimum." The barren-plateau discussion in Section 5.2 now states explicitly that a shallow four-qubit, ≤3-layer sweep cannot and does not probe barren plateaus, which arise at larger qubit counts and circuit depths than used here.

Major Comment 4 - Clinical governance and interpretability advantages require more concrete evidence and supporting evaluation metrics.

Response: We have moderated these claims and clarified the evidence we actually have. The parameter-count argument supports auditability potential, not demonstrated interpretability. We have removed any implication that the model was shown to be interpretable and added an explicit statement in Section 5.3: "We did not formally evaluate interpretability (no attribution or SHAP study was performed), so we make no interpretability-advantage claim." This is also listed as an explicit limitation in Section 5.5 and as a future direction in Section 5.6. All clinical governance language is now framed as conditional on external validation.

Major Comment 5 - Additional statistical significance testing needed between HQCNN and classical baselines. Overlapping confidence intervals currently limit the comparative claims.

Response: Addressed jointly with Reviewer B's Comment 1. The full statistical analysis is now in Section 3.4.6 and reported in Table 2. The Wilcoxon signed-rank test (primary) and the paired t-test (secondary) were applied across the five shared folds for each comparison, with the Holm–Bonferroni correction. No comparison survives correction (all adjusted p ≥ 0.625). The overlapping CIs noted by the reviewer are confirmed by the formal tests: HQCNN: 96.49% [94.06, 98.92]; SVM: 96.14% [93.88, 98.39]; matched MLP: 95.08% [92.83, 97.33]. All comparative claims have been rewritten to reflect this: the HQCNN is described throughout as competitive with, not significantly superior to, the classical baselines. A dedicated 95% CI figure (Figure 4) has been added showing all six models with overlapping intervals and the subtitle "Differences not significant after Holm–Bonferroni (all adj. p ≥ 0.625)."

Round 2

Reviewer 2 Report

Comments and Suggestions for Authors

Publication is recommended

Back to TopTop