Review Reports
- George G. Makiev 1,2,*,
- Igor V. Samoylenko 1 and
- Tigran G. Gevorkyan 1
- et al.
Reviewer 1: Anonymous Reviewer 2: Qianqian Song
Round 1
Reviewer 1 Report
Comments and Suggestions for Authors- The title is clear and represents the work designed.
- The Simple Summary is clear but should acknowledge limitations such as false positives and the need for further validation.
- The abstract reports promising pooled results but underemphasizes high heterogeneity and low positive predictive values that limit clinical interpretation.
- The introduction is very brief and would benefit from a more detailed discussion on the role of AI tools in cancer detection and risk prediction. Incorporating recent evidence on AI applications in oncology, such as the study reported in 10.1002/bdd.70016, would help strengthen the scientific context and better justify the relevance and novelty of using AI-based approaches for early cancer detection.
- The introduction lacks a clear discussion of how this review differs from previous studies and does not sufficiently frame EHR-based AI models as risk-stratification tools rather than direct screening or diagnostic solutions.
- The search strategy is broad and includes multiple databases; however, the exact search strings and compliance with PRISMA or PRISMA-DTA guidelines should be stated more clearly to improve reproducibility.
- The lack of a formal risk-of-bias or quality assessment makes it harder to trust the pooled results.
- Pooling sensitivity and specificity despite very high heterogeneity limits the reliability of statistical estimates.
- The discussion overemphasizes model performance without sufficiently addressing the impact of high heterogeneity, low positive predictive values, and the practical challenges of implementing these AI models in real-world clinical settings.
- The conclusion should better highlight the need for further validation before clinical use.
Author Response
Comments 1: The title is clear and represents the work designed.
Response 1: -
Comments 2: The Simple Summary is clear but should acknowledge limitations such as false positives and the need for further validation.
Response 2: The following sentence has been added to the Simple Summary.
Comments 3: The abstract reports promising pooled results but underemphasizes high heterogeneity and low positive predictive values that limit clinical interpretation.
Response 3: The following sentence has been added to the abstract.
Comments 4: The introduction is very brief and would benefit from a more detailed discussion on the role of AI tools in cancer detection and risk prediction. Incorporating recent evidence on AI applications in oncology, such as the study reported in 10.1002/bdd.70016, would help strengthen the scientific context and better justify the relevance and novelty of using AI-based approaches for early cancer detection.
Response 4: We incorporated the source texts into the Introduction. As an example, we will cite other sources of literature.
Comments 5: The introduction lacks a clear discussion of how this review differs from previous studies and does not sufficiently frame EHR-based AI models as risk-stratification tools rather than direct screening or diagnostic solutions.
Response 5: We have supplemented the Introduction section.
Comments 6: The search strategy is broad and includes multiple databases; however, the exact search strings and compliance with PRISMA or PRISMA-DTA guidelines should be stated more clearly to improve reproducibility.
Response 6: We have made some additions to the text. All the listed search terms were used directly to search the databases.
Comments 7: The lack of a formal risk-of-bias or quality assessment makes it harder to trust the pooled results.
Response 7: We have supplemented the text by adding more information on the limitations of our analysis, and we have also included an assessment of the methodological quality of the studies using the Newcastle-Ottawa Scale.
Comments 8: Pooling sensitivity and specificity despite very high heterogeneity limits the reliability of statistical estimates.
Response 8: Given the limited number of studies in this field, a complete recalculation appears unfeasible. We have significantly strengthened the caveats in the "Results" (Sections 3.4 and 3.5) and "Discussion" sections to clearly state that the extreme heterogeneity (I² ≥ 99.9%) renders the pooled estimates of Se/Sp unreliable for model selection or setting clinical expectations. We have shifted the emphasis to the variability in this area rather than providing a definitive summary metric. We have also outlined potential future analytical approaches for this field as more data become available.
Comments 9: The discussion overemphasizes model performance without sufficiently addressing the impact of high heterogeneity, low positive predictive values, and the practical challenges of implementing these AI models in real-world clinical settings.
Response 9: We have expanded the Discussion section by adding dedicated paragraphs on: the clinical implications of a low positive predictive value as one of the key barriers to population-wide screening; practical obstacles to implementation in real-world clinical practice; a discussion regarding data heterogeneity and potential errors.
Comments 10: The conclusion should better highlight the need for further validation before clinical use.
Response 10: A sentence has been added to the Conclusion section.
Reviewer 2 Report
Comments and Suggestions for AuthorsThis manuscript presents a timely and relevant systematic review and meta-analysis evaluating the performance of electronic health record based artificial intelligence models for early detection of pancreatic cancer. Specifically, I have some major concerns as below.
1) The reported heterogeneity for sensitivity and specificity is extreme and raises concerns about the appropriateness of pooling these metrics using univariate random-effects models. While the authors acknowledge this limitation, the manuscript would benefit from a more rigorous diagnostic accuracy framework.
2) The included studies vary substantially in prediction horizons (ranging from months to years prior to diagnosis) and outcome definitions. These differences materially affect performance metrics and clinical interpretation. Stratified analyses by prediction window or outcome type would improve interpretability and reduce conceptual heterogeneity.
3) The conclusion that neural networks outperform other models may be overstated. Differences in AUC across algorithms could reflect dataset size, feature engineering, temporal modeling, or validation strategy rather than intrinsic algorithmic superiority. The authors should temper causal interpretations and emphasize that performance differences may be confounded by study-level factors.
4) Most included studies rely on internal validation, with few externally validated models. This significantly limits clinical applicability. The manuscript would benefit from a clearer distinction between internally validated proof-of-concept models and those approaching clinical readiness. Moreover, I highly recommend the authors to include related studies (https://academic.oup.com/jamia/article/31/11/2474/7698330 and https://pmc.ncbi.nlm.nih.gov/articles/PMC10462236/ ) in the discussion section.
Author Response
Comments 1: The reported heterogeneity for sensitivity and specificity is extreme and raises concerns about the appropriateness of pooling these metrics using univariate random-effects models. While the authors acknowledge this limitation, the manuscript would benefit from a more rigorous diagnostic accuracy framework.
Response 1:
Descriptive and qualitative comparisons of the extracted data were used to summarize differences across AI approaches and methodological frameworks. For quantitative synthesis, we employed univariate random-effects meta-analysis for AUC, sensitivity, and specificity using R version 4.5.1 (R Foundation for Statistical Computing) with the metafor package (version 4.8-0). This approach was selected due to its ability to account for between-study variability and its frequent use in prognostic and diagnostic meta-analyses of AI models in oncology.
We acknowledge that bivariate models, such as the Hierarchical Summary Receiver Operating Characteristic (HSROC) model, are generally recommended for meta-analyses of diagnostic accuracy studies, particularly when sensitivity and specificity are correlated and heterogeneity is high. However, due to insufficient reporting of paired sensitivity–specificity data with their confidence intervals in many included studies, as well as variability in threshold reporting across studies, application of HSROC was not feasible in the present analysis.
Therefore, we proceeded with univariate pooling while transparently reporting heterogeneity metrics (I² and Cochran’s Q) and interpreting results with appropriate caution. Where possible, subgroup and meta-regression analyses were conducted to explore sources of heterogeneity. Fixed-effects models were used for meta-regression and within predefined algorithm categories. Publication bias was assessed using funnel plots and Egger’s regression test for analyses including five or more studies.
We added the phrase:
“Given the anticipated high heterogeneity in sensitivity and specificity estimates, we initially employed univariate random-effects meta-analysis. However, we acknowledge that bivariate models (such as HSROC) are more appropriate for diagnostic accuracy meta-analyses and recommend their use in future studies when sufficient data are available.”
There is also an explanation in the Discussion.
Comments 2: The included studies vary substantially in prediction horizons (ranging from months to years prior to diagnosis) and outcome definitions. These differences materially affect performance metrics and clinical interpretation. Stratified analyses by prediction window or outcome type would improve interpretability and reduce conceptual heterogeneity.
Response 2:
We have added a clarification in the Discussion section.
We can also create a summary table with grouping by prediction horizon (<1 year, 1–3 years, >3 years) and by outcome type (binary classification vs. time-to-event), validation based on AUC, Sp, and Se data.
However, the necessity for a separate table is unclear, as this information is already presented in the main Table 1. Conducting a stratified analysis appears challenging due to the limited number of studies that provide sufficient data for such an analysis.
Comments 3: The conclusion that neural networks outperform other models may be overstated. Differences in AUC across algorithms could reflect dataset size, feature engineering, temporal modeling, or validation strategy rather than intrinsic algorithmic superiority. The authors should temper causal interpretations and emphasize that performance differences may be confounded by study-level factors.
Response 3: We have added text to the Discussion section to reduce emphasis on neural networks due to potential confounding factors in the studies.
Comments 4: Most included studies rely on internal validation, with few externally validated models. This significantly limits clinical applicability. The manuscript would benefit from a clearer distinction between internally validated proof-of-concept models and those approaching clinical readiness. Moreover, I highly recommend the authors to include related studies (https://academic.oup.com/jamia/article/31/11/2474/7698330 and https://pmc.ncbi.nlm.nih.gov/articles/PMC10462236/ ) in the discussion section.
Response 4: Information about the table in comment 3.
We have added a paragraph to the Discussion section.
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsMay be accepted for publication.
Reviewer 2 Report
Comments and Suggestions for AuthorsThe authors have fully addressed my concerns.