Artificial Intelligence in Gastrointestinal Wireless Capsule Endoscopy: A Systematic Literature Review and Meta-Analysis
Round 1
Reviewer 1 Report
Comments and Suggestions for AuthorsThe study needs some improvements:
1) The high heterogeneity observed in some analysis should be explored to identify the eventual sources
2) Please provide a reproducible search strategy
3) I don't see in the submission the PRISMA checklist.....
4) Was the meta-analysis registered in PROSPERO? If not, this should be listed as a limitation
5) The authors should mention in the Discussion the great impact of AI in the diagnostic management of small bowel diseases (for example Crohn disease). in this regard cite the relevant review PMID: 36688019 )
Author Response
Dear Reviewer,
Thank you for your constructive comments and valuable suggestions. We have carefully addressed all the points raised and revised the manuscript accordingly. The corresponding changes have been highlighted in the revised version.
Sincerely,
The Authors
|
|
Comment |
Reply |
|
1 |
The high heterogeneity observed in some analysis should be explored to identify the eventual sources. |
Thanks for your comment. The observed heterogeneity across studies can be attributed to several factors, including substantial variation in dataset size and composition, differences in AI model architectures, variability in clinical tasks and target pathologies, and inconsistent validation strategies. For example, included studies ranged from small cohorts to large-scale datasets. This is one of the our findings that shows we need more robust settings and analysis in this domain. We mention this in section 5.2 now. We also added Egger’s analysis and funnel plots to interpret the heterogeneity, please see section 4.9. |
|
2 |
Please provide a reproducible search strategy |
Thanks for you suggestion. Supplementary S1 file is our search queries. |
|
3 |
I don't see in the submission the PRISMA checklist.... |
Thanks for mentioning that. Supplementary S2 file is our PRISMA checklist. |
|
4 |
Was the meta-analysis registered in PROSPERO? If not, this should be listed as a limitation |
Thanks for your suggestion. We have now added this in the limitations section 5.5 of our study. |
|
5 |
The authors should mention in the Discussion the great impact of AI in the diagnostic management of small bowel diseases (for example Crohn disease). in this regard cite the relevant review PMID: 36688019 ) |
Thanks for your suggestion. It is cited to the Introduction, section “The Promise of Artificial Intelligence ” which is related to the mentioned article. |
Reviewer 2 Report
Comments and Suggestions for AuthorsThis was a very comprehensive and broad systematic review including 72 studies applying artificial intelligence (AI) to wireless capsule endoscopy (WCE). The authors reviewed AI performance across multiple diagnostic indications (bleeding, ulcers, polyps, tumors, IBD) and whole GI regions. The authors highlight methodological limitations in existing research and propose recommendations to improve clinical applicability.
I have some concern and suggestion
- Methodological heterogeneity not adequately addressed. While the authors acknowledge high heterogeneity, the meta-analysis pools extremely diverse studies differing in imaging device manufacturers and massively unequal sample sizes (36 to >10 million frames). May consider subgroup analyses or sensitivity analyses removing high-bias or extremely small studies.
- The QUADAS‑2 assessment shows a mixture of low and unclear risk of bias in the Index Test domain. May consider excluding high-bias studies in secondary analysis.
- May consider clarify how outcome definitions such as “bleeding lesion,” “ulcer,” “general abnormality” were harmonized before pooling.
- May added subsection specifically analyzing model performance differences between device types.
- The review correctly notes that most studies do not perform patient-wise analysis, but this critical limitation is not sufficiently emphasized in the Discussion.
Author Response
Dear Reviewer,
Thank you for your constructive comments and valuable suggestions. We have carefully addressed all the points raised and revised the manuscript accordingly. The corresponding changes have been highlighted in the revised version.
Sincerely,
The Authors
|
|
Comment |
Reply |
|
|
This was a very comprehensive and broad systematic review including 72 studies applying artificial intelligence (AI) to wireless capsule endoscopy (WCE). The authors reviewed AI performance across multiple diagnostic indications (bleeding, ulcers, polyps, tumors, IBD) and whole GI regions. The authors highlight methodological limitations in existing research and propose recommendations to improve clinical applicability. I have some concern and suggestion
|
Thanks for your comment. We investigated additional potential subgroup analyses; however, in these cases, each subgroup contained a very small number of studies, making statistical analysis not feasible. In response to your suggestion, we excluded studies with extremely small datasets, and six studies were removed from the analysis. Accordingly, Figures 9 and 10 have been updated. In addition, to evaluate potential bias and heterogeneity, we conducted Egger’s test and generated funnel plots, which are presented in Figures 11 and 12. |
|
|
The QUADAS‑2 assessment shows a mixture of low and unclear risk of bias in the Index Test domain. May consider excluding high-bias studies in secondary analysis. |
Thanks for your comments. We discussed it in our group, and we concluded that those studies could reflect a real limitations in utilizing AI systems in real clinical practice. We also did the bias analysis as funnel plots and Egger’s analysis to cover that. |
|
|
May consider clarify how outcome definitions such as “bleeding lesion,” “ulcer,” “general abnormality” were harmonized before pooling. |
Thanks for your comment. Section 3.4 related to this has been updated and highlighted now. |
|
|
May added subsection specifically analyzing model performance differences between device types. |
Thanks for your suggestion. We evaluated the impact of device types; however, only a small number specified the device, which makes a quantitative subgroup analysis by device type not feasible. Therefore, we decided not to include this analysis, as the findings were too heterogeneous and insufficient to support statistically meaningful conclusions. |
|
|
The review correctly notes that most studies do not perform patient-wise analysis, but this critical limitation is not sufficiently emphasized in the Discussion. |
Thanks for your suggestion. Section 5.1 is now updated and elaborated on that. |
Reviewer 3 Report
Comments and Suggestions for AuthorsThis systematic review and meta-analysis examines artificial intelligence applications in wireless capsule endoscopy for gastrointestinal (GI) disease diagnosis. Searching PubMed, Scopus, Embase, Web of Science, and Google Scholar, the authors identified 72 eligible studies from an initial pool of 1,451 records. Meta-analysis using random-effects models revealed consistently high pooled diagnostic performance, with bleeding and vascular lesion detection achieving the strongest results. The authors identify critical barriers to clinical adoption, including limited external validation, small cohorts, and retrospective designs, culminating in recommendations for future research.
However, I have the following concerns,
(1) Publication bias is not formally assessed. Studies reporting superior performance metrics are more likely to be published, and omitting this analysis means that the pooled estimates—particularly AUC values exceeding 95% across multiple categories—may be inflated. This weakens the credibility of the meta-analytic conclusions and the practical recommendations derived from them. The authors should generate funnel plots for each primary metric (accuracy, sensitivity, specificity, and AUC) within each subgroup where the number of studies permits (generally at least 10 studies). Egger’s regression test should be conducted and the results reported. Where asymmetry is detected, the authors should apply a trim-and-fill correction and present both unadjusted and adjusted pooled estimates, along with a transparent interpretation of the implications.
(2) The authors briefly note that multiple studies used the Kvasir dataset and acknowledge its class imbalance (Section 5: “Multiple studies leveraged the Kvasir dataset […] for training AI models […]. Despite its popularity, the dataset is significantly imbalanced”). However, no systematic analysis is conducted to determine how many of the 72 included studies share this or other common public datasets. When multiple included studies train and test on overlapping data, their results are not statistically independent observations. Pooling such results in a meta-analysis violates a core assumption of the random-effects model and can artificially inflate pooled performance estimates while narrowing confidence intervals. The Kvasir, KID, CAD-CAP, and KvasirCapsule datasets appear repeatedly across the included studies, suggesting this is a substantive issue rather than a minor concern.
Author Response
Dear Reviewer,
Thank you for your constructive comments and valuable suggestions. We have carefully addressed all the points raised and revised the manuscript accordingly. The corresponding changes have been highlighted in the revised version.
Sincerely,
The Authors
|
|
|
|
|
|
(1) Publication bias is not formally assessed. Studies reporting superior performance metrics are more likely to be published, and omitting this analysis means that the pooled estimates—particularly AUC values exceeding 95% across multiple categories—may be inflated. This weakens the credibility of the meta-analytic conclusions and the practical recommendations derived from them. The authors should generate funnel plots for each primary metric (accuracy, sensitivity, specificity, and AUC) within each subgroup where the number of studies permits (generally at least 10 studies). Egger’s regression test should be conducted and the results reported. Where asymmetry is detected, the authors should apply a trim-and-fill correction and present both unadjusted and adjusted pooled estimates, along with a transparent interpretation of the implications. |
Thanks for your suggestion. Egger’s test and funnel plots have now been added for subgroups with n≥10. The explanations for the results of this analysis are presented and highlighted in the last paragraphs of the Results section. It is worth noting that, in this field, most published articles report that they have developed or implemented AI systems that perform promisingly for lesion detection or classification; therefore, performance metrics greater than 90% are commonly observed. This is an issue that we have addressed in the Discussion, as although the reported results are promising, there are other important limitations that need to be considered to ensure reliable and robust performance in real clinical practice. |
|
|
The authors briefly note that multiple studies used the Kvasir dataset and acknowledge its class imbalance (Section 5: “Multiple studies leveraged the Kvasir dataset […] for training AI models […]. Despite its popularity, the dataset is significantly imbalanced”). However, no systematic analysis is conducted to determine how many of the 72 included studies share this or other common public datasets. When multiple included studies train and test on overlapping data, their results are not statistically independent observations. Pooling such results in a meta-analysis violates a core assumption of the random-effects model and can artificially inflate pooled performance estimates while narrowing confidence intervals. The Kvasir, KID, CAD-CAP, and KvasirCapsule datasets appear repeatedly across the included studies, suggesting this is a substantive issue rather than a minor concern. |
Thanks for noticing this issue. Figure 7(d) illustrates the proportion of studies that used public datasets, showing that approximately 25% of studies utilized publicly available data. To address your second point, we re-evaluated the final set of included studies, and the meta analysis has been updated accordingly, with studies affected by this issue excluded from the analysis. Please see Figure 9 and 10.
|
Round 2
Reviewer 1 Report
Comments and Suggestions for AuthorsThe revised manuscript is OK. Thank you!
Reviewer 3 Report
Comments and Suggestions for AuthorsRevision looks good to me.
