Next Article in Journal
Improving the Efficiency of Collaboration Between Humans and Embodied AI Agents in 3D Virtual Environments
Previous Article in Journal
Characteristics of Stratum Disturbance During the Construction of Dual-Line Shield Tunnels with Consideration of Soil Spatial Variability
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Neural Network Architectures in Video Capsule Endoscopy: A Systematic Review and Meta-Analysis on Accuracy and Reading Time Performances

1
Department of Gastroenterology and Endoscopy, Fondazione Poliambulanza Istituto Ospedaliero, 25124 Brescia, Italy
2
Research and Clinical Trials Unit, Fondazione Poliambulanza Istituto Ospedaliero, 25124 Brescia, Italy
3
Center for Endoscopic Research Therapeutics and Training (CERTT), Università Cattolica del Sacro Cuore, 00168 Rome, Italy
4
Digestive Endoscopy Unit, Fondazione Policlinico Universitario Agostino Gemelli IRCCS, 00168 Rome, Italy
5
Department of Emergency, Fondazione Policlinico Universitario Agostino Gemelli IRCCS, Università Cattolica del Sacro Cuore, Largo Gemelli 8, 00168 Rome, Italy
*
Author to whom correspondence should be addressed.
These authors contributed equally to this work.
Joint last authors.
Appl. Sci. 2026, 16(2), 1134; https://doi.org/10.3390/app16021134
Submission received: 23 December 2025 / Revised: 15 January 2026 / Accepted: 19 January 2026 / Published: 22 January 2026
(This article belongs to the Section Computing and Artificial Intelligence)

Abstract

Artificial intelligence (AI) has revolutionized medical image analysis. Several neural network (NN) architectures were developed and applied across the last decade, becoming essential for automated diagnosis and clinical applications. AI based on NNs has become increasingly integrated into gastroenterology, offering new opportunities for automated lesion detection and workflow optimization. Small-bowel capsule endoscopy (SBCE) has benefited substantially from these advances, addressing long-standing challenges such as time-consuming video review and variability among readers. This systematic review and meta-analysis evaluated neural network-based models for lesion detection in SBCE, assessing pooled diagnostic accuracy and the impact of AI on reading time. A total of 44 primary studies were included: 36 validation studies for accuracy and 9 clinical studies for reading time. All NN architectures demonstrated high diagnostic performance, with a pooled accuracy of 95.3% (95% CI: 94.1–96.5%). More recent architectures, including transformer-based and capsule networks, outperformed classical convolutional neural networks (CNNs). AI assistance significantly reduced SBCE reading time, with a pooled mean reduction of 84% compared to standard review. These findings highlight the strong potential of AI to enhance SBCE efficiency and diagnostic reliability.

1. Introduction

Medical imaging represents one of the fields most profoundly transformed by artificial intelligence (AI), with substantial potential to reshape how images are acquired, interpreted, and integrated into clinical decision-making [1]. Deep learning, a subset of machine learning based on multilayered artificial neural networks, has driven much of this progress. Unlike traditional machine learning approaches that rely on handcrafted features, deep learning models automatically learn hierarchical representations directly from raw data, enabling superior performance in complex pattern-recognition tasks such as image analysis. In parallel, a wide variety of neural network architectures have been developed, each optimized for specific computational and clinical objectives [2].
Convolutional neural networks (CNNs) are particularly well suited for image-based tasks, as they process grid-structured data through convolutional filters that extract spatial features at increasing levels of abstraction [3]. CNNs have become the cornerstone of contemporary medical image analysis, enabling accurate modeling of complex, non-linear relationships in high-dimensional data. Over the past decade, several architectures have been successfully applied to clinical imaging. AlexNet introduced key innovations that facilitated the training of deep networks [4], while VGGNet enabled efficient feature extraction through stacked small filters [5]. U-Net was specifically designed for biomedical image segmentation, preserving pixel-level detail through its encoder–decoder structure [6]. ResNet and Desnet addressed training limitations of very deep networks through residual and dense connections, respectively, improving information flow and performance [7]. Lastly, EfficientNet optimized depth, width, and resolution simultaneously, achieving high accuracy with lower computational resources [8]. Collectively, these architectures have been applied across multiple imaging modalities, including CT, MRI, and endoscopic imaging, with continuously improving clinical performance [9].
Gastroenterology has been at the forefront of this transformation, particularly in diagnostic and therapeutic endoscopy. The integration of CNNs has opened new possibilities for real-time image interpretation, automated lesion detection, risk stratification, and clinical decision support. In colonoscopy, Computer-Aided Detection systems have consistently improved adenoma and polyp detection rates, enhancing identification of subtle or easily overlooked precancerous lesions [10,11,12,13]. In contrast, Computer-Aided Diagnosis systems have not yet demonstrated a clear clinical benefit for the resect-and-discard strategy, underscoring the need for further refinement and validation of these technologies [14]. Beyond colorectal applications, AI has increasingly been applied to inflammatory bowel disease management. Machine learning models can predict biologic treatment response, estimate relapse or complication risk, and integrate multimodal clinical data to support personalized care [15,16]. AI-based systems have also demonstrated high accuracy in grading endoscopic and histological disease activity in ulcerative colitis and Crohn’s disease, with potential to reduce interobserver variability and standardize disease assessment [17,18,19].
Small-bowel capsule endoscopy (SBCE) represents another area in which AI has rapidly gained traction. Although SBCE plays a crucial role in diagnosing small-bowel disorders—including suspected SB bleeding, Crohn’s disease, and small-bowel tumors—it has long been limited by several factors: prolonged reading times, substantial interobserver variability, and challenges in objectively assessing bowel cleanliness. To address these limitations, several CNN-based models have been developed, demonstrating encouraging performance in automated assessment of bowel preparation quality, capsule localization and transit analysis, and detection of clinically relevant small-bowel lesions, including ulcers, erosions, vascular abnormalities, and tumors [20]. These advances suggest that AI may enhance diagnostic accuracy, standardize reporting, and substantially reduce the workload associated with SBCE interpretation.
However, existing studies are characterized by heterogeneity in neural network architectures, training datasets, validation methodologies, and reported performance metrics. A systematic synthesis of the available evidence is therefore required.
The aim of this systematic review and meta-analysis is to summarize and evaluate the different neural network models applied in small-bowel capsule endoscopy and to assess their diagnostic performance across the available literature.

2. Materials and Methods

The performed analyses and applied methods were compliant with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines.

2.1. Search Strategy

Studies were identified by searching the electronic databases Pubmed and Scopus (on 13 October 2025) using the following search terms: (“capsule endoscopy”[MeSH Terms] OR (“capsule”[All Fields] AND “endoscopy”[All Fields]) OR “capsule endoscopy”[All Fields] OR (“video”[All Fields] AND “capsule”[All Fields] AND “endoscopy”[All Fields]) OR “video capsule endoscopy”[All Fields]) AND (“deep”[All Fields] AND (“neural networks, computer”[MeSH Terms] OR (“neural”[All Fields] AND “networks”[All Fields] AND “computer”[All Fields]) OR “computer neural networks”[All Fields] OR (“neural”[All Fields] AND “network”[All Fields]) OR “neural network”[All Fields])) in either the title or the abstract. We found n = 142 findings in Pubmed and n = 328 in Scopus (original article and reviews in English language) from January 2016 to October 2025 reporting data on the use of convolutional neural networks for lesion detection in the gastrointestinal field. The research strategy and related findings were independently validated by two reviewers.

2.2. Inclusion and Exclusion Criteria

We included studies that reported the use of neural networks (NNs) for lesion detection in CE. Studies reporting lesion detection accuracy and/or capsule video reading time as outcomes were considered. The accuracy metric was defined as the proportion of true positives (TP) and true negatives (TN) over the total number of cases/lesions [(TP + TN)/Tot lesions]. Two types of studies were included and analyzed: validation studies and clinical studies. Papers focused on the assessment and validation of the detection performance of neural network computational models have been considered as validation studies; meanwhile, papers evaluating clinicians’ performance, in terms of lesion detection and reading time, using AI-based tools powered by NNs have been considered as clinical studies. For validation studies, inclusion was based on reporting the accuracy metric defined above. For clinical studies, those reporting the accuracy metric and/or reading time were included. Studies that did not report any of the specified outcomes were excluded.

2.3. Study Selection and Data Extraction

The searches yielded a total of 142 titles and abstracts that were screened. Of these, 50 were excluded because they were not primary studies, or dealt with topics unrelated to the subject of interest. Additionally, 48 did not meet the inclusion criteria (missing the metric of interest) (Figure 1).
The extracted data included the following: year of publication, country where the study was conducted, whether it was a validation study or a clinical study, sample size, whether the study was retrospective or prospective, whether it was multicenter study, type of VCE (video capsule endoscopy) used, metric calculated, metric value, analyzed segment, reading time, and the main detection outcome.

2.4. Outcome Measures

Two different meta-analyses were conducted, each addressing a distinct outcome. The first meta-analysis included validation studies and aimed to estimate the pooled diagnostic accuracy of AI-assisted video capsule endoscopy. The second meta-analysis focused on clinical studies and evaluated the impact of AI assistance on workflow efficiency by comparing mean VCE reading times between standard interpretation and AI-assisted reading modes.

2.5. Data Analysis

Two different random-effects meta-analyses were conducted with distinct objectives. The first meta-analysis was performed to obtain a pooled estimate of the proportion of correctly identified gastrointestinal lesions, together with 95% confidence intervals (CI). The random-effects model was chosen as a conservative approach, assuming that there would be substantial between-study heterogeneity in patient populations, diagnostic methods, and criteria for lesion classification. Proportions and corresponding variances were calculated for each study using the escalc() function with the measure of proportion set as “PR” [21].
The second meta-analysis was applied to compute a pooled reading time difference between video capsule reading time obtained using a standard clinical approach compared with an artificial intelligence–assisted approach. In detail, a random-effects model was established in terms of using mean difference as the effect measure size using the escalc() function with the measure of effect size set as “MD” [21].
Study heterogeneity was assessed using the Q statistic and the I2 index (defined as percentage of total variability due to heterogeneity). A significant Q value and high (larger than 50) I2 index indicate lack of homogeneity of findings among studies [22]. Categorical characteristics were treated as moderators and compared between subgroups based on country (Europe vs. non-Europe), year of publication (2019–2022 vs. 2023–2025), type and purpose of NN architecture, gastrointestinal tract inspected (colon vs. small bowel), and type of lesion (ulcers vs. other). Publication bias was assessed using Begg’s rank correlation test [14]. All statistical analyses were performed using R version 4.5.1, with the metafor package [21]. The level of statistical significance was set at p < 0.05.

2.6. Quality Assessment

Study quality was assessed using the QUADAS 2 tool for the quality assessment of systematic reviews of diagnostic accuracy [23]. This tool comprises 4 domains: patient selection, index test, reference standard, and flow and timing. Each domain is assessed in terms of risk of bias. Risk-of bias assessments were independently conducted by two reviewers.

3. Results

A total of 44 primary studies were considered and assessed in the meta-analyses (Table 1). Of these, 36 validation studies were included in the meta-analysis, assessing pooled diagnostic accuracy, while 9 clinical studies were included in the meta-analysis, assessing the pooled mean difference in reading time between AI-assisted reading and standard reading; 1 study was included in both meta-analyses.
The included studies were broadly distributed geographically and temporally. Thirteen of the thirty-six validation studies were conducted in Europe, while the remaining studies were carried out across the United States, the Middle East, Asia, and North Africa. With respect to publication year, 23 studies were published between 2019 and 2022, and 13 were published between 2023 and 2025. Studies were almost equally distributed between single-center and multicenter designs and were predominantly retrospective in nature.
Information regarding NN architectures was reported in 32 studies (73%). Among these, CNNs were the most frequently employed architecture (91%) followed by transformer neural networks (TNNs, n = 2) and a capsule network (n = 1). Regarding the purpose of neural networks, 27 studies applied NNs for classification, 4 for object identification, and 1 for image segmentation. Among CNN-based studies, Xception, ResNet, and AlexNet were the most commonly used architectures. Regarding the capsule endoscopy device, 24 studies (55%) used the PillCam capsule and approximately 10% used alternative capsules such as MiroCam, NaviCam, or OMOM, while the remaining studies employed other capsule endoscopy devices.

3.1. Assessment of Quality of Research

As shown in Figure S1, the risk of bias for all included studies was assessed across four predefined domains—patient selection, index test, reference standard, and flow and timing—using a structured quality assessment framework. Among validation studies, the risk of bias was judged to be low in at least 95% of cases across all four domains. This finding largely reflects the experimental nature of these studies, which focused on computational model development and validation and were therefore not subject to patient selection bias.
In contrast, a higher, though still acceptable, risk of bias (approximately 20%) was observed among the nine clinical studies. This was primarily attributable to concerns within the index test domain, mainly related to missing or incompletely reported data. Overall, the risk-of-bias assessment suggests that the included evidence is generally robust, with limited sources of potential bias that were transparently identified and considered in the interpretation of results.

3.2. Meta-Analysis on Accuracy of Neural Networks in Lesion Detection

The pooled estimate of overall detection accuracy was 95.30% (95% CI [94.10–96.50%]). Substantial statistical heterogeneity was observed (Q(36) = 4887.06, p < 0.001; I2 = 99.8) (Figure 2). Across individual studies, reported accuracy ranged from 88% to 100%, indicating consistently high performance of neural network-based systems for lesion detection despite marked heterogeneity. Assessment of publication bias using Begg’s rank correlation test did not suggest the presence of small-study effects.
Subgroup analysis demonstrated that the proportion of correctly detected lesions was significantly higher in studies inspecting the upper GI tract (99.47%, 95% CI [98.69–100.00]) compared to those evaluating the colon (94.8%, 95% CI [91.99–97.61]) (p < 0.001) (Table 2). In addition, studies employing CNN architecture resulted in significantly lower accuracy compared with more recent architectures, such as capsule networks or transformer-based models (93% vs. 98%, p = 0.015).
No significant differences were observed regarding country, year of publication, lesion type, or intended purpose of the neural network. Furthermore, when comparing the different types of neural network models, no significant differences were found between Xception, the most frequently used architecture, and the other architectures included in the analysis.

3.3. Meta-Analysis of Video Capsule Endoscopy Reading Time with and Without Neural Network Assistance

A second meta-analysis was performed to assess the difference in mean reading time of video capsule endoscopy (VCE) between standard interpretation and neural network-assisted reading. The analysis included nine studies reporting reading time outcomes. Four studies provided both mean reading time and standard deviation (SD), while five studies did not report SDs. For studies with missing SDs, values were imputed in accordance with established meta-analytic approaches by applying the inverse formula of the coefficient of variation (CV; SD/mean). A constant CV was assumed, calculated from the pooled SD of the four studies with complete data. This imputation method was used to allow inclusion of all eligible studies in the quantitative synthesis and is reported transparently to facilitate reproducibility. The pooled mean reading time was 50.9 min for standard reading (Figure S2) and 6.97 min for AI-assisted reading (Figure S3). The pooled mean difference between standard and AI-assisted reading was 43.82 min (Figure 3), corresponding to an estimated reduction in reading time of approximately 84% associated with AI assistance.

4. Discussion

In recent years, deep neural networks have demonstrated high accuracy and robustness in medical imaging and have increasingly been recognized as reliable tools for automated classification and segmentation. In GI endoscopy, CNN architectures—using ResNet, Xception, and AlexNet models—have been widely adopted, showing comparable performance due to their ability to extract complex and reproducible patterns. Across studies, reported accuracy, sensitivity, and specificity frequently equal or exceed those achieved by classical manual interpretation, further supporting the potential role of these systems in clinical decision support.
Interestingly, our analysis showed that classical CNN architectures performed worse than more recent neural network models. This finding is consistent with broader developments in computer vision, where purely convolutional backbones are increasingly complemented or replaced by transformer-based architectures [67]. While CNNs are highly effective in capturing local image features, they have limited inherent capacity to model long-range dependencies and complex temporal context, unless specifically engineered for this purpose. In endoscopic video analysis, lesion detection often depends not only on static texture or color changes but also on subtle changes in shape, motion, and contextual information across consecutive frames. Newer architectures, including transformer-based models (e.g., Vision Transformer model) and hybrid spatiotemporal networks (e.g., CapsNet-based approaches), can aggregate information across frames and learn richer representations of video dynamics. These capabilities may explain their improved robustness to motion artifacts, blur, and transient occlusions [68]. Moreover, later-generation models often benefit from pretraining on large-scale image or video datasets, advanced regularization techniques, and more sophisticated training pipelines which may enhance generalizability across different endoscope systems, imaging protocols, and patient populations. The superior performance of these architectures observed in our meta-analysis suggests that GI endoscopy, as a predominantly video-based modality, particularly benefits from models that are designed to exploit temporal information rather than relying solely on frame-level analysis.
Despite substantial heterogeneity among included studies—related to different NN architectures, clinical applications, and evaluated gastrointestinal segments (Q index 4887.06)—the pooled accuracy was very high at 0.95. This finding confirms the overall reliability of AI-based approaches in gastroenterological imaging and reinforces their growing role in routine clinical practice.
Importantly, no significant performance differences were observed between studies conducted in European versus non-European populations, suggesting that CNN-based systems generalize well across diverse ethnic and geographic settings. Similarly, no significant variation was observed between CNNs applied to different capsule endoscopy platforms, indicating that AI tools perform consistently regardless of the capsule manufacturer. These findings underscore the broad feasibility of adopting CNN-based systems across varied clinical settings and device types.
The higher detection rate observed in upper GI studies may reflect technical and procedural advantages, as gastric examinations are generally less affected by inadequate cleansing and residual debris compared with colonic evaluations. However, this subgroup comparison should be interpreted cautiously, as SBCE is dedicated exclusively to small-bowel assessment, and extrapolation to other GI segments is indirect.
Several international societies have proposed standardized quality indicators for SBCE, reflecting a growing emphasis on examination quality and reproducibility [69]. CNN-based systems may play a pivotal role in operationalizing these recommendations by enabling automated, objective, and reproducible image analysis. Integration of AI tools into SBCE workflows has the potential to reduce inter-reader variability and to improve adherence to established quality benchmarks.
A major advantage of implementing CNN-based systems in SBCE is the consistently high diagnostic accuracy reported across a wide range of lesion types. Numerous studies have shown that CNNs can detect mucosal abnormalities with sensitivities and specificities comparable to, or in some cases exceeding, those of expert human readers [35,60]. Importantly, these systems also substantially reduce reading times—often by as much as 80–85%—without compromising diagnostic performance. This reduction is particularly relevant because prolonged image interpretation is associated with reader fatigue, which in turn increases the risk of missed or overlooked findings.
By shortening review times and supporting lesion detection, AI-enabled SBCE interpretation has the potential to alleviate clinician workload and mitigate the cognitive burden historically linked to capsule reading. Consequently, clinicians may allocate more time to clinical decision-making, patient counseling, and integration of endoscopic findings into broader patient management. Consequently, AI adoption has the potential not only to enhance technical diagnostic performance but also to improve workflow efficiency and patient outcomes.
Importantly, further validation through prospective, multicenter studies is required to confirm the clinical effectiveness and generalizability of these technologies. Such studies should evaluate not only diagnostic accuracy but also clinically meaningful outcomes, including reading time, interobserver variability, and impact on patient management. Addressing these aspects will be crucial to enable the safe, standardized, and widespread adoption of AI-assisted capsule endoscopy in daily practice.
This meta-analysis has several limitations. First, there was substantial heterogeneity across studies regarding lesion types, outcome definitions (per-frame, per-lesion, or per-patient metrics), and applied neural network models (CNN, TNN, and others). Although subgroup analyses were performed, residual confounding cannot be excluded. Second, most studies were retrospective and single-center, often relying on internal validation cohorts, which raises concerns about the generalizability of the reported performance, especially for high-complexity models that may be more sensitive to domain shift. Third, publication bias cannot be excluded despite formal assessment, as studies reporting negative or modest results may be underrepresented, while early reports of emerging models may rely on small or highly curated datasets. Fourth, incomplete reporting of methodological details in several studies limited our ability to fully adjust for potential confounders. Lastly, methodological features and technical details related to neural network architectures, as well as information regarding model tuning and pretraining, were highly heterogeneous, inconsistently reported in only a few studies, and entirely absent in the vast majority of the included primary studies. Consequently, our review and meta-analysis is inherently unable to draw evidence-based conclusions on these technical aspects.

5. Conclusions

In summary, this meta-analysis shows that neural network-based lesion detection in video endoscopy currently achieves higher accuracy in the upper GI tract than in the colon, and that more recent architectures outperform traditional CNNs. However, further prospective, multicenter studies with standardized outcome definitions, external validation, and more comprehensive reporting of technical and computational details are required to facilitate safe and effective clinical implementation and to ensure computational reliability.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/app16021134/s1. Figure S1: Quality (in percentage on the considered studies) of the 36 validation studies (panels A–B) and of the nine clinical studies (panels C–D) assessed in the four domains of the QUADAS 2 assessment tool. Figure S2: Forest plot of 9 studies reporting the reading time: reading time estimate (reading time in standard reading) with 95% confidence limit (bars); pooled mean reading time is reported as a diamond. Figure S3: Forest plot of 9 studies reporting the reading time: reading time estimate (reading time in NN-assisted reading) with 95% confidence limit (bars); pooled mean reading time is reported as a diamond. Table S1: Research strategy findings. Reference [70] is cited in the supplementary materials.

Author Contributions

Conceptualization, D.S., C.Z. and C.F.; methodology, C.Z. and C.F.; validation, C.Z. and C.F.; formal analysis, C.Z. and C.F.; investigation, D.S., C.Z. and C.F.; resources C.Z.; data curation, D.S., C.Z. and C.F.; writing—original draft preparation, S.P., L.Z.D.V., G.T., L.G., D.S. and C.Z.; writing—review and editing, D.S., C.Z. and P.C.; supervision, C.F. and C.S. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The raw data supporting the conclusions of this article are available in the published studies included in the meta-analysis and can be accessed through the respective publications.

Acknowledgments

We thank Fondazione Roma for the invaluable support of this scientific research—FR-CEMAD 21–25.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial intelligence
SBCESmall-bowel capsule endoscopy
CNNConvolutional neural network
NNNeural network
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
TPTrue positive
TNTrue negative
VCEVideo capsule endoscopy
CIConfidence intervals
SDStandard deviation
CVCoefficient of variation
VValidation study
CClinical study
RRetrospective study
PProspective study
TNNTransformer neural network
RNNRecurrent neural network
GI TractGastrointestinal tract
Upper GIUpper gastrointestinal
SRStandard reading
AIRArtificial intelligence reading

References

  1. Schmidhuber, J. Annotated History of Modern AI and Deep Learning. arXiv 2022, arXiv:2212.11279. [Google Scholar] [CrossRef]
  2. Zhao, X.; Wang, L.; Zhang, Y.; Han, X.; Deveci, M.; Parmar, M. A review of convolutional neural networks in computer vision. Artif. Intell. Rev. 2024, 57, 99. [Google Scholar] [CrossRef]
  3. LeCun, Y.; Boser, B.; Denker, J.S.; Henderson, D.; Howard, R.E.; Hubbard, W.; Jackel, L.D. Backpropagation Applied to Handwritten Zip Code Recognition. Neural. Comput. 1989, 1, 541–551. [Google Scholar] [CrossRef]
  4. Alom, M.Z.; Taha, T.M.; Yakopcic, C.; Westberg, S.; Sidike, P.; Nasrin, M.S.; Van Essen, B.; Awwal, A.A.S.; Asari, V.K. The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches. arXiv 2018, arXiv:1803.01164. [Google Scholar] [CrossRef]
  5. Sengupta, A.; Ye, Y.; Wang, R.; Liu, C.; Roy, K. Going Deeper in Spiking Neural Networks: VGG and Residual Architectures. Front. Neurosci. 2019, 13, 425055. [Google Scholar] [CrossRef]
  6. Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional Networks for Biomedical Image Segmentation. In Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2015; Volume 9351, pp. 234–241. [Google Scholar] [CrossRef]
  7. Wightman, R.; Touvron, H.; Jégou, H.; Ai, F. ResNet Strikes Back: An Improved Training Procedure in Timm. arXiv 2021, arXiv:2110.00476. [Google Scholar] [CrossRef]
  8. Tan, M.; Le, Q.V. EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, Long Beach, CA, USA, 9–15 June 2019; pp. 10691–10700. [Google Scholar]
  9. Mienye, I.D.; Swart, T.G.; Obaido, G.; Jordan, M.; Ilono, P. Deep Convolutional Neural Networks in Medical Image Analysis: A Review. Information 2025, 16, 195. [Google Scholar] [CrossRef]
  10. Spada, C.; Salvi, D.; Ferrari, C.; Hassan, C.; Barbaro, F.; Belluardo, N.; Grazioli, L.M.; Milluzzo, S.; Olivari, N.; Papparella, L.; et al. A comprehensive RCT in screening, surveillance, and diagnostic AI-assisted colonoscopies (ACCENDO-Colo study). Dig. Liver Dis. 2025, 57, 762–769. [Google Scholar] [CrossRef]
  11. Shaukat, A.; Lichtenstein, D.R.; Somers, S.C.; Chung, D.C.; Perdue, D.G.; Gopal, M.; Colucci, D.R.; Phillips, S.A.; Marka, N.A.; Church, T.R.; et al. Computer-Aided Detection Improves Adenomas per Colonoscopy for Screening and Surveillance Colonoscopy: A Randomized Trial. Gastroenterology 2022, 163, 732–741. [Google Scholar] [CrossRef]
  12. Soleymanjahi, S.; Huebner, J.; Elmansy, L.; Rajashekar, N.; Lüdtke, N.; Paracha, R.; Thompson, R.; Grimshaw, A.A.; Foroutan, F.; Sultan, S.; et al. Artificial Intelligence–Assisted Colonoscopy for Polyp Detection. Ann. Intern. Med. 2024, 177, 1652–1663. [Google Scholar] [CrossRef] [PubMed]
  13. Repici, A.; Spadaccini, M.; Antonelli, G.; Correale, L.; Maselli, R.; Galtieri, P.A.; Pellegatta, G.; Capogreco, A.; Milluzzo, S.M.; Lollo, G.; et al. Artificial intelligence and colonoscopy experience: Lessons from two randomised trials. Gut 2022, 71, 757–765. [Google Scholar] [CrossRef]
  14. Hassan, C.; Rizkala, T.; Mori, Y.; Spadaccini, M.; Misawa, M.; Antonelli, G.; Rondonotti, E.; Dekker, E.; Houwen, B.B.S.L.; Pech, O.; et al. Computer-aided diagnosis for the resect-and-discard strategy for colorectal polyps: A systematic review and meta-analysis. Lancet Gastroenterol. Hepatol. 2024, 9, 1010–1019. [Google Scholar] [CrossRef]
  15. Gui, X.; Bazarova, A.; del Amor, R.; Vieth, M.; de Hertogh, G.; Villanacci, V.; Zardo, D.; Parigi, T.L.; Røyset, E.S.; Shivaji, U.N.; et al. PICaSSO Histologic Remission Index (PHRI) in ulcerative colitis: Development of a novel simplified histological score for monitoring mucosal healing and predicting clinical outcomes and its applicability in an artificial intelligence system. Gut 2022, 71, 889–898. [Google Scholar] [CrossRef]
  16. Cannatelli, R.; Parigi, T.L.; Iacucci, M.; Nardone, O.M.; Tontini, G.E.; Labarile, N.; Buda, A.; Rimondi, A.; Bazarova, A.; Bisschops, R.; et al. A virtual chromoendoscopy artificial intelligence system to detect endoscopic and histologic activity/remission and predict clinical outcomes in ulcerative colitis. Endoscopy 2022, 55, 332–341. [Google Scholar] [CrossRef]
  17. Iacucci, M.; Parigi, T.L.; Del Amor, R.; Meseguer, P.; Mandelli, G.; Bozzola, A.; Bazarova, A.; Bhandari, P.; Bisschops, R.; Danese, S.; et al. Artificial Intelligence Enabled Histological Prediction of Remission or Activity and Clinical Outcomes in Ulcerative Colitis. Gastroenterology 2023, 164, 1180–1188.e2. [Google Scholar] [CrossRef]
  18. Maeda, Y.; Kudo, S.-E.; Ogata, N.; Misawa, M.; Iacucci, M.; Homma, M.; Nemoto, T.; Takishima, K.; Mochida, K.; Miyachi, H.; et al. Evaluation in real-time use of artificial intelligence during colonoscopy to predict relapse of ulcerative colitis: A prospective study. Gastrointest. Endosc. 2022, 95, 747–756.e2. [Google Scholar] [CrossRef]
  19. Labarile, N.; Vitello, A.; Sinagra, E.; Nardone, O.M.; Calabrese, G.; Bonomo, F.; Maida, M.; Iacucci, M. Artificial Intelligence in Advancing Inflammatory Bowel Disease Management: Setting New Standards. Cancers 2025, 17, 2337. [Google Scholar] [CrossRef] [PubMed]
  20. Piccirelli, S.; Salvi, D.; Pugliano, C.L.; Tettoni, E.; Facciorusso, A.; Rondonotti, E.; Mussetto, A.; Fuccio, L.; Cesaro, P.; Spada, C. Unmet Needs of Artificial Intelligence in Small Bowel Capsule Endoscopy. Diagnostics 2025, 15, 1092. [Google Scholar] [CrossRef] [PubMed]
  21. Viechtbauer, W. Conducting Meta-Analyses in R with the metafor Package. J. Stat. Softw. 2010, 36, 1–48. [Google Scholar] [CrossRef]
  22. Higgins, J.P.T.; Thompson, S.G.; Deeks, J.J.; Altman, D.G. Measuring inconsistency in meta-analyses. BMJ 2003, 327, 557–560. [Google Scholar] [CrossRef]
  23. Whiting, P.F.; Rutjes, A.W.S.; Westwood, M.E.; Mallett, S.; Deeks, J.J.; Reitsma, J.B.; Leeflang, M.M.G.; Sterne, J.A.C.; Bossuyt, P.M.M.; QUADAS-2 Group. QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy studies. Ann. Intern. Med. 2011, 155, 529–536. [Google Scholar] [CrossRef]
  24. Afonso, J.; Mascarenhas, M.; Ribeiro, T.; Cardoso, H.; Andrade, P.; Ferreira, J.P.; Saraiva, M.M.; Macedo, G. Deep Learning for Automatic Identification and Characterization of the Bleeding Potential of Enteric Protruding Lesions in Capsule Endoscopy. Gastro Hep Adv. 2022, 1, 835–843. [Google Scholar] [CrossRef]
  25. Afonso, J.; Saraiva, M.M.; Ferreira, J.P.S.; Cardoso, H.; Ribeiro, T.; Andrade, P.; Parente, M.; Jorge, R.N.; Macedo, G. Automated detection of ulcers and erosions in capsule endoscopy images using a convolutional neural network. Med. Biol. Eng. Comput. 2022, 60, 719–725. [Google Scholar] [CrossRef]
  26. Alam, M.J.; Rashid RBin Fattah, S.A.; Saquib, M. RAt-CapsNet: A Deep Learning Network Utilizing Attention and Regional Information for Abnormality Detection in Wireless Capsule Endoscopy. IEEE J. Transl. Eng. Health Med. 2022, 10, 3300108. [Google Scholar] [CrossRef]
  27. Alaskar, H.; Hussain, A.; Al-Aseem, N.; Liatsis, P.; Al-Jumeily, D. Application of Convolutional Neural Networks for Automated Ulcer Detection in Wireless Capsule Endoscopy Images. Sensors 2019, 19, 1265. [Google Scholar] [CrossRef]
  28. Aoki, T.; Yamada, A.; Aoyama, K.; Saito, H.; Fujisawa, G.; Odawara, N.; Kondo, R.; Tsuboi, A.; Ishibashi, R.; Nakada, A.; et al. Clinical usefulness of a deep learning-based system as the first screening on small-bowel capsule endoscopy reading. Dig. Endosc. 2020, 32, 585–591. [Google Scholar] [CrossRef]
  29. Aoki, T.; Yamada, A.; Aoyama, K.; Saito, H.; Tsuboi, A.; Nakada, A.; Niikura, R.; Fujishiro, M.; Oka, S.; Ishihara, S.; et al. Automatic detection of erosions and ulcerations in wireless capsule endoscopy images based on a deep convolutional neural network. Gastrointest. Endosc. 2019, 89, 357–363.e2. [Google Scholar] [CrossRef]
  30. Aoki, T.; Yamada, A.; Kato, Y.; Saito, H.; Tsuboi, A.; Nakada, A.; Niikura, R.; Fujishiro, M.; Oka, S.; Ishihara, S.; et al. Automatic detection of blood content in capsule endoscopy images based on a deep convolutional neural network. J. Gastroenterol. Hepatol. 2020, 35, 1196–1200. [Google Scholar] [CrossRef] [PubMed]
  31. Aoki, T.; Yamada, A.; Oka, S.; Tsuboi, M.; Kurokawa, K.; Togo, D.; Tanino, F.; Teshima, H.; Saito, H.; Suzuki, R.; et al. Comparison of clinical utility of deep learning-based systems for small-bowel capsule endoscopy reading. J. Gastroenterol. Hepatol. 2024, 39, 157–164. [Google Scholar] [CrossRef] [PubMed]
  32. Barash, Y.; Azaria, L.; Soffer, S.; Yehuda, R.M.; Shlomi, O.; Ben-Horin, S.; Eliakim, R.; Klang, E.; Kopylov, U. Ulcer severity grading in video capsule images of patients with Crohn’s disease: An ordinal neural network solution. Gastrointest. Endosc. 2021, 93, 187–192. [Google Scholar] [CrossRef] [PubMed]
  33. Blanes-Vidal, V.; Baatrup, G.; Nadimi, E.S. Addressing priority challenges in the detection and assessment of colorectal polyps from capsule endoscopy and colonoscopy in colorectal cancer screening using machine learning. Acta Oncol. 2019, 58, S29–S36. [Google Scholar] [CrossRef]
  34. de Maissin, A.; Vallée, R.; Flamant, M.; Fondain-Bossiere, M.; Le Berre, C.; Coutrot, A.; Normand, N.; Mouchère, H.; Coudol, S.; Trang, C.; et al. Multi-expert annotation of Crohn’s disease images of the small bowel for automatic detection using a convolutional recurrent attention neural network. Endosc. Int. Open 2021, 9, E1136–E1144. [Google Scholar] [CrossRef] [PubMed]
  35. Ding, Z.; Shi, H.; Zhang, H.; Meng, L.; Fan, M.; Han, C.; Zhang, K.; Ming, F.; Xie, X.; Liu, H.; et al. Gastroenterologist-Level Identification of Small-Bowel Diseases and Normal Variants by Capsule Endoscopy Using a Deep-Learning Model. Gastroenterology 2019, 157, 1044–1054.e5. [Google Scholar] [CrossRef]
  36. Ferreira, J.P.S.; De Mascarenhas Saraiva, M.J.d.Q.e.C.; Afonso, J.P.L.; Ribeiro, T.F.C.; Cardoso, H.M.C.; Andrade, A.P.R.; Parente, M.P.L.; Jorge, R.N.; Lopes, S.I.O.; de Macedo, G.M.G. Identification of Ulcers and Erosions by the Novel PillcamTM Crohn’s Capsule Using a Convolutional Neural Network: A Multicentre Pilot Study. J. Crohns Colitis 2022, 16, 169–172. [Google Scholar] [CrossRef]
  37. Gan, T.; Yang, Y.; Liu, S.; Zeng, B.; Yang, J.; Deng, K.; Wu, J.; Yang, L. Automatic Detection of Small Intestinal Hookworms in Capsule Endoscopy Images Based on a Convolutional Neural Network. Gastroenterol. Res. Pract. 2021, 2021, 5682288. [Google Scholar] [CrossRef]
  38. Ghosh, T.; Chakareski, J. Deep Transfer Learning for Automated Intestinal Bleeding Detection in Capsule Endoscopy Imaging. J. Digit. Imaging 2021, 34, 404–417. [Google Scholar] [CrossRef] [PubMed]
  39. Guo, X.; Pang, L.; Chen, P.; Jiang, Q.; Zhong, Y. Deep ensemble framework with Bayesian optimization for multi-lesion recognition in capsule endoscopy images. Med. Biol. Eng. Comput. 2025, 63, 3037–3052. [Google Scholar] [CrossRef]
  40. Huang, Y.-H.; Lin, Q.; Jin, X.-Y.; Chou, C.-Y.; Wei, J.-J.; Xing, J.; Guo, H.-M.; Liu, Z.-F.; Lu, Y. Classification of pediatric video capsule endoscopy images for small bowel abnormalities using deep learning models. World J. Gastroenterol. 2025, 31, 107601. [Google Scholar] [CrossRef]
  41. Hwang, Y.; Lee, H.H.; Park, C.; Tama, B.A.; Kim, J.S.; Cheung, D.Y.; Chung, W.C.; Cho, Y.; Lee, K.; Choi, M.; et al. Improved classification and localization approach to small bowel capsule endoscopy using convolutional neural network. Dig. Endosc. 2021, 33, 598–607. [Google Scholar] [CrossRef]
  42. İncetan, K.; Celik, I.O.; Obeid, A.; Gokceler, G.I.; Ozyoruk, K.B.; Almalioglu, Y.; Chen, R.J.; Mahmood, F.; Gilbert, H.; Durr, N.J.; et al. VR-Caps: A Virtual Environment for Capsule Endoscopy. Med. Image Anal. 2021, 70, 101990. [Google Scholar] [CrossRef] [PubMed]
  43. Klang, E.; Barash, Y.; Margalit, R.Y.; Soffer, S.; Shimon, O.; Albshesh, A.; Ben-Horin, S.; Amitai, M.M.; Eliakim, R.; Kopylov, U. Deep learning algorithms for automated detection of Crohn’s disease ulcers by video capsule endoscopy. Gastrointest. Endosc. 2020, 91, 606–613.e2. [Google Scholar] [CrossRef]
  44. Kwon, Y.S.; Park, T.Y.; Kim, S.E.; Park, Y.; Lee, J.G.; Lee, S.P.; Kim, K.O.; Jang, H.J.; Yang, Y.J.; Cho, B.-J. Deep learning-based localization and lesion detection in capsule endoscopy for patients with suspected small-bowel bleeding. World J. Gastroenterol. 2025, 31, 106819. [Google Scholar] [CrossRef]
  45. Lafraxo, S.; Souaidi, M.; El Ansari, M.; Koutti, L. Semantic Segmentation of Digestive Abnormalities from WCE Images by Using AttResU-Net Architecture. Life 2023, 13, 719. [Google Scholar] [CrossRef] [PubMed]
  46. Li, L.; Yang, L.; Zhang, B.; Yan, G.; Bao, Y.; Zhu, R.; Li, S.; Wang, H.; Chen, M.; Jin, C.; et al. Automated detection of small bowel lesions based on capsule endoscopy using deep learning algorithm. Clin. Res. Hepatol. Gastroenterol. 2024, 48, 102334. [Google Scholar] [CrossRef] [PubMed]
  47. Li, Q.; Xie, W.M.; Wang, Y.; Qin, K.; Huang, M.; Liu, T.M.; Chen, Z.M.; Chen, L.; Teng, L.; Fang, Y.; et al. A Deep Learning Application of Capsule Endoscopic Gastric Structure Recognition Based on a Transformer Model. J. Clin. Gastroenterol. 2024, 58, 937–943. [Google Scholar] [CrossRef] [PubMed]
  48. Li, X.; Gan, Y.; Duan, D.; Yang, X. Toward automatic and reliable evaluation of human gastric motility using magnetically controlled capsule endoscope and deep learning. Sci. Rep. 2025, 15, 25955. [Google Scholar] [CrossRef]
  49. Nadimi, E.S.; Braun, J.-M.; Schelde-Olesen, B.; Khare, S.; Gogineni, V.C.; Blanes-Vidal, V.; Baatrup, G. Towards full integration of explainable artificial intelligence in colon capsule endoscopy’s pathway. Sci. Rep. 2025, 15, 5960. [Google Scholar] [CrossRef]
  50. Nam, S.J.; Moon, G.; Park, J.H.; Kim, Y.; Lim, Y.J.; Choi, H.S. Deep Learning-Based Real-Time Organ Localization and Transit Time Estimation in Wireless Capsule Endoscopy. Biomedicines 2024, 12, 1704. [Google Scholar] [CrossRef]
  51. Oukdach, Y.; Garbaz, A.; Kerkaou, Z.; El Ansari, M.; Koutti, L.; Papachrysos, N.; El Ouafdi, A.F.; de Lange, T.; Distante, C. Vision transformer distillation for enhanced gastrointestinal abnormality recognition in wireless capsule endoscopy images. J. Med. Imaging 2025, 12, 014505. [Google Scholar] [CrossRef]
  52. Pinto, L.; Figueiredo, I.N.; Figueiredo, P.N. Reducing reading time and assessing disease in capsule endoscopy videos: A deep learning approach. Int. J. Med. Inform. 2025, 195, 105792. [Google Scholar] [CrossRef]
  53. Ribeiro, T.; Saraiva, M.J.M.; Afonso, J.; Cardoso, P.; Mendes, F.; Martins, M.; Andrade, A.P.; Cardoso, H.; Saraiva, M.M.; Ferreira, J.; et al. Design of a Convolutional Neural Network as a Deep Learning Tool for the Automatic Classification of Small-Bowel Cleansing in Capsule Endoscopy. Medicina 2023, 59, 810. [Google Scholar] [CrossRef]
  54. Saraiva, M.J.M.; Afonso, J.; Ribeiro, T.; Ferreira, J.; Cardoso, H.; Andrade, A.P.; Parente, M.; Natal, R.; Saraiva, M.M.; Macedo, G. Deep learning and capsule endoscopy: Automatic identification and differentiation of small bowel lesions with distinct haemorrhagic potential using a convolutional neural network. BMJ Open Gastroenterol. 2021, 8, e000753. [Google Scholar] [CrossRef] [PubMed]
  55. Saraiva, M.J.M.; Afonso, J.; Ribeiro, T.; Cardoso, P.; Mendes, F.; Martins, M.; Andrade, A.P.; Cardoso, H.; Saraiva, M.M.; Ferreira, J.; et al. AI-Driven Colon Cleansing Evaluation in Capsule Endoscopy: A Deep Learning Approach. Diagnostics 2023, 13, 3494. [Google Scholar] [CrossRef] [PubMed]
  56. Saraiva, M.M.; Ribeiro, T.; Afonso, J.; Ferreira, J.P.; Cardoso, H.; Andrade, P.; Parente, M.P.; Jorge, R.N.; Macedo, G. Artificial Intelligence and Capsule Endoscopy: Automatic Detection of Small Bowel Blood Content Using a Convolutional Neural Network. GE Port J. Gastroenterol. 2021, 29, 331–338. [Google Scholar] [CrossRef]
  57. Saraiva, M.M.; Ribeiro, T.; Afonso, J.; Andrade, P.; Cardoso, P.; Ferreira, J.; Cardoso, H.; Macedo, G. Deep Learning and Device-Assisted Enteroscopy: Automatic Detection of Gastrointestinal Angioectasia. Medicina 2021, 57, 1378. [Google Scholar] [CrossRef]
  58. Saraiva, M.M.; Ferreira, J.P.S.; Cardoso, H.; Afonso, J.; Ribeiro, T.; Andrade, P.; Parente, M.P.L.; Jorge, R.N.; Macedo, G. Artificial intelligence and colon capsule endoscopy: Automatic detection of blood in colon capsule endoscopy using a convolutional neural network. Endosc. Int. Open 2021, 9, E1264–E1268. [Google Scholar] [CrossRef]
  59. Mascarenhas, M.; Ribeiro, T.; Afonso, J.; Ferreira, J.P.; Cardoso, H.; Andrade, P.; Parente, M.P.; Jorge, R.N.; Saraiva, M.M.; Macedo, G. Deep learning and colon capsule endoscopy: Automatic detection of blood and colonic mucosal lesions using a convolutional neural network. Endosc. Int. Open 2022, 10, E171–E177. [Google Scholar] [CrossRef]
  60. Spada, C.; Piccirelli, S.; Hassan, C.; Ferrari, C.; Toth, E.; González-Suárez, B.; Keuchel, M.; McAlindon, M.; Finta, Á.; Rosztóczy, A.; et al. AI-assisted capsule endoscopy reading in suspected small bowel bleeding: A multicentre prospective study. Lancet Digit. Health 2024, 6, e345–e353. [Google Scholar] [CrossRef]
  61. Su, Q.; Wang, F.; Chen, D.; Chen, G.; Li, C.; Wei, L. Deep convolutional neural networks with ensemble learning and transfer learning for automated detection of gastrointestinal diseases. Comput. Biol. Med. 2022, 150, 106054. [Google Scholar] [CrossRef]
  62. Xie, X.; Xiao, Y.-F.; Yang, H.; Peng, X.; Li, J.-J.; Zhou, Y.-Y.; Fan, C.-Q.; Meng, R.-P.; Huang, B.-B.; Liao, X.-P.; et al. A new artificial intelligence system for both stomach and small-bowel capsule endoscopy. Gastrointest. Endosc. 2024, 100, 878.e1–878.e14. [Google Scholar] [CrossRef] [PubMed]
  63. Xie, X.; Xiao, Y.-F.; Zhao, X.-Y.; Li, J.-J.; Yang, Q.-Q.; Peng, X.; Nie, X.-B.; Zhou, J.-Y.; Zhao, Y.-B.; Yang, H.; et al. Development and Validation of an Artificial Intelligence Model for Small Bowel Capsule Endoscopy Video Review. JAMA Netw. Open 2022, 5, E2221992. [Google Scholar] [CrossRef]
  64. Xu, T.; Li, Y.-Y.; Huang, F.; Gao, M.; Cai, C.; He, S.; Wu, Z.-X. A Multi-task Neural Network for Image Recognition in Magnetically Controlled Capsule Endoscopy. Dig. Dis. Sci. 2024, 69, 4231–4239. [Google Scholar] [CrossRef]
  65. Yogapriya, J.; Chandran, V.; Sumithra, M.G.; Anitha, P.; Jenopaul, P.; Dhas, C.S.G. Gastrointestinal Tract Disease Classification from Wireless Endoscopy Images Using Pretrained Deep Learning Model. Comput. Math. Methods Med. 2021, 2021, 5940433. [Google Scholar] [CrossRef]
  66. Zhang, R.-Y.; Qiang, P.-P.; Cai, L.-J.; Li, T.; Qin, Y.; Zhang, Y.; Zhao, Y.-Q.; Wang, J.-P. Automatic detection of small bowel lesions with different bleeding risks based on deep learning models. World J. Gastroenterol. 2024, 30, 170–183. [Google Scholar] [CrossRef] [PubMed]
  67. Islam, K. Recent Advances in Vision Transformer: A Survey and Outlook of Recent Work. arXiv 2022, arXiv:2203.01536. [Google Scholar]
  68. Arnab, A.; Dehghani, M.; Heigold, G.; Sun, C.; Lucic, M.; Schmid, C. ViViT: A Video Vision Transformer. In Proceedings of the IEEE International Conference on Computer Vision 2021, Montreal, BC, Canada, 10–17 October 2021; pp. 6816–6826. [Google Scholar] [CrossRef]
  69. Salvi, D.; Piccirelli, S.; Parmigiani, M.; Cesaro, P.; Spada, C. How to measure quality in capsule endoscopy. Best Pract. Res. Clin. Gastroenterol. 2025, 76, 102012. [Google Scholar] [CrossRef] [PubMed]
  70. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef]
Figure 1. PRISMA flow diagram of study selection. * number of records identified from each database. ** records excluded by a human check.
Figure 1. PRISMA flow diagram of study selection. * number of records identified from each database. ** records excluded by a human check.
Applsci 16 01134 g001
Figure 2. Forest plot showing accuracy estimates for the 44 included studies. Boxes represent individual study estimates and horizontal lines indicate 95% confidence intervals. The diamond denotes the pooled accuracy estimate.
Figure 2. Forest plot showing accuracy estimates for the 44 included studies. Boxes represent individual study estimates and horizontal lines indicate 95% confidence intervals. The diamond denotes the pooled accuracy estimate.
Applsci 16 01134 g002
Figure 3. Forest plot of the nine studies reporting capsule endoscopy reading time. The plot shows mean differences in reading time between standard interpretation and neural network-assisted reading, with 95% confidence intervals (horizontal lines). Pooled mean difference is reported as a diamond.
Figure 3. Forest plot of the nine studies reporting capsule endoscopy reading time. The plot shows mean differences in reading time between standard interpretation and neural network-assisted reading, with 95% confidence intervals (horizontal lines). Pooled mean difference is reported as a diamond.
Applsci 16 01134 g003
Table 1. Summary information of the studies included in the meta-analysis.
Table 1. Summary information of the studies included in the meta-analysis.
AuthorYearCountryValidation (V) or Clinical Study (C)Study DesignMulticenter StudyType of Neural NetworkType of ArchitecturePurpose
of NN
Application [Upper GI, Small Bowel, Colon]Type of LesionSample Size PatientsSample Test £Mean SRSD SRMean AIRSD AIRTP + TN
Afonso et al. [24]2022PortugalVPyesXceptionCNNclassificationsmall bowelOther lesion & 4264 4136
Afonso et al. [25]2022PortugalVRnoXceptionCNNclassificationsmall bowel 1226 1172
Alam et al. [26]2022BangladeshV RAt-CapsNetCapsule networkclassificationGI tract 9447 9306
Alaskar et al. [27]2019Saudi ArabiaV AlexNetCNNclassificationGI tractUlcers 105 105
Aoki et al. [28]2020JapanVRnoResNetCNNclassificationcolonOther lesion & 10,208 10,197
Aoki et al. [29]2019JapanVRyes small bowelUlcers 10,440 9480
Aoki et al. [30]2024JapanCRyesResNetCNNclassificationsmall bowel 36 10.15.033.616.8
Aoki et al. [31]2020JapanCRyes small bowel 20 4.82.417.88.9
Barash et al. [32]2021IsraelVRno small bowelUlcers49248 226
Blanes-Vidal et al. [33]2019DenmarkVRnoAlexNetCNNclassificationcolon 1695 1634
de Maissin et al. [34]2021FranceVRyes small bowelUlcers 350 326
Ding et al. [35]2019ChinaCRyes small bowel 697042065.92.296.622.5
Ferreira et al. [36]2022PortugalVRyesXceptionCNNclassificationsmall bowelUlcers 4935 4560
Gan et al. [37]2021ChinaVRnoYOLOCNNobject detectionsmall bowel 10,529 9602
Ghosh et al. [38]2021USAV AlexNetCNNclassificationsmall bowelOther lesion & 96 95
Guo et al. [39]2025ChinaV EfficientNetCNNclassificationGI tract 867 731
Huang et al. [40]2025ChinaV noVGGNetCNNclassificationsmall bowel 458 415
Hwang et al. [41]2020KoreaVRnoVGGNetCNNclassificationsmall bowel 5265760 5577
Incetan et al. [42]2021TurkeyV ResNetCNNclassificationGI tract 800 736
Klang et al. [43]2020IsraelVRno small bowelUlcers 3528 3412
Kwon et al. [44]2025South KoreaCRyesDenseNetCNNclassificationGI tractOther lesion &32328.74.353.926.9
Lafraxo et al. [45]2023MoroccoV U-NetCNNsegmentationGI tractOther lesion & 652 647
Li et al. [46]2024ChinaVRyesViTTNNclassificationupper GI 118 118
Li et al. [47]2025ChinaV upper GI 11 10
Li et al. [48]2024ChinaVRyesYOLOCNNobject detectionsmall bowel 2985.62.833.026.7264
Nadimi et al. [49]2025DenmarkV CartoonX23CNNclassification/object detectioncolonOther lesion & 5838 5137
Nam et al. [50]2024South KoreaV ResNetCNNclassificationGI tract 72 70
Oukdach et al. [51]2025MoroccoV ViTTNNclassificationGI tract 1049 1018
Pinto et al. [52]2025PortugalV AlexNetCNNclassificationGI tract 87010.05.058.029.0
Ribeiro et al. [53]2023PortugalVR RegNet YCNNclassificationsmall bowel 791 729
Saraiva et al. [54]2021PortugalVRnoXceptionCNNclassificationcolonOther lesion & 728 671
Saraiva et al. [55]2023PortugalV yesResNetCNNclassificationcolon 6725 6389
Saraiva et al. [56]2022PortugalVPyesXceptionCNNclassificationcolonOther lesion & 1143 1089
Saraiva et al. [57]2021PortugalVRnoXceptionCNNclassificationsmall bowelOther lesion & 6136 6044
Saraiva et al. [58]2021PortugalVPno small bowel 1348 1285
Saraiva et al. [59]2021PortugalVRnoXceptionCNNclassificationcolonOther lesion & 1165 1125
Spada et al. [60]2024ItalyCPyes small bowel 133 3.83.333.722.9
Su et al. [61]2022ChinaV XceptionCNNclassificationGI tract 1600 1517
Xie et al. [62]2022ChinaCPyes small bowelOther lesion &2927 5.41.551.411.6
Xie et al. [63]2024ChinaVRyes GI tract 342 9.94.980.840.4
Xu et al. [64]2024ChinaV YOLOCNNobject detectionupper GIOther lesion & 208 207
Yogapriya et al. [65]2021IndiaV VGGNetCNNclassificationGI tract 6407 6174
Zhang et al. [66]2024ChinaV noResNetCNNclassificationsmall bowel 70137,287 36,899
& Other lesion: protruding lesion, blood/bleeding, or angiodysplasia AVM; £ Number of images/videos used as a test dataset to evaluate the neural network’s performance. V: validation study; C: clinical study; R: retrospective study; P: prospective study; CNN: convolutional neural network; TNN: transformer neural network; GI Tract: gastrointestinal tract; Upper GI: upper gastrointestinal; SR: standard reading; AIR: artificial intelligence reading.
Table 2. Subgroup analysis for accuracy outcome.
Table 2. Subgroup analysis for accuracy outcome.
Study SubgroupsN of StudiesAccuracy Estimate95% ICp-ValueHeterogeneityPublication Bias
Group HeterogeneityBegg’s Test
I2Qdf (Q)p-ValueTaup-Value
Total360.95[0.94, 0.96] 99.804887.0635<0.0010.1170.322
Country Group
Europe130.94[0.92, 0.96]0.25498.08768.8312<0.001−0.2300.306
Non-Europe230.95[0.94, 0.97] 99.893368.0422<0.0010.1460.345
Year
2019–2022230.95[0.94, 0.96]0.50399.693758.4122<0.0010.2010.188
2023–2025130.94[0.91, 0.97] 99.341092.7312<0.001−0.0250.952
Site Comparison 1
Colon70.94[0.91, 0.97]0.91299.321296.696<0.001−0.0471.000
Small Bowel170.94[0.93, 0.96] 99.692013.5216<0.001−0.0580.776
Site Comparison 2
Colon70.94[0.91, 0.97]<0.00199.321296.696<0.001−0.0471.000
Upper GI30.99[0.98, 1.00] 100.001.3220.515−0.3331.000
Type of Lesion
Ulcers60.94[0.91, 0.96]0.17598.24295.645<0.0010.0661.000
No Ulcers100.96[0.94, 0.98] 99.601099.919<0.0010.0221000
Type of Architecture
CNN260.93[0.89, 0.97]0.01599.973919.2925<0.001−0.0210.895
Other30.98[0.97, 0.99] 88.0011.0220.0040.3331000
Purpose of NN
Classification240.96[0.94, 0.97]0.26599.761974.6523<0.0010.1230.4172
Other50.93[0.88, 0.98] 99.44677.994<0.0010.0001.000
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Salvi, D.; Zani, C.; Spada, C.; Piccirelli, S.; Zileri Dal Verme, L.; Tripodi, G.; Gualtieri, L.; Cesaro, P.; Ferrari, C. Neural Network Architectures in Video Capsule Endoscopy: A Systematic Review and Meta-Analysis on Accuracy and Reading Time Performances. Appl. Sci. 2026, 16, 1134. https://doi.org/10.3390/app16021134

AMA Style

Salvi D, Zani C, Spada C, Piccirelli S, Zileri Dal Verme L, Tripodi G, Gualtieri L, Cesaro P, Ferrari C. Neural Network Architectures in Video Capsule Endoscopy: A Systematic Review and Meta-Analysis on Accuracy and Reading Time Performances. Applied Sciences. 2026; 16(2):1134. https://doi.org/10.3390/app16021134

Chicago/Turabian Style

Salvi, Daniele, Chiara Zani, Cristiano Spada, Stefania Piccirelli, Lorenzo Zileri Dal Verme, Giulia Tripodi, Loredana Gualtieri, Paola Cesaro, and Clarissa Ferrari. 2026. "Neural Network Architectures in Video Capsule Endoscopy: A Systematic Review and Meta-Analysis on Accuracy and Reading Time Performances" Applied Sciences 16, no. 2: 1134. https://doi.org/10.3390/app16021134

APA Style

Salvi, D., Zani, C., Spada, C., Piccirelli, S., Zileri Dal Verme, L., Tripodi, G., Gualtieri, L., Cesaro, P., & Ferrari, C. (2026). Neural Network Architectures in Video Capsule Endoscopy: A Systematic Review and Meta-Analysis on Accuracy and Reading Time Performances. Applied Sciences, 16(2), 1134. https://doi.org/10.3390/app16021134

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop