Next Article in Journal
Ototoxicity Associated with Antineoplastic Agents in the Pediatric Population: An Evidence-Based Review of Auditory Monitoring Strategies and Contemporary Diagnostic Frameworks—Narrative Review
Previous Article in Journal
Multiplanar AS-OCT Detection of Clinically Occult Posterior Gas Bubble Dislocation After DSAEK
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Artificial Intelligence in Gastrointestinal Wireless Capsule Endoscopy: A Systematic Literature Review and Meta-Analysis

1
Institute of Mechanical and Electrical Engineering, University of Southern Denmark, 5230 Odense, Denmark
2
Department of Clinical Research, University of Southern Denmark, 5000 Odense, Denmark
3
Department of Biology, University of Southern Denmark, 5230 Odense, Denmark
*
Author to whom correspondence should be addressed.
Diagnostics 2026, 16(9), 1269; https://doi.org/10.3390/diagnostics16091269
Submission received: 19 March 2026 / Revised: 10 April 2026 / Accepted: 12 April 2026 / Published: 23 April 2026
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)

Abstract

Background: Wireless capsule endoscopy is widely used for diagnosing gastrointestinal diseases, but manual interpretation of capsule videos is time-consuming and can vary between clinicians. Artificial intelligence has been increasingly studied to support capsule analysis and reduce clinical workload. This systematic literature review and meta-analysis summarizes current evidence on artificial intelligence methods applied to wireless capsule endoscopy, with a focus on diagnostic performance, validation strategies, and clinical readiness. Methods: A systematic search was conducted in PubMed, Scopus, Embase, Web of Science, and Google Scholar. Original journal articles were included based on predefined eligibility criteria. The reviewed studies addressed multiple artificial intelligence tasks, including detection, classification, segmentation, and localization of gastrointestinal abnormalities. Results: A total of 72 studies were included. Meta-analysis using random effects models showed high pooled diagnostic performance across clinical indications and gastrointestinal tract locations, with the strongest results reported for bleeding and vascular lesions and more variable performance for inflammatory bowel disease and mixed abnormality categories. The review also identified important clinical and technical barriers that may limit reliability and slow clinical adoption. These included limited external validation, small patient cohorts, retrospective study designs, and inconsistent reporting and evaluation practices. Conclusions: Artificial intelligence methods show strong potential to support wireless capsule endoscopy interpretation. Based on the findings, we propose practical recommendations to improve study design and validation. If these recommendations are applied, future studies may report more robust and reliable results, supporting better translation into clinical workflows.

1. Introduction

1.1. The Global Burden of Gastrointestinal Disorders

The global pattern of Gastrointestinal (GI) diseases is changing rapidly due to a growing wave of cancers, inflammatory conditions, and functional disorders. This trend poses a serious challenge to healthcare systems in many regions. The burden is twofold as developed countries continue to report high rates of GI cancers, while newly industrialized nations are seeing a sharp rise in cases, mainly driven by changes in population structure and lifestyle habits [1,2].
Colorectal Cancer (CRC) and Inflammatory Bowel Disease (IBD) represent the most significant drivers of this expanding burden. CRC remains the third most frequently diagnosed cancer worldwide [2,3], while IBD is rapidly spreading beyond western countries into newly industrialized regions due to urbanization and lifestyle changes [4,5]. Together, these conditions create a critical need for long-term chronic care and regular endoscopic surveillance. This places immense pressure on healthcare systems to improve early diagnosis and match capacity with rising demand.
Digestive diseases also impose a major clinical and financial burden on healthcare systems. In 2019, around 332 million people in Europe were living with digestive conditions, which resulted in 498,000 GI-related deaths and over €20 billion in inpatient care costs annually along with significant indirect losses due to reduced productivity [6]. The rising demand to investigate symptoms like Obscure Gastrointestinal Bleeding (OGIB), iron-deficiency anemia, and chronic abdominal pain has added further pressure on endoscopy procedures [7].

1.2. Wireless Capsule Endoscopy

Wireless Capsule Endoscopy (WCE), introduced in 2000, is a breakthrough imaging technology that has transformed the diagnosis of GI disorders [8]. Unlike conventional endoscopy or radiology, WCE allows for direct visualization of the entire GI tract using a swallowable capsule with a tiny camera, light source, battery, and wireless transmitter (Figure 1). Over the past two decades, WCE has become a core tool in gastroenterology and is now widely used as a first-line method for small bowel investigation. Ongoing advancements in both hardware and software have significantly improved its imaging performance [9]. Modern capsules now include higher-resolution sensors, wider fields of view, and extended battery life, typically operating 8 to 12 h and capturing 50,000 to 60,000 images per exam [10]. Newer models also offer features such as adaptive frame rates, enhanced optics, and broader viewing angles, which contribute to better image quality and more reliable lesion detection. In addition, specialized capsule designs have been introduced for the esophagus, colon, and stomach, expanding the diagnostic reach of WCE [11].
From a clinical perspective, WCE has significantly improved the detection and assessment of GI diseases. Its primary indication is the investigation of OGIB and unexplained iron deficiency anemia. In these cases, WCE enables full visualization of the small bowel mucosa and often reveals underlying lesions such as angioectasias, ulcers, or tumors that were missed by conventional methods [8]. WCE is also a valuable tool in the diagnosis of Crohn’s and Ulcerative Colitis (UC), as it can detect inflammatory lesions such as aphthous ulcers or erosions in the small bowel that are beyond the reach of ileocolonoscopy. This is especially useful for patients with suggestive symptoms but normal findings on standard endoscopy. Additionally, WCE plays an important role in detecting small bowel tumors and in surveillance of polyposis syndromes like Peutz–Jeghers syndrome by offering a noninvasive method to examine the entire intestine for polyps or neoplastic lesions [12]. It is also a convenient alternative for patients who are unable to undergo sedation or who have experienced an incomplete colonoscopy [8]. Notably, recent developments are shifting the field toward pan-enteric capsules that image both the small bowel and colon in a single examination. This supports a broader move toward whole-gut assessment.

1.3. The Limitations of Human Interpretation

WCE generates a high-volume video stream consisting of tens of thousands of frames over several hours that requires manual review. Interpreting a single examination typically takes around 45 min of physician time [13]. Since only a small fraction of frames contain relevant pathology, it is not uncommon for clinicians to overlook subtle lesions or misclassify normal findings. Retrospective studies have reported notable miss rates as human readers failed to detect 5.9% of vascular lesions, 0.5% of ulcers, and up to 18.9% of neoplastic lesions on WCE [14]. These omissions often result from the difficulty of identifying small or subtle abnormalities within largely normal mucosa, as well as from considerable inter-observer variability. For example, a study by the I-CARE Group [15] reported that only 23% of inter-reader comparisons achieved good or near-perfect agreement, which highlights inconsistencies across different reviewers. Diagnostic accuracy is further diminished by fatigue, which can reduce performance after reviewing even one full-length recording [15]. These human limitations underscore the need for assistive technologies that can enhance both efficiency and consistency in WCE interpretation.

1.4. The Promise of Artificial Intelligence

Artificial Intelligence (AI), particularly deep learning, has emerged as a promising approach to overcome the challenges associated with WCE interpretation. AI-based systems can operate as a second reader or even perform an initial screening of WCE data to reduce workload and improve diagnostic accuracy. Several AI-assisted platforms are now capable of automatically detecting frames with potential lesions while skipping those without relevant findings, significantly reducing the manual review burden [16]. For example, proprietary capsule software modes such as QuickView or ExpressView use algorithms to highlight suspected pathology and filter out normal frames. In practice, these tools have reduced the number of frames requiring clinician review by up to 20-fold [13].
Convolutional Neural Networks (CNNs) are among the first successful applications that enable automated detection of common abnormalities such as ulcers, angioectasias, polyps, and bleeding with high accuracy [17,18,19,20,21]. CNN-based algorithms consistently achieve classification accuracies above 90%, with pooled sensitivities between 95% and 98% reported for ulcer and bleeding detection [17]. This level of performance has been reported to be comparable to expert endoscopists in specific studies and may represent an important advance in reducing diagnostic errors. In addition to classification, more advanced computer vision models have been used to localize lesions within individual frames. Single-shot object detectors such as You Only Look Once (YOLO) networks [22] and Single Shot MultiBox Detector (SSD) [23] have been adapted for WCE images to draw bounding boxes around abnormalities and mark their precise locations. These models support real-time detection of lesions such as polyps or bleeding during video playback.
The latest generation of deep learning models, Vision Transformers (ViTs), has shown strong potential in WCE applications. Unlike CNNs, which extract localized features through convolutional filters, ViTs use self-attention mechanisms to capture long-range dependencies and global context across the entire image. This ability to model full-frame relationships offers a distinct advantage in WCE, where lesions often appear subtle, dispersed, or embedded within a broader mucosal background. Recent studies have demonstrated that ViT-based architectures can match or surpass CNNs in detecting GI lesions such as bleeding, ulcers, and polyps [24]. ViTs also show improved robustness to image variability and stronger generalization across datasets and imaging conditions. These strengths make ViTs well-suited for the large-scale, visually complex image streams generated by WCE [24].
In summary, AI methods achieve high performance in detecting GI lesions on WCE. In research studies, their performance often matches or surpasses expert readers. These systems can identify subtle abnormalities that may go unnoticed in manual review and reduce the time required for interpretation.

2. Research Gaps and Objectives

The contributions of this study are threefold:
(1)
Almost all review articles on the use of AI in WCE applications are restricted to specific settings, such as particular diseases, GI locations, or predefined applications. To address this limitation, in the current study we extracted all original journal research articles that applied AI for the analysis of WCE outputs. In other words, our search strategy was not restricted to specific diseases, applications, or GI locations.
(2)
Moreover, almost all studies in this domain have largely focused on analyzing and comparing the performance metric values. However, several critical technical and clinical barriers must be considered to improve the trustworthiness of AI results and to increase the likelihood of successful deployment of AI systems in clinical practice. Therefore, in this study, we go beyond performance metrics and highlight these barriers and provide recommendations for future research.
(3)
Finally, this study presents a more comprehensive meta-analysis by systematically structuring the results across two complementary dimensions including clinical indication–based domains and GI tract locations. The included studies were first grouped into five categories based on their clinical indications, and subsequently reclassified into four categories according to the GI tract locations.

3. Method

This systematic review was conducted in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) [25] guidelines and focused on the questions listed in Table 1.

3.1. Search Strategy and Study Selection

We searched four databases, including PubMed, Scopus, Embase (Ovid), and Web of Science. We also searched for articles on Google Scholar. To find relevant studies, we used a mixure of keywords and filters organized into five categories as AI terms, technical terms, medical terms, document type, publication year, and language. Keywords within each category were connected using OR, while the categories were combined using AND. Table 2 shows the search criteria of this study. The full queries for databases are presented in Supplementary File S1.
We considered different criteria to include or exclude of studies. These conditions are presented in Table 3.

3.2. Data Extraction

Two researchers (AN, AS) individually examined the titles and abstracts using the Covidence tool. Subsequently, they conducted a full-text review, resolving disagreements with a senior researcher (AK), who decided on the inclusion or exclusion of the article. The extracted data was compiled into spreadsheets using the Critical Appraisal and Data Extraction for Systematic Reviews (CHARMS) checklist [26].

3.3. Risk of Bias Assessment

Risk of Bias (ROB) assessment of each study was investigated using Quality Assessment of Diagnostic Accuracy Studies 2 (QUADAS-2) tool [27]. It evaluates four key domains: Patient Selection, Index Test, Reference Standard, and Flow and Timing. QUADAS-2 ensures the reliability and generalizability of study findings by assigning three levels of bias as Low, Unclear, and High to each domain for each study.

3.4. Meta-Analysis

In this study, we conducted a meta-analysis using a random-effects model to synthesize findings across multiple studies while accounting for heterogeneity [28]. We assessed heterogeneity using Cochran’s Q, I2, and τ 2 , ensuring that our pooled results reflect study variability and enhance the robustness of our conclusions. Four commonly reported performance metrics, including accuracy, sensitivity, specificity, and Area Under the Curve (AUC) were considered to quantitatively synthesize the diagnostic performance of AI models across studies.
To perform the meta-analysis, two complementary stratification strategies were considered. First, a clinical indication-based approach, which groups studies according to the targeted GI pathology, was implemented. Then, a GI tract location-based approach, which categorizes studies according to the anatomical region of the GI tract, was investigated.
In the first scenario, studies were categorized into five clinically meaningful groups based on the primary GI pathology. IBD (Crohn’s disease, UC) included studies focusing on IBD activity, severity assessment, or disease-related lesions. Polyp, Tumor, Protruded Lesion comprised studies targeting protruding mucosal abnormalities. Ulcer, Erosion, Inflammation included studies detecting mucosal breaks and inflammatory lesions not specific to IBD. Bleeding, Vascular covered studies addressing GI bleeding and vascular abnormalities. Mixed, General Abnormality referred to studies detecting multiple lesion types or general abnormalities without a single dominant focus.
To ensure consistency across studies with different reporting styles, outcome definitions were harmonized by grouping similar lesion types into these broader categories based on their clinical and visual similarity in capsule endoscopy. This approach reduced variability in outcome definitions and allowed more reliable pooling of performance metrics using random-effects models.
In the second scenario, studies were categorized based on the anatomical region of the GI tract. Four main groups were defined. The Small Bowel category included studies focusing on small bowel capsule endoscopy. The Colon category included studies targeting colonic evaluation. The Entire GI Tract category referred to studies analyzing the full GI tract without restriction to a specific region. The Different Parts of GI Tract category included studies that evaluated multiple anatomical regions separately or in combination. This categorization allowed comparison of AI performance across different GI locations while considering differences in visual features and disease distribution.

4. Results

The initial search across all databases yielded 1451 records after automatic duplicate removal using Covidence. During subsequent manual screening, 119 additional duplicate records were identified that had not been detected by Covidence and were therefore removed. After eligibility screening, full-text assessment, quality appraisal, and reference checking of included studies, a total of 72 studies were finally included. The PRISMA chart of the study selection process is illustrated in Figure 2. The PRISMA checklist is also presented as Supplementary File S2.
The extracted data from final set of studies is presented in two tables in the Appendix A as Table A1 and Table A2. Table A1 summarizes the general characteristics of the studies, while Table A2 presents the AI models and performance metrics reported in the included studies. The complete CHARMS data extraction spreadsheet is also provided as Supplementary File S3.

4.1. Geographical Distribution

Figure 3, derived from Table A1, visualizes the international distribution of studies on AI applications in WCE. As seen, research output is heavily concentrated in East Asia and Europe. China shows the highest research output (n = 14), followed by Portugal (n = 12) and Japan (n = 11). Other countries with notable research activity include South Korea (n = 6), Denmark (n = 5), and Israel (n = 5). Additional research activity is observed in several European and North American countries, including the UK (n = 4), Spain (n = 4), France (n = 4), and the USA (n = 4). Further studies are published from Pakistan and Norway (each n = 3).

4.2. Publication Trend

Figure 4, extracted from Table A1, shows the distribution of included studies by publication year. The number of studies increased steadily after 2019, with a clear rise from 2020 onward and a peak in 2021. Although minor fluctuations are observed in subsequent years, the overall trend indicates growing research interest in AI applications for capsule endoscopy in recent years.

4.3. Cohort Presentation

The number of patients included in each study is one of the important cohort-related aspect. This parameter can indirectly reflect the robustness of the findings [29]. Figure 5, extracted from Table A1, presents the distribution of patient numbers for 28 studies that reported this information. The upper panel shows a histogram with an overlaid Kernel Density Estimation (KDE) curve. The lower panel highlights substantial heterogeneity in cohort sizes, ranging from 21 to 6970 patients. The median and mean numbers of patients across the included studies were 174 and 949, respectively. However, the majority of studies (53.6%) included fewer than 400 patients, indicating that most investigations in this field are small to moderate in size.
In addition to the number of patients, the total number of samples in each dataset is an important indicator of the reliability of AI results. Sample size information was available for 55 studies, and Figure 6 shows the distribution of sample sizes. The upper panel displays a histogram with logarithmic x-axis scaling and an overlaid KDE curve illustrating the distribution of sample sizes. The lower panel presents a box plot with logarithmic scaling that reveals extreme variability spanning seven orders of magnitude, from 36 to 10,000,000 samples. The median and mean numbers of samples across the included studies were 17,640 and 434,058, respectively. The distribution shows a substantial right skew, driven by a small number of large-scale studies using extensive retrospective databases, while most studies (over 70%) analyzed fewer than 50,000 capsule endoscopy samples.

4.4. GI Tract Locations

Based on Table A1 and as shown in Figure 7a, the majority of included studies focused on the small bowel (55.6%), followed by entire GI tract (20.6%), colon (14.3%), mixed regions (7.9%), and stomach (1.6%). This distribution indicates a notable emphasis on small bowel research compared to other regions of the GI tract.

4.5. Applications

AI models have been applied to WCE across four primary tasks, including classification, segmentation, object detection, and localization (Figure 7b and Table A2). Among these, classification is the most widely explored application, accounting for 65.1% of studies. In classification tasks, a key application is the automatic filtering of abnormal images from the large number of frames generated during WCE procedures. Since a single examination can produce tens of thousands of images, AI-based methods help reduce the clinician’s workload by identifying frames that contain potential abnormalities such as bleeding, ulcers, and polyps.
Segmentation accounts for 12.7% of the included studies and provides pixel-level delineation of anatomical structures and lesions. This supports quantitative assessment such as lesion area and boundaries. It can aid severity grading and treatment planning. Object detection, reported in 14.3% of studies, localizes abnormalities using bounding boxes. Compared with segmentation, it is less precise but faster and often sufficient for rapid screening and triage of suspicious frames.
Finally, localization, which accounts for 7.9% of the studies, focuses on mapping detected abnormalities to their corresponding anatomical locations within the GI tract. This application is particularly important for guiding further diagnostic and therapeutic interventions, as it enables precise targeting of lesions identified by WCE during follow-up endoscopic procedures.

4.6. AI Models

Based on Table A2 and Figure 7c, standard CNNs, such as ResNet, VGG, DenseNet, Xception, and Inception, were the most frequently used models, accounting for approximately 49% of the included studies. Many of these CNN-based methods incorporate transfer learning, feature-fusion layers, or customized modules.
Object detection frameworks accounted for approximately 32% of the included studies, combining both one-stage and two-stage detectors. Within this group, single-pass methods such as YOLO (v3, v5, v8) and SSD were commonly used because of their ability to localize lesions, including polyps and angiodysplasias, with high speed and accuracy.
Transformer-based models, including ViTs and TimeSformer, were used in approximately 6% of the studies. It indicates growing interest in self-attention mechanisms for modeling the spatial and temporal characteristics of WCE data. Hybrid or custom architectures, such as Capsule Networks (CapsNet), Symmetric Positive Definite Network (SPDNet), and proprietary ensemble models, accounted for about 11% of the included studies. These models often incorporated specialized attention mechanisms or domain-adversarial training to address challenges such as data imbalance and domain shift. In contrast, only one study employed a classical machine learning algorithm.

4.7. Study Design and Validation Strategies

As seen in Table A1, most studies have not provided information about their cohort. The main information about the patient population in some studies is age and sex statistics. These variables are primarily presented as mean and median values, although a significant proportion (approximately 75%) did not provide age data explicitly (Figure 7d).
The gender column reflects the percentage of male or female participants. The percentage of males in the included studies ranged from approximately 45% to 73%. However, a large portion of the studies (approximately 75%) did not explicitly state the male or female proportion (Figure 7d). Studies with reported male dominance often cited percentages exceeding 60% that highlights a trend toward a predominantly male sample in most datasets [30,31,32,33,34].
As seen in Table A2, the included studies utilized a variety of performance metrics, including sensitivity, specificity, accuracy, precision, Area Under the Curve (AUC), recall, Positive Predictive Value (PPV), Negative Predictive Value (NPV), and Matthews Correlation Coefficient (MCC) to evaluate their AI models. Among these, sensitivity, specificity, accuracy, and AUC were the most commonly reported metrics.
Validation and study design practices are summarized in Figure 7d. Only about 20% of studies explicitly reported the use of cross-validation, while external validation was reported in just 7% of studies. In addition, approximately 25% of studies relied on public datasets, and only a small proportion (about 5%) employed a prospective study design. The limited use of cross-validation, external validation, and prospective designs raises concerns about model robustness and generalizability across diverse clinical settings.

4.8. ROB Assessment

Figure 8 presents the summary of ROB analysis. The complete ROB assessment for each included study is provided in Supplementary File S4. The results show that Patient Selection, Reference Standard, and Flow and Timing domains have a low ROB, with most studies falling into this category. However, the Index Test (AI Model) domain shows a mix of low and unclear ROB, which indicates that, despite some studies following rigorous methods, others had unclear or poorly reported details regarding the AI model’s performance.

4.9. Meta-Analysis

The results of the clinical indication–based scenario are presented in Figure 9. The statistics for this approach is provided as Supplementary File S5.
Across all categories, the pooled estimates were consistently high, indicating strong diagnostic performance of AI-based systems in WCE. The highest pooled accuracy and AUC values were observed in the Bleeding, Vascular category (accuracy = 96.91%; AUC = 98.31%), that reflect the visually distinctive nature of bleeding-related findings and their suitability for automated detection. Similarly, studies focusing on Ulcer, Erosion, Inflammation (accuracy = 94.88%, AUC = 97.93%) and Polyp, Tumor, Protruded Lesion (accuracy = 94.17%; AUC = 95.50%) demonstrated robust pooled performance across all metrics, with narrow confidence intervals, which suggests reliable and reproducible results despite inter-study heterogeneity.
In contrast, the IBD (Crohn’s, UC) category showed slightly lower pooled estimates, particularly for AUC (92.04%), with wider confidence intervals for some metrics like AUC and specificity. Similarly, the Mixed, General Abnormality category showed moderate pooled estimates with sensitivity and specificity (sensitivity = 94.49%, specificity = 94.34%). This may reflect the greater phenotypic variability and diagnostic complexity of these conditions, as well as differences in annotation strategies and reference standards across studies. Nevertheless, the overall high pooled performance across all categories supports the effectiveness of AI approaches for diverse clinical applications in capsule endoscopy.
Similarly, the statistics for the second scenario is provided as Supplementary File S6. As seen in Figure 10, across all GI regions, the pooled estimates are consistently high, indicating strong diagnostic performance of AI-based methods in capsule endoscopy. Studies focusing on the Small Bowel formed the largest subgroup and demonstrated stable pooled values across all metrics (accuracy = 94.55%, sensitivity = 93.93%, AUC = 96.92%, specificity = 95.69%), which reflects the advanced development and widespread clinical use of capsule endoscopy in this region.
Studies analyzing the Entire GI tract (accuracy = 94.93%, AUC = 92.04%), Colon (accuracy = 95.29%, sensitivity = 94.73%, specificity = 97.24%, AUC = 97.06%), and Different Parts of GI Tract (accuracy = 95.40%, sensitivity = 93.49%, specificity = 94.17%, AUC = 94.40%) also showed strong pooled performance, although with slightly wider confidence intervals in some metrics, likely due to increased methodological and clinical heterogeneity. Overall, these results suggest that AI models achieve reliable and generalizable performance across different GI locations, while regional characteristics may still contribute to variability in diagnostic outcomes.
Publication bias was assessed using Egger’s linear regression test ( α = 0.10 ) [87] and the Duval & Tweedie trim-and-fill method [88] for all metric–subgroup combinations with at least 10 studies. A summary heatmap of Egger’s test p-values across all subgroups is presented in Figure 11, where colour intensity reflects the degree of asymmetry. Subgroups with fewer than 10 studies were not formally tested and are shown as grey cells. Also, cells labelled as (!) for significant asymmetry ( p < 0.10 ) and (ok) for no significant asymmetry. Subgroups with fewer than 10 studies are shown as grey cells annotated with their study count and were excluded from formal testing.
Funnel plots were constructed by plotting study-level estimates against the inverted standard error, with the 95% pseudo-confidence interval shaded and the DerSimonian-Laird pooled estimate indicated by a solid vertical line (Figure 12). Five combinations met the inclusion criteria, including accuracy, sensitivity, and specificity within the Bleeding, Vascular category, and accuracy and sensitivity within Mixed, General Abnormality. As seen in Figure 12, significant funnel plot asymmetry was observed in the Bleeding, Vascular subgroup for all three metrics (accuracy p = 0.020 , sensitivity p = 0.023 , specificity p = 0.002 ), whereas no significant asymmetry was detected in Mixed, General Abnormality (accuracy p = 0.112 , sensitivity p = 0.214 ), as also illustrated in the funnel plots. Notably, the observed asymmetry was in a conservative direction, with smaller studies reporting lower performance than larger studies. The Duval & Tweedie trim-and-fill method imputed additional studies on the higher-performance side that results in slightly increased adjusted estimates compared with the original values. The adjusted estimates were 99.13% versus 97.40% for accuracy, 98.86% versus 96.04% for sensitivity, and 99% versus 96.87% for specificity. These results indicate that the pooled estimates are unlikely to be inflated by publication bias. It is also worth noting that due to the high diagnostic performance observed (>95% in most subgroups), the funnel plot analysis is subject to a ceiling effect.

5. Discussion

This systematic review highlights the strong diagnostic performance of AI models in capsule endoscopy while uncovering critical limitations in validation practices, dataset diversity, and clinical readiness.
Many prior reviews have constrained their scope to narrow disease categories such as celiac disease [89] or IBD [90], excluding broader AI applications. Others focus exclusively on specific lesion types like protruded lesions or bleeding sources [17,91,92,93,94]. However, in this study, we considered broader domain to cover more AI-related studies on the WCE applications.
Although many studies have investigated the use of AI across different WCE applications, challenges related to real-world deployment remain largely underexplored. In of the Discussion, we address the limitations of the current research landscape and outline the technical and clinical barriers that must be addressed in future studies to achieve robust and reliable AI results and to support effective integration into real-world clinical practice.

5.1. Datasets

First, one of the main factors affecting the quality and trustworthiness of AI models is the quality and diversity of the datasets [95]. However, many included studies were based on a limited number of patients and samples. More precisely, as shown in Figure 5, the median number of patients in the included studies is 174, which is relatively low for sensitive applications such as polyp detection, bleeding detection, and related tasks.
Another important issue is that almost none of the included studies explicitly reported whether samples were considered and analyzed in a patient-wise manner. In most cases, all images acquired by WCE were pooled, annotated, and used in the training and testing processes. The absence of patient-wise analysis introduces a methodological concern related to data leakage and temporal correlation within WCE data. Since consecutive frames from the same examination are highly similar, random image-level splitting can result in nearly identical samples appearing in both training and testing sets that can lead to artificially inflating performance metrics. This issue is particularly critical in WCE, where thousands of frames originate from a single patient could increase the likelihood of overlap between datasets. As a result, reported accuracy and sensitivity may reflect memorization of patient-specific patterns rather than true generalization. Moreover, this approach does not reflect real-world clinical settings, where personalized models that analyze and make decisions at the individual patient level are crucial for reliable clinical deployment [96].
As seen in Table A1, most included studies showed inconsistent or missing reporting of cohort characteristics, including age, sex, comorbidities, and medication use, with only a limited number of studies providing basic cohort characteristics [34,45]. In 16 studies that provided age data, the median age was 53.5 years, indicating a bias toward older patient populations. These issues reduce the interpretability and clinical relevance of the findings. For example, patients taking blood thinners or antiplatelet drugs may present bleeding patterns that are harder to detect. Without information on medication use, AI models may overlook subtle signs. Likewise, conditions such as liver disease or bleeding disorders can alter lesion appearance and make detection more difficult.
Another technical aspect is the imbalanced data, that influence the training and evaluation of AI models. Multiple studies leveraged the Kvasir dataset [97] for training AI models [47,66,67,81,98,99,100,101]. Despite its popularity, the dataset is significantly imbalanced. For example, while the Normal Clean Mucosa category contains over 30,000 images, abnormal classes are drastically underrepresented with Polyp with only 55 images, Blood–Hematin with 12 images, and Erythema with around 159 images. Many available samples in the datasets are also the same abnormality captured at different times, angles, or distances, which reduces the number of truly distinct abnormal cases. This redundancy can cause AI models to overfit to majority classes and perform poorly in detecting rare but clinically important abnormalities.
Moreover, our results reveal a significant under-representation of certain GI regions and conditions. The majority of studies focus on the small bowel, while regions like the esophagus, stomach, and specific segments of the colon receive considerably less attention. Similarly, common conditions such as bleeding and polyps dominate the research landscape, whereas some complex conditions, such as specific inflammatory diseases, are rarely explored. This imbalance restricts the generalizability of AI models.
Recommendations: Future research should focus on collecting larger and more diverse datasets that cover a broad range of GI abnormalities, with better representation across lesion types and patient groups. In addition, developing a comprehensive public benchmark dataset, similar to COCO in computer vision [102], would support standardized evaluation and improve the reproducibility of research studies. Moreover, future studies should implement patient-level data collection to have more robust and reliable datasets.

5.2. AI Analysis

Although several reviews report promising results, such as sensitivity and specificity values exceeding 90% for the detection of polyps and hemorrhagic lesions [93,94,103], several aspects should be considered when interpreting and drawing conclusions from AI model results. First, our ROB analysis (Figure 8) shows that many of the included studies have unclear or high ROB, particularly in domains related to AI models and methodologies. This indicates that, in most studies, important methodological and AI-related aspects were either not reported or not adequately addressed. Moreover, the observed heterogeneity across studies can be attributed to several factors, including substantial variation in dataset size and composition, differences in AI model architectures, variability in clinical tasks and target pathologies, and inconsistent validation strategies. Included studies ranged from small cohorts to large-scale datasets and employed diverse methodologies such as classification, detection, and segmentation. To further investigate potential sources of this heterogeneity, subgroup analyses were performed based on clinical indication and gastrointestinal location. These stratified analyses enabled a more nuanced assessment of AI performance across different disease categories and anatomical regions, highlighting variations that may account for inconsistencies in pooled estimates and offering more clinically relevant insights into model applicability across diverse settings.
External validation is essential to assess the robustness of AI models in real clinical settings. Without external validation, high performance metrics may not translate reliably to clinical practice.
Several studies report very high performance without external validation. For example, CNN models for detecting obscure GI bleeding have achieved sensitivity and specificity above 98% [56], but their generalizability across different institutions, patient populations, and WCE devices remains uncertain. In contrast, studies that include external validation provide stronger evidence of clinical applicability. Ali et al. [37] conducted a multi-center study on polyp detection and segmentation with external validation and reported a Dice score of 0.82 and accuracy of 98%. Similarly, Xie et al. [33] validated their model across multiple independent centers and achieved a detection rate of 95.9% and sensitivity of 98.8%.
Recommendation: We encourage researchers in this domain to apply their methodology in multi-center settings and evaluate the performance of their models on unseen data from completely different cohorts. AI explainability and interpretability techniques can also be used to increase clinicians’ trust in AI models. In addition, given the rapid advances in Large Language Models (LLMs) and their promising performance across various applications [104], it is worth evaluating their potential in this field.

5.3. Linking WCE and Conventional Endoscopy

WCE and conventional endoscopy differ substantially in image acquisition. In conventional endoscopy, physicians control illumination, focus, and viewing angle, resulting in more consistent image quality. In contrast, WCE images are acquired passively and often show variations in lighting, resolution, and orientation. As a result, the same lesion may appear differently across the two modalities, which complicates follow-up and confirmation.
This difference limits AI models trained on a single modality. Several domain adaptation approaches have been proposed to address this gap, including Cycle-Consistent Generative Adversarial Networks (CycleGAN), which can translate WCE images to a conventional endoscopy style without paired data [105], and Domain Adversarial Neural Networks (DANN), which aim to learn domain-invariant features [106]. Metric learning approaches may also help align lesions across modalities but require paired datasets, which remain limited [107]. Despite their potential, most of these methods lack validation in real clinical settings, and the scarcity of high-quality paired data remains a key challenge for integrating AI across the full diagnostic workflow.
Recommendation: Greater integration between WCE and conventional endoscopy is needed through harmonized annotations and AI models that can operate across both modalities. This would support coherent diagnostic workflows.

5.4. Comparison with Expert Interpretation

A key limitation in current AI research on WCE images is the inconsistent comparison between AI performance and that of human experts. Benchmarking AI systems against experienced gastroenterologists is important to assess clinical relevance and added value. Although some studies report improved performance when AI is used as a support tool, such as higher sensitivity in detecting small bowel vascular lesions [54], many studies report only standalone performance metrics. In addition, variability in human interpretation of WCE images has been reported [15], highlighting the need for standardized comparisons to better understand the clinical value of AI systems.
Recommendation: Researchers should include direct comparisons with human experts using standardized protocols, and evaluate both diagnostic performance and clinical utility, such as workload and reading time.

5.5. Strengths and Limitations

Our review has several strengths. We conducted a comprehensive search across multiple databases using a broad set of keywords, ensuring that we captured a wide range of relevant studies. We also applied QUADAS-2 tool for quality assessment to evaluate the methodological quality and ROB in each study. Additionally, by performing a meta-analysis of the studies, we provide an overview of AI’s diagnostic performance in WCE. However, the main limitation of this review is the high variability among the included studies. Differences in study design, sample sizes, and reporting make it difficult to combine the results of the meta-analysis accurately, which may affect the reliability of its findings. Additionally, this meta-analysis was not prospectively registered in PROSPERO, which may introduce potential risk of reporting bias.

6. Conclusions

This systematic review demonstrates that AI has strong potential to improve diagnostic accuracy and efficiency in WCE applications. In contrast to previous reviews that focused on specific diseases, applications, or limited GI regions, our study provides a comprehensive synthesis across multiple AI application domains and GI tract locations, which means it offers a broader and more integrative perspective on the field.
Although many included studies report high sensitivity, specificity, and overall diagnostic accuracy, our analysis identifies several persistent challenges that limit real-world deployment. These include imbalanced datasets, limited patient and sample sizes, incomplete method reporting, ignoring patient-wise model development, and the lack of external and prospective validation in most studies. These observations highlight the need for more uniform methodologies and clinically oriented study designs in future research. Moreover, direct evaluation against experienced gastroenterologists is essential to determine whether AI models provide meaningful added value in real diagnostic workflows.
Furthermore, our meta-analysis confirms the consistently high diagnostic performance of AI models across different clinical indications and GI regions. However, this work emphasizes that methodological rigor, standardized reporting, and robust validation strategies are essential to improve the reliability, generalizability, and clinical adoption of AI-based systems in WCE applications.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/diagnostics16091269/s1, This manuscript has six Supplementary Files. Supplementary File S1 presents the search queries. Supplementary File S2 includes the PRISMA checklist and diagram. Supplementary File S3 contains the CHARMS-based data extraction table for all included studies. Supplementary File S4 provides the full risk of bias assessment using the QUADAS-2 tool. Supplementary Files S5 and S6 present the detailed results of the clinical indication-based and location-based meta analyses, respectively.

Author Contributions

Conceptualization: A.S. and A.N.; methodology: A.S. and A.N.; software: A.S. and A.N.; validation: A.S., A.K. and A.N.; formal analysis: A.S., A.K. and A.N.; data curation: A.S. and A.N.; data extraction: A.S. and A.N.; writing—original draft preparation: A.S., A.K. and A.N.; writing and editing: A.S., A.K. and A.N.; reviewing: A.S., A.K. and A.N., visualization: A.S. and A.N.; supervision: A.K., A.N. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

This study is a systematic literature review incorporating 72 original research studies published by other scholars. All cited works are listed in the References section. Additionally, we have made all extracted data and relevant information available as Supplementary Files. Further inquiries can be directed to the corresponding author.

Conflicts of Interest

The authors declare no conflicts of interest.

Appendix A. Characteristics of Included Studies

Table A1. Characteristics of included studies.
Table A1. Characteristics of included studies.
First AuthorYearCountryRegion of GI TractTime PeriodSamplesPatients TypeAgeSexData SourceStudy DesignSetting
Teng Zhou [86]2017ChinaSmall bowel2008–200911 celiac disease patients and 10 controls, analyzed via capsule endoscopyPatients with celiac disease and controls50.5 (females mean), 44 (males mean)45.45% maleColumbia University Medical CenterRetrospectiveClinical practice
Dimitris K. Iakovidis [80]2018Greece, UKEntire GI tract2015–201810,698 imagesPatients undergoing endoscopy--Two public datasets (MICCAI and KID)RetrospectiveAnalysis
Zhen Ding [77]2019ChinaSmall bowel2016–20186970 patientsPatients with small-bowel diseases and normal variants--Collected by authors from 77 medical centersRetrospectiveClinical practice
Romain Leenhardt [30]2019FranceSmall bowel2011–20182946 images with lesions, 600 gastrointestinal angiectasia framesPatients with gastrointestinal bleeding72 (mean)61% maleCAD-CAP databaseRetrospectiveAnalysis
Amit Kumar Kundu [61]2019BangladeshEntire GI tract20192350 images (450 bleeding, 1900 non-bleeding)Patients with gastrointestinal abnormalities--Public datasetRetrospectiveAnalysis
Tomonori Aoki [51]2019JapanSmall bowel2009–2018115 patients (training), 65 patients (validation), 5360 training images, 10,440 validation imagesPatients with gastrointestinal abnormalities63 (training set mean)54% male (training dataset)University of Tokyo HospitalRetrospectiveClinical practice
Sen Wang [31]2019ChinaEntire GI tract-1416 videosUlcer patients-73% male30 hospitals and 100 medical examination centers, ChinaRetrospectiveClinical practice
Libin Lan [108]2019ChinaSmall bowel-7381 wireless capsule endoscopy imagesPatients with gastrointestinal abnormalities--Datasets from Jinshan Science & Technology and Given ImagingRetrospectiveAnalysis
U Deding [109]2020DenmarkColon2016–201897 patientsPatients with incomplete optical colonoscopy64.5 (median)25 % maleOdense University Hospital (Denmark)ProspectiveClinical practice
Eyal Klang [45]2020IsraelSmall bowel-17,640 images from 49 patientsPatients with known Crohn’s disease and control subjects28 (mean)50% maleDepartment of Gastroenterology at Sheba Medical CenterRetrospectiveClinical practice
Keita Otani [79]2020JapanSmall bowel2009–2019167 patients; 398 erosions/ulcers, 538 vascular lesions, 4590 tumors, 34,437 normal imagesPatients with gastrointestinal abnormalities63.6 (mean)55% maleTokyo HospitalRetrospectiveClinical practice
Yixuan Yuan [40]2020Hong KongColon20197200 images from 80 patientsPatients with gastrointestinal abnormalities--Not mentionedRetrospectiveAnalysis
Akiyoshi Tsuboi [73]2020JapanSmall bowel-141 patients (training), 28 patients (validation)Patients with diagnosed small-bowel angioectasia using capsule endoscopy68.5 (validation set mean)54% male (validation set)Hiroshima University Hospital, University of Tokyo Hospital, Sendai Kousei HospitalRetrospectiveClinical practice
Tomonori Aoki [60]2020JapanSmall bowel2009–201527,847 images (training), 10,208 images (test)Patients with gastrointestinal abnormalities53.4 (mean)56% maleThe University of Tokyo Hospital (Japan)RetrospectiveClinical practice
Hiroaki Saito [32]2020JapanSmall bowel2009–201830,584 images for training; 17,507 for testingPatients with and without small bowel protruding lesions60.1 (mean)65.8% maleSendai Kousei Hospital (Japan), The University of Tokyo (Japan), Hiroshima University Hospital (Japan), Given Imaging (Israel)RetrospectiveClinical practice
Charles Houdeville [82]2021FranceSmall bowel20211200 framesPatients with gastrointestinal bleeding--CAD-CAPRetrospectiveAnalysis
Tonmoy Ghosh [69]2021USASmall bowel-2350 images (450 bleeding, 1900 non-bleeding)Patients with suspected intestinal bleeding--endoscopy.org and KID public datasetsRetrospectiveAnalysis
Chao Liao [110]2021ChinaSmall bowel-50,000 frames across 9 videosPatients undergoing gastrointestinal motility assessment--Jinshan Science & Technology CompanyRetrospectiveClinical practice
Jürgen Herp [111]2021DenmarkColon-42 patients, 84 videosPatients with colonoscopy follow-up need--Odense University HospitalRetrospectiveClinical practice
Andrea Caroppo [59]2021ItalySmall bowel-2352 images (303 with bleeding)Patients with gastrointestinal bleeding--Public KID DatasetRetrospectiveAnalysis
Furqan Rustam [58]2021Pakistan, South KoreaEntire GI tract-5000 images from 33 patientsPatients with gastrointestinal tract infections--Sheikh Zayed Hospital, Rahim Yar KhanRetrospectiveClinical practice
Ji Xia [68]2021ChinaStomach2014–2018697 patients, 1,023,955 capsule endoscopy imagesPatients with gastrointestinal abnormalities--Changhai Hospital, ShanghaiRetrospectiveClinical practice
Yunseob Hwang [50]2021Republic of KoreaSmall bowel2007–2019526 mall bowel capsule endoscopy videos (training), 5760 independent images (validation)Patients undergoing mall bowel capsule endoscopy--Multiple hospitals in KoreaRetrospectiveClinical practice
Miguel José Mascarenhas Saraiva [85]2021PortugalSmall bowel2015–20204319 patients, 53,555 capsule endoscopy imagesPatients with gastrointestinal abnormalities--São João University Hospital and ManopH Gastroenterology ClinicRetrospectiveClinical practice
Tomonori Aoki [112]2021JapanSmall bowel2009–2018379 patients, 66,028 capsule endoscopy images for trainingPatients undergoing small-bowel capsule endoscopy62.3 (test set mean)57% male (test set)University of Tokyo, Hiroshima University Hospital, Sendai Kosei HospitalRetrospectiveClinical practice
Miguel José Mascarenhas Saraiva [84]2021PortugalColon2010–20203,387,259 frames from 24 colon capsule endoscopy exams, 3640 images for training/validationPatients undergoing colon capsule endoscopy for detection of colonic lesions--São João University Hospital, Porto, PortugalRetrospectiveClinical practice
Yiftach Barash [44]2021IsraelSmall bowel-17,640 images (7391 ulcers, 10,249 normal)Patients with Crohn’s disease--Sheba Medical Center, Tel Hashomer (Israel), Medtronic Dublin (Ireland)RetrospectiveClinical practice
Miguel Mascarenhas Saraiva [57]2021PortugalSmall bowel2020–20216740 images (training and validation)Patients undergoing DAE for suspected mid-gastrointestinal bleeding--São João University Hospital, Porto, PortugalRetrospectiveClinical practice
Eyal Klang [43]2021IsraelSmall bowel2009–201827,892 images (1942 strictures; 14,266 normal mucosa; 11,684 ulcers)Crohn’s Disease patients--Sheba Medical Center, Tel Hashomer (Israel), Medtronic Dublin (Ireland)RetrospectiveClinical practice
Samir Jain [67]2021India, Norway, Czech Republic, MalaysiaSmall bowel2018–2021Combined KID dataset; images with augmentationPatients with gastrointestinal anomalies--KID datasetRetrospectiveAnalysis
Miguel Mascarenhas Saraiva [56]2022PortugalSmall bowel2015–20201229 patients; 22,095 framesPatients with suspected gastrointestinal bleeding--São João University Hospital, Porto, PortugalRetrospectiveClinical practice
Guillem Pascual [39]2022Spain, DenmarkSmall intestine, colon2016–2021Generic: 1,185,033 frames, Polyp dataset: 248,136 frames, CAD-CAP dataset: 1800 imagesMixed patients--CAD-CAP, Internal datasetRetrospectiveAnalysis
João Afonso [49]2022PortugalSmall bowel2015–20201483 wireless capsule endoscopy exams, 6130 frames (4233 containing ulcers and erosions)Patients with inflammatory bowel disease--São João University Hospital, Porto, PortugalRetrospectiveClinical practice
Jun-Xiao Zhou [113]2022ChinaSmall bowel2016–2019277 polyp images from 480 patientsPatients with gastrointestinal disorders--Guangzhou First People’s Hospital, CVC-Colon, CVC-ClinicRetrospectiveClinical practice
Md. Jahin Alam [66]2022Bangladesh, USAEntire GI tract2016–201847,238 imagesPatients with gastrointestinal abnormalities--Kvasir-Capsule public datasetRetrospectiveAnalysis
Naoki Higuchi [42]2022JapanColon2018–2020739,021 imagesPatients with moderate or mild ulcerative colitis--Hirosaki University Hospital, JapanProspectiveClinical practice
João Pedro Sousa Ferreira [48]2022PortugalSmall bowel and Colon2017–20208085 imagesPatients with Crohn’s Disease--São João University Hospital and ManopH Gastroenterology Clinic, PortugalRetrospectiveClinical practice
Saqib Mahmood [47]2022PakistanEntire GI tract2016–201847,238 labeled frames from 117 videosPatients with GI tract abnormalities--Kvasir-Capsule datasetRetrospectiveAnalysis
Meryem Souaidi [114]2022MoroccoEntire GI tract-WCE, PillCam-COLON, CVC-ClinicDB, ETIS-Larib, PASCAL VOC, COCOPatients with gastrointestinal polyp indications--PillCam-COLON, CVC-ClinicDB, ETIS-Larib, PASCAL VOC, COCORetrospectiveAnalysis
Miguel Mascarenhas [38]2022PortugalColon2010–2020124 patients, 5715 colon capsule endoscopy framesPatients with protruding lesions detected via colon capsule endoscopy--São João University Hospital and ManopH Gastroenterology ClinicRetrospectiveClinical practice
Tom Kratter [46]2022IsraelSmall bowel, Colon-33,100 capsule endoscopy imagesPatients with Crohn’s disease and normal mucosa--Sheba Medical Center, Tel Hashomer (Israel), Medtronic Dublin (Ireland)RetrospectiveClinical practice
Xia Xie [33]2022ChinaSmall bowel2012–20215825 small bowel capsule endoscopy examinationsPatients undergoing small bowel capsule endoscopy49.8 (mean)60.9% male51 medical centers, ChinaRetrospectiveClinical practice
SangYup Oh [75]2023South KoreaSmall bowel2002–20221,431,344 images (across 260 cases)Mixed gastrointestinal patients--Dongguk University Ilsan Hospital, South KoreaRetrospectiveClinical practice
Raizy Kellerman [41]2023IsraelSmall bowel2011–2021101 patientsNewly diagnosed Crohn’s patients27 (median)46.5% maleSheba Medical Center, IsraelRetrospectiveClinical practice
Mehrdokht Bordbar [34]2023IranEntire GI tract-29 patients, 14,691 framesMixed gastrointestinal patients53.5 (mean)66% maleNamazi Hospital, Shiraz, IranRetrospectiveClinical practice
Sharib Ali [37]2023UK, Norway, France, Italy, EgyptColon-300 patients, 8037 framesPatients undergoing colonoscopy for colorectal cancer screening--Public PolypGen datasetRetrospectiveAnalysis
Ye Chu [55]2023ChinaSmall bowel2014–202012,403 imagesPatients with angiodysplasias undergoing capsule endoscopy--Ruijin Hospital, ChinaRetrospectiveClinical practice
Sofia A. Athanasiou [72]2023GreeceEntire GI tract2019–20206016 framesPatients with bleeding, angiodysplasias, hemangiomas, or lesions predisposing to bleeding--Kapodistrian University of Athens, Attikon University Hospital, Laikon University Hospital, Aristotle University of ThessalonikiProspectiveClinical practice
Akihiko Sumioka [115]2023JapanSmall bowel2011–202126 patientsPrimary small-bowel follicular lymphoma patients60.9 (mean)50% maleHiroshima University Hospital, JapanRetrospectiveClinical practice
Javeria Naz [65]2023Pakistan, Norway, UK, Saudi ArabiaEntire GI tract2016–2018>30,000 imagesPatients with gastrointestinal abnormalities--Kvasir-V1 and other pulic sourcesRetrospectiveAnalysis
Ayako Nakada [74]2023JapanSmall bowel2009–2019651 patients, >10 million wireless capsule endoscopy imagesPatients with obscure gastrointestinal bleeding, small intestine tumors, abdominal symptoms--Nine hospitals across JapanRetrospectiveAnalysis
Tomonori Aoki [116]2024JapanSmall bowel2009–201936 wireless capsule endoscopy videos, 43 lesionsPatients with abnormal lesions in wireless capsule endoscopy--University of Tokyo, Hiroshima University Hospital, Sendai Kosei Hospital, and Japan Medical AssociationRetrospectiveClinical practice
Ali Sahafi [22]2024Denmark, PolandColon-2500 imagesPatients with gastrointestinal abnormalities--KID DatasetRetrospectiveAnalysis
Rui-Ya Zhang [54]2024ChinaSmall bowel2013–2023111,861 imagesPatients with gastrointestinal issues--Shanxi Provincial Hospital, ChinaRetrospectiveClinical practice
Tsedeke Temesgen Habe [81]2024FinlandEntire GI tract2016–201847,238 images (class-annotated)Patients with gastrointestinal abnormalities--Kvasir-Capsule public datasetRetrospectiveAnalysis
Mohamed Achraf Belabbes [100]2024MoroccoEntire GI tract2023–2024Several datasets: PillCam COLON, Kvasir-SEG, ETIS-Larib, CVC-ClinicDBPatients with suspected polyps--PillCam COLON, Kvasir-SEG, ETIS-Larib, CVC-ClinicDBRetrospectiveAnalysis
Kyung Seok Choi [53]2024South KoreaSmall bowel2018–2020103 paitient (1,037,286 images)Patients with suspected small-bowel bleeding (SSBB)64.7 (mean)62.2% maleConsecutive patient data from Yeouido St. Mary’s and Seoul St. Mary’s HospitalsRetrospectiveClinical practice
Akihito Yokote [70]2024JapanSmall bowel2014–2021954 patients, 18,481 imagesPatients undergoing small bowel capsule endoscopy--Kyushu University HospitalRetrospectiveClinical practice
Dong Liu [98]2024ChinaColon2016–20183 public datasets: 1000 polyp images (Kvasir-SEG), 590 images (Kvasir-Instrument), 55 images (KvasirCapsule-SEG)Patients with gastrointestinal abnormalities--Kvasir-Capsule public datasetRetrospectiveAnalysis
Xia Xie [76]2024ChinaStomach and Small bowel2021–20221069 training, 342 validationPatients with gastrointestinal abnormalities42 (median)45.61% maleThree hospitals in Chongqing (China), University of Southern Denmark, Jinshan Science & Technology (China)RetrospectiveClinical practice
Seung-Joo Nam [83]2024Republic of KoreaStomach, Small bowel, Colon2018–2021126 paitient (2,392,462 images)Patients with gastrointestinal abnormalities--Kangwon National University Hospital (South Korea), Dongguk University Ilsan Hospital (South Korea)RetrospectiveClinical practice
Lan Li [64]2024ChinaSmall bowel2016–20221452 patients for training; 298 for testingPatients with small bowel lesions--Zhejiang UniversityRetrospectiveClinical practice
Esmaeil S. Nadimi [36]2024DenmarkLarge Intestine (Colon)20215838 polyp & 5573 normal images (augmented) for recognition; 5838 for size; 144 (49 neoplastic, 95 non-neoplastic) for characterizationFIT-positive screening participants undergoing colon capsule endoscopy--Odense University Hospital and University of Southern DenmarkRetrospectiveAnalysis
Yeong Seok Kwon [117]2025KoreaSmall Bowel2011–202228,279 small-bowel images, 32 full-length SBCE videos for external validationPatients undergoing small-bowel capsule endoscopy for suspected bleeding--Chuncheon & Dongtan Sacred Heart (training); Yeungnam & Ewha Univ. Hospitals (external)ProspectiveClinical validation
Miguel José Mascarenhas Saraiva [35]2025PortugalSmall and Large Intestine2021–2023191,455 frames extracted from 1245 CE/CCE examsPatients undergoing capsule or colon capsule endoscopy for diagnostic evaluation--CHU São João & ManopH Clinic (Portugal)RetrospectiveAnalysis
Patrícia Andrade [118]2025Multiple (Europe & USA)Small Intestine2021–2024259 SBCE exams (PillCam SB3 and Olympus EC-10)Crohn’s disease patients undergoing small-bowel capsule endoscopy--Multiple hospitals across Europe and the USARetrospectiveClinical practice
Miguel Mascarenhas Saraiva [119]2025Portugal, Spain, Brazil, USAEntire GI tract-330 capsule endoscopy videosPatients undergoing capsule endoscopy across Portugal, Spain, Brazil and the USA--Multi-centre dataset using PillCam, Olympus and OMOM systemsProspectiveClinical practice
Maxime Le Floch [78]2025GermanySmall Intestine2011–202380 VCE recordings; 3,513,539 frames annotatedPatients undergoing video capsule endoscopy--University Hospital Carl Gustav Carus, GermanyRetrospectiveAnalysis
Miguel Martins [63]2025Portugal, SpainEsophagus and Stomach2021–202359,482 frames extracted from 774 capsule endoscopy proceduresPatients undergoing capsule endoscopy of the oesophagus and stomach--São João University Hospital, Portugal & La Princesa University Hospitals, SpainRetrospectiveAnalysis
Charles Houdeville [52]2025France, NetherlandsSmall Intestine2019–2021 148 patients; 1525 imagesPatients with suspected small-bowel bleeding undergoing SBCETraining: 72 (mean); Validation: 78 (mean)52 % maleSorbonne University Hospital, Paris & Radboud University Medical Center, NijmegenRetrospectiveAnalysis
Jian Chen [62]2025ChinaSmall Intestine-34,799 images from multiple public datasetsPatients undergoing small-bowel capsule endoscopy with various lesions--Multi institutional datasets from Chinese hospitals and public sourcesRetrospectiveAnalysis
Bernardo Rosa [71]2025PortugalEntire GI tract2024100 patientsPatients undergoing panenteric capsule endoscopy for suspected mid-to-lower GI bleeding66.5 (median)35 % maleHospital da Senhora da Oliveira & São João University Hospital, PortugalProspectiveClinical practice
Table A2. AI-related characteristics of included studies.
Table A2. AI-related characteristics of included studies.
First AuthorGastrointestinal ProblemApplicationAI ModelsPerformance MetricsCross ValidationExternal Validation
Teng Zhou [86]Detection of celiac diseaseImage classificationGoogLeNetAcc: 100%, Sens: 100%, Spec: 100%7-fold cross-validation performedNo
Dimitris K. Iakovidis [80]Detecting and localizing gastrointestinal anomaliesImage classification and anomaly localizationWeakly supervised CNN with saliency detectionClassification AUC (mean): 88.85%, localization (mean): AUC: 0.8610-fold cross-validationYes
Zhen Ding [77]Detection of inflammation, ulcer, polyps, lymphangiectasia, bleeding, vascular disease, protruding lesionImage classificationResNetPer-patient (Sens: 99.88%, NPV: 99.77%), Per-lesion (Sens: 99.90%, NPV: 99.77%)NoNo
Romain Leenhardt [30]Detection of GI angiectasiaImage classification and segmentationCNNSens: 100%, Spec: 96%, Prec: 96.15%, NPV: 100%NoNo
Amit Kumar Kundu [61]Detection and localization of bleeding regionsImage classification and bleeding zone segmentationSupport vector machineAcc: 96.77%, Sens: 97.55%, Spec: 96.59%10-fold cross-validationNo
Tomonori Aoki [51]Detection of erosions and ulcerations in the small bowelImage classificationSSDAcc: 90.8%, Sens: 88.2%, Spec: 90.9%, AUC: 0.958NoNo
Sen Wang [31]Ulcers detectionImage classificationResNet34Acc: 92.05%, Sens: 91.64%, Spec: 92.42%, prec: 92.37%, F1 score 91.99%, AUC: 0.975-fold cross-validationNo
Libin Lan [108]Detection of bleeding, polyps, and tumorsImage classificationCascadeProposal (based on Fast R-CNN)Performance average: overall mAP: 72.17%, per-Class mAP: 70.54%, region proposal methods mAP: 69.40%NoNo
U Deding [109]Polyp detectionImage localization and trackingAI-based tracking algorithmAcc (mean): 77%NoNo
Eyal Klang [45]Ulcers detection in Crohn’s diseaseImage classificationXceptionAcc: 96%, Sens: 94.8%, Spec: 97%, Prec: 95.8%, NPV: 96.4%, AUC: 0.99%5-fold cross-validationNo
Keita Otani [79]Detection of small bowel lesionsImage classification and object detectionRetinaNet and SSDRetinaNet (best model) performance average: AUC: 0.935-fold cross-validationYes
Yixuan Yuan [40]Polyp detectionImage classificationDenseNet-UDCSAcc: 93.19%, Sens: 95.12%, Prec: 94.56%, F1 score: 94.84%10-fold cross-validationNo
Akiyoshi Tsuboi [73]Detection of small-bowel angioectasiaImage classificationSSDSens: 98.8%, Spec: 98.4%, AUC: 0.99NoYes
Tomonori Aoki [60]Blood content detectionImage classificationResNet50Acc: 99.89%, Sens: 96.63%, Spec: 99.96%, Prec: 98.06%, NPV: 99.93%, AUC: 0.99, Sens: 96.63%NoNo
Hiroaki Saito [32]Detection of polyps, nodules, epithelial tumors, submucosal tumors, and venous structuresImage classificationSSDSens: 90.7%, Spec: 79.8%, AUC: 0.91NoNo
Charles Houdeville [82]Angiectasias detectionImage classificationAxaroPerformance average: Acc: 99.7%, Prec: 98.2%, NPV: 96.9%NoYes
Tonmoy Ghosh [69]Detection of bleedingImage classification, bleeding zone segmentationAlexNet, SegNetBleeding classification (F1 score: 98.49%, Sens: 97.51%, Spec: 99.88%, Prec: 99.50%, Acc: 99.44NoNo
Chao Liao [110]GI motility assessmentRegion-based in consecutive WCE framesResNet-50)Percentage of correct key-points: 66.49%NoNo
Jürgen Herp [111]Screening for polyps, inflammatory diseases, and tumorsCapsule localization and path reconstructionFeature Point Tracking with filteringPath difference: 4 ± 0.7 cm, Classification Acc: 86%, Section Acc: 92%NoYes
Andrea Caroppo [59]Bleeding and lesion classificationImage classificationFusion of VGG19, InceptionV3, and ResNet50Performance average: Acc: 96.95%, Sens: 97.26%, Spec: 96.41%, Prec: 97.79%, F1 score: 97.52%10-fold cross-validationNo
Furqan Rustam [58]Bleeding detectionImage classificationMobileNet+custom CNNAcc: 99%, Prec: 100%, Recall: 99.4%, F1 score: 99.7%10-fold cross-validationNo
Ji Xia [68]Detection of various gastric lesionsImage classificationResNet-34 and Faster-RCNNPerformance average: Acc: 81.17%, Sens: 94.89%, Spec: 68.77%, Prec: 55.57%, NPV: 99.90%, AUC: 0.84NoNo
Yunseob Hwang [50]Classification and localization of small-bowel lesionsImage classification and lesion localizationVGGNetPerformance average: Acc: 96.73%, Sens: 96.34%, Spec: 96.04%, Prec: 96.11%, NPV: 96.41%, AUC: 0.9910-fold cross-validationYes
Miguel José Mascarenhas Saraiva [85]Detection small bowel lesionsImage classificationXceptionAcc: 98%, Sens: 87.8%, Spec: 99.4%, NPV: 99.4%NoNo
Tomonori Aoki [112]Detection of mucosal breaks, angioectasia, Protruding lesions, and bleedingImage classificationSSD and ResNet50Performance average: Acc: 92.60%, Sens: 99.36%, Spec: 81.18%NoNo
Miguel José Mascarenhas Saraiva [84]Protruding polyps, epithelial tumors, submucosal tumors, and nodesImage classificationXceptionAcc: 92.2%, Sens: 90.7%, Spec: 92.6%, Prec: 79.2%, NPV: 96.9%, AUC: 0.97NoNo
Yiftach Barash [44]Ulcer severity in Crohn’s diseaseImage classification and ulcer severity gradingOrdinal CNN (ResNet based)Performance average: Acc: 77.13%, Sens: 66.20%, Spec: 84.57%, F1 score: 67.5%, AUC: 0.825-fold cross-validationNo
Miguel Mascarenhas Saraiva [57]Detection of GI angioectasiaImage classificationXceptionAcc: 95.3%, Sens: 88.5%, Spec: 97.1%, Prec: 88.8%, NPV: 97%, AUC: 0.98,NoNo
Eyal Klang [43]Detection of intestinal strictures in Crohn’s DiseaseImage classifcationEfficientNetB5Performance average: Acc: 86.2%, Sens: 92%, Spec: 89%, Prec: 55%, NPV: 99%, F1 score: 69%, AUC: 0.9310-fold cross-validationNo
Samir Jain [67]Localization and detection of inflammatory, polyp, vascular, and normal mucosaImage classificationWCENet (Custom CNN, Grad-CAM++, SegNet)Performance average: Acc: 98%, Sens: 98%, Spec: 98%, Prec: 98%, F1 score: 0.98, AUC: 0.99, localization IoU: 0.605-fold cross-validationNo
Miguel Mascarenhas Saraiva [56]Obscure gastrointestinal bleedingImage classificationXceptionSens: 98.6%, Spec: 98.9%, Acc: 98.5%, Prec: 98.7%, AUC: 1.0NoNo
Guillem Pascual [39]Polyp detectionImage classificationResNet with SimCLR, TLBAAcc: 92.77%, AUC: 0.955-fold cross-validationNo
João Afonso [49]Detection of ulcers and erosionsImage classification with bleeding riskXceptionAcc: 95.6%, Sens: 90.8%, Spec: 97.1%, Prec: 93.4%, NPV: 97.1%NoNo
Jun-Xiao Zhou [113]Polyp segmentationImage segmentationEnsemble with SegNet, U-Net, Attention-UNet, ResNet-UNet, HarDMSEGGuangzhou First People’s Hospital dataset (IoU: 0.55, Dice: 0.65), CVC datasets (IoU: 0.84, Dice: 0.90)NoNo
Md. Jahin Alam [66]Detection of ulcer, erosion, bleeding, lymphangiectasia, and angiectasiaImage classificationRAt-CapsNetBinary (Acc: 98.51%, Prec: 99.02%, Sens: 98.45%, F1 score: 98.73%), 3-class Acc: 97.92%, 4-class: 95.65%NoNo
Naoki Higuchi [42]Ulcerative colitisImage classification for severity scoringResNet50Acc: 97.3%NoNo
João Pedro Sousa Ferreira [48]Detection of ulcers and erosions in Crohn’s diseaseImage classificationXceptionAcc: 92.4%, Sens: 90.0%, Spec: 96.0%, Prec: 96.6%, NPV: 99.5%, AUC: 1NoNo
Saqib Mahmood [47]Detection of peptic Ulcer and other GI disordersImage classificationGI Disease-Detection Network with BL-SMOTEAcc: 98.9%, Prec: 98.9%, Sens: 98.8% AUC: 0.99, F1 score: 98.9%NoNo
Meryem Souaidi [114]Polyp detectionObject detectionMP-FSSDAverage mAP: 93.4%5-fold cross-validationYes
Miguel Mascarenhas [38]Protruding polyps, epithelial tumors, and subepithelial lesionsImage classificationXceptionAcc: 95.3%, Sens: 90.0%, Spec: 99.1%, AUC: 0.993-fold cross-validationNo
Tom Kratter [46]Detection and grading of ulcerImage classification and gradingEfficientNetB4Performance average: Acc: 87.87%, F1 score: 97.8%, AUC: 0.975-fold cross-validationNo
Xia Xie [33]Detection of small bowel abnormalitiesImage classificationEfficientNet and YOLOSens: 93.45%NoYes
SangYup Oh [75]Detection of small bowel lesionsVideo classificationTransformer-based VWCE-NetSens: 95.1%, Spec: 83.4%NoNo
Raizy Kellerman [41]Predicting the need for biological therapy in Crohn’s DiseaseVideo classificationTimeSformerAcc: 81%, Prec: 81%, Sens: 75%, Spec: 84% AUC: 0.86%5-fold cross-validationNo
Mehrdokht Bordbar [34]Detection of ulcers, bleeding, polyps, and erosions, and vascular abnormalitiesImage classification3D-CNNAcc: 99.20%, Sens: 98.92%, Prec: 99.51%, NPV: 98.92%, F1 score: 99.21%NoYes
Sharib Ali [37]Polyp detection and segmentationImage classification and segmentationFCN-8s, U-Net, PSPNet, DeepLabV3+, ResNet-UNetSingle frame: (Acc: 98%, Dice score: 0.82, Prec: 92%, Sens: 81%), Sequence-based (Acc: 97%, Dice score: 0.71, Prec: 90%, Sens: 73NoYes
Ye Chu [55]Segmentation angiodysplasiasImage segmentationResNet50 with feature fusionAcc: 99%, mIOU: 0.69, PPV: 94.27%, NPV: 98.74%NoNo
Sofia A. Athanasiou [72]Organ boundary detectionImage classificationCNNAcc: 95.56%, Sens: 91.82%,NoYes
Akihiko Sumioka [115]Disease surveillance evaluation of primary small-bowel follicular lymphomaImage classification and disease surveillanceEfficientDetAcc: 85.6%, Sens: 81.2%, Spec: 88.6%, AUC: 0.91NoNo
Javeria Naz [65]Detection of GI abnormalitiesImage classificationXcepNet23 and ResNet18Average performance (Acc: 99.62%, Sens: 99.34%, Spec: 99.94%, Prec: 99.34%, F1 score: 100%)5-fold cross-validationNo
Ayako Nakada [74]Detection of ulcerations, vascular lesions, and tumorsObject detectionRevised RetinaNetPerformance average: Acc: 93.87%, Sens: 89.10%, Spec: 94.73%, IoU: 82.33%, AUC: 0.995-fold cross-validationNo
Tomonori Aoki [116]Detection of mucosal breaks, angioectasia, protruding lesionsImage classificationSSD, EfficientDetSSD Sens: 67%, EfficientDet Sens: 79%NoNo
Ali Sahafi [22]Polyps segmentationImage segmentationYOLO-V8Prec: 98%, Recall: 97.9%, mAP: 97.4%, Dice score: 97.94%NoNo
Rui-Ya Zhang [54]Various small bowel lesions with bleeding risksImage classification and lesion detectionResNet-50 and YOLO-V5Acc: 98.96%, Sens: 99.17%, Spec: 99.92%, AUC: 0.98NoNo
Tsedeke Temesgen Habe [81]Detection of polyps, angiectasia, erosions, ulcers, and bleedingObject detection and classificationSSD300, EfficientDet, Faster R-CNN, RetinaNet, YOLOv3, RTMDet, Faster R-CNN+ResNetBest performance: Prec: 99.67%, Sens: 99.67%, Spec: 99.97%, F1 score: 99.67%5-fold cross-validationNo
Mohamed Achraf Belabbes [100]Polyp detectionObject detection and localizationSPDNet (SSD based)Performance average: mAP: 91.7%5-fold cross-validationYes
Kyung Seok Choi [53]Detection of small-bowel lesionsImage classificationVGGNetAcc: 96%, Sens: 96.5%, Spec: 95.8%, AUC: 0.99NoNo
Akihito Yokote [70]Detection of small-bowel lesionsImage classification and object detectionYOLOv5Sens: 91%, Prec: 67.6%. F1 score: 77.6%, AUC: 0.983-fold cross-validationNo
Dong Liu [98]Segmentation of colonic polypsImage segmentationNA-SegFormerPerformance average: Dice Score: 90.54%, Acc: 93.04%, IoU: 85.71%, Prec: 86.84%, Recall: 95.64%5-fold cross-validationYes
Xia Xie [76]Detection of gastric and small bowel lesionsImage classificationEfficientNet and YOLOPerformance average: Acc: 96.93%, Sens: 97.77%, Spec: 99.80%NoNo
Seung-Joo Nam [83]Real-time organ localization and transit time estimationLocalization and transit timeHybrid CNNAverage classification performance (Acc: 97.1%, F1 score: 97.1%, Sens: 96.87%, Spec: 98.33%), transit time performance (gastric transit time error: 4.3 ± 9.7 min, small bowel transit time error: 24.7 ± 33.8 min )NoNo
Lan Li [64]Detection of small-bowel lesionsImage classification and localizationYOLOv5Performance average: Acc: 92.76%, Sens: 93.86%, Spec: 91.54%NoYes
Esmaeil S. Nadimi [36]Colonic polyp detection, size estimation and neoplastic classificationRecognition and characterisationNasNetLarge-based recognition; VGG16-based classifierRecognition: Sensitivity 99.9%, Specificity 99.4%, NPV 99.8%; Characterisation: Sensitivity 82%, Specificity 80%, Accuracy 81%NoNo
Yeong Seok Kwon [117]Obscure gastrointestinal bleeding (erosions/ulcers, angiodysplasia, bleeding)Lesion classificationDenseNet201 ensemble with post filterAccuracy 99.4% (erosions/ulcers), 99.8% (angiodysplasia), 99.9% (bleeding); Sensitivity 90.1% (erosions/ulcers), 100% (angiodysplasia & bleeding); Specificity 99.7%, 99.8%, 99.9%NoYes
Miguel Jose Mascarenhas Saraiva [35]Protruding lesions (polyps, epithelial tumours, subepithelial lesions)Image classificationModified ResNetSensitivity: 79.7%, Specificity 96.5%, PPV 81.5%, NPV 96.0%, Accuracy 93.7%YesNo
Patricia Andrade [118]Ulcers and erosions in Crohn’s diseaseDetectionConvolutional neural networkAI review: Sensitivity 90.2%, Specificity 84.4%, PPV 76.1%, NPV 94.0%, Accuracy 86.5%NoYes
Miguel Mascarenhas Saraiva [119]Pleomorphic lesions across entire GI tractDetectionDeep neural networkSensitivity 97.5%, Detection rate 96.1% (605/635 lesions)NoYes
Maxime Le Floch [78]Multi-label classification of technical artefacts, view quality, anatomical segments and pathologiesMulti-task classificationResNet-50Anatomical segments: Micro F1 0.89, Macro F1 0.71, Macro accuracy 0.81, Micro accuracy 0.89YesNo
Miguel Martins [63]Pleomorphic oesophageal and gastric lesionsDetectionCNN with cross-covariance transformer (XCiT)Test set: Sensitivity 92.2%, Specificity 95.1%, PPV 79.3%, NPV 98.3%, Accuracy 94.6%YesNo
Charles Houdeville [52]Small bowel capsule endoscopy lesion detectionClassificationRandom forestSpecificity 91.1%, AUC 0.873, Accuracy 84.2%NoYes
Jian Chen [62]Multiple small bowel lesionsMulti-task classificationFocalNet (transformer-based)Precision 88.12%, Sensitivity 85.69%, F1 85.84%, Specificity 98.58%, Accuracy 85.69%, AUC 0.98NoYes
Bernardo Rosa [71]Potential haemorrhagic lesions across entire GI tractDetectionAI PCE methodSensitivity 95.2%, Specificity 97.3%, PPV 98.4%, NPV 92.3%, AUROC 0.963NoYes
Acc: Accuracy, AUC: Area Under the Curve, ANN: Artificial Neural Network, BL-SMOTE: Borderline Synthetic Minority Over-sampling Technique, CE: Capsule Endoscopy, CNN: Convolutional Neural Network, COCO: Common Objects in Context, GI: Gastrointestinal, IoU: Intersection over Union, mAP: Mean Average Precision, mIOU: Mean Intersection over Union, MAE: Mean Absolute Error, NPV: Negative Predictive Value, PASCAL VOC: Visual Object Classes, PPV: Positive Predictive Value, Pr: Precision, Sens: Sensitivity, SimCLR: Simple Framework for Contrastive Learning of Visual Representations, SGD: Stochastic Gradient Descent, SPDNet: Symmetric Positive Definite Network, Spec: Specificity, SSD: Single Shot Multibox Detector, TLBA: Triplet Loss-Based Approach, WCE: Wireless Capsule Endoscopy.

References

  1. Zhang, Y.; Chu, X.; Wang, L.; Yang, H. Global patterns in the epidemiology, cancer risk, and surgical implications of inflammatory bowel disease. Gastroenterol. Rep. 2024, 12, goae053. [Google Scholar] [CrossRef]
  2. Li, M.; Cao, S.; Xu, R.H. Global trends and epidemiological shifts in gastrointestinal cancers: Insights from the past four decades. Cancer Commun. 2025, 45, 774–788. [Google Scholar] [CrossRef] [PubMed]
  3. Yao, L.; Nie, J.; Kong, L.; Shao, S.; Xu, X. The global epidemiology of gastrointestinal cancer burden attributable to dietary risks: A systematic analysis of the Global Burden of Disease Study 2021. Front. Nutr. 2025, 12, 1677735. [Google Scholar] [CrossRef] [PubMed]
  4. Caron, B.; Honap, S.; Peyrin-Biroulet, L. Epidemiology of inflammatory bowel disease across the ages in the era of advanced therapies. J. Crohn’s Colitis 2024, 18, ii3–ii15. [Google Scholar] [CrossRef] [PubMed]
  5. Liu, Y.; Li, J.; Yang, G.; Meng, D.; Long, X.; Wang, K. Global burden of inflammatory bowel disease in the elderly: Trends from 1990 to 2021 and projections to 2051. Front. Aging 2024, 5, 1479928. [Google Scholar] [CrossRef]
  6. Rose, T.C.; Pennington, A.; Kypridemos, C.; Chen, T.; Subhani, M.; Hanefeld, J.; Ricciardiello, L.; Barr, B. Analysis of the burden and economic impact of digestive diseases and investigation of research gaps and priorities in the field of digestive health in the European Region-White Book 2: Executive summary. United Eur. Gastroenterol. J. 2022, 10, 657–662. [Google Scholar] [CrossRef]
  7. Peery, A.F.; Dellon, E.S.; Lund, J.; Crockett, S.D.; McGowan, C.E.; Bulsiewicz, W.J.; Gangarosa, L.M.; Thiny, M.T.; Stizenberg, K.; Morgan, D.R.; et al. Burden of gastrointestinal disease in the United States: 2012 update. Gastroenterology 2012, 143, 1179–1187. [Google Scholar] [CrossRef]
  8. Rey, J.F. Gastric examination by guided capsule endoscopy: A new era. Lancet Gastroenterol. Hepatol. 2021, 6, 879–880. [Google Scholar] [CrossRef]
  9. Oka, P.; McAlindon, M.; Sidhu, R. Capsule endoscopy-a non-invasive modality to investigate the GI tract: Out with the old and in with the new? Expert Rev. Gastroenterol. Hepatol. 2022, 16, 591–599. [Google Scholar] [CrossRef]
  10. Habe, T.T.; Haataja, K.; Toivanen, P. Precision enhancement in wireless capsule endoscopy: A novel transformer-based approach for real-time video object detection. Front. Artif. Intell. 2025, 8, 1529814. [Google Scholar] [CrossRef]
  11. Rosa, B.; Cotter, J. Capsule endoscopy and panendoscopy: A journey to the future of gastrointestinal endoscopy. World J. Gastroenterol. 2024, 30, 1270. [Google Scholar] [CrossRef]
  12. Pennazio, M.; Rondonotti, E.; Despott, E.J.; Dray, X.; Keuchel, M.; Moreels, T.; Sanders, D.S.; Spada, C.; Carretero, C.; Valdivia, P.C.; et al. Small-bowel capsule endoscopy and device-assisted enteroscopy for diagnosis and treatment of small-bowel disorders: European Society of Gastrointestinal Endoscopy (ESGE) Guideline–Update 2022. Endoscopy 2023, 55, 58–95. [Google Scholar] [CrossRef]
  13. Spada, C.; Piccirelli, S.; Hassan, C.; Ferrari, C.; Toth, E.; González-Suárez, B.; Keuchel, M.; McAlindon, M.; Finta, Á; Rosztóczy, A.; et al. AI-assisted capsule endoscopy reading in suspected small bowel bleeding: A multicentre prospective study. Lancet Digit. Health 2024, 6, e345–e353. [Google Scholar] [CrossRef] [PubMed]
  14. Morera, H.; Warman, R.; Anudu, A.; Uche, C.; Radosavljevic, I.; Reddy, N.; Kayastha, A.; Baviriseaty, N.; Mhaskar, R.; Borkowski, A.A.; et al. Reduction of video capsule endoscopy reading times using deep learning with small data. Algorithms 2022, 15, 339. [Google Scholar] [CrossRef]
  15. Cortegoso Valdivia, P.; Deding, U.; Bjørsum-Meyer, T.; Baatrup, G.; Fernández-Urién, I.; Dray, X.; Boal-Carvalho, P.; Ellul, P.; Toth, E.; Rondonotti, E.; et al. Inter/intra-observer agreement in video-capsule endoscopy: Are we getting it all wrong? A systematic review and meta-analysis. Diagnostics 2022, 12, 2400. [Google Scholar] [CrossRef] [PubMed]
  16. Vats, A.; Mohammed, A.; Pedersen, M. From labels to priors in capsule endoscopy: A prior guided approach for improving generalization with few labels. Sci. Rep. 2022, 12, 15708. [Google Scholar] [CrossRef]
  17. Soffer, S.; Klang, E.; Shimon, O.; Nachmias, N.; Eliakim, R.; Ben-Horin, S.; Kopylov, U.; Barash, Y. Deep learning for wireless capsule endoscopy: A systematic review and meta-analysis. Gastrointest. Endosc. 2020, 92, 831–839. [Google Scholar] [CrossRef]
  18. Eker, B.U.; Solak, F.Z. Robust, explainable, and statistically validated gastrointestinal image analysis using modern deep learning architectures. Signal Image Video Process. 2025, 19, 1265. [Google Scholar] [CrossRef]
  19. Oğuz, F.E.; Alkan, A.; Klepaczko, A.; Strumillo, P.; İspiroğlu, M. Artificial intelligence supported colonoscopy bowel preparation assessment: A video-based approach. Signal Image Video Process. 2025, 19, 946. [Google Scholar] [CrossRef]
  20. Bchir, O.; Ben Ismail, M.M.; AlZahrani, N. Multiple bleeding detection in wireless capsule endoscopy. Signal Image Video Process. 2019, 13, 121–126. [Google Scholar] [CrossRef]
  21. Da Rio, L.; Spadaccini, M.; Parigi, T.L.; Gabbiadini, R.; Dal Buono, A.; Busacca, A.; Maselli, R.; Fugazza, A.; Colombo, M.; Carrara, S.; et al. Artificial intelligence and inflammatory bowel disease: Where are we going? World J. Gastroenterol. 2023, 29, 508. [Google Scholar] [CrossRef] [PubMed]
  22. Sahafi, A.; Koulaouzidis, A.; Lalinia, M. Polypoid lesion segmentation using YOLO-V8 network in wireless video capsule endoscopy images. Diagnostics 2024, 14, 474. [Google Scholar] [CrossRef] [PubMed]
  23. Souaidi, M.; Lafraxo, S.; Kerkaou, Z.; El Ansari, M.; Koutti, L. A multiscale polyp detection approach for gi tract images based on improved densenet and single-shot multibox detector. Diagnostics 2023, 13, 733. [Google Scholar] [CrossRef] [PubMed]
  24. Kusters, C.H.; Jaspers, T.J.; Boers, T.G.; Jong, M.R.; Jukema, J.B.; Fockens, K.N.; de Groof, A.J.; Bergman, J.J.; van der Sommen, F.; De With, P.H. Will Transformers change gastrointestinal endoscopic image analysis? A comparative analysis between CNNs and Transformers, in terms of performance, robustness and generalization. Med. Image Anal. 2025, 99, 103348. [Google Scholar] [CrossRef]
  25. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef]
  26. Moons, K.G.; de Groot, J.A.; Bouwmeester, W.; Vergouwe, Y.; Mallett, S.; Altman, D.G.; Reitsma, J.B.; Collins, G.S. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: The CHARMS checklist. PLoS Med. 2014, 11, e1001744. [Google Scholar] [CrossRef]
  27. Whiting, P.F.; Rutjes, A.W.; Westwood, M.E.; Mallett, S.; Deeks, J.J.; Reitsma, J.B.; Leeflang, M.M.; Sterne, J.A.; Bossuyt, P.M. QUADAS-2 Group*. QUADAS-2: A revised tool for the quality assessment of diagnostic accuracy studies. Ann. Intern. Med. 2011, 155, 529–536. [Google Scholar] [CrossRef]
  28. DerSimonian, R.; Laird, N. Meta-analysis in clinical trials. Control. Clin. Trials 1986, 7, 177–188. [Google Scholar] [CrossRef]
  29. Naemi, A.; Schmidt, T.; Mansourvar, M.; Ebrahimi, A.; Wiil, U.K. Quantifying the impact of addressing data challenges in prediction of length of stay. BMC Med. Inform. Decis. Mak. 2021, 21, 298. [Google Scholar] [CrossRef]
  30. Leenhardt, R.; Vasseur, P.; Li, C.; Saurin, J.C.; Rahmi, G.; Cholet, F.; Becq, A.; Marteau, P.; Histace, A.; Dray, X.; et al. A neural network algorithm for detection of GI angiectasia during small-bowel capsule endoscopy. Gastrointest. Endosc. 2019, 89, 189–194. [Google Scholar] [CrossRef]
  31. Wang, S.; Xing, Y.; Zhang, L.; Gao, H.; Zhang, H. Deep convolutional neural network for ulcer recognition in wireless capsule endoscopy: Experimental feasibility and optimization. Comput. Math. Methods Med. 2019, 2019, 7546215. [Google Scholar] [CrossRef]
  32. Saito, H.; Aoki, T.; Aoyama, K.; Kato, Y.; Tsuboi, A.; Yamada, A.; Fujishiro, M.; Oka, S.; Ishihara, S.; Matsuda, T.; et al. Automatic detection and classification of protruding lesions in wireless capsule endoscopy images based on a deep convolutional neural network. Gastrointest. Endosc. 2020, 92, 144–151. [Google Scholar] [CrossRef] [PubMed]
  33. Xie, X.; Xiao, Y.F.; Zhao, X.Y.; Li, J.J.; Yang, Q.Q.; Peng, X.; Nie, X.B.; Zhou, J.Y.; Zhao, Y.B.; Yang, H.; et al. Development and validation of an artificial intelligence model for small bowel capsule endoscopy video review. JAMA Netw. Open 2022, 5, e2221992. [Google Scholar] [CrossRef] [PubMed]
  34. Bordbar, M.; Helfroush, M.S.; Danyali, H.; Ejtehadi, F. Wireless capsule endoscopy multiclass classification using three-dimensional deep convolutional neural network model. BioMed. Eng. Online 2023, 22, 124. [Google Scholar] [CrossRef] [PubMed]
  35. Saraiva, M.J.M.; Almeida, M.J.; Martins, M.; Afonso, J.; Ribeiro, T.; Cardoso, P.M.M.S.; Mendes, F.M.C.S.; Mota, J.; Andrade, A.P.; Cardoso, H.; et al. Deep learning and capsule endoscopy: Automatic panendoscopic detection of protruding lesions. BMJ Open Gastroenterol. 2025, 12, e001655. [Google Scholar] [CrossRef]
  36. Nadimi, E.S.; Braun, J.M.; Schelde-Olesen, B.; Khare, S.; Gogineni, V.C.; Blanes-Vidal, V.; Baatrup, G. Towards full integration of explainable artificial intelligence in colon capsule endoscopy’s pathway. Sci. Rep. 2025, 15, 5960. [Google Scholar] [CrossRef]
  37. Ali, S.; Jha, D.; Ghatwary, N.; Realdon, S.; Cannizzaro, R.; Salem, O.E.; Lamarque, D.; Daul, C.; Riegler, M.A.; Anonsen, K.V.; et al. A multi-centre polyp detection and segmentation dataset for generalisability assessment. Sci. Data 2023, 10, 75. [Google Scholar] [CrossRef]
  38. Mascarenhas, M.; Afonso, J.; Ribeiro, T.; Cardoso, H.; Andrade, P.; Ferreira, J.P.; Saraiva, M.M.; Macedo, G. Performance of a deep learning system for automatic diagnosis of protruding lesions in colon capsule endoscopy. Diagnostics 2022, 12, 1445. [Google Scholar] [CrossRef]
  39. Pascual, G.; Laiz, P.; García, A.; Wenzek, H.; Vitrià, J.; Seguí, S. Time-based self-supervised learning for Wireless Capsule Endoscopy. Comput. Biol. Med. 2022, 146, 105631. [Google Scholar] [CrossRef]
  40. Yuan, Y.; Qin, W.; Ibragimov, B.; Zhang, G.; Han, B.; Meng, M.Q.H.; Xing, L. Densely connected neural network with unbalanced discriminant and category sensitive constraints for polyp recognition. IEEE Trans. Autom. Sci. Eng. 2019, 17, 574–583. [Google Scholar] [CrossRef]
  41. Kellerman, R.; Bleiweiss, A.; Samuel, S.; Margalit-Yehuda, R.; Aflalo, E.; Barzilay, O.; Ben-Horin, S.; Eliakim, R.; Zimlichman, E.; Soffer, S.; et al. Spatiotemporal analysis of small bowel capsule endoscopy videos for outcomes prediction in Crohn’s disease. Ther. Adv. Gastroenterol. 2023, 16, 17562848231172556. [Google Scholar] [CrossRef]
  42. Higuchi, N.; Hiraga, H.; Sasaki, Y.; Hiraga, N.; Igarashi, S.; Hasui, K.; Ogasawara, K.; Maeda, T.; Murai, Y.; Tatsuta, T.; et al. Automated evaluation of colon capsule endoscopic severity of ulcerative colitis using ResNet50. PLoS ONE 2022, 17, e0269728. [Google Scholar] [CrossRef]
  43. Klang, E.; Grinman, A.; Soffer, S.; Margalit Yehuda, R.; Barzilay, O.; Amitai, M.M.; Konen, E.; Ben-Horin, S.; Eliakim, R.; Barash, Y.; et al. Automated detection of Crohn’s disease intestinal strictures on capsule endoscopy images using deep neural networks. J. Crohn’s Colitis 2021, 15, 749–756. [Google Scholar] [CrossRef] [PubMed]
  44. Barash, Y.; Azaria, L.; Soffer, S.; Yehuda, R.M.; Shlomi, O.; Ben-Horin, S.; Eliakim, R.; Klang, E.; Kopylov, U. Ulcer severity grading in video capsule images of patients with Crohn’s disease: An ordinal neural network solution. Gastrointest. Endosc. 2021, 93, 187–192. [Google Scholar] [CrossRef] [PubMed]
  45. Klang, E.; Barash, Y.; Margalit, R.Y.; Soffer, S.; Shimon, O.; Albshesh, A.; Ben-Horin, S.; Amitai, M.M.; Eliakim, R.; Kopylov, U. Deep learning algorithms for automated detection of Crohn’s disease ulcers by video capsule endoscopy. Gastrointest. Endosc. 2020, 91, 606–613. [Google Scholar] [CrossRef] [PubMed]
  46. Kratter, T.; Shapira, N.; Lev, Y.; Mauda, O.; Moshkovitz, Y.; Shitrit, R.; Konyo, S.; Ukashi, O.; Dar, L.; Shlomi, O.; et al. Deep learning multi-domain model provides accurate detection and grading of mucosal ulcers in different capsule endoscopy types. Diagnostics 2022, 12, 2490. [Google Scholar] [CrossRef]
  47. Mahmood, S.; Fareed, M.M.S.; Ahmed, G.; Dawood, F.; Zikria, S.; Mostafa, A.; Jilani, S.F.; Asad, M.; Aslam, M. A robust deep model for classification of peptic ulcer and other digestive tract disorders using endoscopic images. Biomedicines 2022, 10, 2195. [Google Scholar] [CrossRef]
  48. Ferreira, J.P.S.; de Mascarenhas Saraiva, M.J.d.Q.e.C.; Afonso, J.P.L.; Ribeiro, T.F.C.; Cardoso, H.M.C.; Ribeiro Andrade, A.P.; de Mascarenhas Saraiva, M.N.G.; Parente, M.P.L.; Natal Jorge, R.; Lopes, S.I.O.; et al. Identification of ulcers and erosions by the novel Pillcam™ Crohn’s capsule using a convolutional neural network: A multicentre pilot study. J. Crohn’s Colitis 2022, 16, 169–172. [Google Scholar] [CrossRef]
  49. Afonso, J.; Saraiva, M.M.; Ferreira, J.P.S.; Cardoso, H.; Ribeiro, T.; Andrade, P.; Parente, M.; Jorge, R.N.; Macedo, G. Automated detection of ulcers and erosions in capsule endoscopy images using a convolutional neural network. Med. Biol. Eng. Comput. 2022, 60, 719–725. [Google Scholar] [CrossRef]
  50. Hwang, Y.; Lee, H.H.; Park, C.; Tama, B.A.; Kim, J.S.; Cheung, D.Y.; Chung, W.C.; Cho, Y.S.; Lee, K.M.; Choi, M.G.; et al. Improved classification and localization approach to small bowel capsule endoscopy using convolutional neural network. Dig. Endosc. 2021, 33, 598–607. [Google Scholar] [CrossRef]
  51. Aoki, T.; Yamada, A.; Aoyama, K.; Saito, H.; Tsuboi, A.; Nakada, A.; Niikura, R.; Fujishiro, M.; Oka, S.; Ishihara, S.; et al. Automatic detection of erosions and ulcerations in wireless capsule endoscopy images based on a deep convolutional neural network. Gastrointest. Endosc. 2019, 89, 357–363. [Google Scholar] [CrossRef] [PubMed]
  52. Houdeville, C.; Souchaud, M.; Leenhardt, R.; Goltstein, L.C.; Velut, G.; Beaumont, H.; Dray, X.; Histace, A. Toward automated small bowel capsule endoscopy reporting using a summarizing machine learning algorithm: The SUM UP study. Clin. Res. Hepatol. Gastroenterol. 2025, 49, 102509. [Google Scholar] [CrossRef] [PubMed]
  53. Choi, K.S.; Park, D.; Kim, J.S.; Cheung, D.Y.; Lee, B.I.; Cho, Y.S.; Kim, J.I.; Lee, S.; Lee, H.H. Deep learning in negative small-bowel capsule endoscopy improves small-bowel lesion detection and diagnostic yield. Dig. Endosc. 2024, 36, 437–445. [Google Scholar] [CrossRef] [PubMed]
  54. Zhang, R.Y.; Qiang, P.P.; Cai, L.J.; Li, T.; Qin, Y.; Zhang, Y.; Zhao, Y.Q.; Wang, J.P. Automatic detection of small bowel lesions with different bleeding risks based on deep learning models. World J. Gastroenterol. 2024, 30, 170. [Google Scholar] [CrossRef]
  55. Chu, Y.; Huang, F.; Gao, M.; Zou, D.W.; Zhong, J.; Wu, W.; Wang, Q.; Shen, X.N.; Gong, T.T.; Li, Y.Y.; et al. Convolutional neural network-based segmentation network applied to image recognition of angiodysplasias lesion under capsule endoscopy. World J. Gastroenterol. 2023, 29, 879. [Google Scholar] [CrossRef]
  56. Mascarenhas Saraiva, M.; Ribeiro, T.; Afonso, J.; Ferreira, J.P.; Cardoso, H.; Andrade, P.; Parente, M.P.; Jorge, R.N.; Macedo, G. Artificial intelligence and capsule endoscopy: Automatic detection of small bowel blood content using a convolutional neural network. GE-Port. J. Gastroenterol. 2022, 29, 331–338. [Google Scholar] [CrossRef]
  57. Mascarenhas Saraiva, M.; Ribeiro, T.; Afonso, J.; Andrade, P.; Cardoso, P.; Ferreira, J.; Cardoso, H.; Macedo, G. Deep Learning and Device-Assisted Enteroscopy: Automatic Detection of Gastrointestinal Angioectasia. Medicina 2021, 57, 1378. [Google Scholar] [CrossRef]
  58. Rustam, F.; Siddique, M.A.; Siddiqui, H.U.R.; Ullah, S.; Mehmood, A.; Ashraf, I.; Choi, G.S. Wireless capsule endoscopy bleeding images classification using CNN based model. IEEE Access 2021, 9, 33675–33688. [Google Scholar] [CrossRef]
  59. Caroppo, A.; Leone, A.; Siciliano, P. Deep transfer learning approaches for bleeding detection in endoscopy images. Comput. Med. Imaging Graph. 2021, 88, 101852. [Google Scholar] [CrossRef]
  60. Aoki, T.; Yamada, A.; Kato, Y.; Saito, H.; Tsuboi, A.; Nakada, A.; Niikura, R.; Fujishiro, M.; Oka, S.; Ishihara, S.; et al. Automatic detection of blood content in capsule endoscopy images based on a deep convolutional neural network. J. Gastroenterol. Hepatol. 2020, 35, 1196–1200. [Google Scholar] [CrossRef]
  61. Kundu, A.K.; Fattah, S.A. Probability density function based modeling of spatial feature variation in capsule endoscopy data for automatic bleeding detection. Comput. Biol. Med. 2019, 115, 103478. [Google Scholar] [CrossRef]
  62. Chen, J.; Wang, H.; Zhang, Z.; Xia, K.; Gao, F.; Xu, X.; Wang, G. Development and Validation of a Multi-Task Artificial Intelligence-Assisted System for Small Bowel Capsule Endoscopy. Int. J. Gen. Med. 2025, 18, 2521–2536. [Google Scholar] [CrossRef]
  63. Martins, M.; Mascarenhas, M.J.; Almeida, M.J.; Afonso, J.; Ribeiro, T.; Cardoso, P.; Mendes, F.; Mota, J.; Andrade, P.; Cardoso, H.; et al. A ubiquitous and interoperable deep learning model for automatic detection of pleomorphic gastroesophageal lesions. Sci. Rep. 2025, 15, 22889. [Google Scholar] [CrossRef]
  64. Li, L.; Yang, L.; Zhang, B.; Yan, G.; Bao, Y.; Zhu, R.; Li, S.; Wang, H.; Chen, M.; Jin, C.; et al. Automated detection of small bowel lesions based on capsule endoscopy using deep learning algorithm. Clin. Res. Hepatol. Gastroenterol. 2024, 48, 102334. [Google Scholar] [CrossRef]
  65. Naz, J.; Sharif, M.I.; Sharif, M.I.; Kadry, S.; Rauf, H.T.; Ragab, A.E. A comparative analysis of optimization algorithms for gastrointestinal abnormalities recognition and classification based on ensemble XcepNet23 and ResNet18 features. Biomedicines 2023, 11, 1723. [Google Scholar] [CrossRef]
  66. Alam, M.J.; Rashid, R.B.; Fattah, S.A.; Saquib, M. Rat-capsnet: A deep learning network utilizing attention and regional information for abnormality detection in wireless capsule endoscopy. IEEE J. Transl. Eng. Health Med. 2022, 10, 3300108. [Google Scholar] [CrossRef] [PubMed]
  67. Jain, S.; Seal, A.; Ojha, A.; Yazidi, A.; Bures, J.; Tacheci, I.; Krejcar, O. A deep CNN model for anomaly detection and localization in wireless capsule endoscopy images. Comput. Biol. Med. 2021, 137, 104789. [Google Scholar] [CrossRef] [PubMed]
  68. Xia, J.; Xia, T.; Pan, J.; Gao, F.; Wang, S.; Qian, Y.Y.; Wang, H.; Zhao, J.; Jiang, X.; Zou, W.B.; et al. Use of artificial intelligence for detection of gastric lesions by magnetically controlled capsule endoscopy. Gastrointest. Endosc. 2021, 93, 133–139. [Google Scholar] [CrossRef] [PubMed]
  69. Ghosh, T.; Chakareski, J. Deep transfer learning for automated intestinal bleeding detection in capsule endoscopy imaging. J. Digit. Imaging 2021, 34, 404–417. [Google Scholar] [CrossRef]
  70. Yokote, A.; Umeno, J.; Kawasaki, K.; Fujioka, S.; Fuyuno, Y.; Matsuno, Y.; Yoshida, Y.; Imazu, N.; Miyazono, S.; Moriyama, T.; et al. Small bowel capsule endoscopy examination and open access database with artificial intelligence: The SEE-artificial intelligence project. DEN Open 2024, 4, e258. [Google Scholar] [CrossRef]
  71. Rosa, B.; Saraiva, M.J.M.; Afonso, J.; Gonçalves, T.C.; Mendes, F.; Moreira, M.J.; Martins, M.; de Castro, F.D.; Ribeiro, T.; Cardoso, P.; et al. Artificial intelligence-assisted versus conventional reading in pan-intestinal capsule endoscopy for suspected mid-lower gastrointestinal bleeding: A retrospective analysis of a prospective cohort. BMJ Open Gastroenterol. 2025, 12, e001906. [Google Scholar] [CrossRef] [PubMed]
  72. Athanasiou, S.A.; Sergaki, E.S.; Polydorou, A.A.; Polydorou, A.A.; Stavrakakis, G.S.; Afentakis, N.M.; Vardiambasis, I.O.; Zervakis, M.E. Revealing the Boundaries of Selected Gastro-Intestinal (GI) Organs by Implementing CNNs in Endoscopic Capsule Images. Diagnostics 2023, 13, 865. [Google Scholar] [CrossRef] [PubMed]
  73. Tsuboi, A.; Oka, S.; Aoyama, K.; Saito, H.; Aoki, T.; Yamada, A.; Matsuda, T.; Fujishiro, M.; Ishihara, S.; Nakahori, M.; et al. Artificial intelligence using a convolutional neural network for automatic detection of small-bowel angioectasia in capsule endoscopy images. Dig. Endosc. 2020, 32, 382–390. [Google Scholar] [CrossRef] [PubMed]
  74. Nakada, A.; Niikura, R.; Otani, K.; Kurose, Y.; Hayashi, Y.; Kitamura, K.; Nakanishi, H.; Kawano, S.; Honda, T.; Hasatani, K.; et al. Improved Object Detection Artificial Intelligence Using the Revised RetinaNet Model for the Automatic Detection of Ulcerations, Vascular Lesions, and Tumors in Wireless Capsule Endoscopy. Biomedicines 2023, 11, 942. [Google Scholar] [CrossRef]
  75. Oh, S.; Oh, D.; Kim, D.; Song, W.; Hwang, Y.; Cho, N.; Lim, Y.J. Video Analysis of Small Bowel Capsule Endoscopy Using a Transformer Network. Diagnostics 2023, 13, 3133. [Google Scholar] [CrossRef]
  76. Xie, X.; Xiao, Y.F.; Yang, H.; Peng, X.; Li, J.J.; Zhou, Y.Y.; Fan, C.Q.; Meng, R.P.; Huang, B.B.; Liao, X.P.; et al. A new artificial intelligence system for both stomach and small bowel capsule endoscopy. Gastrointest. Endosc. 2024, 100, 878-e1. [Google Scholar] [CrossRef]
  77. Ding, Z.; Shi, H.; Zhang, H.; Meng, L.; Fan, M.; Han, C.; Zhang, K.; Ming, F.; Xie, X.; Liu, H.; et al. Gastroenterologist-level identification of small-bowel diseases and normal variants by capsule endoscopy using a deep-learning model. Gastroenterology 2019, 157, 1044–1054. [Google Scholar] [CrossRef]
  78. Le Floch, M.; Wolf, F.; McIntyre, L.; Weinert, C.; Palm, A.; Volk, K.; Herzog, P.; Kirk, S.H.; Steinhäuser, J.L.; Stopp, C.; et al. Galar-a large multi-label video capsule endoscopy dataset. Sci. Data 2025, 12, 828. [Google Scholar] [CrossRef]
  79. Otani, K.; Nakada, A.; Kurose, Y.; Niikura, R.; Yamada, A.; Aoki, T.; Nakanishi, H.; Doyama, H.; Hasatani, K.; Sumiyoshi, T.; et al. Automatic detection of different types of small-bowel lesions on capsule endoscopy images using a newly developed deep convolutional neural network. Endoscopy 2020, 52, 786–791. [Google Scholar] [CrossRef]
  80. Iakovidis, D.K.; Georgakopoulos, S.V.; Vasilakakis, M.; Koulaouzidis, A.; Plagianakos, V.P. Detecting and locating gastrointestinal anomalies using deep learning and iterative cluster unification. IEEE Trans. Med. Imaging 2018, 37, 2196–2210. [Google Scholar] [CrossRef]
  81. Habe, T.T.; Haataja, K.; Toivanen, P. Efficiency meets Accuracy: Benchmarking Object Detection Models for Pathology Detection in Wireless Capsule Endoscopy. IEEE Access 2024, 12, 126793–126817. [Google Scholar] [CrossRef]
  82. Houdeville, C.; Souchaud, M.; Leenhardt, R.; Beaumont, H.; Benamouzig, R.; McAlindon, M.; Grimbert, S.; Lamarque, D.; Makins, R.; Saurin, J.C.; et al. A multisystem-compatible deep learning-based algorithm for detection and characterization of angiectasias in small-bowel capsule endoscopy. A proof-of-concept study. Dig. Liver Dis. 2021, 53, 1627–1631. [Google Scholar] [CrossRef]
  83. Nam, S.J.; Moon, G.; Park, J.H.; Kim, Y.; Lim, Y.J.; Choi, H.S. Deep Learning-Based Real-Time Organ Localization and Transit Time Estimation in Wireless Capsule Endoscopy. Biomedicines 2024, 12, 1704. [Google Scholar] [CrossRef] [PubMed]
  84. Saraiva, M.; Ferreira, J.; Cardoso, H.; Afonso, J.; Ribeiro, T.; Andrade, P.; Parente, M.; Jorge, R.; Macedo, G. Artificial intelligence and colon capsule endoscopy: Development of an automated diagnostic system of protruding lesions in colon capsule endoscopy. Tech. Coloproctology 2021, 25, 1243–1248. [Google Scholar] [CrossRef] [PubMed]
  85. Saraiva, M.J.M.; Afonso, J.; Ribeiro, T.; Ferreira, J.; Cardoso, H.; Andrade, A.P.; Parente, M.; Natal, R.; Saraiva, M.M.; Macedo, G. Deep learning and capsule endoscopy: Automatic identification and differentiation of small bowel lesions with distinct haemorrhagic potential using a convolutional neural network. BMJ Open Gastroenterol. 2021, 8, e000753. [Google Scholar] [CrossRef] [PubMed]
  86. Zhou, T.; Han, G.; Li, B.N.; Lin, Z.; Ciaccio, E.J.; Green, P.H.; Qin, J. Quantitative analysis of patients with celiac disease by video capsule endoscopy: A deep learning method. Comput. Biol. Med. 2017, 85, 1–6. [Google Scholar] [CrossRef]
  87. Egger, M.; Smith, G.D.; Schneider, M.; Minder, C. Bias in meta-analysis detected by a simple, graphical test. BMJ 1997, 315, 629–634. [Google Scholar] [CrossRef]
  88. Duval, S.; Tweedie, R. Trim and fill: A simple funnel-plot–based method of testing and adjusting for publication bias in meta-analysis. Biometrics 2000, 56, 455–463. [Google Scholar] [CrossRef]
  89. Molder, A.; Balaban, D.V.; Jinga, M.; Molder, C.C. Current evidence on computer-aided diagnosis of celiac disease: Systematic review. Front. Pharmacol. 2020, 11, 341. [Google Scholar] [CrossRef]
  90. Pal, P.; Pooja, K.; Nabi, Z.; Gupta, R.; Tandan, M.; Rao, G.V.; Reddy, N. Artificial intelligence in endoscopy related to inflammatory bowel disease: A systematic review. Indian J. Gastroenterol. 2024, 43, 172–187. [Google Scholar] [CrossRef]
  91. Kim, H.J.; Gong, E.J.; Bang, C.S.; Lee, J.J.; Suk, K.T.; Baik, G.H. Computer-aided diagnosis of gastrointestinal protruded lesions using wireless capsule endoscopy: A systematic review and diagnostic test accuracy meta-analysis. J. Pers. Med. 2022, 12, 644. [Google Scholar] [CrossRef]
  92. Mohan, B.P.; Khan, S.R.; Kassab, L.L.; Ponnada, S.; Chandan, S.; Ali, T.; Dulai, P.S.; Adler, D.G.; Kochhar, G.S. High pooled performance of convolutional neural networks in computer-aided diagnosis of GI ulcers and/or hemorrhage on wireless capsule endoscopy images: A systematic review and meta-analysis. Gastrointest. Endosc. 2021, 93, 356–364. [Google Scholar] [CrossRef] [PubMed]
  93. Bang, C.S.; Lee, J.J.; Baik, G.H. Computer-aided diagnosis of gastrointestinal ulcer and hemorrhage using wireless capsule endoscopy: Systematic review and diagnostic test accuracy meta-analysis. J. Med. Internet Res. 2021, 23, e33267. [Google Scholar] [CrossRef] [PubMed]
  94. Qin, K.; Li, J.; Fang, Y.; Xu, Y.; Wu, J.; Zhang, H.; Li, H.; Liu, S.; Li, Q. Convolution neural network for the diagnosis of wireless capsule endoscopy: A systematic review and meta-analysis. Surg. Endosc. 2022, 36, 16–31. [Google Scholar] [CrossRef] [PubMed]
  95. Naemi, A.; Tashk, A.; Azar, A.S.; Samimi, T.; Tavassoli, G.; Mohasefi, A.B.; Khanshan, E.N.; Najafabad, M.H.; Tarighi, V.; Wiil, U.K.; et al. Applications of Artificial Intelligence for Metastatic Gastrointestinal Cancer: A Systematic Literature Review. Cancers 2025, 17, 558. [Google Scholar] [CrossRef]
  96. Naemi, A.; Schmidt, T.; Mansourvar, M.; Wiil, U.K. Personalized Predictive Models for Identifying Clinical Deterioration Using LSTM in Emergency Departments. In Proceedings of the EFMI-STC, Virtual Event, 26–27 November 2020; pp. 152–156. [Google Scholar]
  97. Smedsrud, P.H.; Thambawita, V.; Hicks, S.A.; Gjestang, H.; Nedrejord, O.O.; Næss, E.; Borgli, H.; Jha, D.; Berstad, T.J.D.; Eskeland, S.L.; et al. Kvasir-Capsule, a video capsule endoscopy dataset. Sci. Data 2021, 8, 142. [Google Scholar] [CrossRef]
  98. Liu, D.; Lu, C.; Sun, H.; Gao, S. NA-segformer: A multi-level transformer model based on neighborhood attention for colonoscopic polyp segmentation. Sci. Rep. 2024, 14, 22527. [Google Scholar] [CrossRef]
  99. Obayya, M.; Al-Wesabi, F.N.; Maashi, M.; Mohamed, A.; Hamza, M.A.; Drar, S.; Yaseen, I.; Alsaid, M.I. Modified salp swarm algorithm with deep learning based gastrointestinal tract disease classification on endoscopic images. IEEE Access 2023, 11, 25959–25967. [Google Scholar] [CrossRef]
  100. Belabbes, M.A.; Oukdach, Y.; Souaidi, M.; Koutti, L.; Charfi, S. Advancements in Polyp Detection: A Developed Single Shot Multibox Detector Approach. IEEE Access 2024, 12, 19199–19215. [Google Scholar] [CrossRef]
  101. Yang, J.; Chang, L.; Li, S.; He, X.; Zhu, T. WCE polyp detection based on novel feature descriptor with normalized variance locality-constrained linear coding. Int. J. Comput. Assist. Radiol. Surg. 2020, 15, 1291–1302. [Google Scholar] [CrossRef]
  102. Lin, T.Y.; Maire, M.; Belongie, S.; Hays, J.; Perona, P.; Ramanan, D.; Dollár, P.; Zitnick, C.L. Microsoft coco: Common objects in context. In Proceedings of the European Conference on Computer Vision; Springer: Berlin/Heidelberg, Germany, 2014; pp. 740–755. [Google Scholar]
  103. Moen, S.; Vuik, F.E.; Kuipers, E.J.; Spaander, M.C. Artificial intelligence in colon capsule endoscopy—A systematic review. Diagnostics 2022, 12, 1994. [Google Scholar] [CrossRef]
  104. Naemi, A.; Sahafi, A. Benchmarking Large Language Models for MIMIC-IV Clinical Note Summarization. J. Healthc. Inform. Res. 2026, 10, 95–115. [Google Scholar] [CrossRef]
  105. Zhu, J.Y.; Park, T.; Isola, P.; Efros, A.A. Unpaired image-to-image translation using cycle-consistent adversarial networks. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2017; pp. 2223–2232. [Google Scholar]
  106. Ganin, Y.; Ustinova, E.; Ajakan, H.; Germain, P.; Larochelle, H.; Laviolette, F.; March, M.; Lempitsky, V. Domain-adversarial training of neural networks. J. Mach. Learn. Res. 2016, 17, 1–35. [Google Scholar]
  107. Sushma, B.; Aparna, P. Summarization of wireless capsule endoscopy video using deep feature matching and motion analysis. IEEE Access 2020, 9, 13691–13703. [Google Scholar] [CrossRef]
  108. Lan, L.; Ye, C.; Wang, C.; Zhou, S. Deep convolutional neural networks for WCE abnormality detection: CNN architecture, region proposal and transfer learning. IEEE Access 2019, 7, 30017–30032. [Google Scholar] [CrossRef]
  109. Deding, U.; Herp, J.; Havshoei, A.; Kobaek-Larsen, M.; Buijs, M.; Nadimi, E.; Baatrup, G. Colon capsule endoscopy versus CT colonography after incomplete colonoscopy. Application of artificial intelligence algorithms to identify complete colonic investigations. United Eur. Gastroenterol. J. 2020, 8, 782–789. [Google Scholar] [CrossRef] [PubMed]
  110. Liao, C.; Wang, C.; Bai, J.; Lan, L.; Wu, X. Deep learning for registration of region of interest in consecutive wireless capsule endoscopy frames. Comput. Methods Programs Biomed. 2021, 208, 106189. [Google Scholar] [CrossRef]
  111. Herp, J.; Deding, U.; Buijs, M.M.; Kroijer, R.; Baatrup, G.; Nadimi, E.S. Feature point tracking-based localization of colon capsule endoscope. Diagnostics 2021, 11, 193. [Google Scholar] [CrossRef]
  112. Aoki, T.; Yamada, A.; Kato, Y.; Saito, H.; Tsuboi, A.; Nakada, A.; Niikura, R.; Fujishiro, M.; Oka, S.; Ishihara, S.; et al. Automatic detection of various abnormalities in capsule endoscopy videos by a deep learning-based system: A multicenter study. Gastrointest. Endosc. 2021, 93, 165–173. [Google Scholar] [CrossRef]
  113. Zhou, J.X.; Yang, Z.; Xi, D.H.; Dai, S.J.; Feng, Z.Q.; Li, J.Y.; Xu, W.; Wang, H. Enhanced segmentation of gastrointestinal polyps from capsule endoscopy images with artifacts using ensemble learning. World J. Gastroenterol. 2022, 28, 5931. [Google Scholar] [CrossRef]
  114. Souaidi, M.; El Ansari, M. A new automated polyp detection network MP-FSSD in WCE and colonoscopy images based fusion single shot multibox detector and transfer learning. IEEE Access 2022, 10, 47124–47140. [Google Scholar] [CrossRef]
  115. Sumioka, A.; Tsuboi, A.; Oka, S.; Kato, Y.; Matsubara, Y.; Hirata, I.; Takigawa, H.; Yuge, R.; Shimamoto, F.; Tada, T.; et al. Disease surveillance evaluation of primary small-bowel follicular lymphoma using capsule endoscopy images based on a deep convolutional neural network (with video). Gastrointest. Endosc. 2023, 98, 968–976. [Google Scholar] [CrossRef]
  116. Aoki, T.; Yamada, A.; Oka, S.; Tsuboi, M.; Kurokawa, K.; Togo, D.; Tanino, F.; Teshima, H.; Saito, H.; Suzuki, R.; et al. Comparison of clinical utility of deep learning-based systems for small-bowel capsule endoscopy reading. J. Gastroenterol. Hepatol. 2024, 39, 157–164. [Google Scholar] [CrossRef] [PubMed]
  117. Kwon, Y.S.; Park, T.Y.; Kim, S.E.; Park, Y.; Lee, J.G.; Lee, S.P.; Kim, K.O.; Jang, H.J.; Yang, Y.J.; Cho, B.J. Deep learning-based localization and lesion detection in capsule endoscopy for patients with suspected small-bowel bleeding. World J. Gastroenterol. 2025, 31, 106819. [Google Scholar] [CrossRef]
  118. Andrade, P.; Mascarenhas, M.; Mendes, F.; Rosa, B.; Cardoso, P.; Afonso, J.; Ribeiro, T.; Martins, M.; Mota, J.; Almeida, M.J.; et al. AI-Assisted Capsule Endoscopy for Detection of Ulcers and Erosions in Crohn’s Disease: A Multicenter Validation Study. Clin. Gastroenterol. Hepatol. 2025, in press. [Google Scholar] [CrossRef]
  119. Mascarenhas Saraiva, M.; Ferreira, J.; Afonso, J.; Mendes, F.; Sonnier, W.P.; Rosa, B.; Ribeiro, T.; Gonçalves, T.C.; Martins, M.; Campelo, P.; et al. Real-Life Clinical Validation of Artificial Intelligence-Assisted Detection and Differentiation of Pleomorphic Lesions in Capsule Endoscopy. Am. J. Gastroenterol. 2025, 121, 993–1000. [Google Scholar] [CrossRef]
Figure 1. WCE detecting GI disorders.
Figure 1. WCE detecting GI disorders.
Diagnostics 16 01269 g001
Figure 2. PRISMA diagram.
Figure 2. PRISMA diagram.
Diagnostics 16 01269 g002
Figure 3. Geographical distribution of included studies.
Figure 3. Geographical distribution of included studies.
Diagnostics 16 01269 g003
Figure 4. Publication trend of included studies.
Figure 4. Publication trend of included studies.
Diagnostics 16 01269 g004
Figure 5. Distribution of number of patients across included studies.
Figure 5. Distribution of number of patients across included studies.
Diagnostics 16 01269 g005
Figure 6. Distribution of the number of patients across included studies.
Figure 6. Distribution of the number of patients across included studies.
Diagnostics 16 01269 g006
Figure 7. Distribution of included studies based on (a) GI tract location, (b) Application type, (c) AI architecture, (d) Validation strategy and dataset features.
Figure 7. Distribution of included studies based on (a) GI tract location, (b) Application type, (c) AI architecture, (d) Validation strategy and dataset features.
Diagnostics 16 01269 g007
Figure 8. ROB analysis of included studies using the QUADAS-2 tool.
Figure 8. ROB analysis of included studies using the QUADAS-2 tool.
Diagnostics 16 01269 g008
Figure 9. Results of meta-analysis by clinical indication: (a) Accuracy, (b) Sensitivity, (c) AUC, and (d) Specificity. Included studies: Saraiva et al. [35], Nadimi et al. [36], Ali et al. [37], Mascarenhas et al. [38], Pascual et al. [39], Yuan et al. [40], Kellerman et al. [41], Higuchi et al. [42], Klang et al. [43], Barash et al. [44], Klang et al. [45], Kratter et al. [46], Mahmood et al. [47], Ferreira et al. [48], Afonso et al. [49], Hwang et al. [50], Wang et al. [31], Aoki et al. [51], Houdeville et al. [52], Choi et al. [53], Zhang et al. [54], Chu et al. [55], Saraiva et al. [56], Saraiva et al. [57], Rustam et al. [58], Caroppo et al. [59], Aoki et al. [60], Kundu et al. [61], Chen et al. [62], Martins et al. [63], Li et al. [64], Naz et al. [65], Bordbar et al. [34], Alam et al. [66], Jain et al. [67], Xia et al. [68], Ghosh et al. [69], Sahafi et al. [22], Saito et al. [32], Yokote et al. [70], Rosa et al. [71], Athanasiou et al. [72], Tsuboi et al. [73], Leenhardt et al. [30], Xie et al. [33], Nakada et al. [74], Oh et al. [75], Xie et al. [76], Ding et al. [77], Floch et al. [78], Kellerman et al. [41], Otani et al. [79], Iakovidis et al. [80], Habe et al. [81], Houdeville et al. [82].
Figure 9. Results of meta-analysis by clinical indication: (a) Accuracy, (b) Sensitivity, (c) AUC, and (d) Specificity. Included studies: Saraiva et al. [35], Nadimi et al. [36], Ali et al. [37], Mascarenhas et al. [38], Pascual et al. [39], Yuan et al. [40], Kellerman et al. [41], Higuchi et al. [42], Klang et al. [43], Barash et al. [44], Klang et al. [45], Kratter et al. [46], Mahmood et al. [47], Ferreira et al. [48], Afonso et al. [49], Hwang et al. [50], Wang et al. [31], Aoki et al. [51], Houdeville et al. [52], Choi et al. [53], Zhang et al. [54], Chu et al. [55], Saraiva et al. [56], Saraiva et al. [57], Rustam et al. [58], Caroppo et al. [59], Aoki et al. [60], Kundu et al. [61], Chen et al. [62], Martins et al. [63], Li et al. [64], Naz et al. [65], Bordbar et al. [34], Alam et al. [66], Jain et al. [67], Xia et al. [68], Ghosh et al. [69], Sahafi et al. [22], Saito et al. [32], Yokote et al. [70], Rosa et al. [71], Athanasiou et al. [72], Tsuboi et al. [73], Leenhardt et al. [30], Xie et al. [33], Nakada et al. [74], Oh et al. [75], Xie et al. [76], Ding et al. [77], Floch et al. [78], Kellerman et al. [41], Otani et al. [79], Iakovidis et al. [80], Habe et al. [81], Houdeville et al. [82].
Diagnostics 16 01269 g009
Figure 10. Results of meta-analysis by GI location: (a) Accuracy, (b) Sensitivity, (c) AUC, and (d) Specificity. Included studies in order of appearance: Alam et al. [66], Pascual et al. [39], Nadimi et al. [36], Ali et al. [37], Mascarenhas et al. [38], Higuchi et al. [42], Yuan et al. [40], Martins et al. [63], Saraiva et al. [35], Nam et al. [83], Naz et al. [65], Bordbar et al. [34], Kratter et al. [46], Xia et al. [68], Rustam et al. [58], Wang et al. [31], Kundu et al. [61], Chen et al. [62], Houdeville et al. [52], Li et al. [64], Choi et al. [53], Zhang et al. [54], Chu et al. [55], Kellerman et al. [41], Ferreira et al. [48], Afonso et al. [49], Saraiva et al. [56], Jain et al. [67], Klang et al. [43], Saraiva et al. [57], Barash et al. [44], Hwang et al. [50], Caroppo et al. [59], Ghosh et al. [69], Aoki et al. [60], Klang et al. [45], Aoki et al. [51], Sahafi et al. [22], Saravia et al. [84], Rosa et al. [71], Xie et al. [33], Athanasiou et al. [72], Yokote et al. [70], Nakada et al. [74], Oh et al. [75], Saravia et al. [56], Saravia et al. [57], Saravia et al. [85], Tsuboi et al. [73], Saito et al. [32], Aoki et al. [51], Leenhardt et al. [30], Xie et al. [76], Ding et al. [77], Iakovidis et al. [80], Otani et al. [79], Zhou et al. [86], Habe et al. [81], Houdeville et al. [82].
Figure 10. Results of meta-analysis by GI location: (a) Accuracy, (b) Sensitivity, (c) AUC, and (d) Specificity. Included studies in order of appearance: Alam et al. [66], Pascual et al. [39], Nadimi et al. [36], Ali et al. [37], Mascarenhas et al. [38], Higuchi et al. [42], Yuan et al. [40], Martins et al. [63], Saraiva et al. [35], Nam et al. [83], Naz et al. [65], Bordbar et al. [34], Kratter et al. [46], Xia et al. [68], Rustam et al. [58], Wang et al. [31], Kundu et al. [61], Chen et al. [62], Houdeville et al. [52], Li et al. [64], Choi et al. [53], Zhang et al. [54], Chu et al. [55], Kellerman et al. [41], Ferreira et al. [48], Afonso et al. [49], Saraiva et al. [56], Jain et al. [67], Klang et al. [43], Saraiva et al. [57], Barash et al. [44], Hwang et al. [50], Caroppo et al. [59], Ghosh et al. [69], Aoki et al. [60], Klang et al. [45], Aoki et al. [51], Sahafi et al. [22], Saravia et al. [84], Rosa et al. [71], Xie et al. [33], Athanasiou et al. [72], Yokote et al. [70], Nakada et al. [74], Oh et al. [75], Saravia et al. [56], Saravia et al. [57], Saravia et al. [85], Tsuboi et al. [73], Saito et al. [32], Aoki et al. [51], Leenhardt et al. [30], Xie et al. [76], Ding et al. [77], Iakovidis et al. [80], Otani et al. [79], Zhou et al. [86], Habe et al. [81], Houdeville et al. [82].
Diagnostics 16 01269 g010
Figure 11. (a) Summary of publication bias assessment across all subgroups, and (b) Comparison of unadjusted versus trim-and-fill-adjusted pooled estimates.
Figure 11. (a) Summary of publication bias assessment across all subgroups, and (b) Comparison of unadjusted versus trim-and-fill-adjusted pooled estimates.
Diagnostics 16 01269 g011
Figure 12. Funnel plots for subgroups with n 10 , (a) Accuracy for Bleeding, Vascular, (b) Accuracy for Mixed, General Abnormality, (c) Sensitivity for Bleeding, Vascular, (d) Sensitivity for Mixed, General Abnormality, (e) Specificity for Bleeding, Vascular.
Figure 12. Funnel plots for subgroups with n 10 , (a) Accuracy for Bleeding, Vascular, (b) Accuracy for Mixed, General Abnormality, (c) Sensitivity for Bleeding, Vascular, (d) Sensitivity for Mixed, General Abnormality, (e) Specificity for Bleeding, Vascular.
Diagnostics 16 01269 g012
Table 1. Research questions related to AI applications in WCE.
Table 1. Research questions related to AI applications in WCE.
QuestionDescription
Q1Which AI models have been used for WCE across different diagnostic tasks?
Q2How do AI models perform across studies based on reported performance metrics?
Q3How do dataset characteristics (size, diversity, imbalance, and annotation quality) influence model performance and generalizability?
Q4What barriers currently limit the clinical adoption of AI in WCE, and how can future research address these barriers?
Table 2. Search criteria of this study.
Table 2. Search criteria of this study.
GroupSearch Criteria
G1—AI keywordsArtificial intelligence, Machine learning, Learning algorithms, Deep learning, Unsupervised machine learning, Supervised learning, Image classification, Object detection, Segmentation, Localization
G2—Technical keywordsCapsule endoscopy, Wireless capsule endoscopy, Video capsule endoscopy
G3—Medical keywordsGastrointestinal, Stomach, Small bowel, Colon, Esophagus, Polyp, Ulcer, Bleeding, Inflammation, Crohn’s disease, Lesion, Celiac, Angiectasias
G4—Document typeEnglish journal articles
G5—Publication year1 January 2000–1 January 2026
G6—Final resultG1 AND G2 AND G3 AND G4 AND G5
Table 3. Inclusion and exclusion criteria of this study.
Table 3. Inclusion and exclusion criteria of this study.
Inclusion CriteriaExclusion Criteria
- AI studies on WCE images.- Studies using traditional statistical models.
- Cohorts comprising relevant patient groups for WCE.- The primary focus is not AI applications on WCE images.
- Journal articles written in English.- Non-English and non-journal articles.
- Review studies.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Sahafi, A.; Koulaouzidis, A.; Naemi, A. Artificial Intelligence in Gastrointestinal Wireless Capsule Endoscopy: A Systematic Literature Review and Meta-Analysis. Diagnostics 2026, 16, 1269. https://doi.org/10.3390/diagnostics16091269

AMA Style

Sahafi A, Koulaouzidis A, Naemi A. Artificial Intelligence in Gastrointestinal Wireless Capsule Endoscopy: A Systematic Literature Review and Meta-Analysis. Diagnostics. 2026; 16(9):1269. https://doi.org/10.3390/diagnostics16091269

Chicago/Turabian Style

Sahafi, Ali, Anastasios Koulaouzidis, and Amin Naemi. 2026. "Artificial Intelligence in Gastrointestinal Wireless Capsule Endoscopy: A Systematic Literature Review and Meta-Analysis" Diagnostics 16, no. 9: 1269. https://doi.org/10.3390/diagnostics16091269

APA Style

Sahafi, A., Koulaouzidis, A., & Naemi, A. (2026). Artificial Intelligence in Gastrointestinal Wireless Capsule Endoscopy: A Systematic Literature Review and Meta-Analysis. Diagnostics, 16(9), 1269. https://doi.org/10.3390/diagnostics16091269

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop