Next Article in Journal
Intestinal Barrier: Mechanisms of Disruption and Strategies for Restoration in Ulcerative Colitis
Previous Article in Journal
The Intestinal Microbiota Profile of Patients with Colon Cancer in Southern Peru: An Exploratory Regional Analysis
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Artificial Intelligence in Helicobacter Pylori Infection: Diagnostic Applications and Emerging Treatment-Related Predictive Uses—A Systematic Review

by
Esteban Zavaleta-Monestel
1,*,
Yennifer Villagra-Hernandez
2,
Jeaustin Mora-Jiménez
3,
Jorge Arturo Villalobos-Madriz
3,
Carolina Rojas-Chinchilla
3,
José Andrés Castro-Gamboa
4,
Luis Guillermo Herrera-Jiménez
4,
Sebastián Arguedas-Chacón
1 and
Christian Campos-Núñez
5
1
Health Research Department, Clinica Biblica, San José 1307-1000, Costa Rica
2
School of Pharmacy, Faculty of Health Sciences, Universidad Latina de Costa Rica, San José 11501, Costa Rica
3
Pharmacy Department, Clinica Biblica, San José 1307-1000, Costa Rica
4
Faculty of Pharmacy, University of Costa Rica, San José 11501-2060, Costa Rica
5
Gastroenterology Service, Clinica Biblica, San José 1307-1000, Costa Rica
*
Author to whom correspondence should be addressed.
Gastrointest. Disord. 2026, 8(2), 23; https://doi.org/10.3390/gidisord8020023
Submission received: 26 March 2026 / Revised: 8 May 2026 / Accepted: 12 May 2026 / Published: 16 May 2026

Abstract

Background: Artificial intelligence (AI) has shown growing potential in the diagnosis of H. pylori infection, particularly through automated analysis of endoscopic images. Emerging studies have also explored treatment-related predictive applications, although this evidence remains limited. The aim of this systematic review was to synthesize current evidence on the use of AI in H. pylori infection, with the primary emphasis on diagnosis and secondary consideration of predictive therapeutic applications. Methods: A systematic review was conducted in accordance with PRISMA 2020 guidelines through searches in PubMed, ScienceDirect, EBSCO, and the Cochrane Library, including articles published between 2020 and 2025. Six studies that employed deep learning or machine learning models, primarily convolutional neural networks and predictive classifiers, were selected. Results: Artificial intelligence models showed consistent diagnostic performance, with accuracies ranging from 79.2% to 94%, sensitivities from 62.5% to 96%, and specificities from 79.4% to 93.4%. Convolutional neural network-based systems generally demonstrated diagnostic performance comparable to or better than that of human endoscopists, particularly among less experienced operators. Limited evidence suggests a possible role for artificial intelligence in predicting treatment failure; however, this finding is based on a single included study. Conclusions: Artificial intelligence appears to be a promising complementary tool for the diagnosis of H. pylori infection, particularly in endoscopic imaging. However, evidence regarding treatment-related and resistance-related applications remains limited and indirect, and these potential uses should therefore be considered preliminary.

1. Introduction

Helicobacter Pylori (H. pylori) is a Gram-negative, spiral-shaped, flagellated bacterium that colonizes the gastrointestinal tract. Its infection affects more than half of the global population [1]. Its adaptive capacity enables survival in the acidic stomach environment and functions as a facultative intracellular bacterium in innate immune cells, which may explain the difficulty in its eradication [2]. H. pylori is recognized as a key pathogenic factor in chronic active gastritis, peptic ulcers, gastric mucosa-associated lymphoid tissue (MALT) lymphoma, and gastric cancer [3]. Furthermore, recent studies have linked its infection with a wide range of pathologies including neurological, dermatological, hematological, cardiovascular, metabolic, and hepatobiliary diseases [4].
Artificial intelligence (AI) is based on the development of trained logical algorithms that enable machines to make decisions in specific situations based on general rules. AI has potential to optimize processes, improve medical care, and increase diagnostic accuracy. One of its most relevant applications is the design of algorithms capable of analyzing endoscopic images and detecting gastritis associated with H. pylori. Several studies have demonstrated that machine learning models can identify the infection with high sensitivity and specificity [5,6,7].
Within AI, machine learning (ML) comprises algorithms that learn from data and improve their performance without being explicitly programmed for each specific task. Deep learning (DL), a more advanced branch of ML, is based on multilayered artificial neural networks capable of automatically extracting complex features from large datasets. Among DL architectures, convolutional neural networks (CNNs) are particularly relevant in gastroenterology because they are highly effective for medical image analysis, especially in endoscopy, where subtle mucosal patterns and textural changes must be recognized. In the context of H. pylori infection, these models can assist in real-time endoscopic diagnosis, classification of infection status, and prediction of treatment failure or antibiotic resistance when combined with clinical data [8,9,10,11].
Given these capabilities, AI has emerged as a promising complementary tool in H. pylori infection, particularly for endoscopic diagnosis. Emerging studies have also explored treatment-related predictive applications, such as the prediction of eradication failure; however, this evidence remains limited. Previous meta-analyses have mainly focused on artificial intelligence for the endoscopic diagnosis of H. pylori infection. However, the field has continued to evolve, with newer multicenter validation studies, real-time endoscopic systems, and early models aimed at predicting treatment failure and supporting therapeutic decision-making. Therefore, the aim of this systematic review was to synthesize recent evidence on AI applications in H. pylori infection, with the primary emphasis on endoscopic diagnosis and secondary consideration of treatment-related predictive uses. Given the limited number of studies on therapeutic prediction and the heterogeneity of models and outcomes, this review was conducted as a narrative synthesis rather than a pooled quantitative meta-analysis.

2. Materials and Methods

This systematic review was conducted in accordance with the recommendations of the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guideline [12], and the PRISMA checklist can be viewed in the Supplementary Materials at the end of the article.

2.1. Protocol Registration

This protocol was registered in the PROSPERO (International Prospective Register of Systematic Reviews) database under the number CRD420261290071. The registration can be publicly viewed at the following link: https://www.crd.york.ac.uk/PROSPERO/view/CRD420261290071 (accessed on 10 May 2026).

2.2. Search Strategy

The literature search was performed between 20 July and 19 August 2025 in PubMed, ScienceDirect, EBSCO, and the Cochrane Library as detailed in Appendix A, Table A1. Eligible studies were limited to articles published from January 2020 to August 2025. This time restriction was predefined because this review focused on contemporary artificial intelligence models, whose architectures, training strategies, and clinical applicability have evolved substantially in recent years. Therefore, prioritizing the most recent literature was considered appropriate to better reflect the current state of AI-based applications in H. pylori management. The search strategy included a combination of key terms and Boolean operators (AND and OR), which were adjusted to the particularities of each database. Terms such as “artificial intelligence,” “machine learning,” “deep learning,” “Helicobacter Pylori,” “diagnosis,” “resistance,” “treatment,” and “follow-up” were used, covering the main clinical phases of interest. In the case of PubMed and the Cochrane Library, MeSH descriptors were also used to improve the sensitivity and specificity of the strategy. The filters applied were adapted to the methodological design criteria and age of publications, prioritizing clinical trials, observational studies, and retrospective studies published in recent years.

2.3. Eligibility Criteria

This systematic review was conducted following the guidelines established by the PRISMA 2020 declaration and was structured according to the PICO question. Adult patients with a confirmed diagnosis or clinical suspicion of H. pylori infection were included as the population (P). The intervention (I) corresponded to the application of artificial intelligence models, such as machine learning, deep learning, or neural networks. The comparator (C) consisted of conventional diagnostic methods such as standard endoscopy, the urea breath test, and histology. The evaluated outcomes (O) were diagnostic performance metrics (accuracy, sensitivity, and specificity), comparative performance against conventional diagnostic methods, and, when available, predictive performance for treatment failure. The review question was formulated as follows: in adults with confirmed or suspected H. pylori infection, how accurately can artificial intelligence models support endoscopic diagnosis and, secondarily, predict treatment failure compared with conventional approaches?

2.4. Inclusion Criteria

Original studies published in the last 5 years, in English or Spanish, evaluating adults (≥18 years) with a confirmed diagnosis or clinical suspicion of H. pylori were included. Studies must apply artificial intelligence (machine learning, deep learning, neural networks, or other algorithms) during diagnosis, treatment selection, or clinical follow-up and report relevant performance metrics or clinical outcomes. Retrospective, prospective, and cohort designs, as well as randomized clinical trials with full text available, were accepted.

2.5. Exclusion Criteria

Studies in individuals under 18 years of age, in animals, or in vitro, as well as those not using AI as a main component were excluded. Reviews, meta-analyses, editorials, letters, conference abstracts, protocols without results, and case reports were not considered. Publications in languages other than English or Spanish, without a valid comparator, or studies without access to full text were also excluded.

2.6. Screening and Selection Process

Two investigators independently reviewed the title and abstract of each reference obtained in the searches in order to collect and evaluate the full texts of the relevant studies that could meet the inclusion criteria. They also conducted a manual search of references from original articles and pertinent reviews.

2.7. Risk-of-Bias Assessment

Data collection was performed using a predetermined table designed prior to the evaluation of selected articles. Two independent reviewers assessed the methodological quality of eligible studies. For all studies, the Joanna Briggs Institute (JBI) Critical Appraisal Tool was used, employing versions corresponding to each study design. Any discrepancy between reviewers was resolved by consensus or with the intervention of a third evaluator.

2.8. Statistical Analysis

Due to substantial heterogeneity in AI model architecture, input data, dataset composition, validation strategy, comparator type, and reported outcomes, a quantitative meta-analysis was considered inappropriate. Instead, a structured narrative synthesis was conducted. Studies were grouped into two clinically distinct categories: (1) diagnostic models, mainly based on endoscopic imaging, and (2) predictive models for treatment failure, mainly based on clinical variables. Within each group, findings were compared descriptively according to study design, validation setting, comparator, and reported performance metrics, with an emphasis on the range and consistency of accuracy, sensitivity, and specificity values.

3. Results

3.1. Study Selection

A total of 61 records were identified through searches of PubMed, ScienceDirect, EBSCO, and the Cochrane Library (Figure 1). After removal of duplicates (n = 5), 56 titles and abstracts were screened, of which 39 were excluded for not meeting the eligibility criteria. Seventeen full-text articles were subsequently assessed for eligibility, and 11 were excluded. Ultimately, six studies were included in the systematic review. Although the final number of included studies was limited, this likely reflects both the still-emerging body of clinically validated AI research in H. pylori infection and the strict eligibility criteria applied in this review, particularly those related to adult clinical populations, full-text availability, and the presence of an appropriate comparator or clinical validation.

3.2. Characteristics of Included Studies

The six selected studies included randomized clinical trials, multicenter prospective investigations, and validation studies in different clinical contexts (Table 1). Most focused on the application of convolutional neural networks for diagnosis from endoscopic images, while one evaluated a real-time deep learning model and another explored machine learning algorithms to predict treatment failure in H. pylori infection. Overall, parameters such as the accuracy, sensitivity, and specificity of artificial intelligence systems were analyzed, as well as their performance compared to conventional diagnostic methods.

3.3. Primary Outcomes

3.3.1. Accuracy

The included studies demonstrated high diagnostic accuracy of artificial intelligence systems applied to H. pylori diagnosis, with values that mostly exceeded 80% (Table 2). Zou et al. developed an AI-assisted endoscopy system called EfficientNet-B0, achieving accuracies of 89.6% in the test phase and 92.8% in a multicenter clinical trial [13]. Shen et al. evaluated the CADSS-HP system, which is based on ResNet34, with an accuracy of 89.9% during real-time endoscopies [14]. Seo et al. applied an Inception-v3 network, with values between 87% and 94% in internal and external validations, confirming model consistency [15]. Similarly, Li et al. developed the IDEA-HP system, which is based on deep learning, with an accuracy of 85.3%, making it comparable to that of expert endoscopists [16]. Finally, Nakashima et al. used a 22-layer LCI-CAD model, obtaining accuracies between 79.2% and 84.2% according to infection category [17].

3.3.2. Sensitivity

In terms of sensitivity, all models reported high performance in H. pylori detection. The EfficientNet-B0 system developed by Zou et al. (2024) [13] stood out for its superior values in both the experimental phase and the AI-assisted group, while CADSS-HP, which was evaluated by Shen et al. (2023) [14], reached similar figures, surpassing endoscopists without assistance. The Inception-v3 model studied by Seo et al. (2023) [15] showed high and constant sensitivity in internal and external validations, as did the IDEA-HP system designed by Li et al. (2023) [16], which maintained stable performance. Nakashima et al. (2020) [17] reported variable values with the LCI-CAD model, and Jiang et al. (2024) [18] recorded a sensitivity close to 80% in predicting treatment failure, reflecting its clinical applicability.

3.3.3. Specificity

According to Table 2, the analyzed models demonstrated adequate capacity to correctly identify negative H. pylori cases. EfficientNet-B0, developed by Zou et al. (2024) [13], obtained the highest values, followed by the CADSS-HP system of Shen et al. (2023) [14], which maintained an appropriate balance between sensitivity and specificity. The Inception-v3 model applied by Seo et al. (2023) [15] and the IDEA-HP system designed by Li et al. (2023) [16] showed consistent performance across different validation stages. Likewise, the LCI-CAD model proposed by Nakashima et al. (2020) [17] reached high values according to infection type, while the Extra-Trees Classifier developed by Jiang et al. (2024) [18] presented specificities close to 80%, confirming adequate overall performance.

3.4. Secondary Outcomes

3.4.1. Artificial Intelligence Versus Human Endoscopists

According to Table 3, artificial intelligence systems showed superior performance compared to endoscopists, especially in those with less experience. Zou et al. (2024) [13], using the EfficientNet-B0 model, reported an accuracy of 92.8% in AI-assisted endoscopists versus 75.6% in non-assisted ones, along with higher sensitivity (91.8%) and specificity (93.4%) values. Similarly, Shen et al. (2023) [14] achieved an accuracy of 89.9% and a sensitivity of 91.5% with the CADSS-HP system, surpassing results obtained by human endoscopists (83.8% and 78.3%, respectively). Li et al. (2023) [16] with the IDEA-HP model demonstrated performance comparable to specialists but superior to less experienced endoscopists, with an accuracy of 84%, a sensitivity of 82%, and a specificity of 86%. Collectively, these findings reinforce the utility of AI as a diagnostic support tool in H. pylori detection.

3.4.2. Artificial Intelligence Versus Standard Diagnostic Tests

Shen et al. (2023) [14] showed that the CADSS-HP system achieved a sensitivity score comparable to that of the urea breath test (91.5% vs. 95.2%) and higher than that of histology (91.5% vs. 82%). In addition, because the system provides real-time results during endoscopy, it may help reduce the need for gastric biopsies (Table 3). Several studies, however, did not include direct head-to-head comparisons with human endoscopists or standard diagnostic tests, as their primary aim was model development or validation rather than formal comparative assessment. Therefore, although standalone performance metrics support the technical potential of these models, the level of clinical validation is stronger in studies that incorporated explicit comparators than in those reporting only internal or external validation performance.

3.5. Methodological Quality

JBI Critical Appraisal Tool Evaluation

The JBI Critical Appraisal Checklist for Diagnostic Test Accuracy Studies, comprising ten domains, was used to assess methodological quality and risk of bias (Figure 2). None of the included diagnostic studies met the criteria for low methodological quality; all fulfilled between 8 and 10 domains and were therefore classified as having high methodological quality. In addition, the study by Zou et al. [13] included a multicenter randomized phase, which was evaluated separately using the JBI Critical Appraisal Checklist for Randomized Controlled Trials, consisting of 13 domains. This phase met seven of the 13 domains, indicating high methodological quality with a low-to-moderate risk of bias (Figure 3).
Across the diagnostic accuracy studies, the domains most frequently rated as unclear were those related to the reference standard, particularly its validity and consistent application across participants, as well as blinding procedures. One study also raised concerns regarding whether the reference standard had been clearly specified. In the randomized phase of Zou et al.’s study, the main methodological limitations were associated with blinding, with additional uncertainty in allocation concealment and certain outcome assessment domains. Overall, the included studies were considered to have acceptable methodological quality.

4. Discussion

4.1. Diagnostic Applications of AI in H. pylori Infection

The studies included in this review consistently showed high diagnostic performance of AI models for H. pylori detection. Accuracies generally ranged from 85% to 94%, while sensitivity and specificity exceeded 80% in most studies. Zou et al. (2024) and Shen et al. (2023) reported accuracies and sensitivities close to 92%. Seo et al. (2023) found values between 86% and 96%, and Li et al. (2023) also reported performance above 80%, supporting the clinical utility of convolutional neural network-based systems during endoscopy [13,14,15,16].
Nakashima et al. (2020) likewise reported accuracy, sensitivity, and specificity values above 75% in some categories, suggesting that AI may help distinguish among active infection, post-eradication status, and the absence of infection, with potential implications for gastric cancer risk stratification [17]. Based on the six studies formally included in this systematic review, the available evidence supports AI as a promising complementary tool for the endoscopic diagnosis of H. pylori infection. By improving diagnostic accuracy and reducing operator dependence, AI-assisted systems may serve as a useful second opinion during endoscopic evaluation [19]. Consistent with this, Zou et al. (2024), Shen et al. (2023), and Li et al. (2023) reported improved accuracy, sensitivity, and specificity with AI assistance, reinforcing its value for clinical decision-making [13,14,16,20].

4.2. Treatment-Related and Predictive Applications

Evidence for treatment-related and predictive applications remains limited. Only one included study evaluated treatment failure prediction, and none directly demonstrated a reduction in antimicrobial resistance; therefore, these implications should be interpreted cautiously. Jiang et al. [18] validated an Extra-Trees classifier for predicting failure of clarithromycin-based triple therapy, with sensitivities and specificities close to 80%. The main predictive factors were age, the interval between prior antibiotic use and therapy initiation, and the type of therapeutic regimen. These findings suggest that AI may eventually help identify patients at a higher risk of treatment failure and support more individualized therapeutic decisions [18].

4.3. Additional Contextual Literature Beyond the Formally Included Studies

The following studies were not part of the six articles formally included in the PRISMA-based systematic review. They are presented only to provide broader context regarding recent developments in the field and should not be interpreted as part of the primary evidence base of this review. An additional retrospective study reported an accuracy of 88%, a sensitivity of 93%, and a specificity of 80% with a CNN-based system [21]. These results are comparable to those reported by Zou et al. (2024), Seo et al. (2023), and Shen et al. (2023), supporting the stability and usability of CNN architectures across different endoscopic image datasets [13,14,15].
A recent multicenter study reported similar findings, showing that artificial intelligence achieved significantly better performance than endoscopists regardless of experience level, with an accuracy of 85.3%, a sensitivity of 85.7%, and a specificity of 85.1% [22]. In addition, AI may help professionals in resource-limited settings evaluate gastric health more quickly and consistently from routine endoscopic images, potentially reducing diagnostic gaps [23].
Although the eligibility framework allowed treatment-related and resistance-related AI applications, the final PRISMA-based sample mainly consisted of clinically oriented studies using endoscopic images or patient-level clinical variables. Studies based on genomic sequencing or molecular resistance prediction were therefore discussed as contextual literature when they were not part of the final included sample under the prespecified eligibility criteria, particularly the requirement for direct clinical applicability and valid comparators or clinical validation. These molecular models should therefore be interpreted as emerging complementary evidence rather than as part of the primary evidence base synthesized in this review.
At the molecular level, other deep learning models based on genomic sequencing achieved 100% sensitivity and specificity in predicting clarithromycin resistance. This approach integrated multiple genetic variants, enabling identification of complex resistance patterns and more accurate anticipation of possible therapeutic failures [24]. Another machine learning-based model also demonstrated the ability to recommend personalized treatments according to individual patient characteristics, such as age, sex, allergies, country of origin, and prior therapeutic history, achieving eradication rates close to 94%, compared with 88% obtained with standard therapies [25].
Overall, the available evidence remains predominantly focused on diagnostic applications of artificial intelligence, particularly endoscopic image-based detection of H. pylori. In contrast, evidence regarding therapeutic or predictive applications remains scarce, and conclusions related to treatment-oriented uses of AI should therefore be considered preliminary.

4.4. Authors’ Critical Perspective on AI in H. pylori Infection

Although the available evidence is encouraging, the application of AI in H. pylori infection remains in a transitional phase between technical feasibility and routine clinical implementation. Most published studies have been conducted in controlled settings, often in high-expertise centers, and have primarily focused on image-based diagnosis. As a result, their reported performance may not be fully generalizable to community hospitals, lower-volume centers, or populations with different disease prevalence, endoscopic equipment, and local antimicrobial resistance patterns. In this context, high diagnostic accuracy alone should not be considered sufficient to support universal adoption, as external validity, reproducibility, workflow integration, and operator-dependent variability remain critical considerations.
From our perspective, AI should be regarded as a complementary clinical support tool rather than a substitute for physician expertise. Its most immediate value appears to lie in standardizing endoscopic interpretation, supporting less experienced endoscopists, and potentially facilitating more personalized therapeutic strategies when clinical, epidemiological, and genomic variables are incorporated into predictive models. However, before widespread implementation can be recommended, further research should prioritize prospective multicenter validation, model explainability, cost-effectiveness, regulatory considerations, and the impact of AI-assisted strategies on patient-centered outcomes such as eradication rates, adverse events, and reductions in unnecessary antibiotic exposure. Thus, although AI represents one of the most promising innovations in this field, stronger real-world evidence is still needed to determine whether improved algorithmic performance consistently translates into better clinical care.

4.5. Limitations

This systematic review has several limitations. First, only six studies were included, which limits the generalizability of the findings. Second, the included studies were heterogeneous in clinical objectives, artificial intelligence models, sample sizes, imaging modalities, comparators, and validation methods, which precluded quantitative synthesis and limited standardized comparisons across studies. Third, the available evidence was overwhelmingly diagnostic, with only one study addressing treatment failure prediction; therefore, conclusions regarding treatment-oriented applications should be considered preliminary. Fourth, Embase was not included in the search strategy, and the use of broad predefined AI terms without a wider set of specific architectures or molecular- or biomarker-related terms may have reduced retrieval of potentially relevant studies. Fifth, the final included evidence base did not comprehensively represent molecular or genomic resistance-prediction studies, which further narrows the scope of this review with respect to AI applications in chemoresistance mechanisms. Finally, not all studies reported confidence intervals or measures of statistical significance, limiting the robustness of interpretation.

5. Conclusions

The available evidence suggests that artificial intelligence is a promising complementary tool for the diagnosis of H. pylori infection, particularly in endoscopic settings, where several models have shown high accuracy, sensitivity, and specificity. In contrast, evidence for therapeutic and predictive applications remains scarce, as only one included study evaluated treatment failure prediction. Therefore, although artificial intelligence may have future utility in supporting treatment-related decision-making, its role in this area remains preliminary and should be interpreted with caution. Further studies are needed to determine whether AI-based approaches can meaningfully contribute to therapeutic optimization and antimicrobial resistance-related outcomes in routine clinical practice.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/gidisord8020023/s1. Table S1: PRISMA 2020 checklist.

Author Contributions

Research: Y.V.-H., J.A.C.-G. and L.G.H.-J.; methodology: Y.V.-H., J.A.C.-G., L.G.H.-J. and E.Z.-M.; project management: E.Z.-M., S.A.-C. and J.M.-J.; supervision: E.Z.-M., S.A.-C. and C.C.-N.; verification: J.A.V.-M. and C.R.-C.; validation: E.Z.-M., J.M.-J., J.A.V.-M., C.R.-C. and S.A.-C.; writing—original draft: Y.V.-H., J.A.C.-G. and L.G.H.-J.; writing—review and editing: E.Z.-M., J.M.-J., J.A.C.-G., L.G.H.-J., S.A.-C. and C.C.-N. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Conflicts of Interest

The authors declare no conflicts of interest. For transparency, Esteban Zavaleta-Monestel, Sebastián Arguedas-Chacón, Jeaustin Mora-Jiménez, Jorge Villalobos-Madrigal, Carolina Rojas-Chinchilla, and Christian Campos-Núñez are affiliated with Clinica Biblica.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
CADSS-HPComputer-Aided Decision Support System for Helicobacter Pylori
CNNsConvolutional Neural Networks
DLDeep Learning
H. pyloriHelicobacter Pylori
IDEA-HPIntelligent Detection Endoscopic Assistant for Helicobacter Pylori
JBIJoanna Briggs Institute
LCI-CADLinked Color Imaging Computer-Aided Diagnosis
MALTMucosa-Associated Lymphoid Tissue
MLMachine Learning
PICOPopulation, Intervention, Comparator, Outcome
PRISMAPreferred Reporting Items for Systematic Reviews and Meta-Analyses
PROSPEROInternational Prospective Register of Systematic Reviews
RCTRandomized Controlled Trial
UBTUrea Breath Test

Appendix A

Table A1. Summary of databases used, keywords, strategies applied, and corresponding filters in the initial article collection process.
Table A1. Summary of databases used, keywords, strategies applied, and corresponding filters in the initial article collection process.
DatabaseSearch StrategyDatabase Filters AppliedNumber of Potential Articles Identified
PubMed/MEDLINE(“artificial intelligence” [MeSH Terms] OR “machine learning” [Title/Abstract] OR “deep learning” [Title/Abstract]) AND (“Helicobacter Pylori” [MeSH Terms] OR “H. pylori” [Title/Abstract]) AND (“diagnosis” [Title/Abstract] OR “antibiotic resistance” [Title/Abstract] OR “treatment” [Title/Abstract] OR “follow-up” [Title/Abstract])Publication year: 2020–2025; randomized Controlled trial; observational study; controlled clinical trial; clinical trial; clinical study; multicenter study6
ScienceDirect(“artificial intelligence” OR “machine learning” OR “deep learning”) AND (“Helicobacter Pylori” OR “H. pylori”) AND (“diagnosis” OR “resistance” OR “treatment” OR “follow-up”)Publication year: 2020–2025; research articles13
EBSCO(“artificial intelligence” OR “machine learning” OR “deep learning”) AND (“Helicobacter Pylori” OR “H. pylori”) AND (“diagnosis” OR “resistance” OR “treatment” OR “follow-up”)Last 5 years; English language; full text available40
Cochrane Library#1 MeSH: [Artificial Intelligence]; #2 MeSH: [Helicobacter Pylori]; #3 diagnosis OR treatment OR resistance OR follow-up; #4 #1 AND #2 AND #3Trials/reviews; English language; last 5 years2
Note: In the Cochrane Library search strategy, “#” indicates the numbered search line used to combine previous search sets.

References

  1. De Pardo Ghetti, E.M. Helicobacter Pylori: Un problema actual. Gac. Médica Boliv. 2013, 36, 108–111. [Google Scholar]
  2. Diaconu, S.; Predescu, A.; Moldoveanu, A.; Pop, C.S.; Fierbințeanu-Braticevici, C. Helicobacter Pylori infection: Old and new. J. Med. Life 2017, 10, 112–117. [Google Scholar]
  3. Aumpan, N.; Mahachai, V.; Vilaichone, R. Management of Helicobacter Pylori infection. JGH Open 2023, 7, 3–15. [Google Scholar] [CrossRef] [PubMed]
  4. Brito, B.B.D.; Silva, F.A.F.D.; Soares, A.S.; Pereira, V.A.; Santos, M.L.C.; Sampaio, M.M.; Neves, P.H.M.; de Melo, F.F. Pathogenesis and clinical management of Helicobacter Pylori gastric infection. World J. Gastroenterol. 2019, 25, 5578–5589. [Google Scholar] [CrossRef]
  5. Avila-Tomás, J.F.; Mayer-Pujadas, M.A.; Quesada-Varela, V.J. La inteligencia artificial y sus aplicaciones en medicina I: Introducción antecedentes a la IA y robótica. Atención Primaria 2020, 52, 778–784. [Google Scholar] [CrossRef] [PubMed]
  6. Aneiros-Fernández, J.; Montero Pavón, P.; García Gómez, N.; Prian, R.M.P.; García, I.S.; Ortiz, A.I.R.; Castro, R.L.; Casado-Sánchez, C.; Turrión, V.S.; Luna, A.; et al. Rapid and Efficient Screening of Helicobacter Pylori in Gastric Samples Stained with Warthin–Starry Using Deep Learning. Diagnostics 2025, 15, 1085. [Google Scholar] [CrossRef]
  7. Addissouky, T.A.; Wang, Y.; El Sayed, I.E.T.; Baz, A.E.; Ali, M.M.A.; Khalil, A.A. Recent trends in Helicobacter Pylori management: Harnessing the power of AI and other advanced approaches. Beni-Suef Univ. J. Basic Appl. Sci. 2023, 12, 80. [Google Scholar] [CrossRef]
  8. Aracena, C.; Villena, F.; Arias, F.; Dunstan, J. Aplicaciones de aprendizaje automático en salud. Rev. Médica Clínica Las Condes 2022, 33, 568–575. [Google Scholar] [CrossRef]
  9. González-Pérez, Y.; Montero Delgado, A.; Martinez Sesmero, J.M. Acercando la inteligencia artificial a los servicios de farmacia hospitalaria. Farm. Hosp. 2024, 48, S35–S44. [Google Scholar] [CrossRef] [PubMed]
  10. Sharma, N.; Sharma, R.; Jindal, N. Machine Learning and Deep Learning Applications-A Vision. Glob. Transit. Proc. 2021, 2, 24–28. [Google Scholar] [CrossRef]
  11. Purwono, P.; Ma’arif, A.; Rahmaniar, W.; Fathurrahman, H.I.K.; Frisky, A.Z.K.; Haq, Q.M.U. Understanding of Convolutional Neural Network (CNN): A Review. Int. J. Robot. Control Syst. 2023, 2, 739–748. [Google Scholar] [CrossRef]
  12. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef]
  13. Zou, P.-Y.; Zhu, J.-R.; Zhao, Z.; Mei, H.; Zhao, J.-T.; Sun, W.-J.; Wang, G.-H.; Chen, D.-F.; Fan, L.-L.; Lan, C.-H. Development and application of an artificial intelligence-assisted endoscopy system for diagnosis of Helicobacter Pylori infection: A multicenter randomized controlled study. BMC Gastroenterol. 2024, 24, 335. [Google Scholar] [CrossRef]
  14. Shen, Y.; Chen, A.; Zhang, X.; Zhong, X.; Ma, A.; Wang, J.; Wang, X.; Zheng, W.; Sun, Y.; Yue, L.; et al. Real-Time Evaluation of Helicobacter Pylori Infection by Convolution Neural Network During White-Light Endoscopy: A Prospective, Multicenter Study (With Video). Clin. Transl. Gastroenterol. 2023, 14, e00643. [Google Scholar] [CrossRef]
  15. Seo, J.Y.; Hong, H.; Ryu, W.-S.; Kim, D.; Chun, J.; Kwak, M.-S. Development and validation of a convolutional neural network model for diagnosing Helicobacter Pylori infections with endoscopic images: A multicenter study. Gastrointest. Endosc. 2023, 97, 880–888.e2, Erratum in Gastrointest. Endosc. 2024, 99, 482. https://doi.org/10.1016/j.gie.2024.01.036. [Google Scholar] [CrossRef] [PubMed]
  16. Li, Y.-D.; Wang, H.-G.; Chen, S.-S.; Yu, J.-P.; Ruan, R.-W.; Jin, C.-H.; Chen, M.; Jin, J.-Y.; Wang, S. Assessment of Helicobacter Pylori infection by deep learning based on endoscopic videos in real time. Dig. Liver Dis. 2023, 55, 649–654. [Google Scholar] [CrossRef]
  17. Nakashima, H.; Kawahira, H.; Kawachi, H.; Sakaki, N. Endoscopic three-categorical diagnosis of Helicobacter Pylori infection using linked color imaging and deep learning: A single-center prospective study (with video). Gastric Cancer 2020, 23, 1033–1040. [Google Scholar] [CrossRef]
  18. Jiang, F.; Lui, T.K.L.; Ju, C.; Guo, C.; Cheung, K.S.; Lau, W.C.Y.; Leung, W.K. Machine learning models in predicting failure of Helicobacter Pylori treatment: A two country validation study. Helicobacter 2024, 29, e13051. [Google Scholar] [CrossRef]
  19. Bang, C.S.; Lee, J.J.; Baik, G.H. Artificial Intelligence for the Prediction of Helicobacter Pylori Infection in Endoscopic Images: Systematic Review and Meta-Analysis of Diagnostic Test Accuracy. J. Med. Internet Res. 2020, 22, e21983. [Google Scholar] [CrossRef]
  20. Wen, Y.; Huang, Y.; Liu, Y.; Zhang, S.; Liu, Z.; Hui, C.; Wang, Y. Artificial intelligence for the diagnosis of Helicobacter Pylori infection in endoscopic and pathological tissues images: A systematic review and meta-analysis. Intell.-Based Med. 2025, 11, 100244. [Google Scholar] [CrossRef]
  21. Lin, C.-H.; Hsu, P.-I.; Tseng, C.-D.; Chao, P.-J.; Wu, I.-T.; Ghose, S.; Shih, C.-A.; Lee, S.-H.; Ren, J.-H.; Shie, C.-B.; et al. Application of artificial intelligence in endoscopic image analysis for the diagnosis of a gastric cancer pathogen-Helicobacter Pylori infection. Sci. Rep. 2023, 13, 13380. [Google Scholar] [CrossRef]
  22. Hu, Y.; Xu, J.; Huang, L.; Zheng, Z.; Zhao, J.; Chen, T.; Liu, J.; Zhang, F.; Ding, X.; Pan, J.; et al. Artificial intelligence–assisted endoscopic diagnosis system for diagnosing Helicobacter Pylori infection: A multicenter study. BMC Med. 2025, 23, 540. [Google Scholar] [CrossRef] [PubMed]
  23. Chiang, T.-H.; Hsu, Y.-N.; Chen, M.-H.; Chen, Y.-R.; Cheng, H.-C.; Chen, M.-J.; Lee, F.-J.; Chang, C.-Y.; Chang, C.-C.; Bair, M.-J.; et al. A rural-to-center artificial intelligence model for diagnosing Helicobacter Pylori infection and premalignant gastric conditions using endoscopy images captured in routine practice. Endoscopy 2025, 58, 343–354. [Google Scholar] [CrossRef] [PubMed]
  24. Yu, J.; Jia, Y.; Yu, Q.; Lin, L.; Li, C.; Chen, B.; Zhong, P.; Lin, X.; Li, H.; Sun, Y.; et al. Deciphering complex antibiotic resistance patterns in Helicobacter Pylori through whole genome sequencing and machine learning. Front. Cell. Infect. Microbiol. 2024, 13, 1306368. [Google Scholar] [CrossRef]
  25. Higgins, K.; Nyssen, O.P.; Southern, J.; Laponogov, I.; Veselkov, D.; Gisbert, J.P.; Kanonnikoff, T.F.; Veselkov, K. The Helicobacter Pylori AI-clinician harnesses artificial intelligence to personalise H. pylori treatment recommendations. Nat. Commun. 2025, 16, 6472. [Google Scholar] [CrossRef] [PubMed]
Figure 1. PRISMA 2020 flow diagram of the study selection process. Arrows indicate the flow of records through identification, screening, eligibility assessment, and inclusion.
Figure 1. PRISMA 2020 flow diagram of the study selection process. Arrows indicate the flow of records through identification, screening, eligibility assessment, and inclusion.
Gastrointestdisord 08 00023 g001
Figure 2. Methodological quality assessment of studies using the JBI Critical Appraisal Tool for Diagnostic Test Accuracy Studies [13,14,15,16,17,18]. Figure created using Microsoft Excel for Microsoft 365, version [insert version], Microsoft Corporation, Redmond, WA, USA.
Figure 2. Methodological quality assessment of studies using the JBI Critical Appraisal Tool for Diagnostic Test Accuracy Studies [13,14,15,16,17,18]. Figure created using Microsoft Excel for Microsoft 365, version [insert version], Microsoft Corporation, Redmond, WA, USA.
Gastrointestdisord 08 00023 g002
Figure 3. Methodological quality assessment of one study using the JBI Critical Appraisal Tool for Randomized Controlled Trials. Figure created using Microsoft Excel for Microsoft 365, version [insert version], Microsoft Corporation, Redmond, WA, USA.
Figure 3. Methodological quality assessment of one study using the JBI Critical Appraisal Tool for Randomized Controlled Trials. Figure created using Microsoft Excel for Microsoft 365, version [insert version], Microsoft Corporation, Redmond, WA, USA.
Gastrointestdisord 08 00023 g003
Table 1. Clinical characteristics and main findings of the included studies.
Table 1. Clinical characteristics and main findings of the included studies.
StudyArtificial Intelligence ApplicationStudy DesignPopulation SizeSummary
Zou et al., 2024 [13]Convolutional neural network (EfficientNet-B0) for artificial intelligence-assisted endoscopy in the diagnosis of Helicobacter PyloriDiagnostic accuracy, randomized controlled trial, and prospective multicenter study952 (training), 411 (internal validation), and 160 (external validation)The study developed and implemented an artificial intelligence-assisted endoscopy system for the diagnosis of Helicobacter Pylori infection, demonstrating higher diagnostic accuracy, sensitivity, and specificity compared with unaided endoscopists.
Shen et al., 2023 [14]Convolutional neural network (ResNet34) for real-time diagnosis of Helicobacter Pylori using the Computer-Aided Decision Support System for Helicobacter Pylori (CADSS-HP)Prospective multicenter diagnostic accuracy study456 (prospective evaluation)The study evaluated CADSS-HP, a computer-aided system based on convolutional neural networks for the detection of Helicobacter Pylori during white-light endoscopy. Patient outcomes demonstrated high sensitivity and specificity, outperforming endoscopic diagnosis and showing comparable performance to the urea breath test, with potential to replace gastric biopsies.
Seo et al., 2023 [15]Convolutional neural network for the diagnosis of Helicobacter Pylori using endoscopic imagesValidation study, diagnostic accuracy, and prospective multicenter study191 (prospective cohort)The study validated a convolutional neural network model for the diagnosis of Helicobacter Pylori infection using endoscopic images, achieving robust performance in both internal and external validations, particularly in distinguishing never-infected patients from those with prior infections.
Li et al., 2023 [16]Deep learning with the Intelligent Detection Endoscopic Assistant for Helicobacter Pylori (IDEA-HP) for real-time evaluation of Helicobacter PyloriDiagnostic accuracy and prospective single-center study639 (development), 201 (testing), and 418 (RCT)The study developed and evaluated IDEA-HP, a deep learning-based system that assesses Helicobacter Pylori infection from real-time endoscopic videos. The model achieved diagnostic accuracy comparable to expert endoscopists and superior to novice endoscopists, supporting its potential as a clinical decision-support tool.
Nakashima et al., 2020 [17]Deep convolutional neural network (22 layers) for three-category Helicobacter Pylori infection status classification using Linked Color Imaging Computer-Aided Diagnosis (LCI-CAD)Prospective single-center diagnostic accuracy study84,609 (training), 27,736 (internal validation), and 18,050 (external validation)This study implemented a computer-aided diagnosis system based on linked color imaging and deep learning, making it capable of classifying Helicobacter Pylori infection status into three categories, achieving high diagnostic accuracy and demonstrating potential application in gastric cancer screening programs.
Jiang et al., 2024 [18]Machine learning (Extra-Trees classifier) for prediction of Helicobacter Pylori treatment failureValidation study515 (total)This study evaluated machine learning algorithms to predict failure of clarithromycin-containing Helicobacter Pylori eradication therapy, identifying the Extra-Trees classifier as the most effective model based on clinical variables such as time interval between antibiotic use, age, and type of triple therapy regimen.
Table 2. Performance of artificial intelligence systems in H. pylori infection.
Table 2. Performance of artificial intelligence systems in H. pylori infection.
StudyType of Artificial Intelligence SystemAccuracy (%)Sensitivity (%)Specificity (%)
Zou et al., 2024 [13]Convolutional neural network (EfficientNet-B0)89.6% (AI system) and 92.8% (AI-assisted group)90.9% (AI system) and 91.8% (AI-assisted group)88.9% (AI system) and 93.4% (AI-assisted group)
Shen et al., 2023 [14]Convolutional neural network (ResNet34 and CADSS-HP)89.9%91.5%88.8%
Seo et al., 2023 [15]Convolutional neural network (Inception-v3)Range: 87–94%Range: 86–96%Range: 79–90%
Li et al., 2023 [16]Convolutional neural network (IDEA-HP)85.3%83.3%85.8%
Nakashima et al., 2020 [17]Deep convolutional neural network (22 layers; LCI-CAD)Range: 79.2–84.2% per categoryRange: 62.5–92.5% per categoryRange: 80–92.5% per category
Jiang et al., 2024 [18]Extra-Trees classifier (machine learning)NRRange: 79.6–80.1% per categoryRange: 79.4–80.2% per category
Table 3. Comparative performance of artificial intelligence use versus conventional diagnostic methods in Helicobacter Pylori.
Table 3. Comparative performance of artificial intelligence use versus conventional diagnostic methods in Helicobacter Pylori.
StudyArtificial Intelligence vs. Human EndoscopistsArtificial Intelligence vs. Standard Diagnostic Test
Zou et al., 2024 [13]Accuracy of 92.8% compared to 75.6%.
Sensitivity of 91.8% compared to 78.6%.
Specificity of 93.4% compared to 74.5%.
Not reported
Shen et al., 2023 [14]Accuracy of 89.9% compared to 83.8%.
Sensitivity of 91.5% compared to 78.3%.
Sensitivity of 91.5% compared to 95.2% in the urea breath test (UBT).
Sensitivity of 91.5% compared to 82% in histopathology.
Specificity of 88.8% compared to 99.6% in UBT or histopathology.
Seo et al., 2023 [15]Not reportedNot reported
Li et al., 2023 [16]Accuracy of 84% compared to 83.6% for experts and 74% for beginners.
Sensitivity of 82.0% compared to 82.4% and 67.2%.
Specificity of 86.0% compared to 84.8% and 80.8%.
Not reported
Nakashima et al., 2020 [17]The diagnostic accuracy of artificial intelligence was shown to be comparable to that of experienced endoscopists.Not reported
Jiang et al., 2024 [18]Not reportedNot reported
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Zavaleta-Monestel, E.; Villagra-Hernandez, Y.; Mora-Jiménez, J.; Villalobos-Madriz, J.A.; Rojas-Chinchilla, C.; Castro-Gamboa, J.A.; Herrera-Jiménez, L.G.; Arguedas-Chacón, S.; Campos-Núñez, C. Artificial Intelligence in Helicobacter Pylori Infection: Diagnostic Applications and Emerging Treatment-Related Predictive Uses—A Systematic Review. Gastrointest. Disord. 2026, 8, 23. https://doi.org/10.3390/gidisord8020023

AMA Style

Zavaleta-Monestel E, Villagra-Hernandez Y, Mora-Jiménez J, Villalobos-Madriz JA, Rojas-Chinchilla C, Castro-Gamboa JA, Herrera-Jiménez LG, Arguedas-Chacón S, Campos-Núñez C. Artificial Intelligence in Helicobacter Pylori Infection: Diagnostic Applications and Emerging Treatment-Related Predictive Uses—A Systematic Review. Gastrointestinal Disorders. 2026; 8(2):23. https://doi.org/10.3390/gidisord8020023

Chicago/Turabian Style

Zavaleta-Monestel, Esteban, Yennifer Villagra-Hernandez, Jeaustin Mora-Jiménez, Jorge Arturo Villalobos-Madriz, Carolina Rojas-Chinchilla, José Andrés Castro-Gamboa, Luis Guillermo Herrera-Jiménez, Sebastián Arguedas-Chacón, and Christian Campos-Núñez. 2026. "Artificial Intelligence in Helicobacter Pylori Infection: Diagnostic Applications and Emerging Treatment-Related Predictive Uses—A Systematic Review" Gastrointestinal Disorders 8, no. 2: 23. https://doi.org/10.3390/gidisord8020023

APA Style

Zavaleta-Monestel, E., Villagra-Hernandez, Y., Mora-Jiménez, J., Villalobos-Madriz, J. A., Rojas-Chinchilla, C., Castro-Gamboa, J. A., Herrera-Jiménez, L. G., Arguedas-Chacón, S., & Campos-Núñez, C. (2026). Artificial Intelligence in Helicobacter Pylori Infection: Diagnostic Applications and Emerging Treatment-Related Predictive Uses—A Systematic Review. Gastrointestinal Disorders, 8(2), 23. https://doi.org/10.3390/gidisord8020023

Article Metrics

Back to TopTop