1. Introduction
Artificial-intelligence systems for dental caries detection have been evaluated using bitewing, periapical, and panoramic radiographs, cone-beam computed tomography, intraoral photographs, fluorescence images, and three-dimensional scans. Systematic reviews and primary diagnostic-accuracy studies have reported promising but variable sensitivity and specificity [
1,
2,
3,
4,
5].
Clinical interpretation is limited because the studies differ in modality, lesion threshold, observational unit, reference standard, sampling frame, and validation design. Accuracy for enamel-only proximal lesions on bitewings is not exchangeable with tooth-level panoramic classification, photograph-based cavitation detection, or patch-level computational tasks.
Standalone algorithm accuracy and clinician performance with AI assistance are also different questions [
6,
7,
8]. The first evaluates algorithm output against a reference standard; the second asks whether AI changes human interpretation, errors, decisions, or outcomes. Cost-effectiveness and patient benefit require further evidence beyond diagnostic discrimination.
This review therefore maps the clinical composition, methodological limitations, external validation, and study-level accuracy of the available evidence. Reports with exact, internally coherent 2 × 2 data are shown as a descriptive subset rather than pooled across non-exchangeable clinical questions. Controlled clinician-plus-AI evidence is summarized separately.
1.1. Research Questions
Primary question: What study-level diagnostic-accuracy evidence is reported for standalone AI systems that detect dental caries in clinically acquired human dental data?
Secondary question: What controlled evidence evaluates changes in clinician performance when AI assistance is added?
1.2. Objective
To describe the clinical composition, risk of bias, external validation, study-level sensitivity and specificity, and certainty of evidence for standalone AI caries detection, and to summarize controlled clinician-plus-AI evidence separately.
2. Materials and Methods
2.1. Study Design and Registration
This systematic review of diagnostic test accuracy was conducted and reported in accordance with PRISMA 2020 and PRISMA-DTA [
9,
10]. The protocol was prospectively registered in PROSPERO (CRD420251232014) before title/abstract screening and data extraction.
2.2. Eligibility Criteria
Eligible studies evaluated a defined standalone AI test or clinician interpretation with AI assistance on clinically acquired human dental data against a study-defined clinical, operative, histological, or expert image-based reference standard and reported at least one quantitative diagnostic result. The target condition and observational unit were recorded as defined by each study.
Reports were retained in the systematic review even when a complete contingency table was unavailable. A descriptive exact-data subset required integer TP, FP, FN, and TN values reported directly or uniquely reconstructable from internally consistent denominators and unrounded measures, together with a clinically interpretable patient, tooth, surface, or complete-image unit. Non-integer pseudocells, incompatible denominators, patch-only units, and unresolved dependent computational subsamples were excluded from this subset.
Animal, phantom, simulated, and purely in vitro studies were excluded, as were reports without a caries outcome, defined reference standard, or quantitative diagnostic result, and non-original publication types. Standalone AI and clinician-plus-AI interpretation were treated as separate index-test constructs.
2.3. Information Sources and Search Strategy
PubMed, Scopus, Web of Science Core Collection, Embase, and the Cochrane Library were searched through 10 June 2026 without database-level publication-type, study-design, subject-area, category, or disease-mapping restrictions. Exact syntax, interfaces, dates, coverage, and raw yields are provided in
Supplementary File S1.
IEEE Xplore was listed in the protocol but was not searched because the required institutional interface/API was unavailable. Trial registers, preprint servers, organizational websites, and other gray-literature sources were not searched, and authors were not contacted for missing cells. These choices may have reduced coverage and were incorporated into the certainty assessment.
2.4. Study Selection Criteria
Database exports were imported into Rayyan (Systems, Inc., Cambridge, MA, USA) [
11]. Before human screening, the review team applied deterministic rules using bibliographic fields, record types, keyword searches, labels, exclusion reasons, and bulk actions. Rule families targeted clearly non-primary publication types and records whose indexed fields or title/abstract placed them outside dental-caries diagnostic evaluation. Rayyan relevance ranking and AI suggestions were not used.
This process marked 130,245 records ineligible before duplicate human screening. The retained audit trail does not contain the complete bulk-excluded record set, rule-specific counts, or a human-screened validation sample. Its false-negative rate, therefore, cannot be estimated. Records with ambiguous metadata or possible relevance were intended to be retained, but this safeguard cannot substitute for a validation audit.
LPV and MSV independently screened the remaining 3248 titles/abstracts and assessed the 35 retrieved full texts. Disagreements were resolved through discussion, with JM available for adjudication. The blinded pre-consensus export was not retained; Cohen’s kappa and percentage agreement cannot be reconstructed, and the screening stage is not independently auditable at the reviewer-decision level.
2.5. Data Extraction
Two reviewers used a standardized, piloted form to extract study design, country, setting, modality, target condition, observational unit, reference standard, model, dataset origin, development/test separation, external validation, clinical role, reported diagnostic measures, and exact contingency cells when available. Disagreements were resolved by discussion and, if required, third-reviewer adjudication. Missing cells were not imputed.
Author-reported limitations were retained only as an ancillary verbatim corpus in the
Supplementary Materials. They were not treated as diagnostic outcomes, observed implementation effects, or evidence of translational readiness.
2.6. Risk of Bias Assessment
QUADAS-3 version 1.2 was applied without modification at the selected-estimate level [
12,
13]. Separate synthesis questions and ideal-study definitions were specified for standalone AI and clinician-plus-AI interpretation. The pre-populated master assessment was independently checked against the corresponding source reports by ACG and JM after calibration. This was source verification of a common matrix, not two de novo blinded assessments from blank forms. Final judgments were consensus verified; no inter-reviewer reliability statistic is claimed. Complete responses, rationales, source locations, verification records, and final judgments are provided in
Supplementary File S4.
2.7. Data Synthesis and Certainty of Evidence
2.7.1. Standalone-AI Evidence
Study characteristics and reported diagnostic measures were synthesized descriptively. For the 12 reports with exact, coherent 2 × 2 data, sensitivity and specificity were calculated and displayed as study-level point estimates. No bivariate summary, HSROC curve, likelihood ratio, diagnostic odds ratio, predictive value, or cross-modality pooled estimate was calculated because the reports did not evaluate a common clinical question.
Most source datasets included multiple teeth, surfaces, images, or lesions per participant without cluster-adjusted variance. Exact binomial intervals would therefore imply independence that was not supported by the study designs. The main display consequently emphasizes point estimates and ranges; raw cells and analytic limitations remain available in
Supplementary File S3.
2.7.2. Clinician-Plus-AI Evidence
Controlled reader studies were summarized using the authors’ reported without-AI and with-AI results. They were not combined with standalone algorithm estimates or pooled with each other because their reader assignment, case allocation, lesion thresholds, and reported outcomes differed.
2.7.3. Certainty of Evidence
Certainty was assessed for (1) standalone-AI study-level sensitivity and specificity and (2) incremental clinician performance with AI. A GRADE-DTA structure considered risk of bias, inconsistency, indirectness, imprecision, and publication bias. Domain judgments and reasons are provided in
Supplementary Table S4.
2.7.4. Reporting Bias
No funnel-plot or asymmetry test was performed because the descriptive exact-data subset was small and clinically heterogeneous. Risk of missing evidence was instead considered qualitatively from the search coverage and the unvalidated bulk-exclusion process.
3. Results
3.1. Study Selection
The searches identified 137,274 records: PubMed 31,663; Scopus 37,582; Web of Science 13,836; Embase 53,523; and Cochrane Library 670. Rule-based preprocessing marked 130,245 records ineligible, and 3781 duplicates were removed. Of 3248 records screened by two reviewers, 3213 were excluded. All 35 full texts were retrieved; six were excluded (irrelevant outcome, n = 2; unsuitable design, n = 4). A total of 29 reports representing 28 unique studies were included [
4,
14,
15,
16,
17,
18,
19,
20,
21,
22,
23,
24,
25,
26,
27,
28,
29,
30,
31,
32,
33,
34,
35,
36,
37,
38,
39,
40,
41].
The automation box in
Figure 1 follows the PRISMA 2020 flow structure, but the magnitude of the bulk exclusion and absence of a retained validation sample are review-level limitations rather than routine deduplication.
3.2. Risk of Bias
QUADAS-3 was applied to 28 distinct numerical estimates (
Figure 2). Arsiwala-Scheppach et al. [
19] was an ancillary report of the Mertens et al. [
33] trial and supplied no separate diagnostic-accuracy estimate. Final consensus risk of bias was high for all 28 estimates. The analysis domain was high risk in 25 (89.3%), primarily because clustered teeth, surfaces, images, or patches were analyzed as independent and because exclusions, missing observations, or indeterminate results were incompletely handled.
Participants were high risk in 22 estimates (78.6%); index test was high risk in 5 (17.9%); and target condition was high risk in 6 (21.4%). Overall applicability concern was high in 17 estimates (60.7%), low in 8 (28.6%), and based on insufficient information in 3 (10.7%). Detailed judgments are in
Supplementary Files S2 and S4.
3.3. Clinical Composition of the Evidence
The 29 reports represented 28 studies published from 2021 to 2026 across 16 countries or country groups. Imaging modality, target condition, observational unit, and reference standard varied substantially. Only five studies used an independent external dataset.
Table 1 makes the principal composition and evidence constraints visible in the main text; full report-level characteristics are provided in
Supplementary File S3.
3.4. Study-Level Standalone-AI Accuracy
Twelve standalone-AI reports supplied exact, internally coherent contingency cells for a clinically interpretable unit. The subset comprised four bitewing, two panoramic, three photographic or QLF, one CBCT, one 3D intraoral scan, and one mixed-radiograph study. Five estimates used surfaces, five used teeth, and two used images; only one used external validation.
Sensitivity ranged from 0.360 to 0.940 and specificity from 0.700 to 0.983 (
Figure 3). These values describe an availability subset selected partly by reporting completeness. No pooled operating point was calculated. The main figure omits binomial confidence intervals because the required participant-level clustering information was generally unavailable and independent-observation intervals could be anti-conservative.
Certainty for standalone-AI sensitivity and specificity was very low because of very serious risk of bias and indirectness, serious inconsistency and imprecision, and concern about missing evidence (
Supplementary Table S4).
3.5. Clinician Interpretation with AI Assistance
Devlin et al. [
27] randomized 23 dentists to read the same 24 bitewings without or with AssistDent prompts. Sensitivity was 0.443 without AI and 0.758 with AI, while specificity was 0.963 and 0.854, respectively; both differences were statistically significant. Mertens et al. [
33] used a cluster-randomized crossover design. Mean sensitivity increased from 0.72 without AI to 0.81 with AI and AUC increased from 0.85 to 0.89; specificity did not differ significantly, but AI assistance increased invasive treatment decisions.
The two studies therefore suggest that AI can increase reader sensitivity, but they do not establish net clinical benefit. Their designs, case allocation, outcomes, and tradeoffs differed, and neither evaluated patient outcomes in routine care. Certainty for incremental clinician benefit was very low.
4. Discussion
This review found a broad and methodologically fragile evidence base. All 28 assessed estimates had high overall risk of bias, only 5 studies used external validation, and the 12 reports with exact contingency data spanned sensitivity from 0.360 to 0.940 and specificity from 0.700 to 0.983. These reports did not estimate one clinically exchangeable quantity, so a cross-modality pooled operating point was not retained.
Controlled reader evidence is more informative for human–AI interaction than standalone algorithm accuracy. Devlin and Mertens both reported sensitivity gains with AI, but one showed lower specificity and the other more invasive treatment decisions. Improved detection is therefore not equivalent to improved care.
4.1. Comparison with Previous Evidence
The numerical estimates reported by previous meta-analyses depend strongly on scope. Abbott et al. [
1] pooled seven mixed clinical-image and radiographic studies and reported sensitivity of 0.76 and specificity of 0.91. Ammar and Kuhnisch [
3] restricted the question to bitewing caries detection and reported sensitivity of 0.87 and specificity of 0.89 across five studies. Carvalho et al. [
42], focusing on approximal caries on bitewings, reported pooled sensitivity and specificity of approximately 0.94 and 0.91. These summary values fall within the wide study-level ranges observed here, but direct comparison is inappropriate because the included modalities, lesion thresholds, units, and contingency-data rules differ.
The stricter exact-data rule reduced the quantitative denominator but did not create a homogeneous clinical question. Its contribution is to expose how much published evidence cannot support verifiable contingency analysis and how strongly study composition governs any apparent average.
4.2. Sources of Dispersion and Uncertainty
Target conditions ranged from enamel-only lesions to cavitated or any caries; observational units ranged from surfaces and teeth to images and patches; and reference standards ranged from single-examiner image labels to multi-expert consensus or clinical examination. These differences can change both sensitivity and specificity and were not tested as causal moderators.
Clustering is a particularly important limitation. Many reports treated multiple teeth, surfaces, lesions, or images from the same participant as independent observations. Without participant-level data or cluster-adjusted variances, conventional binomial intervals and hierarchical meta-analysis can overstate precision. Retrospective exclusion of technically difficult images and internal testing from the same centers or devices further limits transportability.
4.3. Clinical Utility and Human–AI Evaluation
The reader studies indicate that AI prompts can change diagnostic behavior, not merely reproduce standalone algorithm performance. Higher sensitivity may be clinically valuable for early lesions, but false-positive prompts and more invasive management can offset that benefit. Paired or crossover evaluations should account for repeated readers, repeated cases, treatment thresholds, and downstream decisions.
Cost-effectiveness was outside this review’s diagnostic-accuracy eligibility criteria. Schwendicke et al. [
43], a linked economic evaluation excluded from the systematic review, reported similar modeled tooth retention and costs with and without AI and substantial uncertainty. We therefore make no within-review claim about cost-effectiveness.
Future evaluations should use locked models, prespecified thresholds, independent multicenter and multivendor data, clinically meaningful lesion definitions, and participant-level analyses. Reader studies should randomize or counterbalance reading order, include washout where appropriate, and report diagnostic errors, confidence, management decisions, reading time, workload, equity, unintended consequences, and patient outcomes. STARD-AI, DECIDE-AI, and CONSORT-AI provide relevant reporting frameworks [
44,
45,
46].
4.4. Strengths and Limitations
Strengths include explicit separation of standalone and clinician-plus-AI constructs, transparent estimate-selection rules, exact contingency checks, study-level presentation, unmodified QUADAS-3, and structured certainty assessment.
The largest limitation is the selection audit. Rule-based preprocessing removed 130,245 of 137,274 records before duplicate human screening. Because the excluded record set, rule-level counts, and validation sample were not retained, the false-negative rate cannot be estimated. The screening export also does not preserve independent pre-consensus decisions. These deficiencies materially reduce auditability and contributed to very-low certainty.
Coverage was further limited by omission of IEEE Xplore, registers, preprints, organizational websites, and gray literature. Authors were not contacted, so the exact-data subset is an availability sample influenced by reporting completeness. The QUADAS-3 process involved independent source verification of a shared pre-populated matrix rather than de novo duplicate coding, and no reliability statistic is claimed.
All assessed estimates were at high overall risk of bias, and 25 had high analysis-domain risk. Cluster-adjusted uncertainty was generally unavailable; external validation was sparse; and the small, clinically diverse evidence base did not support meta-regression, formal reporting-bias tests, or a pooled operating point.
5. Conclusions
Published accuracy estimates for standalone AI caries detection vary widely across non-exchangeable modalities, lesion thresholds, observational units, reference standards, and validation designs. With high risk of bias in all assessed estimates and very-low certainty, the evidence does not support a transferable clinical performance benchmark.
Two controlled reader studies suggest that AI assistance can increase sensitivity, but specificity and treatment-intensity tradeoffs remain and net clinical benefit has not been established. Future research should prioritize locked-model external validation and controlled prospective human–AI studies that measure treatment decisions, workflow, and patient-relevant outcomes.
Supplementary Materials
The following supporting information can be downloaded at:
https://www.mdpi.com/article/10.3390/diagnostics16182995/s1. Supplementary File S1: Final Database Search Strategies and Rayyan-Assisted Preprocessing; Supplementary File S2: Revised Complementary Tables and Figures; Supplementary File S3: Study Characteristics, Estimate Selection, and Analytic Disposition; Supplementary File S4: Complete QUADAS-3 v1.2 Assessment; PRISMA_2020_Checklist_Completed. References [
4,
11,
12,
13,
14,
15,
16,
17,
18,
19,
20,
21,
22,
23,
24,
25,
26,
27,
28,
29,
30,
31,
32,
33,
34,
35,
36,
37,
38,
39,
40,
41] are cited in the Supplementary Materials and are included in the main reference list.
Author Contributions
A.M.C.G.: contributed to conception, design, data acquisition, and interpretation; drafted and critically revised the manuscript. I.C.S.G.: contributed to conception, data acquisition, and interpretation; drafted and critically revised the manuscript. L.P.V.: contributed to design and data interpretation; drafted and critically revised the manuscript. M.S.V.: contributed to design and data interpretation; drafted and critically revised the manuscript. J.J.M.: contributed to the design and data interpretation; drafted and critically revised the manuscript. All authors have read and agreed to the published version of the manuscript.
Funding
This review received no specific grant or other financial or non-financial support from any funding agency, commercial entity, or not-for-profit organization. No funder or sponsor had any role in the design, conduct, analysis, interpretation, manuscript preparation, or decision to submit the review.
Institutional Review Board Statement
This study is a systematic review based exclusively on previously published data. No new human participants, personal data, or clinical interventions were involved. Therefore, ethical approval and informed consent were not required.
Informed Consent Statement
Not applicable.
Data Availability Statement
Acknowledgments
OpenAI ChatGPT 5.6 Sol was used to assist with English-language editing. It was not used to make final eligibility decisions or to extract primary study data autonomously. All AI-assisted outputs, calculations, interpretations, citations, and manuscript changes were critically reviewed and verified by the authors, who take full responsibility for the accuracy, integrity, and final content.
Conflicts of Interest
The authors declare that they have no conflicts of interest related to the conduct, analysis, or reporting of this study.
References
- Abbott, L.P.; Saikia, A.; Anthonappa, R.P. Artificial Intelligence Platforms in Dental Caries Detection: A Systematic Review and Meta-Analysis. J. Evid. Based Dent. Pract. 2025, 25, 102077. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Albano, D.; Galiano, V.; Basile, M.; Di Luca, F.; Gitto, S.; Messina, C.; Cagetti, M.G.; Del Fabbro, M.; Tartaglia, G.M.; Sconfienza, L.M. Artificial intelligence for radiographic imaging detection of caries lesions: A systematic review. BMC Oral Health 2024, 24, 274. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ammar, N.; Kühnisch, J. Diagnostic performance of artificial intelligence-aided caries detection on bitewing radiographs: A systematic review and meta-analysis. Jpn. Dent. Sci. Rev. 2024, 60, 128–136. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Estai, M.; Tennant, M.; Gebauer, D.; Brostek, A.; Vignarajan, J.; Mehdizadeh, M.; Saha, S. Evaluation of a deep learning system for automatic detection of proximal surface dental caries on bitewing radiographs. Oral Surg. Oral Med. Oral Pathol. Oral Radiol. 2022, 134, 262–270. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Luke, A.M.; Rezallah, N.N.F. Accuracy of artificial intelligence in caries detection: A systematic review and meta-analysis. Head Face Med. 2025, 21, 24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hung, K.; Montalvao, C.; Tanaka, R.; Kawai, T.; Bornstein, M.M. The use and performance of artificial intelligence applications in dental and maxillofacial radiology: A systematic review. Dentomaxillofac. Radiol. 2019, 49, 20190107. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khanagar, S.B.; Al-ehaideb, A.; Maganur, P.C.; Vishwanathaiah, S.; Patil, S.; Baeshen, H.A.; Sarode, S.C.; Bhandi, S. Developments, application, and performance of artificial intelligence in dentistry—A systematic review. J. Dent. Sci. 2021, 16, 508–522. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schwendicke, F.; Samek, W.; Krois, J. Artificial Intelligence in Dentistry: Chances and Challenges. J. Dent. Res. 2020, 99, 769–774. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- McInnes, M.D.F.; Moher, D.; Thombs, B.D.; McGrath, T.A.; Bossuyt, P.M.; PRISMA-DTA Group. Preferred Reporting Items for a Systematic Review and Meta-analysis of Diagnostic Test Accuracy Studies: The PRISMA-DTA Statement. JAMA 2018, 319, 388–396, Correction in JAMA 2019, 322, 2026. https://doi.org/10.1001/jama.2019.18307. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ouzzani, M.; Hammady, H.; Fedorowicz, Z.; Elmagarmid, A. Rayyan—A web and mobile app for systematic reviews. Syst. Rev. 2016, 5, 210. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Davenport, C.F.; Rutjes, A.W.S.; Mallett, S.; Tomlinson, E.; Yang, B.; Holmes, J.; Westwood, M.E.; Takwoingi, Y.; Reitsma, J.B.; Hyde, C.; et al. QUADAS-3 Explanation and Elaboration: Guidance for Quality Assessment of Diagnostic Test Accuracy Studies. Ann. Intern. Med. 2026, 179, e2504943. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Whiting, P.F.; Tomlinson, E.; Rutjes, A.W.S. QUADAS-3: A Revised Tool for the Quality Assessment of Diagnostic Test Accuracy Studies. Ann. Intern. Med. 2026, 179, 548–555. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Aarti-Rajambigai, M.; Ahmed, F.; Mohammed Harris, S.; Rath, J.; Varghese, M.; Elango, P.A. Role of artificial intelligence in teledentistry diagnosis and treatment planning. Bioinformation 2026, 22, 1974–1977. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chaudhari, A.Y.; Birwadkar, P.; Joshi, S.; Verma, Y.; Sindgi, R. Classification of periapical dental X-ray using the YOLOv8 deep learning model. MethodsX 2025, 15, 103721. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rahimi, H.; Naeimi, S.M.; Darvish, S.; Nazemi Salman, B.; Razzaghi, P.; Luchian, I.; Budala, D.G. Detection of Pediatric Dental Caries in Panoramic Radiograph Using Deep Learning: A Benchmark Study on MD-OPG. Sensors 2026, 26, 2481. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abdelaziz, S.; Yasser, I.; Amgad, M.; Al Qudah, M.; Kunnath Menon, R. Evaluation of the diagnostic reliability of an AI prototype in detecting clinical features from dental photographs—Original research. Front. Dent. Med. 2026, 7, 1814876. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ahmed, W.M.; Azhari, A.A.; Alnakhli, R.; Alsulami, Y.; Merdad, Y.; Abdelrazek, M.; Sahlol, A.T. Development and evaluation of an artificial intelligence (AI) model for detecting dental caries from 3D intraoral scans. J. Prosthet. Dent. 2025, 134, 1115.e1–1115.e8. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Arsiwala-Scheppach, L.T.; Castner, N.J.; Rohrer, C.; Mertens, S.; Kasneci, E.; Cejudo Grano de Oro, J.E.; Schwendicke, F. Impact of artificial intelligence on dentists’ gaze during caries detection: A randomized controlled trial. J. Dent. 2024, 140, 104793. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bayraktar, Y.; Ayan, E. Diagnosis of interproximal caries lesions with deep convolutional neural network in digital bitewing radiographs. Clin. Oral Investig. 2022, 26, 623–632. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Biçengil, K.; Kurt, A.; Naralan, M.E.; Okumuş, İ. Deep Learning-Based Dental Caries Diagnosis on Panoramic Radiographies: Performance of YOLOv8 Versus Human Observers. Diagnostics 2026, 16, 1150. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Biltekin, H.; Geduk, G.; Altan, A.; Karasu, S. Evaluation of deep learning systems in detection of dental caries on panoramic radiography. Am. J. Dent. 2025, 38, 163–168. [Google Scholar] [PubMed]
- Caldwell, J.; Parekh, K.; Crowther, B.; Gohel, C.; Pileggi, R.; Garcia, A.I.; Ghorbanifarajzadeh, M.; Dolan, T.A.; Gohel, A. Performance evaluation of AI-based caries detection technology and its educational training module: A dual-phase investigation. Front. Dent. Med. 2026, 6, 1741855. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cao, L.; van Nistelrooij, N.; Chaves, E.T.; Bergé, S.; Cenci, M.S.; Xi, T.; Loomans, B.; Vinayahalingam, S. Automated chart filing on bitewings using deep learning: Enhancing clinical diagnosis in a multi-center study. J. Dent. 2025, 161, 105919. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, S.; Wu, W.; Chen, P.; Zhou, G.; Yang, Y.; Shen, B.; Hu, J.; Ma, J. Automated Tooth Detection and Caries Identification in CBCT With Deep Learning. Int. Dent. J. 2026, 76, 109508. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, X.; Guo, J.; Ye, J.; Zhang, M.; Liang, Y. Detection of Proximal Caries Lesions on Bitewing Radiographs Using Deep Learning Method. Caries Res. 2022, 56, 455–463. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Devlin, H.; Williams, T.; Graham, J.; Ashley, M. The ADEPT study: A comparative study of dentists’ ability to detect enamel-only proximal caries in bitewing radiographs with and without the use of AssistDent artificial intelligence software. Br. Dent. J. 2021, 231, 481–485. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fadilah, R.P.N.; Rikmasari, R.; Akbar, S.; Setiawan, A.S. Examination of new clinical dental caries in school children using real intra oral photos with artificial intelligence model YOLO-V8x. BMC Oral Health 2026, 26, 164. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Giannakopoulos, K.; Kavadella, A.; Paraskevis, D.; Arhakis, A.; Makrygiannakis, M.A.; Kaklamanos, E.G. Evaluating AI diagnostic accuracy in approximal dental caries detection on bitewing radiographs. Clin. Oral Investig. 2026, 30, 207. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kang, J.-Y.; Kim, H.-K. Feasibility of multi-class dental caries detection using deep learning–based smartphone images: A pilot prospective study. Front. Oral Health 2026, 7, 1805448. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, Z.; Li, J.; Wang, S.; Liu, W.; Wang, W.; Lin, H.; Pang, L. A Multimodal Model for Caries Screening Using Intraoral Images and Questionnaires. Int. Dent. J. 2026, 76, 109420. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ma, Y.; Al-Aroomi, M.A.; Zheng, Y.; Ren, W.; Liu, P.; Wu, Q.; Liang, Y.; Jiang, C. Application of Mask R-CNN for automatic recognition of teeth and caries in cone-beam computerized tomography. BMC Oral Health 2025, 25, 927. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mertens, S.; Krois, J.; Cantu, A.G.; Arsiwala, L.T.; Schwendicke, F. Artificial intelligence for caries detection: Randomized trial. J. Dent. 2021, 115, 103849. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Orhan, K.; Aktuna Belgin, C.; Manulis, D.; Golitsyna, M.; Bayrak, S.; Aksoy, S.; Sanders, A.; Önder, M.; Ezhov, M.; Shamshiev, M.; et al. Determining the reliability of diagnosis and treatment using artificial intelligence software with panoramic radiographs. Imaging Sci. Dent. 2023, 53, 199–208. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Parekh, A.; Jagtap, R.; Samata, Y.; Kotha, P.; Sanjana, M.; Tretto, P.; Roach, M.D.; Jaju, P.; Friedel, A.; Feinberg, M. Deep learning-based segmentation of caries, implants, fixed prosthesis, and restorations on bitewing radiographs: A retrospective study. Sci. Prog. 2026, 109, 00368504261439717. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Park, E.Y.; Cho, H.; Kang, S.; Jeong, S.; Kim, E.-K. Caries detection with tooth surface segmentation on intraoral photographic images using deep learning. BMC Oral Health 2022, 22, 573. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Park, E.Y.; Jeong, S.; Kang, S.; Cho, J.; Cho, J.-Y.; Kim, E.-K. Tooth caries classification with quantitative light-induced fluorescence (QLF) images using convolutional neural network for permanent teeth in vivo. BMC Oral Health 2023, 23, 981. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ren, L.; Chen, J. A Dual-Branch Deep Learning Framework with Explainability for Dental Caries Classification Using Intra-Oral Photographs and Radiographs. J. Imaging 2026, 12, 207. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- van Nistelrooij, N.; Jurkáček, P.; Runge, J.; Do, W.; El Ghoul, K.; Xi, T.; Cenci, M.S.; Loomans, B.A.C.; Vinayahalingam, S. Development and Validation of AI System for Tooth Detection and Diagnosis in Dental Radiographs. Int. Dent. J. 2026, 76, 109576. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wang, L.; Li, Z. Transformer-based intelligent detection model for early dental caries in panoramic radiographs. Sci. Rep. 2026, 16, 3507. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, J.-W.; Fan, J.; Zhao, F.-B.; Ma, B.; Shen, X.-Q.; Geng, Y.-M. Diagnostic accuracy of artificial intelligence-assisted caries detection: A clinical evaluation. BMC Oral Health 2024, 24, 1095. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Carvalho, B.K.G.; Nolden, E.-L.; Wenning, A.S.; Kiss-Dala, S.; Agócs, G.; Róth, I.; Kerémi, B.; Géczi, Z.; Hegyi, P.; Kivovics, M. Diagnostic accuracy of artificial intelligence for approximal caries on bitewing radiographs: A systematic review and meta-analysis. J. Dent. 2024, 151, 105388. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schwendicke, F.; Mertens, S.; Cantu, A.G.; Chaurasia, A.; Meyer-Lueckel, H.; Krois, J. Cost-effectiveness of AI for caries detection: Randomized trial. J. Dent. 2022, 119, 104080. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, X.; Cruz Rivera, S.; Moher, D.; Calvert, M.J.; Denniston, A.K.; Chan, A.-W.; Darzi, A.; Holmes, C.; Yau, C.; Ashrafian, H.; et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: The CONSORT-AI extension. Nat. Med. 2020, 26, 1364–1374. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Sounderajah, V.; Guni, A.; Liu, X.; Collins, G.S.; Karthikesalingam, A.; Markar, S.R.; Golub, R.M.; Denniston, A.K.; Shetty, S.; Moher, D.; et al. The STARD-AI reporting guideline for diagnostic accuracy studies using artificial intelligence. Nat. Med. 2025, 31, 3283–3289, Correction in Nat. Med. 2026. https://doi.org/10.1038/s41591-026-04570-9. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vasey, B.; Nagendran, M.; Campbell, B.; Clifton, D.A.; Collins, G.S.; Denaxas, S.; Denniston, A.K.; Faes, L.; Geerts, B.; Ibrahim, M.; et al. Reporting guideline for the early-stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. Nat. Med. 2022, 28, 924–933, Correction in Nat. Med. 2022, 28, 2218. https://doi.org/10.1038/s41591-022-01951-8. [Google Scholar] [CrossRef] [PubMed]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |