Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study
Abstract
1. Introduction
2. Materials and Methods
2.1. Study Design and Reference Standards
2.2. Chatbot Systems Evaluated
- ▪
- ChatGPT powered by GPT-4o (OpenAI, San Francisco, CA, USA);
- ▪
- Gemini 2.5 Pro (Google, Mountain View, CA, USA);
- ▪
- DeepSeek Chat powered by DeepSeek-V3-0324 (DeepSeek AI, Hangzhou, China);
- ▪
- ScholarGPT (academic version built on OpenAI’s GPT-4 architecture; OpenAI, San Francisco, CA, USA);
- ▪
- MedGebra GPT-4 (clinically oriented decision-support model trained on Saudi Arabian medical guidelines, developed by Medgebra Chatbots for Medical Professionals, accessed via https://medical-chatbot.medgebra.com/, accessed on 13 September 2026).
2.3. Question Development
2.4. Content Validity Assessment
2.5. Guideline Concordance Mapping
2.6. Administration Protocol
2.7. Response Scoring and Rater Procedure
2.8. Statistical Analysis
3. Results
3.1. Overall Model Effects
3.2. Temporal Effects
3.3. Short-Term Response Consistency
4. Discussion
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial intelligence |
| LLM | Large language model |
| AAE | American Association of Endodontists |
| ESE | European Society of Endodontology |
| CVR | Content Validity Ratio |
| CHART | Chatbot Assessment Reporting Tool |
| GQS | Global Quality Score |
References
- Dilsizian, S.E.; Siegel, E.L. Artificial intelligence in medicine and cardiac imaging: Harnessing big data and advanced computing to provide personalized medical diagnosis and treatment. Curr. Cardiol. Rep. 2014, 16, 441. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jha, S.; Topol, E.J. Adapting to artificial intelligence. JAMA 2016, 316, 2353. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Basu, K.; Sinha, R.; Ong, A.; Basu, T. Artificial intelligence: How is it changing medical sciences and its future? Indian J. Dermatol. 2020, 65, 365. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lee, D.-W.; Kim, S.-Y.; Jeong, S.-N.; Lee, J.-H. Artificial intelligence in fractured dental implant detection and classification: Evaluation using dataset from two dental hospitals. Diagnostics 2021, 11, 233. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Wuttisarnwattana, P.; Wongsapai, M.; Theppitak, S.; Ittichaicharoen, J.; Warin, K.; Thanathornwong, B.; Suebnukarn, S. Precise Identification of Oral Cancer Lesions Using Artificial Intelligence. Stud. Health Technol. Inform. 2024, 316, 1096–1097. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Bunyarit, S.S.; Nambiar, P.; Naidu, M.; Asif, M.K.; Poh, R.Y.Y. Dental age estimation of Malaysian Indian children and adolescents: Applicability of Chaillet and Demirjian’s modified method using artificial neural network. Ann. Hum. Biol. 2022, 49, 192–199. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Aminoshariae, A.; Kulild, J.; Nagendrababu, V. Artificial intelligence in endodontics: Current applications and future directions. J. Endod. 2021, 47, 1352–1357. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Endres, M.G.; Hillen, F.; Salloumis, M.; Sedaghat, A.R.; Niehues, S.M.; Quatela, O.; Hanken, H.; Smeets, R.; Beck-Broichsitter, B.; Rendenbach, C.; et al. Development of a Deep Learning Algorithm for Periapical Disease Detection in Dental Radiographs. Diagnostics 2020, 10, 430. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Pauwels, R.; Brasil, D.M.; Yamasaki, M.C.; Jacobs, R.; Bosmans, H.; Freitas, D.Q.; Haiter-Neto, F. Artificial intelligence for detection of periapical lesions on intraoral radiographs: Comparison between convolutional neural networks and human observers. Oral Surg. Oral Med. Oral Pathol. Oral Radiol. 2021, 131, 610–616. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ekert, T.; Krois, J.; Meinhold, L.; Elhennawy, K.; Emara, R.; Golla, T.; Schwendicke, F. Deep learning for the radiographic detection of apical lesions. J. Endod. 2019, 45, 917–922.e5. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Suárez, A.; Díaz-Flores García, V.; Algar, J.; Gómez Sánchez, M.; Llorente de Pedro, M.; Freire, Y. Unveiling the ChatGPT phenomenon: Evaluating the consistency and accuracy of endodontic question answers. Int. Endod. J. 2024, 57, 108–113. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Thorat, V.A.; Rao, P.; Joshi, N.; Talreja, P.; Shetty, A. The Role of Chatbot GPT Technology in Undergraduate Dental Education. Cureus 2024, 16, e54193. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Qutieshat, A.; Al Rusheidi, A.; Al Ghammari, S.; Alarabi, A.; Salem, A.; Zelihic, M. Comparative analysis of diagnostic accuracy in endodontic assessments: Dental students vs. artificial intelligence. Diagnosis 2024, 11, 259–265. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Alhaidry, H.M.; Fatani, B.; Alrayes, J.O.; Almana, A.M.; Alfhaed, N.K. ChatGPT in Dentistry: A Comprehensive Review. Cureus 2023, 15, e38317. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mohammad-Rahimi, H.; Setzer, F.C.; Aminoshariae, A.; Dummer, P.M.H.; Duncan, H.F.; Nosrat, A. Artificial intelligence chatbots in endodontic education-Concepts and potential applications. Int. Endod. J. 2025, 59, 999–1012. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abd-Alrazaq, A.; AlSaad, R.; Alhuwail, D.; Ahmed, A.; Healy, P.M.; Latifi, S.; Aziz, S.; Damseh, R.; Alabed Alrazak, S.; Sheikh, J. Large Language Models in Medical Education: Opportunities, Challenges, and Future Directions. JMIR Med. Educ. 2023, 9, e48291. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- American Association of Endodontists. Guide to Clinical Endodontics, 6th ed.; American Association of Endodontists: Chicago, IL, USA, 2019; Available online: https://www.aae.org/specialty/download/guide-to-clinical-endodontics/ (accessed on 5 May 2025).
- European Society of Endodontology. ESE Publishes S3 Level Clinical Practice Guidelines. Available online: https://www.e-s-e.eu/news-related-to-ese/ese-news/ese-publishes-s3-level-clinical-practice-guidelines/ (accessed on 5 May 2025).
- Fontenele, R.C.; Jacobs, R. Unveiling the power of artificial intelligence for image-based diagnosis and treatment in endodontics: An ally or adversary? Int. Endod. J. 2025, 58, 155–170. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Eggmann, F.; Weiger, R.; Zitzmann, N.U.; Blatz, M.B. Implications of large language models such as ChatGPT for dental medicine. J. Esthet. Restor. Dent. 2023, 35, 1098–1102. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kung, T.H.; Cheatham, M.; Medenilla, A.; Sillos, C.; De Leon, L.; Elepaño, C.; Madriaga, M.; Aggabao, R.; Diaz-Candido, G.; Maningo, J.; et al. Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models. PLoS Digit. Health 2023, 2, e0000198. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Karaca, B.; Çakmak, Y.E.; Erkal, D. Clinical Relevance of Large Language Models in Endodontics: Diagnostic Appropriateness Based on 50 Simulated Case Scenarios. Aust. Endod. J. 2026, 52, 130–138. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abdulrab, S.; Abada, H.; Mashyakhy, M.; Mostafa, N.; Alhadainy, H.; Halboub, E. Performance of 4 artificial intelligence chatbots in answering endodontic questions. J. Endod. 2025, 51, 602–608. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Özbay, Y.; Erdoğan, D.; Dinçer, G.A. Evaluation of the performance of large language models in clinical decision-making in endodontics. BMC Oral Health 2025, 25, 648. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Büyüközer Özkan, H.; Doğan Çankaya, T.; Kölüş, T. The Impact of Language Variability on Artificial Intelligence Performance in Regenerative Endodontics. Healthcare 2025, 13, 1190. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- de Araújo, L.P.; Moreno, L.B.; de Araújo, B.C.C.; Chaves, E.T.; Botero, T.M.; Romero, V.H.D. From Evidence-based Endodontics to Generative AI: A Comparative Study of 11 Large Language Models. J. Endod. 2026, 52, 1010–1015. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- The CHART Collaborative. Reporting guideline for chatbot health advice studies: The CHART statement. JAMA Netw. Open 2025, 8, e2530220. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lawshe, C.H. A Quantitative Approach to Content Validity. Pers. Psychol. 1975, 28, 563–575. [Google Scholar] [CrossRef] [Scilit]
- Ayre, C.; Scally, A.J. Critical values for Lawshe’s content validity ratio: Revisiting the original methods of calculation. Meas. Eval. Couns. Dev. 2014, 47, 79–86. [Google Scholar] [CrossRef] [Scilit]
- Bernard, A.; Langille, M.; Hughes, S.; Rose, C.; Leddin, D.; Veldhuyzen van Zanten, S. A systematic review of patient inflammatory bowel disease information resources on the World Wide Web. Am. J. Gastroenterol. 2007, 102, 2070–2077. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef] [Scilit]
- Erkal, D.; Felek, T.; Butean, O.P.; Er, K. Dens invaginatus as a diagnostic challenge: Evaluating large language models against expert endodontic reasoning. BMC Oral Health 2025, 25, 1552. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Platz, J.J.; Bryan, D.S.; Naunheim, K.S.; Ferguson, M.K. Chatbot reliability in managing thoracic surgical clinical scenarios. Ann. Thorac. Surg. 2024, 118, 275–281. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chatzopoulos, G.S.; Koidou, V.P.; Tsalikis, L.; Kaklamanos, E.G. Large language models in periodontology: Assessing their performance in clinically relevant questions. J. Prosthet. Dent. 2025, 134, 2328–2336. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Giannakopoulos, K.; Kavadella, A.; Aaqel Salim, A.; Stamatopoulos, V.; Kaklamanos, E.G. Evaluation of the Performance of Generative AI Large Language Models ChatGPT, Google Bard, and Microsoft Bing Chat in Supporting Evidence-Based Dentistry: Comparative Mixed Methods Study. J. Med. Internet Res. 2023, 25, e51580. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Günay, S.; Öztürk, A.; Yiğit, Y. The accuracy of Gemini, GPT-4, and GPT-4o in ECG analysis: A comparison with cardiologists and emergency medicine specialists. Am. J. Emerg. Med. 2024, 84, 68–73. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tokgöz Kaplan, T.; Cankar, M. Evidence-based potential of generative artificial intelligence large language models on dental avulsion: ChatGPT versus Gemini. Dent. Traumatol. 2025, 41, 178–186. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Balel, Y. Can ChatGPT be used in oral and maxillofacial surgery? J. Stomatol. Oral Maxillofac. Surg. 2023, 124, 101471. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kofos, G.; Fardi, A.; Lillis, T.; Ioannis, F.; Dabarakis, N. Evaluation of artificial intelligence conversational models in providing information on dental implants: A comparative analysis of ChatGPT, Gemini and MedGebra. J. Eval. Clin. Pract. 2025, 31, e70304. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hakami, Z.; Saheb, S.A.K.; Bawazeer, O.A. Orthodontic knowledge assessment: A comparison of five AI Chatbots. Saudi Dent. J. 2026, 38, 20. [Google Scholar] [CrossRef] [Scilit] [PubMed]




| Effect | Test Statistic (Day/Session) | df (Day/Session) | p-Value (Day/Session) | Interpretation |
|---|---|---|---|---|
| Model | 29.613/25.931 | 2.72/2.87 | <0.001/<0.001 | Significant model effect in both analyses |
| Time | 2.997/3.426 | 1.87/1.44 | 0.054/0.048 | No day effect; significant session effect |
| Model × Time | 1.221/2.434 | 5.19/5.22 | 0.295/0.030 | No day interaction; significant session interaction |
| Comparison | Model Effect | Day Effect | Model × Day | Interpretation | |||
|---|---|---|---|---|---|---|---|
| Raw p | Bonferroni-Corrected p | Raw p | Bonferroni-Corrected p | Raw p | Bonferroni-Corrected p | ||
| ChatGPT-4o—Gemini 2.5 Pro | 0.789 | 1.000 | 0.067 | 0.201 | 0.479 | 1.000 | No significant difference |
| ChatGPT-4o—DeepSeek-V3-0324 | 0.617 | 1.000 | 0.034 | 0.101 | 0.567 | 1.000 | No significant difference |
| ChatGPT-4o—ScholarGPT | 0.628 | 1.000 | 0.071 | 0.212 | 0.053 | 0.160 | No significant difference |
| ChatGPT-4o—MedGebra | <0.001 | <0.001 | 0.168 | 0.505 | 0.392 | 1.000 | ChatGPT-4o performed significantly higher |
| Gemini 2.5 Pro—DeepSeek-V3-0324 | 0.917 | 1.000 | 0.247 | 0.740 | 0.369 | 1.000 | No significant difference |
| Gemini 2.5 Pro—ScholarGPT | 0.929 | 1.000 | 0.230 | 0.689 | 0.050 | 0.149 | No significant difference |
| Gemini 2.5 Pro—MedGebra | <0.001 | <0.001 | 0.600 | 1.000 | 0.315 | 0.945 | Gemini 2.5 Pro performed significantly higher |
| DeepSeek-V3-0324—ScholarGPT | 1.000 | 1.000 | 0.065 | 0.196 | 0.212 | 0.636 | No significant difference |
| DeepSeek-V3-0324—MedGebra | <0.001 | <0.001 | 0.187 | 0.562 | 0.722 | 1.000 | DeepSeek-V3-0324 performed significantly higher |
| ScholarGPT—MedGebra | <0.001 | <0.001 | 0.078 | 0.235 | 0.743 | 1.000 | ScholarGPT performed significantly higher |
| Comparison | Effect | Test Statistic | df | p-Value | Bonferroni-Corrected p |
|---|---|---|---|---|---|
| ChatGPT-4o—Gemini 2.5 Pro | Model effect | 0.072 | 1.00 | 0.789 | 1.000 |
| Day effect | 2.773 | 1.84 | 0.067 | 0.201 | |
| Model × Day | 0.704 | 1.78 | 0.479 | 1.000 | |
| ChatGPT-4o—DeepSeek-V3-0324 | Model effect | 0.250 | 1.00 | 0.617 | 1.000 |
| Day effect | 3.431 | 1.94 | 0.034 | 0.101 | |
| Model × Day | 0.489 | 1.56 | 0.567 | 1.000 | |
| ChatGPT-4o—ScholarGPT | Model effect | 0.235 | 1.00 | 0.628 | 1.000 |
| Day effect | 2.733 | 1.81 | 0.071 | 0.212 | |
| Model × Day | 2.957 | 1.95 | 0.053 | 0.160 | |
| ChatGPT-4o—MedGebra | Model effect | 51.781 | 1.00 | <0.001 | <0.001 |
| Day effect | 1.812 | 1.75 | 0.168 | 0.505 | |
| Model × Day | 0.894 | 1.65 | 0.392 | 1.000 | |
| Gemini 2.5 Pro—DeepSeek-V3-0324 | Model effect | 0.011 | 1.00 | 0.917 | 1.000 |
| Day effect | 1.401 | 1.85 | 0.247 | 0.740 | |
| Model × Day | 0.988 | 1.90 | 0.369 | 1.000 | |
| Gemini 2.5 Pro—ScholarGPT | Model effect | 0.0078 | 1.00 | 0.929 | 1.000 |
| Day effect | 1.471 | 1.99 | 0.230 | 0.689 | |
| Model × Day | 3.007 | 1.99 | 0.050 | 0.149 | |
| Gemini 2.5 Pro—MedGebra | Model effect | 43.173 | 1.00 | <0.001 | <0.001 |
| Day effect | 0.471 | 1.76 | 0.600 | 1.000 | |
| Model × Day | 1.084 | 1.30 | 0.315 | 0.945 | |
| DeepSeek-V3-0324—ScholarGPT | Model effect | 0.000 | 1.00 | 1.000 | 1.000 |
| Day effect | 2.790 | 1.86 | 0.065 | 0.196 | |
| Model × Day | 1.564 | 1.74 | 0.212 | 0.636 | |
| DeepSeek-V3-0324—MedGebra | Model effect | 59.258 | 1.00 | <0.001 | <0.001 |
| Day effect | 1.693 | 1.78 | 0.187 | 0.562 | |
| Model × Day | 0.291 | 1.78 | 0.722 | 1.000 | |
| ScholarGPT—MedGebra | Model effect | 51.931 | 1.00 | <0.001 | <0.001 |
| Day effect | 2.577 | 1.92 | 0.078 | 0.235 | |
| Model × Day | 0.244 | 1.67 | 0.743 | 1.000 |
| AI Model | Timing Comparison | Weighted κ | SE | Z | p-Value | 95% CI |
|---|---|---|---|---|---|---|
| ChatGPT-4o | G1–G2 | 0.606 | 0.148 | 4.150 | <0.001 | 0.316–0.896 |
| G1–G3 | 0.662 | 0.155 | 4.327 | <0.001 | 0.358–0.966 | |
| G2–G3 | 0.630 | 0.131 | 4.486 | <0.001 | 0.374–0.886 | |
| Gemini 2.5 Pro | G1–G2 | 0.730 | 0.115 | 4.995 | <0.001 | 0.504–0.956 |
| G1–G3 | 0.717 | 0.127 | 4.695 | <0.001 | 0.469–0.966 | |
| G2–G3 | 0.689 | 0.118 | 4.742 | <0.001 | 0.457–0.921 | |
| DeepSeek-V3-0324 | G1–G2 | 0.579 | 0.142 | 4.015 | <0.001 | 0.301–0.857 |
| G1–G3 | 0.520 | 0.157 | 3.547 | <0.001 | 0.212–0.829 | |
| G2–G3 | 0.645 | 0.134 | 4.571 | <0.001 | 0.381–0.908 | |
| ScholarGPT | G1–G2 | 0.458 | 0.146 | 3.443 | 0.001 | 0.171–0.745 |
| G1–G3 | 0.504 | 0.161 | 3.633 | <0.001 | 0.188–0.821 | |
| G2–G3 | 0.592 | 0.146 | 4.167 | <0.001 | 0.306–0.879 | |
| MedGebra | G1–G2 | 0.739 | 0.098 | 5.873 | <0.001 | 0.547–0.932 |
| G1–G3 | 0.438 | 0.149 | 3.566 | <0.001 | 0.147–0.730 | |
| G2–G3 | 0.240 | 0.177 | 1.924 | 0.054 | −0.107–0.587 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
İlgenli, I.; Avcı, E.; Köse, T. Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study. Dent. J. 2026, 14, 598. https://doi.org/10.3390/dj14090598
İlgenli I, Avcı E, Köse T. Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study. Dentistry Journal. 2026; 14(9):598. https://doi.org/10.3390/dj14090598
Chicago/Turabian Styleİlgenli, Ilgın, Ezgi Avcı, and Timur Köse. 2026. "Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study" Dentistry Journal 14, no. 9: 598. https://doi.org/10.3390/dj14090598
APA Styleİlgenli, I., Avcı, E., & Köse, T. (2026). Accuracy and Short-Term Consistency of Generative AI Chatbots in Guideline-Based Endodontic Decision Support: A Multi-Model Benchmarking Study. Dentistry Journal, 14(9), 598. https://doi.org/10.3390/dj14090598

