Evaluating Large Language Models for Diagnostic Accuracy and Health Information Quality in Oral Mucosal Diseases
Abstract
1. Introduction
2. Methods
2.1. Study Design and Case Selection
2.2. Search Strategy
- Irrelevant content;
- Non-health-related blogs;
- Videos;
- Broken or paywalled links;
- Non-English content.
2.3. Diagnostic Accuracy Assessment
- Directly matched the confirmed diagnosis;
- Had the correct diagnosis within the top three differential diagnoses.
2.4. Information Quality Assessment
- Section 1: Reliability (Q1–8);
- Section 2: Quality of treatment information (Q9–15);
- Section 3: Overall rating (Q16).
2.5. Readability Assessment
- FRES = 206.835 − (1.015 × ASL) − (84.6 × ASW);
- FKRGL = (0.39 × ASL) + (11.8 × ASW) − 15.59.
- 90–100 = very easy;
- 80–89 = easy;
- 70–79 = fairly easy;
- 60–69 = standard;
- 50–59 = fairly difficult;
- 30–49 = difficult;
- 0–29 = very difficult.
2.6. Statistical Analysis
3. Results
3.1. Readability and Quality of Information
3.2. Diagnostic Accuracy
3.3. Predictive Value Analysis
4. Discussion
4.1. Implications for Clinical and Consumer Use
4.2. Limitations and Future Research
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- GBD 2021 Oral Disorders Collaborators. Trends in the global, regional, and national burden of oral conditions from 1990 to 2021: A systematic analysis for the Global Burden of Disease Study 2021. Lancet 2025, 405, 897–910. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- da Silva, K.D.; de Oliveira, W.L.; Sarkis-Onofre, R.; Aitken-Saavedra, J.P.; Demarco, F.F.; Correa, M.B.; Tarquinio, S.B.C. Prevalence of oral mucosal lesions in population-based studies: A systematic review of the methodological aspects. Community Dent. Oral Epidemiol. 2019, 47, 431–440. [Google Scholar] [CrossRef] [Scilit]
- Feng, J.; Zhou, Z.; Shen, X.; Wang, Y.; Shi, L.; Wang, Y.; Hu, Y.; Sun, H.; Liu, W. Prevalence and distribution of oral mucosal lesions: A cross-sectional study in Shanghai, China. J. Oral Pathol. Med. 2015, 44, 490–494. [Google Scholar] [CrossRef] [Scilit]
- Do, L.G.; Spencer, A.J.; Dost, F.; Farah, C.S. Oral mucosal lesions: Findings from the Australian National Survey of Adult Oral Health. Aust. Dent. J. 2014, 59, 114–120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, X.; Wang, B.; Zhang, Z.; Han, H.; Liu, Z.; Li, S.; Yu, L.; Zhang, P. Global, regional and national burden of lip and oral cavity cancer from 1990 to 2021 and projections to 2036. BMC Oral Health 2025, 25, 1779. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Vaira, L.A.; Lechien, J.R.; Maniaci, A.; De Vito, A.; Mayo-Yáñez, M.; Troise, S.; Consorti, G.; Chiesa-Estomba, C.M.; Cammaroto, G.; Radulesco, T.; et al. Diagnostic Performance of ChatGPT-4o in Analyzing Oral Mucosal Lesions: A Comparative Study with Experts. Medicina 2025, 61, 1379. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Idrees, M.; Halimi, R.; Gadiraju, S.; Frydrych, A.M.; Kujan, O. Clinical competency of dental health professionals and students in diagnosing oral mucosal lesions. Oral Dis. 2024, 30, 3108–3116. [Google Scholar] [CrossRef] [Scilit]
- Zachariadis, C.B.; Leligou, H.C. Harnessing Artificial Intelligence for Automated Diagnosis. Information 2024, 15, 311. [Google Scholar] [CrossRef] [Scilit]
- Gul, S.; Erdemir, I.; Hanci, V.; Aydogmus, E.; Erkoc, Y.S. How artificial intelligence can provide information about subdural hematoma: Assessment of readability, reliability, and quality of ChatGPT, BARD, and Perplexity responses. Medicine 2024, 103, e38009. [Google Scholar] [CrossRef] [Scilit]
- Zhou, M.; Pan, Y.; Zhang, Y.; Song, X.; Zhou, Y. Evaluating AI-generated patient education materials for spinal surgeries: Comparative analysis of readability and DISCERN quality across ChatGPT and deepseek models. Int. J. Med. Inform. 2025, 198, 105871. [Google Scholar] [CrossRef] [Scilit]
- Silva, T.P.; Andrade-Bortoletto, M.F.S.; Ocampo, T.S.C.; Alencar-Palha, C.; Bornstein, M.M.; Oliveira-Santos, C.; Oliveira, M.L. Performance of a commercially available generative pre-trained transformer (GPT) in describing radiolucent lesions in panoramic radiographs and establishing differential diagnoses. Clin. Oral Investig. 2024, 28, 204. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yu, S.; Sun, W.; Mi, D.; Jin, S.; Wu, X.; Xin, B.; Zhang, H.; Wang, Y.; Sun, X.; He, X. Artificial intelligence diagnosing of oral lichen planus: A comparative study. Bioengineering 2024, 11, 1159. [Google Scholar] [CrossRef] [Scilit]
- Pradhan, P. Accuracy of ChatGPT 3.5, 4.0, 4o and Gemini in diagnosing oral potentially malignant lesions based on clinical case reports and image recognition. Med. Oral Patol. Oral Cirugia Bucal 2025, 30, e1–e10. [Google Scholar] [CrossRef] [Scilit]
- Lorenzo-Pouso, A.; Perez-Sayans, M.; Kujan, O.; Castelo-Baz, P.; Chamorro-Petronacci, C.; Garcia-Garcia, A.; Blanco-Carrion, A. Patient-centered web-based information on oral lichen planus: Quality and readability. Med. Oral Patol. Oral Cirugia Bucal 2019, 24, e461–e467. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Weiss, B.D. Health Literacy: A Manual for Clinicians; American Medical Association Foundation: Chicago, IL, USA, 2003. [Google Scholar]
- Cotugna, N.; Vickery, C.E.; Carpenter-Haefele, K.M. Evaluation of literacy level of patient education pages in health-related journals. J. Community Health 2005, 30, 213–219. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Charnock, D.; Shepperd, S.; Needham, G.; Gann, R. DISCERN: An instrument for judging the quality of written consumer health information on treatment choices. J. Epidemiol. Community Health 1999, 53, 105–111. [Google Scholar] [CrossRef] [Scilit]
- Jansen, B.; Mullane, S.; Tan, B.; Joseph, B.; Junaid, M.; Balasubramaniam, R.; Frydrych, A.; Kujan, O. Patient-Centred Web-Based Information on Head and Neck Squamous Cell Carcinoma: Quality and Readability. Oral Dis. 2025. [Google Scholar] [CrossRef] [Scilit]
- Özdemir, Ö.T.; Kavan, M.Y.; Güven, Y. Evaluation of the readability, quality, and accuracy of AI chatbot responses to questions about deleterious oral habits. BMC Oral Health 2025, 25, 1812. [Google Scholar] [CrossRef] [Scilit]
- Dinc, M.T.; Bardak, A.E.; Bahar, F.; Noronha, C. Comparative Analysis of Large Language Models in Clinical Diagnosis: Performance Evaluation Across Common and Complex Medical Cases. JAMIA Open 2025, 8, ooaf055. [Google Scholar] [CrossRef] [Scilit]
- Thirunavukarasu, A.J.; Ting, D.S.J.; Elangovan, K.; Gutierrez, L.; Tan, T.F.; Ting, D.S.W. Large language models in medicine. Nat. Med. 2023, 29, 1930–1940. [Google Scholar] [CrossRef] [Scilit]
- Musheyev, D.; Pan, A.; Gross, P.; Kamyab, D.; Kaplinsky, P.; Spivak, M.; Bragg, M.A.; Loeb, S.; Kabarriti, A.E. Readability and information quality in cancer information from a free vs paid chatbot. JAMA Netw. Open 2024, 7, e2422275. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Lopez-Jornet, P.; Camacho-Alonso, F. The quality of patient-oriented Internet information on oral lichen planus: A pilot study. J. Eval. Clin. Pract. 2010, 16, 883–886. [Google Scholar] [CrossRef] [Scilit]
- Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.J.; Madotto, A.; Fung, P. Survey of hallucination in natural language generation. ACM Comput. Surv. 2023, 55, 248. [Google Scholar] [CrossRef] [Scilit]
- Vimalraj, S.; Sekaran, S. ChatGPT: Empowering dentistry with future possibilities. Oral Oncol. 2023, 144, 106496. [Google Scholar] [CrossRef] [Scilit]
- Milne-Ives, M.; de Cock, C.; Lim, E.; Shehadeh, M.H.; de Pennington, N.; Mole, G.; Normando, E.; Meinert, E. The effectiveness of artificial intelligence conversational agents in health care: Systematic review. J. Med. Internet Res. 2020, 22, e20346. [Google Scholar] [CrossRef] [Scilit]
- Haver, H.L.; Gupta, A.K.; Ambinder, E.B.; Bahl, M.; Oluyemi, E.T.; Jeudy, J.; Yi, P.H. Evaluating the use of ChatGPT to accurately simplify patient-centered information about breast cancer prevention and screening. Radiol. Imaging Cancer 2024, 6, e230086. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hancı, V.; Ergün, B.; Gül, Ş.; Uzun, Ö.; Erdemir, İ.; Hancı, F.B. Assessment of Readability, Reliability, and Quality of ChatGPT®, BARD®, Gemini®, Copilot®, Perplexity® responses on palliative care. Medicine 2024, 103, e39305. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Tursunbayeva, A.; Renkema, M. Artificial intelligence in health care: Implications for the job design of healthcare professionals. Asia Pac. J. Hum. Resour. 2023, 61, 845–887. [Google Scholar] [CrossRef] [Scilit]
- Guimaraes, G.R.; Figueiredo, R.G.; Silva, C.S.; Arata, V.; Contreras, J.C.Z.; Gomes, C.M.; Tiraboschi, R.B.; Junior, J.B. Diagnosis in bytes: Comparing the diagnostic accuracy of Google and ChatGPT 3.5 as an educational support tool. Int. J. Environ. Res. Public Health 2024, 21, 580. [Google Scholar] [CrossRef] [Scilit]


| Platform | DISCERN (Q16) | DISCERN IQRange | FRES | FRES IQRange | FKGL | FKGL Range | College-Level (%) | Professional Level (%) |
|---|---|---|---|---|---|---|---|---|
| Bing | 1.46 | 1.75–1.00 | 67.63 | 63.00–73.00 | 6.80 | 5.67–7.31 | 0 | 0 |
| Yahoo | 1.42 | 1.00–1.00 | 67.63 | 60.00–75.50 | 6.89 | 5.77–7.61 | 0 | 0 |
| 1.61 | 2.00–1.00 | 64.33 | 59.00–69.00 | 7.57 | 6.40–8.46 | 2 | 0 | |
| ChatGPT 3.5 | 2.73 | 3.00–2.25 | 43.58 | 38.00–49.75 | 12.31 | 11.27–13.35 | 26 | 4 |
| ChatGPT 4.5 (Plus) | 3.2 | 3.00–3.50 | 41.04 | 38.50–48.75 | 14.05 | 11.85–14.11 | 37 | 8 |
| Microsoft Copilot | 2.96 | 3.00–3.00 | 44.08 | 37.25–50.75 | 10.67 | 9.99–11.71 | 7 | 1 |
| Claude (Sonnet 4.5) | 2.92 | 2.95–3.12 | 52.15 | 36.00–64.44 | 11.10 | 10.66–12.34 | 9 | 3 |
| Grok | 2.84 | 2.86–3.04 | 56.42 | 42.00–62.10 | 7.88 | 6.84–10.68 | 7 | 2 |
| DeepSeek v3.1 | 2.99 | 2.92–3.00 | 54.25 | 36.64–66.40 | 8.68 | 7.42–10.80 | 8 | 4 |
| Platform | Benign Oral Lesions | Malignant Lesions | OPMDs | Oral Infections | Reactive Oral Lesions |
|---|---|---|---|---|---|
| ChatGPT 4.5 | 90 * | 85 * | 94 * | 89 * | 87 * |
| DeepSeek v3.1 | 74 | 68 | 75 | 73 | 69 |
| Claude (Sonnet 4.5) | 72 | 66 | 72 | 66 | 63 |
| ChatGPT 3.5 | 70 | 63 | 73 | 68 | 65 |
| Copilot | 65 | 60 | 67 | 64 | 61 |
| Grok | 61 | 56 | 65 | 60 | 57 |
| 55 | 50 | 58 | 53 | 50 | |
| Bing | 35 | 28 | 40 | 33 | 30 |
| Yahoo | 25 | 18 | 28 | 22 | 20 |
| Platform | Benign Oral Lesions (PPV/NPV) | Malignant Lesions (PPV/NPV) | OPMDs (PPV/NPV) | Oral Infections (PPV/NPV) | Reactive Oral Lesions (PPV/NPV) |
|---|---|---|---|---|---|
| ChatGPT 4.5 | 92/88 | 70/92 | 85/90 | 80/85 | 91/92 |
| DeepSeek v3.1 | 65/50 | 60/80 | 72/76 | 75/82 | 72/70 |
| Claude (Sonnet 4.5) | 60/45 | 55/78 | 68/74 | 70/80 | 66/64 |
| ChatGPT 3.5 | 55/30 | 30/75 | 55/68 | 60/73 | 60/58 |
| Copilot | 60/35 | 45/80 | 65/70 | 68/77 | 55/53 |
| Grok | 55/30 | 50/85 | 60/75 | 60/80 | 60/60 |
| 50/25 | 45/80 | 60/70 | 65/75 | 50/45 | |
| Bing | 25/15 | 20/50 | 30/40 | 25/45 | 28/25 |
| Yahoo | 12/8 | 8/15 | 15/20 | 10/18 | 12/10 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Iacob, M.; Qawas, A.; Balasubramaniam, R.; Frydrych, A.M.; Kujan, O. Evaluating Large Language Models for Diagnostic Accuracy and Health Information Quality in Oral Mucosal Diseases. J. Pers. Med. 2026, 16, 129. https://doi.org/10.3390/jpm16030129
Iacob M, Qawas A, Balasubramaniam R, Frydrych AM, Kujan O. Evaluating Large Language Models for Diagnostic Accuracy and Health Information Quality in Oral Mucosal Diseases. Journal of Personalized Medicine. 2026; 16(3):129. https://doi.org/10.3390/jpm16030129
Chicago/Turabian StyleIacob, Melisa, Ayham Qawas, Ramesh Balasubramaniam, Agnieszka M. Frydrych, and Omar Kujan. 2026. "Evaluating Large Language Models for Diagnostic Accuracy and Health Information Quality in Oral Mucosal Diseases" Journal of Personalized Medicine 16, no. 3: 129. https://doi.org/10.3390/jpm16030129
APA StyleIacob, M., Qawas, A., Balasubramaniam, R., Frydrych, A. M., & Kujan, O. (2026). Evaluating Large Language Models for Diagnostic Accuracy and Health Information Quality in Oral Mucosal Diseases. Journal of Personalized Medicine, 16(3), 129. https://doi.org/10.3390/jpm16030129

