Next Article in Journal
68Ga-NY104 PET/CT in the Differential Diagnosis of FDG-Negative Renal Masses: A Two-Case Illustration of Clear Cell Carcinoma Versus Renal Hemangioma
Previous Article in Journal
Lung Ultrasound Versus Chest Radiography for Acute Heart Failure: Impact of Heart Failure History and Pleural Effusion
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Assessment of AI-Driven Large Language Models for Orthodontic Aesthetic Scoring Using the IOTN-AC

Department of Orthodontics, Faculty of Dentistry, Zonguldak Bulent Ecevit University, Zonguldak 67600, Türkiye
*
Author to whom correspondence should be addressed.
Diagnostics 2025, 15(23), 3048; https://doi.org/10.3390/diagnostics15233048
Submission received: 8 October 2025 / Revised: 21 November 2025 / Accepted: 25 November 2025 / Published: 29 November 2025
(This article belongs to the Section Machine Learning and Artificial Intelligence in Diagnostics)

Abstract

Background/Objectives: The aim of this study was to evaluate the accuracy of aesthetic assessments performed by artificial intelligence (AI)-based large language models (LLMs) using the Aesthetic Component of the Index of Orthodontic Treatment Need (IOTN-AC), which is widely applied to determine the need for orthodontic treatment. Methods: A total of 150 frontal intraoral photographs from patients in the permanent dentition, scored from 1 to 10 on the IOTN-AC, were assessed by two AI-based LLMs (ChatGPT-5 and ChatGPT-5 Pro). Two experienced clinicians independently scored all photographs, with one evaluator’s scores used as the reference (κ = 0.91, ICC = 0.88). Model performance was analyzed by comparing IOTN-AC scores and treatment need classifications. In addition, performance parameters such as accuracy, precision, specificity, and sensitivity were evaluated. Statistical analyses included Spearman correlation, Cohen’s Kappa, ICC, Mean Absolute Error (MAE), Wilcoxon signed-rank test, and Bland–Altman analysis. Results: Both models demonstrated positive and significant correlations with the reference values for scoring and classification (p < 0.001). Compared to GPT-5 Pro, the GPT-5 model exhibited superior performance, with a lower error rate (MAE = 1.47) and higher classification accuracy (66.7%). Bland–Altman analysis showed that most predictions fell within the 99% confidence interval, and regression analysis revealed no systematic bias (p > 0.05). Conversely, the models failed to achieve consistently high performance in each of the performance parameters. Conclusions: The findings revealed that although AI-based LLMs are promising, statistical accuracy alone is insufficient for safe clinical use, and they should demonstrate consistently high performance across all parameters.
Keywords: orthodontic treatment need; occlusal indices; Index of Orthodontic Treatment Need; aesthetic component; artificial intelligence; large language model orthodontic treatment need; occlusal indices; Index of Orthodontic Treatment Need; aesthetic component; artificial intelligence; large language model

Share and Cite

MDPI and ACS Style

Yıldırım, A.; Cicek, O. Assessment of AI-Driven Large Language Models for Orthodontic Aesthetic Scoring Using the IOTN-AC. Diagnostics 2025, 15, 3048. https://doi.org/10.3390/diagnostics15233048

AMA Style

Yıldırım A, Cicek O. Assessment of AI-Driven Large Language Models for Orthodontic Aesthetic Scoring Using the IOTN-AC. Diagnostics. 2025; 15(23):3048. https://doi.org/10.3390/diagnostics15233048

Chicago/Turabian Style

Yıldırım, Ahmet, and Orhan Cicek. 2025. "Assessment of AI-Driven Large Language Models for Orthodontic Aesthetic Scoring Using the IOTN-AC" Diagnostics 15, no. 23: 3048. https://doi.org/10.3390/diagnostics15233048

APA Style

Yıldırım, A., & Cicek, O. (2025). Assessment of AI-Driven Large Language Models for Orthodontic Aesthetic Scoring Using the IOTN-AC. Diagnostics, 15(23), 3048. https://doi.org/10.3390/diagnostics15233048

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop