Reproducibility of Mandibular Cortical Index Classification Among Dental Examiners and a Single Generative AI Platform: An Observer Agreement Study
Abstract
1. Introduction
2. Materials and Methods
2.1. Study Design
2.2. Patient Selection and Data Sources
2.3. DPR Imaging Data
2.4. Evaluation of the Mandibular Cortical Index (MCI)
- Class 1 (C1): The endosteal margin of the cortex is even and sharp on both sides.
- Class 2 (C2): The endosteal margin shows semilunar defects (lacunar resorption) or forms endosteal cortical residues.
- Class 3 (C3): The cortical layer is heavy with endosteal residues and is clearly porous.
2.5. Examiners: Dentists
- A male periodontist (PER-M28; 28 years of clinical experience)
- A male endodontist (END-M33; 33 years)
- A female prosthodontist (PRO-F29; 29 years)
- A male prosthodontist (PRO-M19; 19 years)
- A male dental radiologist (RAD-REF; 14 years), who served as the reference examiner
- A male general practitioner (GP-M02; 2 years)
- A female postgraduate resident (PGR-F01; 1 year)
- A male postgraduate resident (PGR-M01; 1 year)
2.6. Examiners: Generative AI
2.7. Statistical Analysis
3. Results
3.1. Inter-Rater Reliability
3.2. Intra-Examiner Reliability and Evaluation Time
3.3. Discrepancies in Diagnostic Criteria Between Dentists and Generative AI
3.4. Consistency and Persisting Discrepancies in the Second Evaluation Session
3.5. Agreement Between Individual Examiners and the Reference Standard
3.6. Qualitative Analysis of Diagnostic Disagreements
3.7. Specific Classification Bias and the Nature of AI Discrepancies
3.8. Visualization of Inter-Examiner Agreement Using Heatmaps
4. Discussion
4.1. Comparison of Reproducibility and Inter-Rater Agreement
4.2. Learning Effects in Human Examiners
4.3. Limitations
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Taguchi, A.; Tanaka, R.; Kakimoto, N.; Morimoto, Y.; Arai, Y.; Hayashi, T.; Kurabayashi, T.; Katsumata, A.; Asaumi, J.; Japanese Society for Oral and Maxillofacial Radiology. Clinical guidelines for the application of panoramic radiographs in screening for osteoporosis. Oral Radiol. 2021, 37, 189–208. [Google Scholar] [CrossRef] [PubMed]
- Seki, K.; Nagasaki, M.; Yoshino, T.; Yano, M.; Kawamoto, A.; Shimizu, O. Radiographical diagnostic evaluation of mandibular cortical index classification and mandibular cortical width in female patients prescribed antiosteoporosis medication. Diagnostics 2024, 14, 1009. [Google Scholar] [CrossRef] [PubMed]
- Klemetti, E.; Kolmakov, S.; Kröger, H. Pantomography in assessment of the osteoporosis risk group. Eur. J. Oral Sci. 1994, 102, 68–72. [Google Scholar] [CrossRef] [PubMed]
- Curtis, E.M.; Dennison, E.M.; Cooper, C.; Harvey, N.C. Osteoporosis in 2022: Care gaps to screening and personalised medicine. Best Pract. Res. Clin. Rheumatol. 2022, 36, 101754. [Google Scholar] [CrossRef] [PubMed]
- Singer, A.J.; Sharma, A.; Deignan, C.; Borgermans, L. Closing the gap in osteoporosis management: The critical role of primary care in bone health. Curr. Med. Res. Opin. 2023, 39, 387–398. [Google Scholar] [CrossRef] [PubMed]
- Jowitt, N.; MacFarlane, T.; Devlin, H.; Klemetti, E.; Horner, K. The reproducibility of the mandibular cortical index. Dentomaxillofac. Radiol. 1999, 28, 141–144. [Google Scholar] [CrossRef] [PubMed]
- Revilla-León, M.; Gómez-Polo, M.; Vyas, S.; Barmak, A.B.; Özcan, M.; Att, W.; Krishnamurthy, V.R. Artificial intelligence applications in restorative dentistry: A systematic review. J. Prosthet. Dent. 2022, 128, 867–875. [Google Scholar] [CrossRef] [PubMed]
- Revilla-León, M.; Gómez-Polo, M.; Vyas, S.; Barmak, A.B.; Gallucci, G.O.; Att, W.; Krishnamurthy, V.R. Artificial intelligence applications in implant dentistry: A systematic review. J. Prosthet. Dent. 2023, 129, 293–300. [Google Scholar] [CrossRef] [PubMed]
- Liu, Z.; Nalley, A.; Hao, J.; Ai, Q.Y.H.; Yeung, A.W.K.; Tanaka, R.; Hung, K.F. The performance of large language models in dentomaxillofacial radiology: A systematic review. Dentomaxillofac. Radiol. 2025, 54, 613–631. [Google Scholar] [CrossRef] [PubMed]
- Nguyen, V.A.; Vuong, T.Q.T.; Nguyen, V.H. Benchmarking large-language-model vision capabilities in oral and maxillofacial anatomy: A cross-sectional study. PLoS ONE 2025, 20, e0335775. [Google Scholar] [CrossRef] [PubMed]
- World Medical Association. World Medical Association Declaration of Helsinki: Ethical Principles for Medical Research Involving Human Subjects. JAMA 2013, 310, 2191–2194. [Google Scholar] [CrossRef] [PubMed]
- von Elm, E.; Altman, D.G.; Egger, M.; Pocock, S.J.; Gøtzsche, P.C.; Vandenbroucke, J.P.; STROBE Initiative. Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) statement: Guidelines for reporting observational studies. BMJ 2007, 335, 806–808. [Google Scholar] [CrossRef] [PubMed]
- Kottner, J.; Audigé, L.; Brorson, S.; Donner, A.; Gajewski, B.J.; Hróbjartsson, A.; Roberts, C.; Shoukri, M.; Streiner, D.L. Guidelines for Reporting Reliability and Agreement Studies (GRRAS) were proposed. J. Clin. Epidemiol. 2011, 64, 96–106. [Google Scholar] [CrossRef] [PubMed]
- Yasar, F.; Sener, S.; Yesilova, E.; Akgünlü, F. Mandibular cortical index evaluation in masked and unmasked panoramic radiographs. Dentomaxillofac. Radiol. 2009, 38, 86–91. [Google Scholar] [CrossRef] [PubMed]
- Cohen, J. Weighted kappa: Nominal scale agreement with provision for scaled disagreement or partial credit. Psychol. Bull. 1968, 70, 213–220. [Google Scholar] [CrossRef] [PubMed]
- Kanda, Y. Investigation of the freely available easy-to-use software ‘EZR’ for medical statistics. Bone Marrow Transplant. 2013, 48, 452–458. [Google Scholar] [CrossRef] [PubMed]
- Landis, J.R.; Koch, G.G. The measurement of observer agreement for categorical data. Biometrics 1977, 33, 159–174. [Google Scholar] [CrossRef]
- Feinstein, A.R.; Cicchetti, D.V. High agreement but low kappa: I. The problems of two paradoxes. J. Clin. Epidemiol. 1990, 43, 543–549. [Google Scholar] [CrossRef] [PubMed]
- Gwet, K.L. Computing inter-rater reliability and its variance in the presence of high agreement. Br. J. Math. Stat. Psychol. 2008, 61, 29–48. [Google Scholar] [CrossRef] [PubMed]
- Gwet, K.L. Handbook of Inter-Rater Reliability: The Definitive Guide to Measuring the Extent of Agreement Among Raters, 4th ed.; Advanced Analytics, LLC: Gaithersburg, MD, USA, 2014. [Google Scholar]
- Tassoker, M.; Öziç, M.U.; Yuce, F. Comparison of five convolutional neural networks for predicting osteoporosis based on mandibular cortical index on panoramic radiographs. Dentomaxillofac. Radiol. 2022, 51, 20220108. [Google Scholar] [CrossRef] [PubMed]
- Nakamoto, T.; Taguchi, A.; Kakimoto, N. Osteoporosis screening support system from panoramic radiographs using deep learning by convolutional neural network. Dentomaxillofac. Radiol. 2022, 51, 20220135. [Google Scholar] [CrossRef] [PubMed]
- Urooj, B.; Ali, S.; Naqvi, S.K.H.; Xiao, F.; Huang, P.C. Large language models in medical image analysis: A systematic survey and future directions. Biomed. J. 2025, 48, 100932. [Google Scholar] [CrossRef] [PubMed]
- Bhayana, R. Chatbots and large language models in radiology: A practical primer for clinical and research applications. Radiology 2024, 310, e232756. [Google Scholar] [CrossRef] [PubMed]
- Personal Information Protection Commission; Ministry of Health, Labour and Welfare. Guidance on Appropriate Handling of Personal Information by Medical and Long-Term Care Service Providers. Available online: https://www.ppc.go.jp/personalinfo/legal/iryoukaigo_guidance/ (accessed on 26 April 2026).
- Conduah, A.K.; Ofoe, S.; Siaw-Marfo, D. Data privacy in healthcare: Global challenges and solutions. Digit. Health 2025, 11, 20552076251343959. [Google Scholar] [CrossRef] [PubMed]
- Meszaros, J.; Minari, J.; Huys, I. The future regulation of artificial intelligence systems in healthcare services and medical research in the European Union. Front. Genet. 2022, 13, 927721. [Google Scholar] [CrossRef] [PubMed]
- Ng, M.Y.; Helzer, J.; Pfeffer, M.A.; Seto, T.; Hernandez-Boussard, T. Development of secure infrastructure for advancing generative artificial intelligence research in healthcare at an academic medical center. J. Am. Med. Inform. Assoc. 2025, 32, 586–588. [Google Scholar] [CrossRef] [PubMed]
- Ezhov, M.; Gusarev, M.; Golitsyna, M.; Yates, J.M.; Kushnerev, E.; Tamimi, D.; Aksoy, S.; Shumilov, E.; Sanders, A.; Orhan, K. Clinically applicable artificial intelligence system for dental diagnosis with CBCT. Sci. Rep. 2021, 11, 15006. [Google Scholar] [CrossRef] [PubMed]
- Ding, H.; Wu, J.; Zhao, W.; Matinlinna, J.P.; Burrow, M.F.; Tsoi, J.K.H. Artificial intelligence in dentistry—A review. Front. Dent. Med. 2023, 4, 1085251. [Google Scholar] [CrossRef] [PubMed]
- Claman, D.; Sezgin, E. Artificial intelligence in dental education: Opportunities and challenges of large language models and multimodal foundation models. JMIR Med. Educ. 2024, 10, e52346. [Google Scholar] [CrossRef] [PubMed]
- Uribe, S.E.; Maldupa, I.; Schwendicke, F. Integrating generative AI in dental education: A scoping review of current practices and recommendations. Eur. J. Dent. Educ. 2025, 29, 341–355. [Google Scholar] [CrossRef] [PubMed]
- Ghasemi, N.; Rokhshad, R.; Zare, Q.; Shobeiri, P.; Schwendicke, F. Artificial intelligence for osteoporosis detection on panoramic radiography: A systematic review and meta-analysis. J. Dent. 2025, 156, 105650. [Google Scholar] [CrossRef] [PubMed]
- Khadivi, G.; Akhtari, A.; Sharifi, F.; Zargarian, N.; Esmaeili, S.; Ahsaie, M.G.; Shahbazi, S. Diagnostic accuracy of artificial intelligence models in detecting osteoporosis using dental images: A systematic review and meta-analysis. Osteoporos. Int. 2025, 36, 1–19. [Google Scholar] [CrossRef] [PubMed]





| 1st Session | 2nd Session | |
|---|---|---|
| All Examiners (AI + Dentists) | 0.498 | 0.542 |
| [0.441, 0.555] | [0.485, 0.598] | |
| Dentists Only | 0.606 | 0.627 |
| [0.548, 0.664] | [0.571, 0.683] |
| Examiner ID | Weighted Cohen’s κ [95% CI] | Interpretation | Time 1 (min) | Time 2 (min) | Mean Time (min) |
|---|---|---|---|---|---|
| AI-NBLM | 0.987 [0.962, 1.000] | Almost perfect | 2 | 1 | 1.5 |
| END-M33 | 0.685 [0.578, 0.792] | Substantial | 5 | 7 | 6 |
| PER-M28 | 0.64 [0.521, 0.758] | Substantial | 7 | 6 | 6.5 |
| PRO-F29 | 0.645 [0.512, 0.778] | Substantial | 30 | 11 | 20.5 |
| RAD-REF | 0.703 [0.589, 0.817] | Substantial | 6 | 5 | 5.5 |
| PRO-M19 | 0.761 [0.651, 0.871] | Substantial | 10 | 9 | 9.5 |
| GP-M02 | 0.619 [0.466, 0.772] | Substantial | 8 | 6 | 7 |
| PGR-M01 | 0.416 [0.252, 0.581] | Fair | 15 | 15 | 15 |
| PGR-F01 | 0.257 [0.108, 0.406] | Fair | 6 | 8 | 7 |
| Session 1 | Session 2 | ||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Examiner | C1 (n) | C1 (%) | C2 (n) | C2 (%) | C3 (n) | C3 (%) | Examiner | C1 (n) | C1 (%) | C2 (n) | C2 (%) | C3 (n) | C3 (%) |
| AI-NBLM | 40 | 40 | 42 | 42 | 18 | 18 | AI-NBLM | 41 | 41 | 41 | 41 | 18 | 18 |
| RAD-REF | 58 | 58 | 33 | 33 | 9 | 9 | RAD-REF | 54 | 54 | 35 | 35 | 11 | 11 |
| END-M33 | 28 | 28 | 46 | 46 | 26 | 26 | END-M33 | 26 | 26 | 45 | 45 | 29 | 29 |
| PER-M28 | 23 | 23 | 45 | 45 | 32 | 32 | PER-M28 | 28 | 28 | 51 | 51 | 21 | 21 |
| PRO-F29 | 36 | 36 | 54 | 54 | 10 | 10 | PRO-F29 | 40 | 40 | 54 | 54 | 6 | 6 |
| PRO-M19 | 34 | 34 | 54 | 54 | 12 | 12 | PRO-M19 | 32 | 32 | 54 | 54 | 14 | 14 |
| GP-M02 | 43 | 43 | 50 | 50 | 7 | 7 | GP-M02 | 32 | 32 | 55 | 55 | 13 | 13 |
| PGR-M01 | 58 | 58 | 35 | 35 | 7 | 7 | PGR-M01 | 69 | 69 | 24 | 24 | 7 | 7 |
| PGR-F01 | 15 | 15 | 47 | 47 | 38 | 38 | PGR-F01 | 45 | 45 | 37 | 37 | 18 | 18 |
| AI-NBLM vs. RAD-REF (Session 1) | ||||
|---|---|---|---|---|
| AI-NBLM/RAD-REF | C1 | C2 | C3 | Row Total |
| C1 | 22 | 16 | 2 | 40 |
| C2 | 25 | 13 | 4 | 42 |
| C3 | 11 | 4 | 3 | 18 |
| Col Total | 58 | 33 | 9 | 100 |
| AI-NBLM vs. RAD-REF (Session 2) | ||||
| C1 | 25 | 12 | 4 | 41 |
| C2 | 22 | 16 | 3 | 41 |
| C3 | 7 | 7 | 4 | 18 |
| Col Total | 54 | 35 | 11 | 100 |
| Session 1 | ||||
|---|---|---|---|---|
| Dentist | C1 (AI)→C3 (Dentist) | C3 (AI)→C1 (Dentist) | Total (n) | Total (%) |
| RAD-REF | 2 | 11 | 13 | 13 |
| END-M33 | 11 | 2 | 13 | 13 |
| PER-M28 | 15 | 2 | 17 | 17 |
| PRO-F29 | 6 | 6 | 12 | 12 |
| PRO-M19 | 5 | 5 | 10 | 10 |
| GP-M02 | 4 | 6 | 10 | 10 |
| PGR-M01 | 4 | 10 | 14 | 14 |
| PGR-F01 | 15 | 0 | 15 | 15 |
| Session 2 | ||||
| RAD-REF | 4 | 7 | 11 | 11 |
| END-M33 | 15 | 3 | 18 | 18 |
| PER-M28 | 9 | 4 | 13 | 13 |
| PRO-F29 | 4 | 8 | 12 | 12 |
| PRO-M19 | 8 | 4 | 12 | 12 |
| GP-M02 | 5 | 6 | 11 | 11 |
| PGR-M01 | 2 | 10 | 12 | 12 |
| PGR-F01 | 7 | 8 | 15 | 15 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Seki, K.; Kashima, M.; Akiyama, T.; Kobayashi, A.; Dezawa, K.; Takeuchi, Y.; Furuchi, M.; Kamimoto, A. Reproducibility of Mandibular Cortical Index Classification Among Dental Examiners and a Single Generative AI Platform: An Observer Agreement Study. Clin. Pract. 2026, 16, 142. https://doi.org/10.3390/clinpract16080142
Seki K, Kashima M, Akiyama T, Kobayashi A, Dezawa K, Takeuchi Y, Furuchi M, Kamimoto A. Reproducibility of Mandibular Cortical Index Classification Among Dental Examiners and a Single Generative AI Platform: An Observer Agreement Study. Clinics and Practice. 2026; 16(8):142. https://doi.org/10.3390/clinpract16080142
Chicago/Turabian StyleSeki, Keisuke, Minori Kashima, Taiki Akiyama, Atsushi Kobayashi, Ko Dezawa, Yoshimasa Takeuchi, Mika Furuchi, and Atsushi Kamimoto. 2026. "Reproducibility of Mandibular Cortical Index Classification Among Dental Examiners and a Single Generative AI Platform: An Observer Agreement Study" Clinics and Practice 16, no. 8: 142. https://doi.org/10.3390/clinpract16080142
APA StyleSeki, K., Kashima, M., Akiyama, T., Kobayashi, A., Dezawa, K., Takeuchi, Y., Furuchi, M., & Kamimoto, A. (2026). Reproducibility of Mandibular Cortical Index Classification Among Dental Examiners and a Single Generative AI Platform: An Observer Agreement Study. Clinics and Practice, 16(8), 142. https://doi.org/10.3390/clinpract16080142

