Comparative Evaluation of Deep Learning Models for the Classification of Impacted Maxillary Canines on Panoramic Radiographs
Abstract
1. Introduction
- Comparative Analysis: We comprehensively evaluate four distinct CNN architectures (ResNet50, Xception, InceptionV3, and VGG16) for classifying impacted maxillary canines to identify the optimal feature extractor.
- High Performance Verification: We show that VGG16 achieves the highest diagnostic accuracy (99.28%) and a perfect recall rate, setting a new benchmark for this dental anomaly.
- Proof-of-Concept Visualization: Development of a diagnostic interface to visualize model predictions as a feasibility study, with clinical usability testing reserved for future research.
- Dataset Reliability: We use a clinically validated, mildly imbalanced dataset to ensure robust, unbiased model performance.
2. Materials and Methods
2.1. Study Design and Ethical Considerations
2.2. Dataset Acquisition
2.3. Image Preprocessing
2.4. Model Selection and Architecture
2.5. Proposed Methodological Framework and Implementation Details
- A Global Average Pooling 2D layer to reduce the spatial dimensions of the feature maps.
- A Dense (Fully Connected) layer with 1024 units and ReLU activation to interpret the extracted features.
- A Dropout layer with a rate of 0.5 to enforce regularization.
- A final Dense layer with a single unit and Sigmoid activation to output the binary classification probability ().
2.6. Proposed Methodological Framework and Transfer Learning Strategy
- A Global Average Pooling 2D layer to reduce spatial dimensions and minimize overfitting.
- A Dense layer with 512 units and ReLU activation to interpret high-level features.
- A Dropout layer (rate = 0.5) to further prevent overfitting during training.
- A Dense layer with a Sigmoid activation function to output a probability score indicating the presence of an impacted canine.
- ResNet50 (Residual Network) introduces the concept of residual learning, which addresses the vanishing gradient problem commonly encountered in deep neural networks by incorporating identity shortcut connections. This design allows for the effective training of substantially deeper networks without degradation in performance [16]. ResNet50, consisting of 50 layers, has been widely adopted in medical imaging due to its capacity to capture complex hierarchical features.
- Xception (Extreme Inception) is an architecture based on depthwise separable convolutions, which factorize conventional convolutions into spatial and channel-wise operations. This results in a more efficient model with fewer parameters while maintaining high representational power [18]. Xception has demonstrated strong performance in tasks requiring fine-grained feature discrimination, making it suitable for nuanced medical image classification.
- InceptionV3 employs a sophisticated design involving parallel convolutional filters of varying sizes within its inception modules, enabling multi-scale feature extraction [19]. This architecture balances computational efficiency with depth, making it effective for complex image analysis tasks, including lesion detection and organ segmentation in clinical contexts.
- VGG16 is characterized by its simplicity and uniform architecture, utilizing sequential stacks of 3 × 3 convolutional layers followed by max-pooling layers [20,21]. Despite its relatively straightforward design, VGG16 has shown remarkable performance across various domains due to its deep structure (16 weight layers) and ability to capture detailed spatial hierarchies. Its robustness and ease of implementation make it a popular choice in medical imaging applications, particularly when working with smaller datasets.
Proposed Fine-Tuning and Custom Head Architecture
2.7. Training and Validation Strategy
2.8. Evaluation Metrics
3. Results
3.1. Performance Metrics and Statistical Analysis
3.2. Confusion Matrix and Error Analysis
- False Negatives (Missed Diagnosis): VGG16 produced zero false negatives (FN = 0), meaning it did not miss a single case of an impacted canine. In contrast, Xception and InceptionV3 missed 2 cases each, and ResNet50 missed 1 case. In a clinical screening context, false negatives are the most critical error type, as they lead to untreated pathologies.
- False Positives (False Alarm): ResNet50 and Xception exhibited a higher tendency to classify non-impacted teeth as impacted, generating 5 false positives each. VGG16, conversely, generated only a single false positive (FP = 1), demonstrating superior specificity.
3.3. Training Dynamics and Convergence
- VGG16 (Figure 5): Exhibited the most stable learning trajectory. Both training and validation accuracy increased consistently, while validation loss decreased steadily, indicating successful convergence without overfitting.
- Xception: Showed signs of early overfitting, as illustrated in Figure 3. The validation loss began to trend upward early in the process, which triggered the pre-configured Early Stopping mechanism at epoch 8. This is not an error, but rather an intended outcome of the training protocol, designed to prevent overfitting and retain the model at its point of peak generalization. The variation in training epochs across models is therefore a direct consequence of this automated regularization technique, with the Xception model converging on its optimal weights more rapidly than the other models before performance on the validation set began to degrade.
4. Discussion
4.1. Comparison with Prior Literature
4.2. Architectural Performance and Training Dynamics
4.3. Limitations and Sources of Bias
4.4. Clinical Integration and Future Directions
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| CBCT | Cone-Beam Computed Tomography |
| CNN | Convolutional Neural Network |
| DL | Deep Learning |
| FN | False Negatives |
| FP | False Positives |
| FPR | False Positive Rate |
| MRI | Magnetic Resonance Imaging |
| TN | True Negatives |
| TP | True Positives |
Appendix A

References
- Becker, A. The Orthodontic Treatment of Impacted Teeth, 2nd ed.; Informa Healthcare: London, UK, 2007. [Google Scholar]
- Bishara, S.E. Impacted maxillary canines: A review. Am. J. Orthod. Dentofac. Orthop. 1992, 101, 159–171. [Google Scholar] [CrossRef] [Scilit]
- Alqerban, A.; Willems, G.; Bernaerts, C.; Vangastel, J.; Politis, C.; Jacobs, R. Orthodontic treatment planning for impacted maxillary canines using conventional records versus 3D CBCT. Eur. J. Orthod. 2014, 36, 698–707. [Google Scholar] [CrossRef] [Scilit]
- Küçük, D.B.; Imak, A.; Özçelik, S.T.A.; Çelebi, A.; Türkoğlu, M.; Sengur, A.; Koundal, D. Hybrid CNN-Transformer Model for Accurate Impacted Tooth Detection in Panoramic Radiographs. Diagnostics 2025, 15, 244. [Google Scholar] [CrossRef] [Scilit]
- Litjens, G.; Kooi, T.; Bejnordi, B.E.; Setio, A.A.A.; Ciompi, F.; Ghafoorian, M.; van der Laak, J.A.W.M.; van Ginneken, B.; Sánchez, C.I. A survey on deep learning in medical image analysis. Med. Image Anal. 2017, 42, 60–88. [Google Scholar] [CrossRef] [Scilit]
- Esteva, A.; Kuprel, B.; Novoa, R.A.; Ko, J.; Swetter, S.M.; Blau, H.M.; Thrun, S. Dermatologist-level classification of skin cancer with deep neural networks. Nature 2017, 542, 115–118. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Shen, D.; Wu, G.; Suk, H.I. Deep Learning in Medical Image Analysis. Annu. Rev. Biomed. Eng. 2017, 19, 221–248. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Schwendicke, F.; Samek, W.; Krois, J. Artificial Intelligence in Dentistry: Chances and Challenges. J. Dent. Res. 2020, 99, 769–774. [Google Scholar] [CrossRef] [Scilit]
- Mathew, M. Artificial intelligence for oral health care: Applications and future prospects. Br. Dent. J. 2025, 238, 917. [Google Scholar] [CrossRef] [Scilit]
- Singh, V.; Mathur, P.; Thakker, B.; Rao, Z.; Gowdar, I.M.; Butolia, H.K.; Chouhan, D.S. Assessment of AI-Driven Software Accuracy in Diagnosing Oral Lesions Using Radiographic Imaging. J. Pharm. Bioallied Sci. 2025, 17, S1553–S1555. [Google Scholar] [CrossRef] [Scilit]
- Lee, J.H.; Kim, D.H.; Jeong, S.N.; Choi, S.H. Detection and diagnosis of dental caries using a deep learning-based convolutional neural network algorithm. J. Dent. 2018, 77, 106–111. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Abdulkreem, A.; Bhattacharjee, T.; Alzaabi, H.; Alali, K.; Gonzalez, A.; Chaudhry, J.; Prasad, S. Artificial intelligence-based automated preprocessing and classification of impacted maxillary canines in panoramic radiographs. Dentomaxillofac. Radiol. 2024, 53, 173–177. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Aljabri, M.; Aljameel, S.S.; Min-Allah, N.; Alhuthayfi, J.; Alghamdi, L.; Alduhailan, N.; Alfehaid, R.; Alqarawi, R.; Alhareky, M.; Shahin, S.Y.; et al. Canine impaction classification from panoramic dental radiographic images using deep learning models. Inform. Med. Unlocked 2022, 30, 100918. [Google Scholar] [CrossRef] [Scilit]
- Minhas, S.; Wu, T.-H.; Kim, D.-G.; Chen, S.; Wu, Y.-C.; Ko, C.-C. Artificial intelligence for 3D reconstruction from 2D panoramic X-rays to assess maxillary impacted canines. Diagnostics 2024, 14, 196. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- World Medical Association. World Medical Association Declaration of Helsinki: Ethical principles for medical research involving human subjects. JAMA 2013, 310, 2191–2194. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Deng, J.; Dong, W.; Socher, R.; Li, L.J.; Li, K.; Li, F.F. ImageNet: A Large-Scale Hierarchical Image Database. In Proceedings of the 2009 IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 248–255. [Google Scholar]
- Chollet, F. Xception: Deep Learning with Depthwise Separable Convolutions. In Proceedings of the 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA, 21–26 July 2017; pp. 1800–1807. [Google Scholar]
- Szegedy, C.; Vanhoucke, V.; Ioffe, S.; Shlens, J.; Wojna, Z. Rethinking the Inception Architecture for Computer Vision. In Proceedings of the 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA, 27–30 June 2016; pp. 2818–2826. [Google Scholar]
- Simonyan, K.; Zisserman, A. Very Deep Convolutional Networks for Large-Scale Image Recognition. arXiv 2014, arXiv:1409.1556. [Google Scholar]
- Kingma, D.P.; Ba, J. Adam: A Method for Stochastic Optimization. arXiv 2014, arXiv:1412.6980. [Google Scholar]
- Sokolova, M.; Lapalme, G. A systematic analysis of performance measures for classification tasks. Inf. Process. Manag. 2009, 45, 427–437. [Google Scholar] [CrossRef] [Scilit]
- Flügge, T.; Vinayahalingam, S.; van Nistelrooij, N.; Kellner, S.; Xi, T.; van Ginneken, B.; Bergé, S.; Heiland, M.; Kernen, F.; Ludwig, U.; et al. Automated tooth segmentation in magnetic resonance scans using deep learning—A pilot study. Dento Maxillo Facial Radiol. 2025, 54, 12–18. [Google Scholar] [CrossRef] [Scilit]
- Alenezi, O.; Bhattacharjee, T.; Alseed, H.A.; Tosun, Y.I.; Chaudhry, J.; Prasad, S. Evaluating the Efficacy of Various Deep Learning Architectures for Automated Preprocessing and Identification of Impacted Maxillary Canines in Panoramic Radiographs. Int. Dent. J. 2025, 75, 100940. [Google Scholar] [CrossRef] [Scilit]
- Veerabhadrappa, S.K.; Vengusamy, S.; Padarha, S.; Iyer, K.; Yadav, S. Fully automated deep learning framework for detection and classification of impacted mandibular third molars in panoramic radiographs. J. Oral Med. Oral Surg. 2025, 31, 7. [Google Scholar] [CrossRef] [Scilit]
- Abid, A.; Abdalla, A.; Abid, A.; Khan, D.; Alfozan, A.; Zou, J. Gradio: Hassle-Free Sharing and Testing of ML Models in the Wild. arXiv 2019, arXiv:1906.02569. [Google Scholar] [CrossRef] [Scilit]





| Sex | n (%) | Min Age | Max Age | Mean ± SD |
|---|---|---|---|---|
| Female | 384 (55.3%) | 18.33 | 65.58 | 24.48 ± 10.35 |
| Male | 310 (44.7%) | 18.67 | 64.58 | 26.08 ± 11.27 |
| Model | Accuracy (%) | 95% CI | Precision (%) | Recall (%) | F1-Score (%) |
|---|---|---|---|---|---|
| ResNet50 | 94.24 | 90.4–98.1 | 92.55 | 98.86 | 95.60 |
| Xception | 93.53 | 89.5–97.5 | 92.47 | 97.73 | 95.03 |
| InceptionV3 | 95.68 | 92.3–99.0 | 95.56 | 97.73 | 96.63 |
| VGG16 | 99.28 | 94.9–99.9 | 98.88 | 100.00 | 99.43 |
| Model | TN | FP | FN | TP | Total Errors |
|---|---|---|---|---|---|
| ResNet50 | 33 | 5 | 1 | 65 | 6 |
| Xception | 33 | 5 | 2 | 64 | 7 |
| InceptionV3 | 35 | 3 | 2 | 64 | 5 |
| VGG16 | 37 | 1 | 0 | 66 | 1 |
| Study | Target Task | Dataset Size | Core Architecture |
|---|---|---|---|
| Alenezi et al. (2025) [24] | Classification | 182 | GoogLeNet |
| Küçük et al. (2025) [4] | Object Detection | 407 | YOLOv8 + RT-DETR |
| Veerabhadrappa (2025) [25] | Detection (3rd Molar) | 1100 | VGG16 + YOLOv7 |
| Abdulkreem et al. (2024) [12] | Classification | 182 | SqueezeNet |
| Aljabri et al. (2022) [13] | Classification | 416 | Inception V3 |
| Current Study | Classification | 694 | VGG16 |
| Study | Best Model | Accuracy | Recall | F1-Score |
|---|---|---|---|---|
| Küçük et al. (2025) [4] | Hybrid YOLO | – | 99.20% | 96.00% |
| Alenezi et al. (2025) [24] | GoogLeNet | 94.00% | 91.00% | – |
| Alenezi et al. (2025) [24] * | VGG16 | 83.00% | 78.00% | – |
| Veerabhadrappa (2025) [25] | VGG16 | 93.51% | 89.47% | 91.97% |
| Aljabri et al. (2022) [13] | Inception V3 | 92.59% | – | – |
| Current Study | VGG16 | 99.28% | 100.00% | 99.43% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Tokatlı, N.; Erdem, B.; Özcan, M.; Turan Maviş, B.; Şar, Ç.; Özdemir, F. Comparative Evaluation of Deep Learning Models for the Classification of Impacted Maxillary Canines on Panoramic Radiographs. Diagnostics 2026, 16, 219. https://doi.org/10.3390/diagnostics16020219
Tokatlı N, Erdem B, Özcan M, Turan Maviş B, Şar Ç, Özdemir F. Comparative Evaluation of Deep Learning Models for the Classification of Impacted Maxillary Canines on Panoramic Radiographs. Diagnostics. 2026; 16(2):219. https://doi.org/10.3390/diagnostics16020219
Chicago/Turabian StyleTokatlı, Nazlı, Buket Erdem, Mustafa Özcan, Begüm Turan Maviş, Çağla Şar, and Fulya Özdemir. 2026. "Comparative Evaluation of Deep Learning Models for the Classification of Impacted Maxillary Canines on Panoramic Radiographs" Diagnostics 16, no. 2: 219. https://doi.org/10.3390/diagnostics16020219
APA StyleTokatlı, N., Erdem, B., Özcan, M., Turan Maviş, B., Şar, Ç., & Özdemir, F. (2026). Comparative Evaluation of Deep Learning Models for the Classification of Impacted Maxillary Canines on Panoramic Radiographs. Diagnostics, 16(2), 219. https://doi.org/10.3390/diagnostics16020219

