YOLO-TARC: YOLOv10 with Token Attention and Residual Convolution for Small Void Detection in Root Canal X-Ray Images
Abstract
1. Introduction
- (1)
- A YOLOv10-based network is proposed to integrate Token Attention and Residual Convolution (YOLO-TARC) for small void detection in root canal X-ray images, aiming to address the issues of gradual attenuation and inability to focus on the contours of small targets during feature transmission.
- (2)
- A Residual Convolution (ResConv) is designed using residual connections to combine standard convolution and depthwise convolution. This ensures the transmission of discriminative features and effectively preserves high-frequency details at the pixel level.
- (3)
- A novel Token Attention (TokAtt) is proposed. The input features are divided into small local region tokens. An attention mechanism is then used to dynamically adjust the weights of these tokens to focus on key information, significantly improving the ability to attend to small targets.
- (4)
- The proposed YOLOv10-TARC is validated using a private root canal void dataset from a previous study, demonstrating superior overall performance compared to existing state-of-the-art detection methods.
2. Related Work
3. Method
3.1. The ResConv Module
3.2. The TokAtt Module
| Algorithm 1 Pseudocode of TokAtt mechanism. |
|
3.3. The Bounding Box Loss Function
4. Experiments
4.1. Implementation Details
4.2. Comparison Experiments
4.3. Ablation Experiments
5. Conclusions and Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Akhtar, S.; Silk, H.; Savageau, J.A.; Stevens, G.A. Oral health articles in primary care journals: A bibliometric review. J. Public Health Dent. 2025, 85, 84–91. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rudol, J.F.; Niemczyk, W.; Janik, K.; Zawilska, A.; Kępa, M.; Tanasiewicz, M. How to Deal with Pulpitis: An Overview of New Approaches. Dent. J. 2025, 13, 25. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Çelik, B.; Genç, M.Z.; Çelik, M.E. Evaluation of root canal filling length on periapical radiograph using artificial intelligence. Oral Radiol. 2024, 41, 102–110. [Google Scholar] [CrossRef] [Scilit]
- Libonati, A.; Gallusi, G.; Montemurro, E.; Di Taranto, V. Reduction of radiations exposure in endodontics: Comparative analysis of direct (GX S-700, Gendex) and semidirect (VistaScan Mini View, Dürr) digital systems. J. Biol. Regul. Homeost. Agents 2021, 35, 87–94. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Chen, X.; Ma, N.; Xu, T.; Xu, C. Deep learning-based tooth segmentation methods in medical imaging: A review. Proc. Inst. Mech. Eng. Part H J. Eng. Med. 2024, 238, 115–131. [Google Scholar] [CrossRef] [Scilit]
- Duncan, H.F.; Kirkevang, L.L.; Peters, O.A.; El-Karim, I.; Krastl, G.; Del Fabbro, M.; Chong, B.S.; Galler, K.M.; Segura-Egea, J.J.; Kebschull, M. Treatment of pulpal and apical disease: The European Society of Endodontology (ESE) S3-level clinical practice guideline. Int. Endod. J. 2023, 56, 238–295. [Google Scholar] [CrossRef] [Scilit]
- Fletcher, J.G.; Leng, S.; Yu, L.; McCollough, C.H. Dealing with uncertainty in CT images. Radiology 2016, 279, 5–10. [Google Scholar] [CrossRef] [Scilit]
- Kurz, A.; Hauser, K.; Mehrtens, H.A.; Krieghoff-Henning, E.; Hekler, A.; Kather, J.N.; Fröhling, S.; von Kalle, C.; Brinker, T.J. Uncertainty estimation in medical image classification: Systematic review. JMIR Med. Inform. 2022, 10, e36427. [Google Scholar] [CrossRef] [Scilit]
- Huang, L.; Ruan, S.; Xing, Y.; Feng, M. A review of uncertainty quantification in medical image analysis: Probabilistic and non-probabilistic methods. Med. Image Anal. 2024, 97, 103223. [Google Scholar] [CrossRef] [Scilit]
- Wang, A.; Chen, H.; Liu, L.; Chen, K.; Lin, Z.; Han, J.; Ding, G. YOLOv10: Real-Time End-to-End Object Detection. arXiv 2024, arXiv:2405.14458. [Google Scholar]
- Yeerjiang, A.; Wang, Z.; Huang, X.; Zhang, J.; Chen, Q.; Qin, Y.; He, J. YOLOv1 to YOLOv10: A Comprehensive Review of YOLO Variants and Their Application in Medical Image Detection. J. Artif. Intell. Pract. 2024, 7, 112–122. [Google Scholar]
- Ali, M.L.; Zhang, Z. The YOLO Framework: A Comprehensive Review of Evolution, Applications, and Benchmarks in Object Detection. Computers 2024, 13, 336. [Google Scholar] [CrossRef] [Scilit]
- Chai, R.; Tian, N.; Wan, G.; Liu, S.; Zhan, J.; Li, X.; Bian, H.; Gao, C.; Xia, X.; Wang, D.; et al. Automated detection of early-stage osteonecrosis of the femoral head in adult using YOLOv10: Multi-institutional validation. Eur. J. Radiol. 2025, 184, 111983. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Dongmei, Z.; Dongbo, W. Transformers and their application to medical image processing: A review. J. Radiat. Res. Appl. Sci. 2023, 16, 100680. [Google Scholar]
- Muzammul, M.; Li, X. Comprehensive review of deep learning-based tiny object detection: Challenges, strategies, and future directions. Knowl. Inf. Syst. 2025, 67, 3825–3913. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Zuo, G. Small target detection in UAV view based on improved YOLOv8 algorithm. Sci. Rep. 2025, 15, 421. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Zhang, Y.; Liu, H.; Guo, J.; Liu, L.; Gu, J.; Deng, L.; Li, S. A novel small object detection algorithm for UAVs based on YOLOv5. Phys. Scr. 2024, 99, 036001. [Google Scholar] [CrossRef] [Scilit]
- He, X.; Zheng, X.; Hao, X.; Jin, H.; Zhou, X.; Shao, L. Improving small object detection via context-aware and feature-enhanced plug-and-play modules. J. Real-Time Image Process. 2024, 21, 44. [Google Scholar] [CrossRef] [Scilit]
- Tong, K.; Wu, Y. Small object detection using deep feature learning and feature fusion network. Eng. Appl. Artif. Intell. 2024, 132, 107931. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.; Wen, R.; Ma, L. Small object detection model for UAV aerial image based on YOLOv7. Signal Image Video Process. 2023, 18, 2695–2707. [Google Scholar] [CrossRef] [Scilit]
- Wen, Z.; Su, J.; Zhang, Y.; Li, M.; Gan, G.; Zhang, S.; Fan, D. A lightweight small object detection algorithm based on improved YOLOv5 for driving scenarios. Int. J. Multimed. Inf. Retr. 2023, 12, 38. [Google Scholar] [CrossRef] [Scilit]
- Guo, A.; Sun, K.; Zhang, Z. A lightweight YOLOv8 integrating FasterNet for real-time underwater object detection. J. Real-Time Image Process. 2024, 21, 49. [Google Scholar] [CrossRef] [Scilit]
- Luo, Z.; Tian, Y. Infrared Road Object Detection Based on Improved YOLOv8. IAENG Int. J. Comput. Sci. 2024, 51, 252–259. [Google Scholar]
- Sun, D.; Zhang, K.; Zhong, H.; Xie, J.; Xue, X.; Yan, M.; Wu, W.; Li, J. Efficient Tobacco Pest Detection in Complex Environments Using an Enhanced YOLOv8 Model. Agriculture 2024, 14, 353. [Google Scholar] [CrossRef] [Scilit]
- Rehman, K.U.; Li, J.; Yasin, A.; Bilal, A.; Basheer, S.; Ullah, I.; Jabbar, M.K.; Tian, Y. A feature fusion attention-based deep learning algorithm for mammographic architectural distortion classification. IEEE J. Biomed. Health Inform. 2025, 1–12. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Xie, J.; Zeng, Z.; Ma, Y.; Pan, Y.; Wu, X.; Han, X.; Tian, Y. Defect recognition in sonic infrared imaging by deep learning of spatiotemporal signals. Eng. Appl. Artif. Intell. 2024, 133, 108174. [Google Scholar] [CrossRef] [Scilit]
- Zhorif, N.N.; Anandyto, R.K.; Rusyadi, A.U.; Irwansyah, E. Implementation of Slicing Aided Hyper Inference (SAHI) in YOLOv8 to Counting Oil Palm Trees Using High-Resolution Aerial Imagery Data. Int. J. Adv. Comput. Sci. Appl. (IJACSA) 2024, 15, 869–874. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Tian, Y.; Zhang, Z.; Zhang, X.; Du, B.; Zeng, Z. Teeth segmentation from bite-wing X-ray images by integrating nested dual UNet with Swin Transformers. In Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics (SMC), Sarawak, Malaysia, 6–10 October 2024; pp. 4548–4553. [Google Scholar]
- Cao, Z.; Zeng, Z.; Xie, J.; Zhai, H.; Yin, Y.; Ma, Y.; Tian, Y. Diabetic plantar foot segmentation in active thermography using a two-stage adaptive gamma transform and a deep neural network. Sensors 2023, 23, 8511. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep Residual Learning for Image Recognition. arXiv 2015, arXiv:1512.03385. [Google Scholar]
- Guo, Y.; Li, Y.; Wang, L.; Rosing, T. Depthwise Convolution Is All You Need for Learning Multiple Visual Domains. Proc. AAAI Conf. Artif. Intell. 2019, 33, 8368–8375. [Google Scholar] [CrossRef] [Scilit]
- Vaswani, A.; Shazeer, N.M.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention is All you Need. In Proceedings of the Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Niu, Z.; Zhong, G.; Yu, H. A review on the attention mechanism of deep learning. Neurocomputing 2021, 452, 48–62. [Google Scholar] [CrossRef] [Scilit]
- Ma, S.; Xu, Y. MPDIoU: A Loss for Efficient and Accurate Bounding Box Regression. arXiv 2023, arXiv:2307.07662. [Google Scholar]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. SSD: Single Shot MultiBox Detector. In Computer Vision—ECCV 2016; Springer International Publishing: Berlin/Heidelberg, Germany, 2016; pp. 21–37. [Google Scholar]
- Jocher, G. ultralytics/YOLOv5: V3.1—Bug Fixes and Performance Improvements. 2020. Available online: https://github.com/ultralytics/yolov5 (accessed on 10 January 2023).
- Wang, C.Y.; Bochkovskiy, A.; Liao, H.Y.M. YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors. arXiv 2022, arXiv:2207.02696. [Google Scholar]
- Jocher, G.; Chaurasia, A.; Qiu, J. Ultralytics YOLO8. 2023. Available online: https://github.com/ultralytics/ultralytics/tree/main/ultralytics/cfg/models/v8 (accessed on 10 January 2023).
- Wang, C.Y.; Yeh, I.H.; Liao, H. YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information. arXiv 2024, arXiv:2402.13616. [Google Scholar]
- Jocher, G.; Qiu, J. Ultralytics YOLO11. 2024. Available online: https://github.com/ultralytics/ultralytics/tree/main/ultralytics/cfg/models/v11 (accessed on 30 September 2024).






| Method | Recall (%) | Precision (%) | mAP50 (%) | mAP50-95 (%) |
|---|---|---|---|---|
| Faster R-CNN [35] | 59.8 | 62.8 | 61.7 | 24.1 |
| SSD [36] | 58.8 | 61.6 | 56.1 | 18.2 |
| YOLOv5 [37] | 61.3 | 70.7 | 65.9 | 30.7 |
| YOLOv7-tiny [38] | 62.5 | 63.6 | 64.5 | 25.6 |
| YOLOv8m [39] | 67.6 | 74.2 | 69.9 | 31.6 |
| YOLOv9m [40] | 69.6 | 78.1 | 71.8 | 34.0 |
| YOLOv10m (baseline) [10] | 73.8 | 77.4 | 73.3 | 34.3 |
| YOLOv11 [41] | 64.3 | 78.0 | 72.0 | 31.7 |
| YOLO-TARC (ours) | 80.0 | 79.5 | 80.8 | 37.1 |
| YOLOv10m | ResConv | TokAtt | Loss | Rec (%) | Pre (%) | mAP50 (%) | mAP50-95 (%) |
|---|---|---|---|---|---|---|---|
| ✔ | 73.8 | 77.4 | 73.3 | 34.3 | |||
| ✔ | ✔ | 73.5 | 82.4 | 75.3 | 36.5 | ||
| ✔ | ✔ | 77.5 | 81.5 | 78.0 | 37.1 | ||
| ✔ | ✔ | 72.9 | 78.3 | 74.5 | 34.5 | ||
| ✔ | ✔ | ✔ | 79.6 | 80.4 | 79.4 | 37.3 | |
| ✔ | ✔ | ✔ | 77.9 | 81.7 | 78.7 | 36.9 | |
| ✔ | ✔ | ✔ | 74.7 | 81.5 | 76.0 | 36.2 | |
| ✔ | ✔ | ✔ | ✔ | 80.0 | 79.5 | 80.8 | 37.1 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Pan, Y.; Zhang, Z.; Zhang, X.; Zeng, Z.; Tian, Y. YOLO-TARC: YOLOv10 with Token Attention and Residual Convolution for Small Void Detection in Root Canal X-Ray Images. Sensors 2025, 25, 3036. https://doi.org/10.3390/s25103036
Pan Y, Zhang Z, Zhang X, Zeng Z, Tian Y. YOLO-TARC: YOLOv10 with Token Attention and Residual Convolution for Small Void Detection in Root Canal X-Ray Images. Sensors. 2025; 25(10):3036. https://doi.org/10.3390/s25103036
Chicago/Turabian StylePan, Yin, Zhenpeng Zhang, Xueyang Zhang, Zhi Zeng, and Yibin Tian. 2025. "YOLO-TARC: YOLOv10 with Token Attention and Residual Convolution for Small Void Detection in Root Canal X-Ray Images" Sensors 25, no. 10: 3036. https://doi.org/10.3390/s25103036
APA StylePan, Y., Zhang, Z., Zhang, X., Zeng, Z., & Tian, Y. (2025). YOLO-TARC: YOLOv10 with Token Attention and Residual Convolution for Small Void Detection in Root Canal X-Ray Images. Sensors, 25(10), 3036. https://doi.org/10.3390/s25103036

