RAC-RTDETR: A Lightweight, Efficient Real-Time Small-Object Detection Algorithm for Steel Surface Defect Detection
Abstract
1. Introduction
- The ARNet backbone network is a lightweight architecture designed to reduce the impact of noise on model detection performance, significantly enhancing its ability to identify small targets. Compared to the baseline model, ARNet improves detection accuracy while reducing the model’s parameter count by 48.24%.
- A novel AIFI-ASMD module is developed to enhance the model’s ability to interact with features across different scales, improving spatial awareness and long-range dependency modeling.This results in improved detection performance for multi-scale objects. Experimental results on the NEU-DET dataset [9] show that incorporating this module boosts the model’s MAP by 1.58%.
- The Converse2D upsampling module is introduced to replace traditional upsampling methods. By reconstructing the input feature map, it preserves more detailed feature information, enhancing the model’s ability to detect small objects in low-contrast and feature-sparse scenarios. Experimental results show that the incorporation of this module improves the model’s MAP by 0.41%.
- Extensive experiments were conducted on the NEU-DET and GC10-DET datasets [10] using RAC-RTDETR.The experimental results show that, on the NEU-DET dataset, compared with the baseline model, RAC-RTDETR achieved improvements of 3.56% in MAP and 7.96% in FPS, while reducing the number of model parameters and GFLOPs by 36.18% and 40.70%, respectively.On the GC10-DET dataset, the MAP of RAC-RTDETR increased by 3.47%.
2. Related Works
3. Methods
3.1. RTDETR Model
3.2. RAC-RTDETR Model
3.3. ARNet Backbone Network
3.4. AIFI-ASMD Module
3.4.1. The Adaptive Sparse Self-Attention (ASSA) Submodule
3.4.2. The Spatially Enhanced Feedforward Network (SEFN) Submodule
3.4.3. The Multi-Cognitive Visual Adapter (Mona) Submodule
3.4.4. The Dynamic Tanh (DyT) Function
3.5. Converse2D Module
4. Experiment Setup
4.1. Datasets
4.2. Experimental Environment
4.3. Evaluation Metrics
5. Experimental Results and Analysis
5.1. Effectiveness Experiment of the ARnet Backbone Network
5.2. Effectiveness Experiment of the AIFI-ASMD Module
5.3. Effectiveness Experiment of the Converse2D Module
5.4. Comparison Experiment
5.5. Ablation Experiment
5.6. Visual Qualitative Analysis Experiment
5.7. Generalization Experiment
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Xu, H.; Zhang, Z.; Ye, H.; Song, J.; Chen, Y. Efficient Steel Surface Defect Detection via a Lightweight YOLO Framework with Task-Specific Knowledge-Guided Optimization. Electronics 2025, 14, 2029. [Google Scholar] [CrossRef]
- Chen, B.; Zha, J.; Cai, Z.; Wu, M. Predictive modelling of surface roughness in precision grinding based on hybrid algorithm. CIRP J. Manuf. Sci. Technol. 2025, 59, 1–17. [Google Scholar] [CrossRef]
- Ge, J.; Yao, Z.; Wu, M.; Almeida, J.H.S., Jr.; Jin, Y.; Sun, D. Tackling data scarcity in machine learning-based CFRP drilling performance prediction through a Broad Learning System with Virtual Sample Generation (BLS-VSG). Compos. Part B Eng. 2025, 305, 112701. [Google Scholar] [CrossRef]
- Wu, M.; Yao, Z.; Ye, L.; Verbeke, M.; Karsmakers, P.; Reynaerts, D. Geometrical Feature Classification in Electrical Discharge Machining Using In-Process Monitoring and Machine Learning. Procedia CIRP 2025, 137, 462–467. [Google Scholar] [CrossRef]
- Redmon, J.; Divvala, S.; Girshick, R.; Farhadi, A. You only look once: Unified, real-time object detection. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Denver, CO, USA, 3–7 June 2016; pp. 779–788. [Google Scholar]
- Liu, W.; Anguelov, D.; Erhan, D.; Szegedy, C.; Reed, S.; Fu, C.Y.; Berg, A.C. Ssd: Single shot multibox detector. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 11–14 October 2016; Springer: Berlin/Heidelberg, Germany, 2016; pp. 21–37. [Google Scholar]
- Girshick, R. Fast r-cnn. In Proceedings of the IEEE International Conference on Computer Vision, Santiago, Chile, 7–13 December 2015; pp. 1440–1448. [Google Scholar]
- Zhao, Y.; Lv, W.; Xu, S.; Wei, J.; Wang, G.; Dang, Q.; Liu, Y.; Chen, J. Detrs beat yolos on real-time object detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 16965–16974. [Google Scholar]
- Bao, Y.; Song, K.; Liu, J.; Wang, Y.; Yan, Y.; Yu, H.; Li, X. Triplet-graph reasoning network for few-shot metal generic surface defect segmentation. IEEE Trans. Instrum. Meas. 2021, 70, 5011111. [Google Scholar] [CrossRef]
- Lv, X.; Duan, F.; Jiang, J.j.; Fu, X.; Gan, L. Deep metallic surface defect detection: The new benchmark and detection network. Sensors 2020, 20, 1562. [Google Scholar] [CrossRef] [PubMed]
- Krummenacher, G.; Ong, C.S.; Koller, S.; Kobayashi, S.; Buhmann, J.M. Wheel defect detection with machine learning. IEEE Trans. Intell. Transp. Syst. 2017, 19, 1176–1187. [Google Scholar] [CrossRef]
- Liu, X.; Gao, J. Surface defect detection method of hot rolling strip based on improved SSD model. In Proceedings of the International Conference on Database Systems for Advanced Applications, Taipei, Taiwan, 11–14 April 2021; Springer: Berlin/Heidelberg, Germany, 2021; pp. 209–222. [Google Scholar]
- Liang, C.; Wang, Z.Z.; Liu, X.L.; Zhang, P.; Tian, Z.W.; Qian, R.L. SDD-Net: A Steel Surface Defect Detection Method Based on Contextual Enhancement and Multiscale Feature Fusion. IEEE Access 2024, 12, 185740–185756. [Google Scholar] [CrossRef]
- Lu, S.; Liang, Y.; Ren, Z.; Yu, X.; Wang, X. FEP-YOLO: A lightweight steel surface defect detection method for resource-constrained devices. Meas. Sci. Technol. 2025, 36, 076016. [Google Scholar] [CrossRef]
- Sun, W.; Meng, N.; Chen, L.; Yang, S.; Li, Y.; Tian, S. CTL-YOLO: A Surface Defect Detection Algorithm for Lightweight Hot-Rolled Strip Steel Under Complex Backgrounds. Machines 2025, 13, 301. [Google Scholar] [CrossRef]
- Liao, L.; Song, C.; Wu, S.; Fu, J. A novel YOLOv10-based algorithm for accurate steel surface defect detection. Sensors 2025, 25, 769. [Google Scholar] [CrossRef] [PubMed]
- Leng, Y.; Liu, J. Improved faster R-CNN for steel surface defect detection in industrial quality control. Sci. Rep. 2025, 15, 30093. [Google Scholar] [CrossRef] [PubMed]
- Gao, B.; Zhao, H.; Miao, X. A novel multi-model cascade framework for pipeline defects detection based on machine vision. Measurement 2023, 220, 113374. [Google Scholar] [CrossRef]
- Mao, H.; Gong, Y. Steel surface defect detection based on the lightweight improved RT-DETR algorithm. J. Real-Time Image Process. 2025, 22, 28. [Google Scholar] [CrossRef]
- Su, F.; Meng, P. EL-DETR: A lightweight steel surface defect detection model. Meas. Sci. Technol. 2025, 36, 116003. [Google Scholar] [CrossRef]
- Zhou, S.; Cai, Y.; Zhang, Z.; Yin, J. MESC-DETR: An Improved RT-DETR Algorithm for Steel Surface Defect Detection. Electronics 2025, 14, 2232. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar]
- Wang, C.Y.; Yeh, I.H.; Mark Liao, H.Y. Yolov9: Learning what you want to learn using programmable gradient information. In Proceedings of the European Conference on Computer Vision, Milan, Italy, 29 September–4 October 2024; Springer: Berlin/Heidelberg, Germany, 2024; pp. 1–21. [Google Scholar]
- Cai, X.; Lai, Q.; Wang, Y.; Wang, W.; Sun, Z.; Yao, Y. Poly kernel inception network for remote sensing detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 16–22 June 2024; pp. 27706–27716. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Las Vegas, NV, USA, 27–30 June 2016; pp. 770–778. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 18–23 June 2018; pp. 7132–7141. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. Cbam: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 3–19. [Google Scholar]
- Wang, Q.; Wu, B.; Zhu, P.; Li, P.; Zuo, W.; Hu, Q. ECA-Net: Efficient channel attention for deep convolutional neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 11534–11542. [Google Scholar]
- Wu, M.; Yao, Z.; Verbeke, M.; Karsmakers, P.; Gorissen, B.; Reynaerts, D. Data-driven models with physical interpretability for real-time cavity profile prediction in electrochemical machining processes. Eng. Appl. Artif. Intell. 2025, 160, 111807. [Google Scholar] [CrossRef]
- Liu, W.; Lu, H.; Fu, H.; Cao, Z. Learning to upsample by learning to sample. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 6027–6037. [Google Scholar]






















| Methods | MAP/% | Parameters/M | GFLOPs | FPS/s |
|---|---|---|---|---|
| RepNCSPELAN | 76.22 | 9.1 | 26.5 | 153 |
| RepNCSPELAN-SE | 76.13 | 9.4 | 34.4 | 156 |
| RepNCSPELAN-CBAM | 76.35 | 9.3 | 28.4 | 143 |
| RepNCSPELAN-ECA | 76.24 | 9.1 | 27.1 | 151 |
| RepNCSPELAN-CAA (ARNet) | 77.13 | 10.3 | 32.4 | 133 |
| Methods | MAP/% | Parameters/M | GFLOPs | FPS/s |
|---|---|---|---|---|
| ASSA | 74.75 | 20.7 | 57.8 | 104 |
| ASSA-SEFN | 74.84 | 22.2 | 58.3 | 107 |
| ASSA-SEFN-Mona | 75.47 | 22.3 | 58.4 | 109 |
| ASSA-SEFN-Mona-DyT (AIFI-ASMD) | 76.18 | 22.3 | 58.4 | 111 |
| Methods | Preprocessing/ms | Postprocessing/ms | Reasoning/ms |
|---|---|---|---|
| ASSA | 0.287 | 0.236 | 9.627 |
| ASSA-SEFN | 0.270 | 0.247 | 9.378 |
| ASSA-SEFN-Mona | 0.259 | 0.247 | 9.188 |
| ASSA-SEFN-Mona-DyT | 0.264 | 0.259 | 9.033 |
| Methods | MAP/% | Parameters/M | GFLOPs | FPS/s |
|---|---|---|---|---|
| Nearest-Neighbor Interpolation (NNI) | 74.60 | 19.9 | 57.0 | 113 |
| Dynamic Upsampling (DU) | 74.67 | 19.9 | 57.0 | 110 |
| Converse2D | 75.01 | 19.9 | 57.0 | 88 |
| Baseline | |||||
| Category | Precisoion/% | Recall/% | F1-Score/% | MAP/% | MAP50-90/% |
| crazing | 65.57 | 17.60 | 27.72 | 31.94 | 12.51 |
| inclusion | 73.82 | 72.55 | 73.18 | 77.79 | 42.34 |
| patches | 83.32 | 90.27 | 86.66 | 92.50 | 60.34 |
| pitted surface | 75.63 | 72.19 | 73.87 | 81.43 | 47.60 |
| rolled-in scale | 70.45 | 71.67 | 71.05 | 67.88 | 30.13 |
| scratches | 82.28 | 94.87 | 88.43 | 96.07 | 59.72 |
| all | 75.27 | 69.86 | 70.16 | 74.60 | 42.27 |
| RAC-RTDETR | |||||
| crazing | 70.49 (+4.92) | 24.62 (+7.02) | 36.49 (+8.74) | 40.65 (+8.71) | 17.65 (+5.14) |
| inclusion | 77.26 (+3.44) | 77.45 (+4.90) | 77.36 (+4.18) | 82.82 (+5.03) | 46.81 (+4.47) |
| patches | 79.79 (−3.53) | 89.16 (−1.11) | 84.21 (−2.45) | 90.58 (−1.92) | 60.50 (+0.16) |
| pitted surface | 79.55 (+3.92) | 76.74 (+4.55) | 78.12 (+4.25) | 84.04 (+2.61) | 50.72 (+3.12) |
| rolled−in scale | 70.37 (−0.08) | 73.33 (+1.66) | 71.82 (+0.77) | 74.35 (+6.47) | 36.88 (+5.75) |
| scratches | 86.04 (+3.76) | 94.88 (+0.01) | 90.24 (+1.81) | 96.52 (+0.45) | 61.38 (+1.66) |
| all | 77.25 (+1.98) | 72.69 (+2.83) | 73.04 (+2.88) | 78.16 (+3.56) | 45.66 (+3.39) |
| Methods | MAP/% | Parameters/M | GFLOPs | FPS/s |
|---|---|---|---|---|
| Faster RCNN | 74.81 | 28.4 | 94.12 | 51 |
| Yolov5n | 74.15 | 1.9 | 4.9 | 83 |
| Yolov7-tiny | 71.82 | 6.8 | 13.1 | 77 |
| Yolov8n | 76.55 | 3.1 | 8.1 | 96 |
| Yolov10s | 71.36 | 7.2 | 21.4 | 83 |
| Yolo11n | 76.13 | 2.58 | 6.3 | 92 |
| RTDETR-R18 (Baseline) | 74.60 | 19.9 | 57.0 | 113 |
| RTDETR-R34 | 75.84 | 31.1 | 88.8 | 98 |
| RTDETR-R50 | 75.99 | 41.9 | 129.6 | 62 |
| RAC-RTDETR (ours) | 78.16 | 12.7 | 33.8 | 122 |
| Methods | MAP/% | Parameters/M | GFLOPs | FPS/s |
|---|---|---|---|---|
| Baseline | 74.60 | 19.9 | 57.0 | 113 |
| ARNet | 77.13 | 10.3 | 32.4 | 133 |
| Converse2D | 75.01 | 19.9 | 57.0 | 88 |
| AIFI-ASMD | 76.18 | 22.3 | 58.4 | 111 |
| ARNet+AIFI-ASMD | 75.74 | 12.7 | 33.8 | 119 |
| ARNet+Converse2D | 76.45 | 10.3 | 32.4 | 123 |
| AIFI-ASMD + Converse2D | 76.01 | 22.3 | 58.4 | 95 |
| ARNet+AIFI-ASMD + Converse2D (ours) | 78.16 | 12.7 | 33.8 | 122 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Xu, Z.; Wang, N. RAC-RTDETR: A Lightweight, Efficient Real-Time Small-Object Detection Algorithm for Steel Surface Defect Detection. Electronics 2025, 14, 4968. https://doi.org/10.3390/electronics14244968
Xu Z, Wang N. RAC-RTDETR: A Lightweight, Efficient Real-Time Small-Object Detection Algorithm for Steel Surface Defect Detection. Electronics. 2025; 14(24):4968. https://doi.org/10.3390/electronics14244968
Chicago/Turabian StyleXu, Zhenping, and Nengxi Wang. 2025. "RAC-RTDETR: A Lightweight, Efficient Real-Time Small-Object Detection Algorithm for Steel Surface Defect Detection" Electronics 14, no. 24: 4968. https://doi.org/10.3390/electronics14244968
APA StyleXu, Z., & Wang, N. (2025). RAC-RTDETR: A Lightweight, Efficient Real-Time Small-Object Detection Algorithm for Steel Surface Defect Detection. Electronics, 14(24), 4968. https://doi.org/10.3390/electronics14244968

