YOLO-SR: A Modified YOLO Model with Strip Pooling and a Rectangular Self-Calibration Module for Defect Segmentation in Smart Card Surfaces
Abstract
1. Introduction
1.1. Research Background
1.2. Critical Areas and Verification Challenges of Smart Cards
1.3. Objectives and Methodology Overview
2. Related Work
3. Materials and Methods: Dataset and Data Augmentation
4. Materials and Methods: YOLO-SR Model
4.1. Overall Structure of the Proposed Method
4.2. C3K2_SP Module
4.3. RCM
5. Experimental Results and Discussion
5.1. Dataset
5.2. Experiment Setting
5.3. Results and Analysis
5.4. Computational Complexity and Inference Latency
5.5. Domain Discrepancy Between Synthetic and Real Defects
5.6. Ablation on Module Placement
5.6.1. SP Placement
5.6.2. RCM Placement
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Mazzei, D.; Ramjattan, R. Machine Learning for Industry 4.0: A Systematic Review Using Deep Learning-Based Topic Modelling. Sensors 2022, 22, 8641. [Google Scholar] [CrossRef] [PubMed]
- Zhang, L.; Jia, X.; Chang, Q.; Liu, Y.; Zhang, Z.; Cao, Y.; Liu, J.; Yang, Y. The development of Machine Vision and its Applications in Different Industries: A review. Mech. Eng. Adv. 2024, 2, 1746. [Google Scholar] [CrossRef]
- de la Rosa, F.L.; Sánchez-Reolid, R.; Gómez-Sirvent, J.L.; Morales, R.; Fernández-Caballero, A. A Review on Machine and Deep Learning for Semiconductor Defect Classification in Scanning Electron Microscope Images. Appl. Sci. 2021, 11, 9508. [Google Scholar] [CrossRef]
- Zheng, X.; Zheng, S.; Kong, Y.; Chen, J. Recent advances in surface defect inspection of industrial products using deep learning techniques. Int. J. Adv. Manuf. Technol. 2021, 113, 35–58. [Google Scholar] [CrossRef]
- Bártová, B.; Bína, V.; Váchová, L. A PRISMA-driven Systematic Review of Data Mining Methods Used for Defects Detection and Classification in the Manufacturing Industry. Production 2022, 32, e20210097. [Google Scholar] [CrossRef]
- Prabhu, E.; Debby, G. An Efficient Optimization Approach for Steel Surface Flaw Classification using Machine Learning. Int. J. Res. Appl. Sci. Eng. Technol. 2025, 13, 1526–1538. [Google Scholar] [CrossRef]
- Waseem, F.; Menon, S.; Xu, H.; Mondal, D. VizInspect Pro—Automated Optical Inspection (AOI). arXiv 2022, arXiv:2205.13095. [Google Scholar]
- Wang, K.-J.; Fan-Jiang, H.; Lee, Y.-X. A multiple-stage defect detection model by convolutional neural network. Comput. Ind. Eng. 2022, 168, 108096. [Google Scholar] [CrossRef]
- Ma, Y.; Yin, J.; Huang, F.; Li, Q. Surface defect inspection of industrial products with object detection deep networks: A systematic review. Artif. Intell. Rev. 2024, 57, 333. [Google Scholar] [CrossRef]
- Bower, L.A. Automatic Identification Technology (AIT): The Development of Functional Capability and Card Application Matrices. Master’s Thesis, Naval Postgraduate School, Monterey, CA, USA, 1994. [Google Scholar]
- Access Control Council; Identity Council. PIV Card/Reader Challenges with Physical Access Control Systems: A Field Troubleshooting Guide. Available online: https://www.securetechalliance.org/piv-card-reader-challenges-with-physical-access-control-systems-a-field-troubleshooting-guide/ (accessed on 1 October 2025).
- TransFirst, TSYS. Operating Guide for Merchant Card Processing v6-0915. 2015. Available online: https://assets.tsys.com/Assets/TSYS/downloads/merchant/docs/TransFirst_Merchant_Card%20Processing_Operating_Guide_v6-0915.pdf (accessed on 1 October 2025).
- Liang, X.; Sun, J.; Wang, X.; Li, J.; Zhang, L.; Guo, J. Surface weak scratch detection for optical elements based on a multimodal imaging system and a deep encoder–decoder network. J. Opt. Soc. Am. A 2023, 40, 1237–1248. [Google Scholar] [CrossRef] [PubMed]
- Ye, B.; Xue, R.; Wu, Q. A hybrid attention multi-scale fusion network for real-time semantic segmentation. Sci. Rep. 2025, 15, 872. [Google Scholar] [CrossRef] [PubMed]
- Chen, X.; Cao, J.-W.; Wang, Y.-P. Defect detection in ID cards with accurately reconstructed reference image. In Proceedings of the 5th International Conference on Multimedia and Image Processing, Nanjing China, 10–12 January 2020; pp. 18–22. [Google Scholar]
- Lin, H.; Zhan, Y.; Liu, S.; Ke, X.; Chen, Y. A deep learning based bank card detection and recognition method in complex scenes. Appl. Intell. 2022, 52, 15259–15277. [Google Scholar] [CrossRef]
- Ultralytics. YOLO11 Documentation & Software Repository. 2024. Available online: https://docs.ultralytics.com/models/yolo11/ (accessed on 1 October 2025).
- Cai, X.; Ruan, Z.; Sun, H. Bank Card Number Identification Method Based on YOLOv3 and MobileNetv2. J. Comput.-Aided Des. Comput. Graph. 2022, 34, 142–151. [Google Scholar] [CrossRef]
- Sun, G.; You, F. Bank card number recognition system based on deep learning. In Proceedings of the EITCE 2020: 2020 4th International Conference on Electronic Information Technology and Computer Engineering, Online, 6–8 November 2020; pp. 745–749. [Google Scholar]
- Bochkovskiy, A.; Wang, C.-Y.; Liao, H.-Y.M. YOLOv4: Optimal Speed and Accuracy of Object Detection. arXiv 2020, arXiv:2004.10934. [Google Scholar] [CrossRef]
- Ghiasi, G.; Cui, Y.; Srinivas, A.; Qian, R.; Lin, T.-Y.; Cubuk, E.D.; Le, Q.V.; Zoph, B. Simple Copy-Paste Is a Strong Data Augmentation Method for Instance Segmentation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 2918–2928. [Google Scholar]
- Cubuk, E.D.; Zoph, B.; Mane, D.; Vasudevan, V.; Le, Q.V. AutoAugment: Learning Augmentation Strategies from Dat. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 113–123. [Google Scholar]
- Li, C.-L.; Sohn, K.; Yoon, J.; Pfister, T. CutPaste: Self-Supervised Learning for Anomaly Detection and Localization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 9664–9674. [Google Scholar]
- Zhang, G.; Cui, K.; Hung, T.-Y.; Lu, S. Defect-GAN: High-Fidelity Defect Synthesis for Automated Defect Inspection. In Proceedings of the 2021 IEEE Winter Conference on Applications of Computer Vision (WACV), Virtual, 5–9 January 2021; pp. 2524–2534. [Google Scholar]
- Hou, Q.; Zhou, D.; Feng, J. Coordinate Attention for Efficient Mobile Network Design. In Proceedings of the 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 13708–13717. [Google Scholar]
- Wang, B.; Wang, M.; Yang, J.; Luo, H. YOLOv5-CD: Strip steel surface defect detection method based on coordinate attention and a decoupled head. Meas. Sens. 2023, 30, 100909. [Google Scholar] [CrossRef]
- Hou, Q.; Zhang, L.; Cheng, M.-M.; Feng, J. Strip Pooling: Rethinking Spatial Pooling for Scene Parsing. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 4002–4011. [Google Scholar]
- Takikawa, T.; Acuna, D.; Jampani, V.; Fidler, S. Gated-scnn: Gated shape cnns for semantic segmentation. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 5229–5238. [Google Scholar]
- Qin, X.; Zhang, Z.; Huang, C.; Gao, C.; Dehghan, M.; Jagersand, M. BASNet: Boundary-Aware Salient Object Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; pp. 7479–7489. [Google Scholar]
- Feng, Y.; Fan, Z.; Yan, Y.; Jiang, Z.; Zhang, S. MFAFNet: Multi-Scale Feature Adaptive Fusion Network Based on DeepLab V3+ for Cloud and Cloud Shadow Segmentation. Remote. Sens. 2025, 17, 1229. [Google Scholar] [CrossRef]
- Ni, Z.; Chen, X.; Zhai, Y.; Tang, Y.; Wang, Y. Context-Guided Spatial Feature Reconstruction for Efficient Semantic Segmentation. In Computer Vision–ECCV 2024; Springer: Cham, Switzerland, 2024; pp. 239–255. [Google Scholar]
- Sun, B.; Cheng, X. Smoke Detection Transformer: An Improved Real-Time Detection Transformer Smoke Detection Model for Early Fire Warning. Fire 2024, 7, 488. [Google Scholar] [CrossRef]
- Liu, L.; Li, J. MCRS-YOLO: Multi-Aggregation Cross-Scale Feature Fusion Object Detector for Remote Sensing Images. Remote Sens. 2025, 17, 2204. [Google Scholar] [CrossRef]
- Zhang, Y.; Yu, X.; Ji, X. WMC-RTDETR: A Lightweight Tea Disease Detection Model. Front. Plant Sci. 2025, 16, 1574920. [Google Scholar] [CrossRef] [PubMed]
- Hu, F.; Abula, M.; Wang, D.; Li, X.; Yan, N.; Xie, Q.; Zhang, X. Investigation of an Efficient Multi-Class Cotton Leaf Disease Detection Algorithm That Leverages YOLOv11. Plants 2025, 25, 4432. [Google Scholar] [CrossRef] [PubMed]















| Dataset | Class | Original Data | Augmented Data | Total (Train + Val) | Train/Val (9:1) | Test Data |
|---|---|---|---|---|---|---|
| ICChip | damaged | 32 | 968 | 1000 | 900/100 | 30 |
| foreign | 28 | 972 | 1000 | 900/100 | 30 | |
| scratch | 36 | 964 | 1000 | 900/100 | 30 | |
| Trace | 32 | 968 | 1000 | 900/100 | 30 | |
| background | 110 | 890 | 1000 | 900/100 | 30 | |
| SignPlate | damaged | 24 | 976 | 1000 | 900/100 | 30 |
| foreign | 26 | 974 | 1000 | 900/100 | 30 | |
| scratch | 30 | 970 | 1000 | 900/100 | 30 | |
| background | 90 | 910 | 1000 | 900/100 | 30 |
| Category | Environment |
|---|---|
| Hardware | Intel(R) Core(TM) i9-14900KF × 1 RAM 64 GB GeForce RTX 4080 SUPER 16 GB × 1 |
| Software | Windows 10 (10.0.26100) Python 3.10.15 Cuda 12.1 Pytorch 2.5.1 |
| Models | Batch Size | Epoch | Input Image Pixels | Learning Rate |
|---|---|---|---|---|
| Mask R-CNN | 40 | 300 | 640 × 640 | 0.0001 |
| Swin Transformer | 40 | 300 | 640 × 640 | 0.0001 |
| YOLO11n | 40 | 300 | 640 × 640 | 0.0001 |
| YOLO11 + SP | 40 | 300 | 640 × 640 | 0.0001 |
| YOLO11 + RCM | 40 | 300 | 640 × 640 | 0.0001 |
| YOLO-SR | 40 | 300 | 640 × 640 | 0.0001 |
| + | Method | Recall | Precision | Map@0.5 | Map@0.5:0.95 |
|---|---|---|---|---|---|
| ICChip | Mask R-CNN | 0.684 | 0.698 | 0.635 | 0.299 |
| Swin Transformer | 0.758 | 0.727 | 0.714 | 0.328 | |
| YOLO11n | 0.737 | 0.716 | 0.701 | 0.322 | |
| YOLO11 + SP | 0.719 | 0.756 | 0.719 | 0.314 | |
| YOLO11 + RCM | 0.748 | 0.721 | 0.691 | 0.324 | |
| YOLO-SR | 0.800 | 0.781 | 0.787 | 0.335 | |
| SignPlate | Mask R-CNN | 0.541 | 0.729 | 0.623 | 0.347 |
| Swin Transformer | 0.689 | 0.856 | 0.763 | 0.372 | |
| YOLO11n | 0.643 | 0.802 | 0.713 | 0.361 | |
| YOLO11 + SP | 0.645 | 0.843 | 0.766 | 0.370 | |
| YOLO11 + RCM | 0.701 | 0.811 | 0.732 | 0.386 | |
| YOLO-SR | 0.719 | 0.867 | 0.808 | 0.405 |
| Method | Params (M) | Latency (ms/Image) | FPS |
|---|---|---|---|
| YOLO11n | 2.84 | 6.34 | 157 |
| YOLO11 + SP | 3.08 | 10.31 | 97.0 |
| YOLO11 + RCM | 3.58 | 6.75 | 148.1 |
| YOLO-SR | 6.28 | 10.84 | 92.2 |
| Recall | Precision | Map@0.5 | Map@0.5:0.95 | ||
|---|---|---|---|---|---|
| S0 | YOLO11n (baseline) | 0.643 | 0.802 | 0.713 | 0.361 |
| S1 | SP-shallow (backbone) | 0.625 | 0.789 | 0.694 | 0.338 |
| S2 | SP-deep (backbone) | 0.648 | 0.831 | 0.755 | 0.364 |
| S3 | SP-head | 0.632 | 0.827 | 0.763 | 0.352 |
| S4 | SP-deep + head (ours) | 0.645 | 0.843 | 0.766 | 0.370 |
| S5 | SP-shallow + head | 0.605 | 0.772 | 0.685 | 0.333 |
| Recall | Precision | Map@0.5 | Map@0.5:0.95 | ||
|---|---|---|---|---|---|
| R0 | YOLO11n (baseline) | 0.643 | 0.802 | 0.713 | 0.361 |
| R1 | RCM-mid | 0.683 | 0.749 | 0.681 | 0.317 |
| R2 | RCM-top (ours) | 0.701 | 0.811 | 0.732 | 0.386 |
| R3 | RCM-head | 0.607 | 0.715 | 0.654 | 0.302 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/licenses/by/4.0/).
Share and Cite
Yao, T.; Hossain, F.M.F.; Kim, S.-H.; Yoo, K.-H. YOLO-SR: A Modified YOLO Model with Strip Pooling and a Rectangular Self-Calibration Module for Defect Segmentation in Smart Card Surfaces. Appl. Sci. 2025, 15, 12980. https://doi.org/10.3390/app152412980
Yao T, Hossain FMF, Kim S-H, Yoo K-H. YOLO-SR: A Modified YOLO Model with Strip Pooling and a Rectangular Self-Calibration Module for Defect Segmentation in Smart Card Surfaces. Applied Sciences. 2025; 15(24):12980. https://doi.org/10.3390/app152412980
Chicago/Turabian StyleYao, Tianshui, F. M. Fahmid Hossain, Sung-Hoon Kim, and Kwan-Hee Yoo. 2025. "YOLO-SR: A Modified YOLO Model with Strip Pooling and a Rectangular Self-Calibration Module for Defect Segmentation in Smart Card Surfaces" Applied Sciences 15, no. 24: 12980. https://doi.org/10.3390/app152412980
APA StyleYao, T., Hossain, F. M. F., Kim, S.-H., & Yoo, K.-H. (2025). YOLO-SR: A Modified YOLO Model with Strip Pooling and a Rectangular Self-Calibration Module for Defect Segmentation in Smart Card Surfaces. Applied Sciences, 15(24), 12980. https://doi.org/10.3390/app152412980

