GO-PILL: A Geometry-Aware OCR Pipeline for Reliable Recognition of Debossed and Curved Pill Imprints
Abstract
1. Introduction
- We propose a geometry-aware IR-ILR framework that explicitly models the physical morphology of pill imprints. By repurposing centerline extraction for adaptive contrast enhancement and nonlinear rectification, we bridge the structural gap between irregular pill surfaces and generic OCR engines.
- We introduce a text alignment and rectification strategy that transforms curved, diagonal, and irregular imprint arrangements into geometrically normalized representations, thereby reducing distortion and improving OCR accuracy.
- We develop a unified three-stage pipeline—comprising IR, ILR, and OCR—and demonstrate its superior robustness and consistent performance across diverse imprint conditions, validating its effectiveness for real-world clinical applications.
2. Related Works
2.1. Shape-Based Identification Methods
2.2. Imprint-Based Identification Methods
2.3. OCR-Based Imprint Recognition
3. Methodology
3.1. Imprint Refinement
3.1.1. TextSnake-Based Text-Region Extraction
3.1.2. Morphological Contrast Enhancement
3.1.3. Localized Otsu Binarization
3.2. Imprint Localization and Rectification
3.2.1. Centerline Refinement
3.2.2. Text Instance Classification
3.2.3. Text Line Rectification
3.3. Text Recognition
4. Experimental Results
- (1)
- Baseline OCR, which applies OCR directly to raw images;
- (2)
- Non-OCR approaches, which rely on detection or classification rather than text decoding; and
- (3)
- Preprocessed OCR, which incorporates the proposed IR and ILR modules prior to OCR.
4.1. Experimental Setup
4.2. Ablation Study
4.3. Comparative Evaluation
| Category | Recognition Model | Class Type | Precision | Recall | F1-Score |
|---|---|---|---|---|---|
| Baseline OCR | Paddle OCR [25] | All | 80.66 | 41.50 | 54.80 |
| Printed | 96.60 | 81.95 | 88.67 | ||
| Debossed | 76.68 | 35.92 | 48.93 | ||
| Curved | 71.61 | 16.83 | 27.26 | ||
| Diagonal | 80.71 | 59.11 | 68.24 | ||
| Linear | 83.72 | 57.40 | 68.11 | ||
| Easy OCR [40] | All | 81.01 | 35.36 | 49.24 | |
| Printed | 93.91 | 77.98 | 85.21 | ||
| Debossed | 77.22 | 29.52 | 42.71 | ||
| Curved | 68.04 | 19.72 | 30.58 | ||
| Diagonal | 86.21 | 46.47 | 60.39 | ||
| Linear | 85.29 | 45.56 | 59.39 | ||
| TrOCR [36] | All | 59.01 | 53.90 | 56.34 | |
| Printed | 56.56 | 49.82 | 52.98 | ||
| Debossed | 59.40 | 54.43 | 56.81 | ||
| Curved | 63.16 | 51.39 | 56.67 | ||
| Diagonal | 50.11 | 43.31 | 46.46 | ||
| Linear | 60.43 | 63.67 | 62.01 | ||
| MMOCR [12] | All | 71.88 | 74.09 | 72.97 | |
| Printed | 64.71 | 59.57 | 62.03 | ||
| Debossed | 72.48 | 75.52 | 73.97 | ||
| Curved | 50.87 | 52.69 | 51.76 | ||
| Diagonal | 83.36 | 83.83 | 83.60 | ||
| Linear | 85.76 | 85.08 | 85.42 | ||
| Non-OCR approach | YOLOv11 [39] | All | 73.14 | 66.40 | 69.61 |
| Printed | 75.57 | 60.29 | 67.07 | ||
| Debossed | 72.86 | 67.03 | 69.82 | ||
| Curved | 62.23 | 49.85 | 55.36 | ||
| Diagonal | 70.87 | 67.42 | 69.10 | ||
| Linear | 84.06 | 83.49 | 83.77 | ||
| Heo et al. [29] | All | 51.47 | 70.94 | 59.66 | |
| Printed | 72.84 | 88.09 | 79.74 | ||
| Debossed | 48.93 | 68.60 | 57.11 | ||
| Curved | 31.69 | 28.88 | 30.22 | ||
| Diagonal | 38.11 | 56.88 | 45.64 | ||
| Linear | 74.90 | 89.75 | 81.66 | ||
| Preprocessed OCR | Ponte et al. [24] | All | 80.60 | 42.54 | 55.69 |
| Printed | 92.47 | 79.78 | 85.66 | ||
| Debossed | 77.66 | 37.48 | 50.56 | ||
| Curved | 73.75 | 17.63 | 28.46 | ||
| Diagonal | 80.35 | 60.04 | 68.72 | ||
| Linear | 83.01 | 59.00 | 68.97 | ||
| GO-PILL | All | 81.93 | 81.73 | 81.83 | |
| Printed | 96.88 | 89.53 | 93.06 | ||
| Debossed | 80.05 | 80.61 | 80.33 | ||
| Curved | 63.72 | 65.44 | 64.57 | ||
| Diagonal | 87.91 | 85.13 | 86.50 | ||
| Linear | 99.19 | 98.05 | 98.62 |
4.4. Performance Evaluation Under Glare Noise Conditions
5. Discussion
6. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- El Hajj, M.; Asiri, R.; Husband, A.; Todd, A. Medication errors in community pharmacies: A systematic review of the international literature. PLoS ONE 2025, 20, e0322392. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- World Health Organization. Global Burden of Preventable Medication-Related Harm in Health Care: A Systematic Review; World Health Organization: Geneva, Switzerland, 2023; Available online: https://iris.who.int/server/api/core/bitstreams/f208903d-b47d-4a8c-9fac-d8136c2bdbb6/content (accessed on 25 September 2025).
- World Health Organization. Medication Safety for Look-Alike, Sound-Alike Medicines; World Health Organization: Geneva, Switzerland, 2023; Available online: https://iris.who.int/server/api/core/bitstreams/5a2f5a2c-e95c-45b8-ad44-f15c9351f9f1/content (accessed on 25 September 2025).
- Bates, D.; Slight, S. Medication errors: What is their impact? Mayo Clin. Proc. 2014, 89, 1027–1029. [Google Scholar] [CrossRef] [Scilit]
- Wong, Y.; Ng, H.; Leung, K.; Chan, K.; Chan, S.; Loy, C. Development of fine-grained pill identification algorithm using deep convolutional network. J. Biomed. Inform. 2017, 74, 130–136. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ou, Y.; Tsai, A.; Zhou, X.; Wang, J. Automatic drug pills detection based on enhanced feature pyramid network and convolution neural networks. IET Comput. Vis. 2020, 14, 9–17. [Google Scholar] [CrossRef] [Scilit]
- Tan, L.; Huangfu, T.; Wu, L.; Chen, W. Comparison of RetinaNet, SSD, and YOLO v3 for real-time pill identification. BMC Med. Inform. Decis. Mak. 2021, 21, 324. [Google Scholar] [CrossRef] [Scilit]
- Larios Delgado, N.; Usuyama, N.; Hall, A.; Hazen, R.; Ma, M.; Sahu, S.; Lundin, J. Fast and accurate medication identification. npj Digit. Med. 2019, 2, 10. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Nguyen, A.; Pham, H.; Trung, H.; Nguyen, Q.; Truong, T.; Nguyen, P. High accurate and explainable multi-pill detection framework with graph neural network-assisted multimodal data fusion. PLoS ONE 2023, 18, e0291865. [Google Scholar] [CrossRef] [Scilit]
- My, L.N.T.; Le, V.-T.; Vo, T.; Hoang, V.T. A comprehensive review of pill image recognition. Comput. Mater. Contin. 2025, 82, 3693–3740. [Google Scholar] [CrossRef] [Scilit]
- Long, S.; Ruan, J.; Zhang, W.; He, X.; Wu, W.; Yao, C. TextSnake: A flexible representation for detecting text of arbitrary shapes. arXiv 2018, arXiv:1807.01544. [Google Scholar]
- OpenMMLab. MMOCR. Available online: https://github.com/open-mmlab/mmocr (accessed on 14 November 2025).
- Fang, S.; Xie, H.; Wang, Y.; Mao, Z.; Zhang, Y. read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition. arXiv 2021, arXiv:2103.06495. [Google Scholar] [CrossRef] [Scilit]
- National Library of Medicine (US). RxImage Dataset. Available online: https://datadiscovery.nlm.nih.gov/ (accessed on 5 November 2024).
- Kwon, H.; Kim, H.; Lee, S. Pill detection model for medicine inspection based on deep learning. Chemosensors 2022, 10, 4. [Google Scholar] [CrossRef] [Scilit]
- He, K.; Gkioxari, G.; Dollár, P.; Girshick, R. Mask R-CNN. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Venice, Italy, 22–29 October 2017; pp. 2961–2969. [Google Scholar]
- Kim, S.; Park, E.; Kim, J.; Ihm, S. Combination pattern method using deep learning for pill classification. Appl. Sci. 2024, 14, 9065. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G. YOLOv5. Available online: https://github.com/ultralytics/yolov5 (accessed on 8 May 2025).
- Lee, Y.; Park, U.; Jain, A.; Lee, S. Pill-ID: Matching and retrieval of drug pill images. Pattern Recognit. Lett. 2012, 33, 904–910. [Google Scholar] [CrossRef] [Scilit]
- Lowe, D. Distinctive image features from scale-invariant keypoints. Int. J. Comput. Vis. 2004, 60, 91–110. [Google Scholar] [CrossRef] [Scilit]
- Guo, Z.; Zhang, L.; Zhang, D. A completed modeling of local binary pattern operator for texture classification. IEEE Trans. Image Process. 2010, 19, 1657–1663. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Al-Hussaeni, K.; Karamitsos, I.; Adewumi, E.; Amawi, R. CNN-based pill image recognition for retrieval systems. Appl. Sci. 2023, 13, 5050. [Google Scholar] [CrossRef] [Scilit]
- Canny, J. A computational approach to edge detection. IEEE Trans. Pattern Anal. Mach. Intell. 1986, 8, 679–698. [Google Scholar] [CrossRef] [Scilit]
- Ponte, E.; Amparo, X.; Huang, K.; Kwak, D. Automatic pill identification system based on deep learning and image preprocessing. In Proceedings of the 2023 Congress in Computer Science, Computer Engineering, & Applied Computing (CSCE), Las Vegas, NV, USA, 24–27 July 2023; pp. 1–6. [Google Scholar]
- Cui, C.; Sun, T.; Lin, M.; Gao, T.; Zhang, Y.; Liu, J.; Wang, X.; Zhang, Z.; Zhou, C.; Liu, H.; et al. PaddleOCR 3.0 technical report. arXiv 2025, arXiv:2507.05595. [Google Scholar] [CrossRef] [Scilit]
- Suntronsuk, S.; Ratanotayanon, S. Pill image binarization for detecting text imprints. In Proceedings of the 2016 13th International Joint Conference on Computer Science and Software Engineering (JCSSE), Khon Kaen, Thailand, 13–15 July 2016; pp. 1–6. [Google Scholar]
- Otsu, N. A threshold selection method from gray-level histograms. IEEE Trans. Syst. Man Cybern. 1979, 9, 62–66. [Google Scholar] [CrossRef] [Scilit]
- Smith, R. An overview of the Tesseract OCR engine. In Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), Curitiba, Brazil, 23–26 September 2007; Volume 2, pp. 629–633. [Google Scholar]
- Heo, J.; Kang, Y.; Lee, S.; Jeong, D.; Kim, K. An accurate deep learning-based system for automatic pill identification: Model development and validation. J. Med. Internet Res. 2023, 25, e41043. [Google Scholar] [CrossRef] [Scilit]
- Dhivya, A.; Sundaresan, M. Tablet identification using support vector machine-based text recognition and error correction by enhanced n-grams algorithm. IET Image Process. 2020, 14, 1366–1372. [Google Scholar]
- Cortes, C.; Vapnik, V. Support-vector networks. Mach. Learn. 1995, 20, 273–297. [Google Scholar]
- Li, M.; Lv, T.; Chen, J.; Cui, L.; Lu, Y.; Florencio, D.; Zhang, C.; Li, Z.; Wei, F. TrOCR: Transformer-based optical character recognition with pre-trained models. arXiv 2021, arXiv:2109.10282. [Google Scholar] [CrossRef] [Scilit]
- Aluri, M.; Tatavarthi, U. Geometric deep learning for enhancing irregular scene text detection. Rev. D’intelligence Artif. 2024, 38, 115–125. [Google Scholar] [CrossRef] [Scilit]
- Ronneberger, O.; Fischer, P.; Brox, T. U-Net: Convolutional networks for biomedical image segmentation. In Proceedings of the International Conference on Medical Image Computing and Computer-Assisted Intervention (MICCAI), Munich, Germany, 5–9 October 2015; pp. 234–241. [Google Scholar]
- Kipf, T.; Welling, M. Semi-supervised classification with graph convolutional networks. arXiv 2016, arXiv:1609.02907. [Google Scholar]
- Hu, Y.; Dong, B.; Huang, K.; Ding, L.; Wang, W.; Huang, X.; Wang, Q. Scene text recognition via dual-path network with shape-driven attention alignment. ACM Trans. Multimed. Comput. Commun. Appl. 2024, 20, 107. [Google Scholar] [CrossRef] [Scilit]
- Rehman, N.; Haroon, F. Adaptive Gaussian and double thresholding for contour detection and character recognition of two-dimensional area using computer vision. Eng. Proc. 2023, 32, 23. [Google Scholar]
- Sauvola, J.; Pietikäinen, M. Adaptive document image binarization. Pattern Recognit. 2000, 33, 225–236. [Google Scholar] [CrossRef] [Scilit]
- Jocher, G.; Qiu, J. Ultralytics YOLO11. Available online: https://github.com/ultralytics/ultralytics (accessed on 2 December 2025).
- JaidedAI. EasyOCR. Available online: https://github.com/JaidedAI/EasyOCR (accessed on 15 November 2025).















| Parameter | Symbol | Value | Parameter | Symbol | Value |
|---|---|---|---|---|---|
| Minimum text-region area threshold | 125 | Geometric kernel size | |||
| Minimum centerline area threshold | 250 | Angular variance threshold for diagonal text | 15 | ||
| Minimum number of centerline points | 10 | Edge angle threshold for diagonal text | 3 | ||
| Curved-text average angle threshold | 20 | Minimum number of center points (curved) | 12 | ||
| Diagonal text average angle range | , | Minimum number of center points (diagonal) | 5 | ||
| Curved-text angular variance threshold | 3 | Minimum number of center points (linear) | 3 |
| Preprocessing Module | Precision | Recall | F1-Score | |
|---|---|---|---|---|
| IR | ILR | |||
| - | - | 71.88 | 74.09 | 72.97 |
| √ | - | 55.07 | 48.02 | 51.30 |
| - | √ | 54.95 | 34.64 | 42.49 |
| √ | √ | 81.93 | 81.73 | 81.83 |
| Thresholding Technique | Precision | Recall | F1-Score |
|---|---|---|---|
| Adaptive Gaussian [37] | 82.98 | 77.68 | 80.24 |
| Sauvola [38] | 67.47 | 49.06 | 56.81 |
| Otsu [27] | 81.93 | 81.73 | 81.83 |
| Category | Recognition Model | Top-1 | Top-5 | Top-10 |
|---|---|---|---|---|
| Baseline OCR | MMOCR [12] | 36.95 | 61.90 | 69.62 |
| Non-OCR | YOLOv11 [39] | 27.91 | 54.02 | 62.07 |
| Heo et al. [29] | 28.90 | 46.47 | 53.86 | |
| Preprocessed OCR | Ponte et al. [24] | 23.65 | 40.23 | 46.14 |
| GO-PILL | 46.63 | 68.80 | 76.52 |
| Model | Glare Intensity | |||||
|---|---|---|---|---|---|---|
| Level 1 | Level 2 | |||||
| Precision | Recall | F1-Score | Precision | Recall | F1-Score | |
| MMOCR [12] | 76.96 | 69.67 | 73.13 | 69.37 | 57.07 | 62.62 |
| YOLOv11 [39] | 71.41 | 61.28 | 65.96 | 69.80 | 52.03 | 59.62 |
| Heo et al. [29] | 50.41 | 58.87 | 54.31 | 45.12 | 52.02 | 48.32 |
| Ponte et al. [24] | 77.63 | 36.63 | 49.77 | 78.93 | 25.16 | 38.15 |
| GO-PILL | 83.09 | 75.10 | 78.89 | 78.60 | 60.91 | 68.63 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Jo, J.; Yoon, S.; Cho, J. GO-PILL: A Geometry-Aware OCR Pipeline for Reliable Recognition of Debossed and Curved Pill Imprints. Mathematics 2026, 14, 356. https://doi.org/10.3390/math14020356
Jo J, Yoon S, Cho J. GO-PILL: A Geometry-Aware OCR Pipeline for Reliable Recognition of Debossed and Curved Pill Imprints. Mathematics. 2026; 14(2):356. https://doi.org/10.3390/math14020356
Chicago/Turabian StyleJo, Jaehyeon, Sungan Yoon, and Jeongho Cho. 2026. "GO-PILL: A Geometry-Aware OCR Pipeline for Reliable Recognition of Debossed and Curved Pill Imprints" Mathematics 14, no. 2: 356. https://doi.org/10.3390/math14020356
APA StyleJo, J., Yoon, S., & Cho, J. (2026). GO-PILL: A Geometry-Aware OCR Pipeline for Reliable Recognition of Debossed and Curved Pill Imprints. Mathematics, 14(2), 356. https://doi.org/10.3390/math14020356

