CoordAtt-PARSeq: A Coordinate-Attention-Stem- and Focal-Loss-Augmented Permuted Autoregressive Transformer for Robust Fiscal Receipt Text Recognition
Abstract
1. Introduction
- We analyze the limitations of pure Vision-Transformer-based STR models on fiscal receipt images, and attribute the performance gap mainly to the lack of explicit spatial inductive bias and the long-tailed distribution of receipt characters.
- We develop CoordAtt-PARSeq, a lightweight extension of PARSeq-tiny that integrates a CoordAtt stem and Focal Loss. The CoordAtt stem introduces only 999 extra parameters, less than 0.02% of the model size, while preserving a computational profile close to that of the original PARSeq-tiny recognizer.
- We conduct controlled experiments on WildReceipt and SROIE with five strong baselines, including CRNN, ViTSTR, ABINet, TRBA, and PARSeq-tiny. The results show that CoordAtt-PARSeq achieves the best performance on all reported main metrics, with absolute gains of 4.19 and 4.47 percentage points in case-insensitive word accuracy on WildReceipt and SROIE, respectively.
2. Related Work
2.1. Scene Text Recognition
2.2. Receipt OCR and Document Understanding
2.3. Attention Stems and Class-Imbalance Learning
3. Methodology
3.1. Overall Architecture
3.2. CoordAtt Stem
3.3. Focal Loss
3.4. Network Architecture in Detail
4. Experimental Analysis
4.1. Datasets
4.2. Baselines and Evaluation Metrics
4.3. Implementation Details
4.4. Main Results
4.5. Ablation Study
4.6. Sensitivity Analysis of Focal Loss Parameter
5. Discussion
5.1. Qualitative Behaviour
5.2. Error Analysis
5.3. Limitations and Future Work
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Ribeiro, J.; Lima, R.; Eckhardt, T.; Paiva, S. Robotic process automation and artificial intelligence in industry 4.0—A literature review. Procedia Comput. Sci. 2021, 181, 51–58. [Google Scholar] [CrossRef]
- Zhang, Y.; Xiong, F.; Xie, Y.; Fan, X.; Gu, H. The impact of artificial intelligence and blockchain on the accounting profession. IEEE Access 2020, 8, 110461–110477. [Google Scholar] [CrossRef]
- Tang, Q.; Lee, Y.; Jung, H. The industrial application of artificial intelligence-based optical character recognition in modern manufacturing innovations. Sustainability 2024, 16, 2161. [Google Scholar] [CrossRef]
- Baviskar, D.; Ahirrao, S.; Potdar, V.; Kotecha, K. Efficient automated processing of the unstructured documents using artificial intelligence: A systematic literature review and future directions. IEEE Access 2021, 9, 72894–72936. [Google Scholar] [CrossRef]
- Mahadevkar, S.V.; Patil, S.; Kotecha, K.; Soong, L.W.; Choudhury, T. Exploring AI-driven approaches for unstructured document analysis and future horizons. J. Big Data 2024, 11, 92. [Google Scholar] [CrossRef]
- Huang, Z.; Chen, K.; He, J.; Bai, X.; Karatzas, D.; Lu, S.; Jawahar, C. Icdar2019 competition on scanned receipt ocr and information extraction. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR); IEEE: New York, NY, USA, 2019; pp. 1516–1520. [Google Scholar]
- Xu, Y.; Xu, Y.; Lv, T.; Cui, L.; Wei, F.; Wang, G.; Lu, Y.; Florencio, D.; Zhang, C.; Che, W.; et al. Layoutlmv2: Multi-modal pre-training for visually-rich document understanding. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers); Association for Computational Linguistics: Cambridge, MA, USA, 2021; pp. 2579–2591. [Google Scholar]
- Hong, T.; Kim, D.; Ji, M.; Hwang, W.; Nam, D.; Park, S. Bros: A pre-trained language model focusing on text and layout for better key information extraction from documents. Proc. AAAI Conf. Artif. Intell. 2022, 36, 10767–10775. [Google Scholar] [CrossRef]
- Ha, H.T.; Horák, A. Information extraction from scanned invoice images using text analysis and layout features. Signal Process. Image Commun. 2022, 102, 116601. [Google Scholar] [CrossRef]
- Sun, H.; Kuang, Z.; Yue, X.; Lin, C.; Zhang, W. Spatial dual-modality graph reasoning for key information extraction. arXiv 2021, arXiv:2103.14470. [Google Scholar]
- Zhang, H.; Whittaker, E.; Kitagishi, I. Extending TrOCR for text localization-free OCR of full-page scanned receipt images. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2023; pp. 1479–1485. [Google Scholar]
- Shi, B.; Bai, X.; Yao, C. An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition. IEEE Trans. Pattern Anal. Mach. Intell. 2016, 39, 2298–2304. [Google Scholar] [PubMed]
- Baek, J.; Kim, G.; Lee, J.; Park, S.; Han, D.; Yun, S.; Oh, S.J.; Lee, H. What is wrong with scene text recognition model comparisons? dataset and model analysis. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2019; pp. 4715–4723. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Atienza, R. Vision transformer for fast and efficient scene text recognition. In Proceedings of the International Conference on Document Analysis and Recognition; Springer: Cham, Switzerland, 2021; pp. 319–334. [Google Scholar]
- Fang, S.; Xie, H.; Wang, Y.; Mao, Z.; Zhang, Y. Read like humans: Autonomous, bidirectional and iterative language modeling for scene text recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 7098–7107. [Google Scholar]
- Fang, S.; Mao, Z.; Xie, H.; Wang, Y.; Yan, C.; Zhang, Y. Abinet++: Autonomous, bidirectional and iterative language modeling for scene text spotting. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 45, 7123–7141. [Google Scholar] [CrossRef]
- Bautista, D.; Atienza, R. Scene text recognition with permuted autoregressive sequence models. In Proceedings of the European Conference on Computer Vision; Springer: Cham, Switzerland, 2022; pp. 178–196. [Google Scholar]
- Johnson, J.M.; Khoshgoftaar, T.M. Survey on deep learning with class imbalance. J. Big Data 2019, 6, 27. [Google Scholar] [CrossRef]
- Hou, Q.; Zhou, D.; Feng, J. Coordinate attention for efficient mobile network design. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 13713–13722. [Google Scholar]
- Lin, T.Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal loss for dense object detection. In Proceedings of the IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 2980–2988. [Google Scholar]
- LeCun, Y.; Bottou, L.; Bengio, Y.; Haffner, P. Gradient-based learning applied to document recognition. Proc. IEEE 1998, 86, 2278–2324. [Google Scholar] [CrossRef]
- Graves, A.; Fernández, S.; Gomez, F.; Schmidhuber, J. Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks. In Proceedings of the 23rd International Conference on Machine Learning; Association for Computing Machinery: New York, NY, USA, 2006; pp. 369–376. [Google Scholar]
- Liu, Y.; Wang, Y.; Shi, H. A convolutional recurrent neural-network-based machine learning for scene text recognition application. Symmetry 2023, 15, 849. [Google Scholar] [CrossRef]
- Jaderberg, M.; Simonyan, K.; Vedaldi, A.; Zisserman, A. Synthetic data and artificial neural networks for natural scene text recognition. arXiv 2014, arXiv:1406.2227. [Google Scholar]
- Gupta, A.; Vedaldi, A.; Zisserman, A. Synthetic data for text localisation in natural images. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016; pp. 2315–2324. [Google Scholar]
- Mishra, A.; Alahari, K.; Jawahar, C. Scene text recognition using higher order language priors. In Proceedings of the BMVC-British Machine Vision Conference; BMVA: Durham, UK, 2012. [Google Scholar]
- Karatzas, D.; Gomez-Bigorda, L.; Nicolaou, A.; Ghosh, S.; Bagdanov, A.; Iwamura, M.; Matas, J.; Neumann, L.; Chandrasekhar, V.R.; Lu, S.; et al. ICDAR 2015 competition on robust reading. In Proceedings of the 2015 13th International Conference on Document Analysis and Recognition (ICDAR); IEEE: New York, NY, USA, 2015; pp. 1156–1160. [Google Scholar]
- Shi, B.; Yang, M.; Wang, X.; Lyu, P.; Yao, C.; Bai, X. Aster: An attentional scene text recognizer with flexible rectification. IEEE Trans. Pattern Anal. Mach. Intell. 2018, 41, 2035–2048. [Google Scholar] [CrossRef] [PubMed]
- Luo, C.; Jin, L.; Sun, Z. Moran: A multi-object rectified attention network for scene text recognition. Pattern Recognit. 2019, 90, 109–118. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Yu, D.; Li, X.; Zhang, C.; Liu, T.; Han, J.; Liu, J.; Ding, E. Towards accurate scene text recognition with semantic reasoning networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 12113–12122. [Google Scholar]
- Qiao, Z.; Zhou, Y.; Yang, D.; Zhou, Y.; Wang, W. Seed: Semantics enhanced encoder-decoder framework for scene text recognition. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2020; pp. 13528–13537. [Google Scholar]
- Wang, Y.; Xie, H.; Fang, S.; Wang, J.; Zhu, S.; Zhang, Y. From two to one: A new scene text recognizer with visual language modeling network. In Proceedings of the IEEE/CVF International Conference on Computer Vision; IEEE: New York, NY, USA, 2021; pp. 14194–14203. [Google Scholar]
- Wan, Z.; He, M.; Chen, H.; Bai, X.; Yao, C. Textscanner: Reading characters in order for robust scene text recognition. Proc. AAAI Conf. Artif. Intell. 2020, 34, 12120–12127. [Google Scholar] [CrossRef]
- Wei, J.; Zhan, H.; Lu, Y.; Tu, X.; Yin, B.; Liu, C.; Pal, U. Image as a language: Revisiting scene text recognition via balanced, unified and synchronized vision-language reasoning network. Proc. AAAI Conf. Artif. Intell. 2024, 38, 5885–5893. [Google Scholar] [CrossRef]
- Zhao, S.; Quan, R.; Zhu, L.; Yang, Y. Clip4str: A simple baseline for scene text recognition with pre-trained vision-language model. IEEE Trans. Image Process. 2024, 33, 6893–6904. [Google Scholar] [CrossRef]
- He, Y.; Chen, C.; Zhang, J.; Liu, J.; He, F.; Wang, C.; Du, B. Visual semantics allow for textual reasoning better in scene text recognition. Proc. AAAI Conf. Artif. Intell. 2022, 36, 888–896. [Google Scholar] [CrossRef]
- Huang, Y.; Lv, T.; Cui, L.; Lu, Y.; Wei, F. Layoutlmv3: Pre-training for document ai with unified text and image masking. In Proceedings of the 30th ACM International Conference on Multimedia; ACM: New York, NY, USA, 2022; pp. 4083–4091. [Google Scholar]
- Kim, G.; Hong, T.; Yim, M.; Park, J.; Yim, J.; Hwang, W.; Yun, S.; Han, D.; Park, S. Donut: Document understanding transformer without ocr. arXiv 2021, arXiv:2111.15664. [Google Scholar]
- Li, M.; Lv, T.; Chen, J.; Cui, L.; Lu, Y.; Florencio, D.; Zhang, C.; Li, Z.; Wei, F. Trocr: Transformer-based optical character recognition with pre-trained models. Proc. AAAI Conf. Artif. Intell. 2023, 37, 13094–13102. [Google Scholar] [CrossRef]
- Yu, J.M.; Ma, H.J.; Kong, J.L. Receipt Recognition Technology Driven by Multimodal Alignment and Lightweight Sequence Modeling. Electronics 2025, 14, 1717. [Google Scholar] [CrossRef]
- Abdalla, M.; Kasem, M.S.; Mahmoud, M.; Yagoub, B.; Senussi, M.F.; Abdallah, A.; Hun Kang, S.; Kang, H.S. Receiptqa: A question-answering dataset for receipt understanding. Mathematics 2025, 13, 1760. [Google Scholar] [CrossRef]
- Li, D.L.; Lee, S.K.; Liu, Y.T. Printed document layout analysis and optical character recognition system based on deep learning. Sci. Rep. 2025, 15, 23761. [Google Scholar] [CrossRef] [PubMed]
- Singh, L.G.; Middleton, S.E. Tabular context-aware optical character recognition and tabular data reconstruction for historical records: LG Singh, SE Middleton. Int. J. Doc. Anal. Recognit. (IJDAR) 2025, 28, 357–376. [Google Scholar] [CrossRef]
- Heakl, A.; Sohail, M.A.; Ranjan, M.; Elbadry, R.; Ahmad, G.S.; El-Geish, M.; Maher, O.; Shen, Z.; Khan, F.S.; Khan, S. Kitab-bench: A comprehensive multi-domain benchmark for arabic ocr and document understanding. In Findings of the Association for Computational Linguistics: ACL; Association for Computational Linguistics: Vienna, Austria, 2025; pp. 22006–22024. [Google Scholar]
- Wang, S. Development of an automated transformer-based text analysis framework for monitoring fire door defects in buildings. Sci. Rep. 2025, 15, 43910. [Google Scholar] [CrossRef] [PubMed]
- Wang, S. Graph neural network–driven text classification for fire-door defect inspection in pre-completion construction. Sci. Rep. 2025, 15, 44382. [Google Scholar] [CrossRef] [PubMed]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar]
- Hu, J.; Shen, L.; Sun, G. Squeeze-and-excitation networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2018; pp. 7132–7141. [Google Scholar]
- Woo, S.; Park, J.; Lee, J.Y.; Kweon, I.S. Cbam: Convolutional block attention module. In Proceedings of the European Conference on Computer Vision (ECCV); IEEE: New York, NY, USA, 2018; pp. 3–19. [Google Scholar]





| Dataset | Split | Number of Original Receipt Images | Number of Cropped Text Images | LMDB Usage |
|---|---|---|---|---|
| WildReceipt | Training Set | 1267 | 40,363 | Training |
| WildReceipt | Test Set | 472 | 15,201 | Final Evaluation |
| SROIE | Training Set | 626 | 33,618 | Training |
| SROIE | Test Set | 347 | 18,703 | Final Evaluation |
| Model | Params (M) | FLOPs (G) | FPS | Latency (ms/img) |
|---|---|---|---|---|
| PARSeq-tiny | 6.018 | 1.533 | 1108.46 | 0.90 |
| CoordAtt-PARSeq | 6.019 | 1.539 | 1050.82 | 0.95 |
| Model | W.Acc. (IC) ↑ | W.Acc. (Strict) ↑ | NED ↑ | CER ↓ |
|---|---|---|---|---|
| CRNN [12] | 68.43 | 67.80 | 91.00 | 18.00 |
| ViTSTR [15] | 79.12 | 78.62 | 95.18 | 9.86 |
| TRBA [13] | 81.67 | 81.10 | 96.22 | 7.60 |
| ABINet [16] | 81.88 | 81.28 | 96.19 | 7.65 |
| Base (PARSeq-tiny) [18] | 81.88 | 81.40 | 96.44 | 6.94 |
| Ours (CoordAtt-PARSeq) | 86.07 | 85.47 | 97.55 | 5.09 |
| Model | W.Acc. (IC) ↑ | W.Acc. (Strict) ↑ | NED ↑ | CER ↓ |
|---|---|---|---|---|
| CRNN [12] | 64.74 | 63.86 | 87.35 | 27.02 |
| ABINet [16] | 77.81 | 76.92 | 94.21 | 12.19 |
| ViTSTR [15] | 72.22 | 71.35 | 92.25 | 16.63 |
| Base (PARSeq-tiny) [18] | 76.73 | 75.92 | 94.59 | 11.08 |
| TRBA [13] | 78.93 | 78.04 | 94.76 | 11.93 |
| Ours (CoordAtt-PARSeq) | 81.20 | 80.31 | 96.09 | 9.24 |
| Dataset | Variant | W.Acc. (IC) ↑ | W.Acc. (Strict) ↑ | NED ↑ | CER ↓ |
|---|---|---|---|---|---|
| WildReceipt | BASE (PARSeq-tiny) [18] | 81.88 | 81.40 | 96.44 | 6.94 |
| BASE + CoordAtt | 84.21 | 83.61 | 96.98 | 5.91 | |
| BASE + Loss | 85.19 | 84.48 | 97.11 | 5.78 | |
| OURS (CoordAtt-PARSeq) | 86.07 | 85.47 | 97.55 | 5.09 | |
| SROIE | BASE (PARSeq-tiny) [18] | 76.73 | 75.92 | 94.59 | 11.08 |
| BASE + CoordAtt | 78.21 | 77.36 | 95.42 | 9.92 | |
| BASE + Loss | 79.47 | 78.55 | 94.85 | 9.85 | |
| OURS (CoordAtt-PARSeq) | 81.20 | 80.31 | 96.09 | 9.24 |
| W.Acc. (IC) ↑ | W.Acc. (Strict) ↑ | NED ↑ | CER ↓ | |
|---|---|---|---|---|
| (=CE) | 84.21 | 83.61 | 96.98 | 5.91 |
| 85.69 | 85.15 | 97.49 | 5.16 | |
| 85.73 | 85.18 | 97.45 | 5.25 | |
| 86.07 | 85.47 | 97.55 | 5.09 | |
| 85.53 | 84.99 | 97.43 | 5.27 | |
| 85.53 | 84.98 | 97.41 | 5.34 |
| Loss | Class Weight | W.Acc. (IC) ↑ | CER ↓ | |
|---|---|---|---|---|
| CE | None | 0 | 84.21 | 5.91 |
| Weighted CE | Inverse frequency | 0 | 85.14 | 5.47 |
| Focal | None | 2 | 85.56 | 5.33 |
| Weighted Focal | Clipped inverse frequency | 2 | 86.07 | 5.09 |
| Variant | Front-End Module | W.Acc. (IC) ↑ | W.Acc. (Strict) ↑ | NED ↑ | CER ↓ |
|---|---|---|---|---|---|
| Base + Loss | None | 85.19 | 84.48 | 97.11 | 5.78 |
| Shallow Conv | convolutional Stem | 85.03 | 84.36 | 96.99 | 5.86 |
| SE | Squeeze-and-Excitation Stem | 85.61 | 84.95 | 97.32 | 5.40 |
| CBAM | Convolutional Block Attention Module Stem | 85.82 | 85.19 | 97.44 | 5.23 |
| CoordAtt-PARSeq | Coordinate Attention Stem | 86.07 | 85.47 | 97.55 | 5.09 |
| Stratum | Symbol Class | Freq. | CRNN | ABINet | PARSeq-Tiny | Ours |
|---|---|---|---|---|---|---|
| Head (>10% freq.) | Uppercase A–Z | 34.11 | 17.83 | 6.66 | 5.66 | 4.20 |
| Head (>10% freq.) | Lowercase a–z | 31.11 | 22.84 | 10.19 | 9.02 | 6.73 |
| Head (>10% freq.) | Digits 0–9 | 25.25 | 9.97 | 4.13 | 3.94 | 3.07 |
| Body (1–10%) | Decimal point “.” | 3.31 | 8.00 | 5.59 | 4.81 | 4.79 |
| Body (1–10%) | Common punctuation | 2.25 | 51.71 | 21.92 | 26.11 | 11.62 |
| Body (1–10%) | Colon “:” | 1.51 | 18.64 | 5.72 | 6.90 | 4.11 |
| Tail (<1% freq.) | Hyphen “-” | 0.87 | 13.73 | 7.47 | 5.26 | 6.99 |
| Tail (<1% freq.) | Currency symbols | 0.72 | 26.16 | 23.87 | 23.93 | 20.72 |
| Tail (<1% freq.) | Slash “/” | 0.60 | 12.43 | 5.54 | 3.26 | 4.13 |
| Tail (<1% freq.) | Other symbols | 0.27 | 75.34 | 47.94 | 43.47 | 26.39 |
| Overall CER | – | 18.00 | 7.65 | 6.94 | 5.09 |
| Stratum | Symbol Class | Freq. | CRNN | ABINet | PARSeq-Tiny | Ours |
|---|---|---|---|---|---|---|
| Head (>10% freq.) | Uppercase A–Z | 64.08 | 34.47 | 17.18 | 15.57 | 13.19 |
| Head (>10% freq.) | Digits 0–9 | 25.29 | 11.68 | 2.48 | 2.42 | 1.74 |
| Body (1–10%) | Decimal point “.” | 3.41 | 7.42 | 2.82 | 2.11 | 2.00 |
| Body (1–10%) | Common punctuation | 3.33 | 29.86 | 8.57 | 7.82 | 5.32 |
| Body (1–10%) | Colon “:” | 1.83 | 18.88 | 2.82 | 2.72 | 2.07 |
| Tail (<1% freq.) | Hyphen “-” | 0.93 | 8.49 | 3.17 | 1.83 | 1.83 |
| Tail (<1% freq.) | Slash “/” | 0.67 | 25.85 | 6.86 | 5.79 | 5.86 |
| Tail (<1% freq.) | Other symbols | 0.35 | 32.50 | 12.41 | 13.97 | 2.88 |
| Tail (<1% freq.) | Currency symbols | 0.11 | 17.65 | 0.00 | 0.46 | 0.26 |
| Tail (<1% freq.) | Lowercase a–z | 0.01 | 4.90 | 2.90 | 2.70 | 2.32 |
| Overall CER | – | 27.02 | 12.19 | 11.08 | 9.24 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ma, J.; Dai, S.; Li, H.; Jing, H. CoordAtt-PARSeq: A Coordinate-Attention-Stem- and Focal-Loss-Augmented Permuted Autoregressive Transformer for Robust Fiscal Receipt Text Recognition. Appl. Sci. 2026, 16, 7003. https://doi.org/10.3390/app16147003
Ma J, Dai S, Li H, Jing H. CoordAtt-PARSeq: A Coordinate-Attention-Stem- and Focal-Loss-Augmented Permuted Autoregressive Transformer for Robust Fiscal Receipt Text Recognition. Applied Sciences. 2026; 16(14):7003. https://doi.org/10.3390/app16147003
Chicago/Turabian StyleMa, Jiaqi, Shenglin Dai, Hui Li, and Hongjun Jing. 2026. "CoordAtt-PARSeq: A Coordinate-Attention-Stem- and Focal-Loss-Augmented Permuted Autoregressive Transformer for Robust Fiscal Receipt Text Recognition" Applied Sciences 16, no. 14: 7003. https://doi.org/10.3390/app16147003
APA StyleMa, J., Dai, S., Li, H., & Jing, H. (2026). CoordAtt-PARSeq: A Coordinate-Attention-Stem- and Focal-Loss-Augmented Permuted Autoregressive Transformer for Robust Fiscal Receipt Text Recognition. Applied Sciences, 16(14), 7003. https://doi.org/10.3390/app16147003

