Next Article in Journal
PEASE: A PUF-Based Efficient Authentication and Session Establishment Protocol for Machine-to-Machine Communication in Industrial IoT
Previous Article in Journal
A New Decentralized PQ Control for Parallel Inverters in Grid-Tied Microgrids Propelled by SMC-Based Buck–Boost Converters
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions

1
Key Laboratory of China’s Ethnic Languages and Information Technology of Ministry of Education, Northwest Minzu University, Lanzhou 730030, China
2
School of Artificial Intelligence, Chongqing University of Education, Chongqing 400065, China
3
LinkDoc Technology, Beijing 100089, China
*
Author to whom correspondence should be addressed.
Electronics 2022, 11(23), 3919; https://doi.org/10.3390/electronics11233919
Submission received: 28 October 2022 / Revised: 23 November 2022 / Accepted: 25 November 2022 / Published: 27 November 2022
(This article belongs to the Topic Computer Vision and Image Processing)

Abstract

The construction of a character dataset is an important part of the research on document analysis and recognition of historical Tibetan documents. The results of character segmentation research in the previous stage are presented by coloring the characters with different color values. On this basis, the characters are annotated, and the character images corresponding to the annotation are extracted to construct a character dataset. The construction of a character dataset is carried out as follows: (1) text annotation of segmented characters is performed; (2) the character image is extracted from the character block based on the real position information; (3) according to the class of annotated text, the extracted character images are classified to construct a preliminary character dataset; (4) data augmentation is used to solve the imbalance of classes and samples in the preliminary dataset; (5) research on character recognition based on the constructed dataset is performed. The experimental results show that under low-resource conditions, this paper solves the challenges in the construction of a historical Uchen Tibetan document character dataset and constructs a 610-class character dataset. This dataset lays the foundation for the character recognition of historical Tibetan documents and provides a reference for the construction of relevant document datasets.
Keywords: historical Tibetan documents; character annotation; character extraction; data augmentation; character recognition historical Tibetan documents; character annotation; character extraction; data augmentation; character recognition

Share and Cite

MDPI and ACS Style

Zhang, C.; Wang, W.; Zhang, G. Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions. Electronics 2022, 11, 3919. https://doi.org/10.3390/electronics11233919

AMA Style

Zhang C, Wang W, Zhang G. Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions. Electronics. 2022; 11(23):3919. https://doi.org/10.3390/electronics11233919

Chicago/Turabian Style

Zhang, Ce, Weilan Wang, and Guowei Zhang. 2022. "Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions" Electronics 11, no. 23: 3919. https://doi.org/10.3390/electronics11233919

APA Style

Zhang, C., Wang, W., & Zhang, G. (2022). Construction of a Character Dataset for Historical Uchen Tibetan Documents under Low-Resource Conditions. Electronics, 11(23), 3919. https://doi.org/10.3390/electronics11233919

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop