Next Article in Journal
Study on Heat Transfer Performance and Anti-Fouling Mechanism of Ternary Ni-W-P Coating
Next Article in Special Issue
Error Detection for Arabic Text Using Neural Sequence Labeling
Previous Article in Journal
Process Monitoring of Antisolvent Based Crystallization in Low Conductivity Solutions Using Electrical Impedance Spectroscopy and 2-D Electrical Resistance Tomography
Previous Article in Special Issue
Integrated Model for Morphological Analysis and Named Entity Recognition Based on Label Attention Networks in Korean
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

UPC: An Open Word-Sense Annotated Parallel Corpora for Machine Translation Study

School of IT Convergence, University of Ulsan, Ulsan 44610, Korea
*
Authors to whom correspondence should be addressed.
Appl. Sci. 2020, 10(11), 3904; https://doi.org/10.3390/app10113904
Submission received: 28 April 2020 / Revised: 21 May 2020 / Accepted: 30 May 2020 / Published: 4 June 2020
(This article belongs to the Special Issue Machine Learning and Natural Language Processing)

Abstract

Machine translation (MT) has recently attracted much research on various advanced techniques (i.e., statistical-based and deep learning-based) and achieved great results for popular languages. However, the research on it involving low-resource languages such as Korean often suffer from the lack of openly available bilingual language resources. In this research, we built the open extensive parallel corpora for training MT models, named Ulsan parallel corpora (UPC). Currently, UPC contains two parallel corpora consisting of Korean-English and Korean-Vietnamese datasets. The Korean-English dataset has over 969 thousand sentence pairs, and the Korean-Vietnamese parallel corpus consists of over 412 thousand sentence pairs. Furthermore, the high rate of homographs of Korean causes an ambiguous word issue in MT. To address this problem, we developed a powerful word-sense annotation system based on a combination of sub-word conditional probability and knowledge-based methods, named UTagger. We applied UTagger to UPC and used these corpora to train both statistical-based and deep learning-based neural MT systems. The experimental results demonstrated that using UPC, high-quality MT systems (in terms of the Bi-Lingual Evaluation Understudy (BLEU) and Translation Error Rate (TER) score) can be built. Both UPC and UTagger are available for free download and usage.
Keywords: comparative corpus linguistics; Korean-English parallel corpus; word-sense disambiguation; neural machine translation; statistical machine translation comparative corpus linguistics; Korean-English parallel corpus; word-sense disambiguation; neural machine translation; statistical machine translation

Share and Cite

MDPI and ACS Style

Vu, V.-H.; Nguyen, Q.-P.; Shin, J.-C.; Ock, C.-Y. UPC: An Open Word-Sense Annotated Parallel Corpora for Machine Translation Study. Appl. Sci. 2020, 10, 3904. https://doi.org/10.3390/app10113904

AMA Style

Vu V-H, Nguyen Q-P, Shin J-C, Ock C-Y. UPC: An Open Word-Sense Annotated Parallel Corpora for Machine Translation Study. Applied Sciences. 2020; 10(11):3904. https://doi.org/10.3390/app10113904

Chicago/Turabian Style

Vu, Van-Hai, Quang-Phuoc Nguyen, Joon-Choul Shin, and Cheol-Young Ock. 2020. "UPC: An Open Word-Sense Annotated Parallel Corpora for Machine Translation Study" Applied Sciences 10, no. 11: 3904. https://doi.org/10.3390/app10113904

APA Style

Vu, V.-H., Nguyen, Q.-P., Shin, J.-C., & Ock, C.-Y. (2020). UPC: An Open Word-Sense Annotated Parallel Corpora for Machine Translation Study. Applied Sciences, 10(11), 3904. https://doi.org/10.3390/app10113904

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop