Next Article in Journal
A New Finite-Difference Method for Nonlinear Absolute Value Equations
Previous Article in Journal
A Galerkin Finite Element Method for a Nonlocal Parabolic System with Nonlinear Boundary Conditions Arising from the Thermal Explosion Theory
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

A Context-Preserving Tokenization Mismatch Resolution Method for Korean Word Sense Disambiguation Based on the Sejong Corpus and BERT

Department of Software Convergence Engineering, Mokpo National University, Muan 58554, Republic of Korea
Mathematics 2025, 13(5), 864; https://doi.org/10.3390/math13050864
Submission received: 24 January 2025 / Revised: 3 March 2025 / Accepted: 3 March 2025 / Published: 5 March 2025
(This article belongs to the Section E1: Mathematics and Computer Science)

Abstract

The disambiguation of word senses (Word Sense Disambiguation, WSD) plays a crucial role in various natural language processing (NLP) tasks, such as machine translation, sentiment analysis, and information retrieval. Due to the complex morphological structure and polysemy of the Korean language, the meaning of words can change depending on the context, making the WSD problem challenging. Since a single word can have multiple meanings, accurately distinguishing between them is essential for improving the performance of NLP models. Recently, large-scale pre-trained models like BERT and GPT, based on transfer learning, have shown promising results in addressing this issue. However, for languages with complex morphological structures, like Korean, the tokenization mismatch between pre-trained models and fine-tuning data prevents the rich contextual and lexical information learned by the pre-trained models from being fully utilized in downstream tasks. This paper proposes a novel method to address the tokenization mismatch issue during the fine-tuning of Korean WSD, leveraging BERT-based pre-trained models and the Sejong corpus, which has been annotated by language experts. Experimental results using various BERT-based pre-trained models and datasets from the Sejong corpus demonstrate that the proposed method improves performance by approximately 3–5% compared to existing approaches.
Keywords: word sense disambiguation; BERT; Sejong corpus; tokenization; transfer learning word sense disambiguation; BERT; Sejong corpus; tokenization; transfer learning

Share and Cite

MDPI and ACS Style

Jeong, H. A Context-Preserving Tokenization Mismatch Resolution Method for Korean Word Sense Disambiguation Based on the Sejong Corpus and BERT. Mathematics 2025, 13, 864. https://doi.org/10.3390/math13050864

AMA Style

Jeong H. A Context-Preserving Tokenization Mismatch Resolution Method for Korean Word Sense Disambiguation Based on the Sejong Corpus and BERT. Mathematics. 2025; 13(5):864. https://doi.org/10.3390/math13050864

Chicago/Turabian Style

Jeong, Hanjo. 2025. "A Context-Preserving Tokenization Mismatch Resolution Method for Korean Word Sense Disambiguation Based on the Sejong Corpus and BERT" Mathematics 13, no. 5: 864. https://doi.org/10.3390/math13050864

APA Style

Jeong, H. (2025). A Context-Preserving Tokenization Mismatch Resolution Method for Korean Word Sense Disambiguation Based on the Sejong Corpus and BERT. Mathematics, 13(5), 864. https://doi.org/10.3390/math13050864

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop