Abstract
The main challenge of resource-poor languages—namely, the lack of sufficiently large and linguistically informed datasets for training neural models—is addressed in this paper by developing a dataset generation technology based on a Complete Set of Endings (CSE) morphological model for Turkic languages. Building on this technology, we propose a CSE-Guided Framework for morphology-aware statistical tokenization and neural model segmentation, with Kazakh as a case study. Applying the proposed CSE-guided approach to adapt well-known tokenizers for Kazakh leads to measurable reductions in neural model training time (up to approximately 33%) in our experimental setting, primarily due to shorter tokenized sentence lengths. In addition, we extend the SOTA FEMSeg-CRF architecture by incorporating Kazakh vowel–consonant harmony rules at the embedding generation stage. Within the proposed framework, training on a corpus of CSE-generated wordforms results in the FEMSeg_kaz_v2 model, which is evaluated using intrinsic segmentation metrics. Training on a CSE-segmented sentence corpus yields FEMSeg_kaz_v3, which is further assessed using intrinsic, extrinsic, and external evaluation on a manually prepared gold-standard dataset. The paper presents a CSE-guided framework for morphology-aware tokenization and segmentation for Turkic languages, supported by corpus construction, model extensions, and multi-level evaluation. The proposed CSE-Guided Framework can potentially be adapted for other Turkic languages.
1. Introduction
Tokenization in neural models is a crucial preprocessing step that impacts training efficiency. Although popular tokenizers such as Byte Pair Encoding (BPE), WordPiece, and SentencePiece perform well in most languages, they often exhibit limitations when applied to morphologically rich and low-resource languages.
Turkic languages, aside from Turkish, are considered low-resource languages. Turkic languages have a rich, agglutinative word structure. Kazakh, a Turkic language with an agglutinative structure and vowel/consonant harmony, remains insufficiently covered by existing tokenization strategies.
Modern natural language processing systems for Turkic languages typically rely on multilingual pretrained neural models, whose tokenizers often fail to account for the morphological structure of the language. Tokenizers for these models usually produce segmented subword units that do not correspond to the language’s natural morphemes, which can negatively affect downstream application performance. Therefore, research into fine-tuning pre-trained tokenizers for morphologically rich languages such as Kazakh is highly relevant, as morphology-adapted tokenizers can improve performance in downstream applications.
Tokenization quality can be assessed using entropy-based metrics, such as vocabulary size, token count, character count, token compression, bits/token (theoretical), and bits/token (empirical), which serve as intrinsic indicators of statistical efficiency and regularity of subword segmentation. In morphologically rich languages, entropy metrics are sensitive to how subword units reflect recurrent morphological patterns. Preserving natural morpheme boundaries tends to produce more structured token distributions and improved compression properties. Such intrinsic characteristics are expected to correlate with more stable and efficient neural model training. The impact of subword segmentation on downstream performance is evaluated separately through extrinsic experiments.
Our research focuses on two areas. One is the training of widely used statistical tokenizers (SentencePiece BPE and SentencePiece Unigram) for the Kazakh language on CSE model-based datasets and the evaluation of their intrinsic and extrinsic performance in question-answering tasks. The second area is exploring improvements to the FemSeg-CRF neural network model [1] by incorporating vowel/consonant harmony and training it on a large-volume CSE model-based dataset. It should be noted that tokenizers are based on segmentation. Well-known tokenizers are based on the BPE word segmentation algorithm. For morphologically rich, agglutinative Turkic languages, research into new segmentation models is relevant. It can be used in new tokenizers and in other natural language processing tasks.
Our approach is based on the Complete Set of Endings (CSE) model, which provides a rule-based and decision-table framework for segmenting word forms based on inflectional and derivational morphology. We created several segmented corpora using this morphological CSE model. Both research areas are used to train these CSE-based corpora.
The contribution of this work to statistical tokenizer models is the development of a method for building specialized tokenizers for morphologically rich Turkic languages, using Kazakh as a case study, with the following key components:
- •
- We formalize and implement a morphology-aware segmentation algorithm based on the CSE model, using the Kazakh language as an example.
- •
- We provide a segmented Kazakh corpus of 284,707 sentences, together with a reproducible pipeline for training and evaluating tokenizers.
- •
- We demonstrate indicative improvements in token coverage, segmentation accuracy, and entropy-based efficiency compared to standard tokenizers under controlled experimental settings.
The contribution of this work to neural model segmentation includes:
- •
- The development of a corpus of segmented Kazakh wordforms comprising 2,329,377 wordforms.
- •
- The development of the FEMSeg_kaz neural segmentation model, enhanced by incorporating vowel–consonant harmony features.
- •
- The training of FEMSeg_kaz_v2 on the corpus of 2,329,377 segmented wordforms, achieving a boundary accuracy of 99.7% on intrinsic (in-domain) evaluation.
- •
- The training of FEMSeg_kaz_v3 on a corpus of 50,000 segmented sentences, achieving a boundary accuracy of 95.55% on a manually annotated external test set and an average edit distance of 3.48%, providing indicative evidence of improved segmentation quality.
This work proposes a scalable methodology that can be adapted for creating specialized tokenizers and neural segmentation models for other Turkic languages.
The paper is organized as follows: Section 1 introduces the problem of constructing morphology-aware tokenizers and segmentation; Section 2 introduces related works; Section 3 introduces materials and methods; Section 3.1 introduces the CSE Morphological Model; Section 3.2 introduces SentencePiece Tokenizer Training for Kazakh; Section 3.3 introduces Neural Model-Based Kazakh Segmentation; Section 4 introduces Results of our Research; Section 4.1 introduces Intrinsic Evaluation of Morphology-Aware Kazakh Tokenizers Based on Statistical Models; Section 4.2 introduces Training of Morphology-Aware Neural Models FEMSeg_kaz; Section 4.3—Intrinsic Evaluation of Morphology-Aware Kazakh Segmentation based on Neural Models; Section 4.4—External Estimation of Statistical Tokenizers and Neural Segmentation Models; Section 4.5—Extrinsic Evaluation of Morphology-Aware Kazakh Tokenizers Based on Statistical Models; Section 4.6—Extrinsic Evaluation of Morphology-Aware FEMSeg_kaz_v3 Based on a Neural Model; Section 5 introduces Discussion; and Section 6 introduces Conclusions and Future Work.
2. Related Work
Subword tokenization is a preprocessing component of neural network modeling that enables efficient dictionary compression and the processing of out-of-vocabulary words. Widely used methods, such as Byte Pair Encoding (BPE) [2], WordPiece [3], and SentencePiece [4], offer language-independent segmentation strategies based on the statistical frequency of the generated subwords. These approaches are practical for languages with substantial resources and relatively simple morphology. However, they do not consider the morphologically rich structure of Turkic languages, where morphemes encode various grammatical and semantic functions.
In the context of Kazakh, a morphologically rich Turkic language, existing NLP systems typically rely on multilingual pretrained models such as mT5 [5] and MiniLM [6]. Although these models provide baseline performance across tasks, their tokenizers are not tailored to Kazakh’s morphological complexity. This approach often results in arbitrary subword splits and fragmented morphemes, significantly hindering interpretability and potentially reducing the effectiveness of downstream models.
Morphology-aware tokenization has shown promise in other agglutinative and low-resource languages, including Turkish and Finnish [7], and Turkish and Uyghur [8]. The results of these studies show that incorporating linguistic segmentation into text preprocessing reduces vocabulary size and improves model performance. However, these works do not address segmentation and tokenization for the Kazakh language, nor do they address extending the tools used for segmentation to other languages, that is, how to extend the tools for preparing datasets to related languages. They do not provide scalable corpus-generation pipelines.
An essential issue of morphological alignment, namely how often token boundaries align with morpheme boundaries for languages, and how it affects model training and performance, is investigated in [9].
For Turkic languages, the Complete Set of Endings (CSE) morphology model [10] provides a formal framework for encoding inflectional and derivational morphology. It has been applied to machine translation [11], morphological analysis, and the generation of wordforms [12], but has not yet been integrated into tokenizer training pipelines or evaluated using entropy-based metrics.
The paper [8] studies morphological word segmentation for Turkish and Uyghur in Neural Machine Translation. In this study, the authors used morphological analysis tools for Turkish and Uyghur to perform word segmentation.
The paper [13] proposed MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies, and evaluated it on four languages: English, Russian, Hungarian, and Arabic. The limitations of this paper are the absence of tools for preparing segmented datasets for training MorphBPE and the lack of an extrinsic evaluation of the proposed methodology.
The paper [14] presents an optimal BPE segmentation algorithm based on a dynamic programming method that reduces token counts by 3–5%, compared to greedy segmentation. The authors performed intrinsic and extrinsic experiments. In the intrinsic experiments, four languages were used: English, Finnish, Indonesian, and Turkish. The intrinsic experiments used a special metric, the Token Saving Ratio (TSR), which measures the difference in token sequences between the two compared tokenization methods as a percentage. In intrinsic Turkic-language experiments, the TSR metric ranges from 2.62% to 2.90%. In extrinsic experiments for the Turkish language, an accuracy improvement of results by optimal segmentation is 0.76–0.79%.
In the paper [1], the authors propose two morphological segmentation models for Uyghur and Kazakh: (1) a supervised model—the Feature-Enhanced Morphological Segmentation Model (FEMSeg); and (2) an unsupervised model—Masked Morphological Segmentation (MMSeg). For training, word-level Uyghur and Kazakh datasets were used. The authors conducted extensive experiments with the MMSeg, FEMSeg, and FEMSeg-CRF models. For comparison, the authors used the following models: with MMSeg—Morfessor and BPE; with FEMSeg—BiLSTM, BiGRU; and with FEMSeg-CRF—HMM, CRF, BiLSTM-CRF, BiLSTM-ATT-CRF. The metrics accuracy, recall, and F1 are used for evaluation, adapted for morphological segmentation. The results of the experiments show all metrics averaging 92–93%. However, the authors did not use developed models in extrinsic evaluation for downstream tasks.
The paper [15] continues to apply convolutional neural networks for morpheme segmentation. The author proposes a convolutional neural network architecture with left and right convolutions. The left observes the current symbols and some symbols preceding it, while the right observes the current one and the ones following it. Proposed convolutional neural models used for 4 Mexican languages are Mexicanero, Nahuatl, Wixarika, and Yorem Nokki, and North S’ami, a Finno-Ugric language. Experiment results for the metric micro-averaged (per morpheme boundary) boundary F1 range from 62.8 to 80.6 across different languages.
Our study is built by applying the CSE model to generate a segmented Kazakh dataset for training a SentencePiece tokenizer and a neural segmentation model. We propose and evaluate a CSE-guided framework for morphology-aware segmentation and tokenization in Turkic languages, using Kazakh as a case study. In doing so, it contributes to a reproducible methodology for linguistically informed tokenization and segmentation in low-resource settings.
3. Materials and Methods
3.1. CSE Morphological Model
Kazakh, as a member of the Turkic family, is a language with a highly agglutinative morphological structure, in which wordforms are constructed by the systematic concatenation of affixes that encode grammatical features such as number, case, possession, and person [10]. Standard subword tokenization methods often struggle to capture these regularities, leading to fragmented morphemes and reduced semantic coherence.
To address this problem, we use the Complete Set of Endings (CSE) morphological model [10], which formalizes Kazakh morphology as a rule-governed generative system consisting of stems and a limited set of endings.
All Kazakh affixes are divided into two classes: affixes to nominal stems (nouns, adjectives, numbers) and affixes to verb stems (verbs, participles, gerunds, moods, and voices).
The scheme for inferring endings for each class of affixes is considered separately. However, the following four-step procedure is applied uniformly in all cases:
- •
- The determination of possible combinations of elemental affix-type placements;
- •
- The selection of semantically acceptable placements of basic affix types;
- •
- The enumeration of possible ending variants for each semantically acceptable placement;
- •
- The arrangement of endings into a complete set of endings for a given language.
For nominal stems, affixes are grouped into four grammatical categories: plural (K), possessive (T), case (C), and personal (J).
The formula determines the number of placements [10]:
Ank = n!/(n − k)!
Then, the number of placements will be determined as follows:
A41 = 4!/(4 − 1)! = 4; A42 = 4!/(4 − 2)! = 12; A43 = 4!/(4 − 3)! = 24; A44 = 4!/(4 − 4)! = 24.
Combinatorial analysis yields 64 theoretical affix placements, of which only 15 are linguistically valid, constrained by morphophonemic rules and vowel harmony [10].
For instance, the KT placement (Plural–Possessive) generates 30 distinct endings, derived from six plural and five possessive suffixes (Table 1).
Table 1.
Inferring of endings for placement type КТ (Plural—Possessive).
The complete set of inferred endings for Kazakh is 4679 [10].
Each wordform was segmented into morphemes using the rule-based algorithm aligned with the CSE model. The output format follows the structure: stem@@inflectional_suffixes.
Examples:
- •
- балаларымызға→бала@@лар@@ымыз@@ға
- •
- oқытушыдан→oқытушы@@дан@@быз
- •
- жазғыштарымызға→жазғыш@@тар@@ымыз@@ға
Segmented endings were obtained from a parallel mapping file that presents endings and their morpheme segmentation in separate columns.
3.2. SentencePiece Tokenizer Training for Kazakh
3.2.1. SentencePiece Kazakh Architecture
To train a subword tokenizer aligned with Kazakh morphology, we adopted the SentencePiece framework [2], which supports two model types: Byte Pair Encoding (BPE) and the Unigram Language Model. Both approaches are language-independent and operate directly on raw text without requiring token-level annotations:
- •
- BPE merges frequent character sequences iteratively, producing compact subword units.
- •
- Unigram models the likelihood of subword sequences and selects the most probable segmentation.
For our experiments, we used both models, BPE and Unigram. The BPE model is favored for its deterministic behavior. The Unigram model is favored for its remarkable adaptability to the morphemic structure of words. The vocabulary size was set to 50,000, with optional reduction to 16,000 for lightweight applications. For the experiment presented in this article, a dictionary size of 32,000 was used for both tokenizers.
3.2.2. Morphology-Aware Adaptation of SentencePiece for Kazakh
To align token boundaries with morpheme structure, we trained SentencePiece on a segmented corpus generated using the CSE model. Each wordform was decomposed into a stem and inflectional endings, separated by the delimiter “@@”. This segmentation ensured that frequent morphemes were preserved as atomic units during subword training, reducing token-level unpredictability and improving entropy-based compression.
Training SentencePiece for Kazakh used the kazakh_segmented_corpus, which contains 284,707 sentences, obtained using the CSE morphology model tools [16].
3.2.3. Kazakh Text Tokenization Examples
Table 2 shows examples of tokenization before and after training for the Baseline Tokenizer (mT5) and the Morphology-Aware Kazakh Tokenizer.
Table 2.
Examples of tokenization before and after training.
The morphology-aware tokenizer preserves grammatical morphemes as discrete units. Compared to the baseline tokenizer, it yields more consistent segmentation. For example, the word ‘балаларымызға’ has been grammatically segmented as бала@@лар@@ымыз@@ға. Baseline Tokenizer (mT5) segments words as бал@@алар@@ымыз@@ға, while Sentence-Piece_Uni_kaz segments as бала@@лар@@ымыз@@ға. This shows that the Baseline Tokeniser (mT5) made segmentation errors.
3.3. Neural Model-Based Kazakh Segmentation
In this branch, the FEMSeg-CRF neural segmentation model presented in [1] serves as the basis.
3.3.1. Development of a Neural Model FEMSeg_kaz
The developed neural model FEMSeg_kaz is enhanced by incorporating vowel–consonant harmony features and employs character-level sequence labeling for morphological segmentation.
The character-level sequence labeling is performed using the BMES tagging scheme, which is well-suited to capturing the morphemic structure of Turkic languages. BMES defines four tag types:
- •
- B—Begin (beginning of a morpheme),
- •
- M—Middle (internal part of a morpheme),
- •
- E—End (end of a morpheme),
- •
- S—Single (single-character morpheme).
Under this formulation, each character in a word is assigned a BMES tag.
Example.
Word:
Ылғалдандырады (moisturizes)
CSE partitioning:
ылғал | дан | дыр | а | ды
BMES labels:
ы B
л M
ғ M
а M
л E → morphem “ылғал”
д B
а M
н E → morphem “дан”
д B
ы M
р E → morphem “дыр”
а S → single morph
д B
ы E → morphem “ды”.
Boundary metrics.
How does BMES relate to the Boundary metrics dimension TP (True Positive)/FP (False Positive)/FN (False Negative)/TN (True Negative)?
In FEMSeg_kaz, we train a model to predict the BMES for each character.
Boundary metrics are calculated post hoc based on BMES.
A boundary occurs between morphs: E → B; E → S; S → B; S → S.
There is no boundary if: B → M; M → M; M → E. That is:
- •
- TP—the model correctly predicted the transition “end of morph → beginning of next”;
- •
- FP—the model incorrectly predicted a break (E → B) when it was absent;
- •
- FN—the model missed a real break (connects two morphs);
- •
- TN—the model correctly assumes that there is no continuation of the morph.
3.3.2. Architecture Enhancements in FEMSeg_kaz
FEMSeg_kaz follows the general multi-feature segmentation framework of FEMSeg-CRF, combining CNN, BiLSTM, Transformer and CRF layers [1]. However, our architecture extends the original model in three significant ways:
Vowel/Consonant embedding,
Harmony class embedding,
Stem-boundary heuristic.
Linguistically informed architecture FEMSeg_kaz Pipeline presented in Figure 1.
Figure 1.
Linguistically informed architecture of the proposed FEMSeg_kaz model.
We introduce a linguistically informed architecture for FEMSeg_kaz, including three phonological embedding channels: vowel/consonant class, front/back harmony class, and stem-final boundary indicator.
These features encode the core principles of Kazakh vowel/consonant harmony and morphotactics directly into the model.
When concatenated with character embeddings, they provide strong cues to CNN, BiLSTM, and Transformer layers, allowing the model to resolve allomorphic suffix variants, detect stem boundaries, and generalize to unseen word forms.
The FEMSeg_kaz neural segmentation model is trained in two variants: FEMSeg_kaz_v2 and FEMSeg_kaz_v3. FEMSeg_kaz_v2 is trained on a corpus of 2.33 million Kazakh wordforms automatically generated and segmented using the CSE model. In contrast, FEMSeg_kaz_v3 is trained on an independent corpus of 284,707 real-world Kazakh sentences; while the textual data are independent, the segmentation annotations are produced using the same CSE-based framework.
4. Results
4.1. Intrinsic Evaluation of Morphology-Aware Kazakh Tokenizers Based on Statistical Models
To assess the efficiency and linguistic adequacy of our morphology-aware tokenizer, we conducted an intrinsic evaluation on a shared unsegmented Kazakh corpus (“Chunk_099”; 5000 sentences). We compared our trained SentencePiece BPE and Unigram models against standard tokenizers used in mT5 (SentencePiece) and MiniLM (WordPiece). The analysis focuses on three key indicators of tokenization quality: compression, entropy, and information content—each reflecting different dimensions of structural compactness and statistical predictability. It is worth noting that intrinsic evaluation in our experiments reflects consistency with the CSE morphological model, not human linguistic annotation.
4.1.1. Intrinsic Metric Definitions
We used the following metrics, grounded in information-theoretic principles:
- •
- Vocabulary Size. Number of unique subword units learned by the tokenizer.
- •
- Token Count. Total number of tokens produced from the corpus.
- •
- Character Count. Total number of characters in the original corpus before tokenization.
- •
- Compression Ratio. The compression ratio reflects the average number of characters per token and is calculated as:Compression = Characters/Tokens.Higher values indicate more compact tokenization and shorter sequences.
- •
- Theoretical Bits per Token (BPT). Minimum number of bits required to encode each token under uniform distribution:Bitsₜₕₑₒᵣ = log2(V), where V is the vocabulary size.
- •
- Empirical Bits per Token. Entropy of the token distribution:where pᵢ is the empirical probability of token i. Lower values of H indicate more predictable and efficient token usage.H = −∑ᵢ pᵢ · log2(pᵢ),
- •
- Theoretical Information Content. Total information required to encode the corpus, assuming uniform token distribution:where N is the total number of tokens.Iₜₕₑₒᵣ = N · log2(V),
- •
- Empirical Information Content. Total information based on actual token frequencies:Iₑₘₚᵢᵣ = N · H.Both values are expressed in megabits (Mbit) and reflect the overall encoding cost of the corpus.
4.1.2. Source and Segmented Datasets
For the experiments on statistical tokenizer models, we used a corpus of 300,000 parallel sentences in Kazakh and English, collected from various Kazakh websites as part of a scientific grant project (2018–2020) in our research group. We extracted the Kazakh portion of this parallel corpus and further checked for duplicates and single-word sentences, yielding a corpus of 284,707 sentences.
This Kazakh corpus of 284,707 sentences was processed using a segmentation program [16] based on the model described above, yielding a segmented corpus used to train the BPE_Kaz tokenizer via SentencePiece.
A test corpus of 5000 sentences in Kazakh was used to evaluate the morphology-aware Kazakh tokenizer.
For the experiments on neural tokenizer models, we created a corpus of Kazakh segmented wordforms totaling 2,329,377, generated by a computational vowel/consonant harmony model of Kazakh [12], and a dataset of 50,000 sentences from the 284,707-sentence Kazakh corpus.
4.1.3. Results of Intrinsic Evaluation of Tokenizers
The results of the intrinsic evaluation of tokenizers (vocab size: BPE-32000, Unigram-32000) on the unsegmented corpus are shown in Table 3.
Table 3.
Results of intrinsic evaluation of four tokenizers on Unsegmented Corpus 4999 sentences.
4.1.4. Interpretation
Although all tokenizers were trained with a fixed vocabulary size of 32,000 tokens, intrinsic statistics are computed over the observed token distributions in the evaluation corpus, which explains the smaller vocabulary counts reported in Table 3.
- (1)
- Observed Vocabulary Size (unique tokens in 4999 sentences).
The custom Kazakh tokenizers (BPE_Kaz and Unigram_Kaz) have larger observed vocabularies, reflecting richer coverage of Kazakh morphemes. Multilingual baseline tokenizers (mT5, MiniLM) use smaller observed vocabularies, which forces them to break words more aggressively. Custom Kazakh tokenizers capture more Kazakh-specific morphological patterns.
- (2)
- Token Count.
A higher token count leads to longer sequences and more computational time.
BPE_Kaz and Unigram_Kaz produce approximately the exact token count (113 k), suggesting stable, consistent segmentation behavior. This level of granularity lies between the extremes observed in the multilingual baselines: mT5 splits Kazakh words into many smaller subwords (130 k). At the same time, MiniLM tends to merge multiple morphological units into larger tokens (112 k).
- (3)
- Compression Ratio (characters/tokens).
Higher compression keeps to longer, more meaningful tokens—lower compression results in longer sequences. mT5 has far lower compression, reflecting many micro-tokens. MiniLM achieves the highest compression. BPE_Kaz/Unigram_Kaz strikes a balance.
- (4)
- Empirical Entropy (Bits/token, empirical)
Lower entropy causes tokens to repeat more often and carry less unique information. Lower entropy leads to better token predictability.
mT5 has the lowest value: it predicts Kazakh morphology better.
BPE_Kaz has the highest empirical entropy, meaning that tokens are more unique and less repetitive, on average.
4.1.5. Effects of Morphology-Aware Kazakh Tokenisers Based on Statistical Models
- Fewer tokens → shorter training sequences:
BPE_Kaz: 113,142 and Unigram_Kaz: 113,704 tokens versus 130,605 (mT5). This data indicates that texts are, on average, encoded more compactly and that sequences are processed more efficiently. The neural network requires fewer steps to process the same sentence and less training time.
- 2.
- Energy Savings and Green Impact estimation:
Reducing the number of tokens per corpus means fewer matrix–matrix multiplications in the Transformer.
This is expected to reduce computational and energy costs, → reducing the carbon footprint.
- 3.
- Efficient Corpus Compression:
Compression (characters/tokens):
BPE_Kaz = 3.9157 (better)
Unigram_Kaz = 3.8963
mT5 = 3.3921
MiniLM = 3.9306
This data shows that BPE_Kaz and Unigram_Kaz “compress” the text more efficiently, leading to fewer sequences and less computational time.
- 4.
- More balanced entropy:
The theoretical BPT (BPE_Kaz = 13.0289, Unigram_Kaz = 12.9874) is close to mT5 (12.84) and MiniLM (12.69).
The empirical entropy of mT5 (10.12) is less than that of others. That means that tokens in mT5 are more predictable.
4.1.6. Statistical Analysis of Sentence Length for Tokenizers
Table 4 presents the statistical characteristics of sentence length for four tokenizers. The sentence length in tokens after tokenization significantly affects the neural model’s training time, as the maximum sentence length determines the resulting data matrices. For sentences shorter than the maximum, the matrices undergo padding, filling empty elements with padding symbols.
Table 4.
Statistical characteristics of sentence length for four tokenizers.
Statistical characteristics of sentence length can therefore be used to estimate relative computational costs associated with different tokenizers under identical training settings. The experiment processed data from 4999 Kazakh sentences.
Interpretation.
Mean length is the mean length of the sentence in tokens. Median is the length after which half of all sentences are longer. The 90th percentile is the length of a sentence before which there are 90% of the sentences, and the remaining 10% are longer. The 95th percentile is the length of sentences that are 95% of the data. Min is the minimal length of sentences. Max is the maximum length of sentences.
The 95th percentile of token length determines the size of attention matrices and the actual computational cost. Therefore, a tokenizer with a shorter 95th percentile produces smaller attention matrices, supports larger batch sizes, accelerates training, and reduces energy consumption.
The calculated difference in the sentence-length statistical parameters between the mT5/BPE_Kaz and MiniLM/BPE_Kaz pairs shows that MiniLM’s values are closer to BPE_Kaz’s than those of mT5. A detailed analysis of these intrinsic parameters and their relationships with extrinsic data is provided in paragraph 4.5.5.
4.2. Training of Morphology-Aware Neural Models FEMSeg_kaz
Training Characteristics of Morphology-Aware Neural Models: FEMSeg_kaz_v2, FEMSeg_kaz_v3 (10 k), and FEMSeg_kaz_v3 (50 k) are presented in Table 5.
Table 5.
Training Characteristics of Morphology-Aware Neural Models.
4.3. Intrinsic Evaluation of Morphology-Aware Kazakh Segmentation Based on Neural Models
4.3.1. Evaluation of Morphology-Aware Kazakh Segmentation FEMSeg_kaz_v2
The FEMSeg_kaz_v2 model is trained on a corpus of segmented wordforms containing 2,239,377 words. Screenshots of the training results for two epochs are shown below:
Epoch 01: train_loss = 8.7340, val_token_acc = 0.9982, morph_P = 0.9962, morph_R = 0.9955, morph_F1 = 0.9959.
Epoch 02: train_loss = 10.3484, val_token_acc = 0.9983, morph_P = 0.9958, morph_R = 0.9967, morph_F1 = 0.9962.
Interpretation. (1) val_token_acc = 0.9983 (99.83%)—the percentage of characters for which the model predicts the correct BMES tag. (2) morph_P = 0.9958 (99.58%), Precision—the proportion of predicted morpheme boundaries that are correct. (3) morph_R = 0.9967 (99.67%), Recall—the proportion of actual morpheme boundaries that the model successfully identifies. (4) morph_F1 = 0.9962 (99.62%), F1-score—the harmonic mean of Precision and Recall, the overall measure of segmentation quality.
4.3.2. Comparative Evaluation with FEMSeg-CRF
To compare our results with the FEMSeg-CRF model reported in [1], we conducted morphological boundary evaluation using the same BMES sequence-labeling formulation. The FEMSeg-CRF model achieved an F1-score of 92.84% on Kazakh morphological segmentation when trained on a dataset of 20,000 manually segmented wordforms.
A comparative summary of the FEMSeg-CRF and our FEMSeg_kaz_v2 models is presented in Table 6.
Table 6.
Comparison of FEMSeg-CRF and our FEMSeg_kaz_v2 model.
Our FEMSeg_kaz_v2 model, trained on a substantially larger corpus of 2,329,377 CSE-segmented Kazakh wordforms, achieved a morphological boundary F1-score of 99.62%, with precision and recall values exceeding 99.5%. This result reflects the combined effect of (i) the availability of a large-scale, consistently segmented training corpus generated using the CSE morphological model, and (ii) the use of a linguistically enriched segmentation architecture. Importantly, this intrinsic score measures consistency with the CSE segmentation framework, rather than human-annotated linguistic correctness.
4.4. External Estimation of Statistical Tokenizers and Neural Segmentation Models
External estimation of statistical tokenizers MiniLM Standard WordPiece, mT5 Standard SentencePiece, SentencePiece BPE_Kaz, SentencePiece Unigram_Kaz, and neural models’ segmentation FEMSeg_kaz_v2, FEMSeg_kaz_v3 (10 k), FEMSeg_kaz_v3 (50 k) on the external dataset, a manually prepared gold standard of ≈200 sentences (≈15,000 boundary positions).
4.4.1. External Estimation on Boundary Metrics
For external estimation, we used boundary metrics. Key quality metrics of morphemes labeling are:
TP (true positive)—the model detected a boundary, and Gold did, too;
FP (false positive)—the model detected a boundary, but Gold did not;
FN (false negative)—Gold detected a boundary, but the model did not;
TN (true negative)—neither Gold nor the model detected a boundary.
Precision = TP/(TP + FP)—“If the model detected a boundary, how often was it correct?”
Recall = TP/(TP + FN)—“Did it detect all true morpheme boundaries?”
F1 = (2⋅Precision⋅Recall)/(Precision + Recall)—the harmonic mean of precision and recall.
Results of Boundary Metrics (SP-format) on the external, manually prepared gold-standard dataset (≈200 sentences, ≈15,000 positions) for the considered models are presented in Table 7.
Table 7.
Boundary Metrics (SP-format) on the external manually prepared gold-standard dataset (≈200 sentences, ≈15,000 positions).
The boundary-based evaluation on a manually annotated gold-standard set of approximately 200 sentences (≈15,000 positions) shows that the proposed FEMSeg-v3 models achieve higher boundary F1 scores than standard WordPiece, standard SentencePiece, SentencePiece BPE_Kaz, and Unigram_Kaz tokenizers adapted for Kazakh. In particular, FEMSeg-v3 demonstrates substantially higher precision with competitive recall, indicating more accurate placement of morpheme boundaries. These results provide indicative evidence that incorporating morphology-aware segmentation can improve boundary-level segmentation quality under controlled evaluation conditions.
4.4.2. Levenstein Distance External Estimation
Table 8 presents the study on Levenshtein distance estimation for the statistical and neural segmentation models. Levenshtein distance estimation uses a manually annotated gold-standard set of approximately 200 sentences (≈15,000 positions) for statistical tokenizers MiniLM Standard WordPiece, mT5 Standard SentencePiece, SentencePiece BPE_Kaz, SentencePiece Unigram_Kaz, and neural models’ segmentation FEMSeg_kaz_v2, FEMSeg_kaz_v3 (10 k), FEMSeg_kaz_v3 (50 k).
Table 8.
Average Edit Distance for the statistical and neural segmentation models on a manually annotated gold-standard set of approximately 200 sentences (≈15,000 positions).
The results show that morphology-aware FEMSeg-v3 models substantially reduce edit distance, compared to both statistical tokenizers and earlier neural segmentation variants. At the same time, we note that edit distance reflects surface-level segmentation similarity, and should not be interpreted as a complete measure of linguistic correctness. Therefore, these results are interpreted as indicative evidence of improved segmentation quality and are considered complementary to boundary-based metrics and downstream evaluation.
The trends observed in the average edit distance are consistent with the boundary-based evaluation results. Models with lower edit distance also exhibit higher boundary F1 scores, suggesting that both metrics capture complementary aspects of segmentation quality. While boundary metrics explicitly evaluate morpheme boundary placement, edit distance provides a sequence-level similarity measure. Together, these evaluations provide a coherent, mutually reinforcing assessment of segmentation behavior within the limitations of the available gold-standard data.
4.5. Extrinsic Evaluation of Morphology-Aware Kazakh Tokenizers Based on Statistical Models
The extrinsic evaluation is designed as a controlled diagnostic task to assess the impact of different tokenization strategies on semantic representation quality, rather than as a fully optimized end-to-end question-answering system. The goal of this evaluation is not to achieve state-of-the-art QA performance, but to isolate the contribution of morphology-aware tokenization under identical model architectures and training conditions.
4.5.1. Dataset
The Kazakh QA corpus consists of 5944 manually curated question–answer pairs covering the domain of learning programming for secondary school. The corpus is used exclusively for fine-tuning and evaluation within this study. A held-out test subset of 595 QA pairs is used for extrinsic evaluation to prevent overlap between training and evaluation data. While the corpus is limited in size and domain diversity, it serves as a controlled benchmark for analyzing differences induced by tokenizers.
4.5.2. Configurations of Models and Tokenizers
Three different configurations of models and tokenizers were used for comparative analysis:
Standard tokenizer + MiniLM—the basic multilingual paraphrase model, MiniLM-L12-v2, was used with the official HuggingFace tokenizer. This version was taken as the main baseline.
mT5 tokenizer + MiniLM—the MiniLM encoder was retained, and the multilingual subword tokenizer of the Google/mt5-base model was used for tokenization. This approach used the rich mT5 dictionary and demonstrated the advantages of processing Kazakh text in a multilingual environment.
BPE_Kaz (morphology-aware Kazakh tokenizer) + MiniLM—the SentencePiece BPE_Kaz tokenizer, based on BPE and adapted for Kazakh, was used. This tokenizer decomposes words into roots and suffixes at the morpheme level, converts rare forms into compact units, and enables efficient model learning.
When evaluating different tokenization schemes with a fixed encoder architecture, special care must be taken to ensure tokenizer–embedding compatibility. In our experiments, we do not perform a naïve tokenizer swap. Instead, for each replacement tokenizer, we explicitly rebuild the encoder’s input embedding layer to match the tokenizer vocabulary. The pretrained MiniLM encoder architecture and weights are retained, while the embedding matrix is resized to the new vocabulary size. Embeddings are transferred using token–string–based remapping: pretrained embeddings are copied for tokens shared across vocabularies, whereas embeddings for newly introduced tokens are randomly initialized, according to the original embedding distribution. Special tokens ([PAD], [CLS], [SEP], [MASK]) are explicitly defined and validated in the encoder configuration. After alignment, the model is fine-tuned using the same Sentence-Transformer training protocol to adapt the embedding space to the new tokenization, ensuring valid and comparable extrinsic evaluation results.
4.5.3. Model Training
All models were fine-tuned with the same training parameters. The CosineSimilarityLoss function was used during training. During the fine-tuning, the models were trained for three epochs. In all experiments, identical training hyperparameters were used, including a fixed batch size of 8 and the same maximum sequence length of 128 tokens, ensuring a fair and controlled comparison across tokenizers. Training time was measured as the model’s total end-to-end fine-tuning time, rather than per-step or per-epoch timing. In addition, 100 warm-up steps were introduced, allowing the model to gradually increase the learning rate at the initial stage and maintain learning process stability.
Table 9 presents the training time for the three tokenizers used with the MiniLM model.
Table 9.
Model Training Time for three different tokenizers used with the MiniLM model (four runs).
Table 10 presents the training time variability for the three tokenizers. The training time in seconds is used to calculate statistical parameters.
Table 10.
Training time variability across the three tokenizers.
The observed coefficients of variation (3–8%) are within the expected range for GPU-based training and reflect normal run-to-run variability. Since the measured training-time reductions are substantially larger than this variance, the reported efficiency gains can be considered robust and systematic.
For the extrinsic evaluation of the morphology-aware Kazakh tokenizer, we used the Accuracy and BERTScore. The evaluation metrics for dataset 5944 question–answer pairs are shown in Table 11.
Table 11.
Question–Answer Evaluation Results for morphology-aware Kazakh tokenizer.
Although automatic metrics such as BERTScore show limited discrimination across tokenizers, this behavior is expected in short-answer QA settings, where semantic overlap is high. Therefore, these metrics primarily serve as stability indicators, while relative differences are further examined through expert evaluation.
4.5.4. Expert-Based Evaluation Question–Answer Results
For expert assessment, 50 and 200 questions were manually reviewed.
Expert Evaluation Results of answers (correct or incorrect) for 50 questions:
- •
- Standard tokenizer + MiniLM: 26 correct, 24 incorrect—52% accuracy;
- •
- mT5 tokenizer + MiniLM: 23 correct, 27 incorrect—46% accuracy;
- •
- BPE_Kaz + MiniLM: 27 correct, 23 incorrect—54% accuracy.
Expert Evaluation Results of answers (correct or incorrect) for 200 questions:
- •
- Standard tokenizer + MiniLM: 102 correct, 98 incorrect—51% accuracy;
- •
- mT5 tokenizer + MiniLM: 88 correct, 112 incorrect—44% accuracy;
- •
- BPE_Kaz + MiniLM: 105 correct, 95 incorrect—52.5% accuracy.
Due to the limited number of evaluation samples, these results should be interpreted with caution and are intended to provide diagnostic and qualitative insights, rather than statistically conclusive or generalizable comparisons. The expert assessment indicates a consistent, though preliminary, trend: morphology-aware tokenization is associated with a higher proportion of correct answers under controlled experimental conditions.
In this diagnostic setting, the BPE_Kaz configuration achieves the highest expert accuracy, while the mT5-tokenizer-based configuration demonstrates comparatively lower performance. The relatively low absolute expert accuracy (46–54%) reflects the exploratory nature of the task, the morphological complexity of Kazakh, and the semantic variability of acceptable answers. Accordingly, these results should not be interpreted as evidence of strong downstream QA performance, but rather as indicative signals supporting the potential benefit of morphology-aware tokenization.
4.5.5. Estimation of Energy Saving-Green Impact for Morphology-Aware Kazakh Tokeniser
We estimate the potential computational and environmental benefits of the proposed morphology-aware Kazakh tokenizer using both intrinsic (token statistics-based) and extrinsic (training-time-based) analyses.
Intrinsic estimation.
The relative reduction in token count achieved by the morphology-aware tokenizer is:
vs. mT5 tokenizer: (130,605 − 105,641)/130,605 ≈ 19.1%;
vs. MiniLM tokenizer: (112,711 − 105,641)/112,711 ≈ 6.27%.
Beyond average token counts, the upper tail of the sentence-length distribution plays a decisive role in determining computational cost in Transformer models. In practice, the 95th percentile of token length determines the maximum sequence length, the size of attention matrices, and the amount of padding, thereby directly affecting memory usage, batch size, and training speed. A tokenizer with a shorter 95th-percentile length, therefore, produces smaller attention matrices, supports larger batch sizes, and accelerates training.
For the BPE_Kaz tokenizer, the 95th percentile is 39 tokens, whereas for the mT5 tokenizer, it is 49 tokens, under the standard Transformer assumption that attention complexity scales as . This corresponds to a relative attention-cost ratio of , indicating that the mT5 tokenizer incurs approximately 58% higher attention cost. Equivalently, this implies an expected reduction in attention-related computational cost of approximately 37% when using the morphology-aware tokenizer instead of mT5.
For the MiniLM tokenizer, the corresponding ratio is , implying an expected cost reduction of approximately 14%. These intrinsic estimates should be interpreted as approximate lower-bound indicators, rather than exact predictors of training time.
Extrinsic estimation.
Empirical training-time measurements confirm a substantial reduction in wall-clock training time when using the morphology-aware tokenizer (averaged over four runs):
vs. mT5 tokenizer: (8964 s – 5707 s)/8964 s ≈ 36.0%;
vs. MiniLM tokenizer: (8124 s – 5707 s)/8124 s ≈ 29%.
The close correspondence between the 95th-percentile-based theoretical estimate and the observed training-time reduction indicates that upper-tail sequence lengths explain real-world computational savings more effectively than average token statistics. Because Transformer costs grow as a combination of linear and quadratic functions of sequence length, reductions in the longest sequences yield disproportionately large gains in practice.
While the intrinsic token-count reduction is approximately 19% relative to mT5 and 6% relative to MiniLM, and the 95th-percentile-based estimates suggest reductions of approximately 37% and 14%, respectively, the observed training-time savings are 36% and 29%, respectively. The larger discrepancy observed with MiniLM suggests that additional architectural and implementation factors—such as subword-fragmentation patterns, padding efficiency, and kernel-level optimization—also contribute to the measured speedup.
Finally, although energy consumption and carbon emissions were not directly measured, training time is commonly used as a proxy for estimating computational and energy costs under fixed hardware conditions. Since all experiments were conducted using identical hardware and software configurations, the observed reduction in training time suggests a proportional reduction in energy consumption. Accordingly, the reported energy and carbon-footprint reductions should be interpreted as hypothesized environmental benefits inferred from proxy measurements, rather than as direct measurements.
4.6. Extrinsic Evaluation of Morphology-Aware FEMSeg_kaz_v3 Based on Neural Model
This subsection compares the extrinsic evaluation results of the FEMSeg + MiniLM and BPE_Kaz + MiniLM models for morphological segmentation (Table 12). The performance of the BPE_Kaz + MiniLM baseline and the FEMSeg + MiniLM model was evaluated on a held-out TEST set consisting of 595 question–answer pairs. The observed differences are primarily associated with segmentation strategy under identical model architectures and training conditions, although stochastic training effects and dataset limitations may also contribute.
Table 12.
Evaluation results of the FEMSeg_kaz_v3 + MiniLM and BPE_Kaz + MiniLM.
Exact Match Accuracy (Accuracy@1 Exact):
FemSeg_kaz_v3 + MiniLM—0.0857,
BPE_Kaz + MiniLM—0.0840.
Both models showed similar results in verbatim retrieval of the reference response. This low score indicates that Kazakh QA retrieval systems tend to reconstruct responses, rather than directly reproduce them semantically. Therefore, the Exact Match metric is not considered a key metric in this study.
Semantic accuracy (Accuracy@1, Semantic Cosine ≥ 0.85):
FemSeg_kaz_v3 + MiniLM—0.2958,
BPE_Kaz + MiniLM—0.2235.
The FemSeg_kaz_v3 + MiniLM configuration achieves a 32% higher score in semantic assessment under the Cosine Similarity ≥0.85 criterion. This result indicates that morphology-aware segmentation is associated with improved semantic similarity estimates in the embedding space
Given the agglutinative nature of the Kazakh language, such tokenization is likely to facilitate a more consistent representation of complex morphemes. However, the observed improvement should be interpreted as the combined effect of segmentation quality and model-specific factors, including stochastic training dynamics and random initialization-induced variance.
Semantic completeness (Accuracy@1, BERTScore F1 ≥ 0.85):
FemSeg_kaz_v3 + MiniLM—0.8639,
BPE_Kaz + MiniLM—0.8118.
The FemSeg_kaz_v3 configuration also led (+5.2%) in the BERTScore-F1 ≥0.85 criterion. This indicates that the model can return a more detailed, contextually relevant answer. BERTScore provides a more linguistically informed evaluation by accounting for semantic similarity at the embedding level, making the observed difference meaningful.
Average Cosine Similarity:
FemSeg_kaz_v3 + MiniLM—0.6064,
BPE_Kaz + MiniLM—0.5549.
Overall, embeddings produced by the FemSeg_kaz_v3 configuration exhibit a higher cosine similarity, by 0.0515. This indicates that the model tends to retrieve semantically more similar answers to the test questions, since morphological segmentation better preserves the internal structure of words and improves the quality of the embeddings.
Evaluation of the FEMSeg_kaz_v3 + MiniLM and BPE_Kaz + MiniLM on the BERTScore (Precision, Recall, F1) is shown in Table 13.
Table 13.
Comparison of the FEMSeg_kaz_v3 + MiniLM and BPE_Kaz + MiniLM on the BERTScore metrics.
FemSeg_kaz_v3 consistently outperforms on all three metrics. The observed differences are primarily associated with segmentation strategy under identical model architectures and training conditions, although stochastic training effects and dataset limitations may also contribute.
Overall, extrinsic experiments suggest that morphology-aware tokenization and segmentation can positively influence semantic representation quality in Kazakh under controlled conditions. However, the results should be interpreted as indicative rather than conclusive, given the limited scope of the QA corpus and evaluation protocol.
5. Discussion
5.1. Why Morphological Segmentation Improves Tokenization
Morphologically rich languages such as Kazakh pose specific challenges for subword tokenization. Standard statistical tokenization methods often fragment linguistically meaningful morphemes into arbitrary subword units, leading to sparse representations and misalignment between token boundaries and grammatical structure. This effect is particularly pronounced in low-resource settings, where inefficient tokenization can exacerbate data sparsity and training instability.
By incorporating morphology-aware segmentation—specifically, through the Complete Set of Endings (CSE) model—token boundaries are more closely aligned with inflectional and derivational morphemes. This alignment results in several observable effects: (i) a reduction in the total number of tokens through the use of fewer, but more informative, units; (ii) improved character-to-token compression; (iii) lower empirical entropy, reflecting more regular token distributions; and (iv) increased interpretability, as tokens correspond more directly to stems, suffixes, and endings.
Such properties are expected to facilitate neural architectures’ modeling of syntactic and semantic regularities, particularly in resource-constrained scenarios where efficient use of training data is critical. Within this context, the FEMSeg_kaz_v2 and FEMSeg_kaz_v3 models are designed to provide high-precision, rule-consistent segmentation based on the CSE framework, making them well suited for morphology-aware preprocessing in downstream neural modeling for Turkic languages.
5.2. Limitations of Length-Based Intrinsic Metrics
The observed discrepancy between intrinsic length-based statistics and empirical training-time reductions—most notably for MiniLM—highlights the limitations of relying solely on sentence-length measures to estimate computational efficiency. Metrics such as average sequence length and upper percentiles capture coarse properties of tokenized data, but do not fully reflect the internal structure of subword sequences.
In agglutinative languages, WordPiece-based tokenization tends to produce stronger intra-word fragmentation, increasing the number of subword units per word without necessarily causing a proportional increase in overall sentence length. This structural effect influences attention locality, padding efficiency, and memory-access patterns during training, yet remains largely invisible to standard length-based statistics.
These findings suggest that additional intrinsic metrics—such as average subwords per word, continuation–token ratio, or padding efficiency—may be required to better capture tokenizer-induced computational effects. Incorporating such structure-aware metrics represents a promising direction for future analysis, particularly for explaining tokenizer–architecture interactions in models where empirical training-time reductions are not fully predicted by length statistics alone.
5.3. Adaptability to Other Turkic Languages
The proposed CSE-based segmentation framework is designed to be adaptable to other Turkic languages, many of which share key linguistic characteristics, including agglutinative morphology, productive suffixation, vowel harmony systems, and similar affix placement rules. Languages such as Azerbaijani, Kyrgyz, Turkmen, Turkish, Uzbek, Tatar, and Karakalpak exhibit structural similarities that make them suitable candidates for the application of the same formal segmentation principles.
With appropriate language-specific stem lexicons and suffix inventories, the CSE-guided segmentation pipeline may be reused across different Turkic languages, with limited adaptation. This suggests the feasibility of developing unified morphology-aware tokenization strategies for the Turkic language family, enabling cross-lingual comparisons and supporting multilingual NLP systems for under-resourced languages. A systematic empirical validation across multiple Turkic languages is left for future work.
5.4. Integration Challenges with Pretrained Neural Models
Despite the benefits observed in intrinsic and extrinsic evaluations, integrating morphology-aware tokenizers with pretrained neural models presents non-trivial challenges. Pretrained models are typically optimized for their original tokenization schemes, and deviations from these schemes can introduce mismatches between the tokenizer output and the model’s learned representations.
These compatibility issues may affect convergence behavior and downstream performance, and warrant further investigation. A detailed analysis of tokenizer–model co-adaptation, including retraining or partial re-initialization strategies, lies beyond the scope of the present study, and represents an important direction for future research.
6. Conclusions and Future Work
The main challenge of resource-poor languages—namely, the lack of sufficiently large and linguistically informed datasets for training neural models—is addressed in this paper through the development of a dataset generation and segmentation technology based on a CSE morphological model for Turkic languages. Using Kazakh as a case study, we demonstrate that training well-known tokenizers, such as SentencePiece, on morphologically segmented corpora can yield measurable efficiency gains when linguistic knowledge is explicitly incorporated.
In our experimental setting, morphology-aware tokenization reduces the average training time of neural models by up to approximately 33%, primarily by lowering sentence length in tokens through improved character-level compression. This effect is achieved by segmenting Kazakh texts using universal CSE-based morphological segmentation programs.
In addition, we extend the SOTA FEMSeg-CRF architecture by incorporating Kazakh vowel–consonant harmony rules at the embedding generation stage and by training on CSE-segmented corpora, resulting in the FEMSeg_kaz models. Both intrinsic and extrinsic evaluations were conducted for the proposed morphology-aware tokenizers and neural segmentation models. The observed reductions in training time—up to approximately 1.6× in wall-clock measurements—are consistent across repeated runs.
External evaluation experiments were conducted on a manually prepared gold-standard dataset. While limited in size, these results provide indicative evidence that the morphology-aware FEMSeg_kaz_v3 model trained on 50,000 CSE-segmented sentences achieves higher segmentation quality than the considered baseline tokenizers and segmentation models.
The results further suggest that morphology-aware SentencePiece BPE_Kaz and Unigram_Kaz tokenizers offer clear advantages over standard subword models for Kazakh. Reducing the number of tokens for a fixed corpus size leads to shorter input sequences and, consequently, reduced training time and computational cost for Transformer-based models. Under fixed hardware conditions, these efficiency gains are expected to correlate with lower energy consumption; however, the environmental impact should be interpreted as inferred, rather than directly measured.
Additionally, intrinsic analysis shows higher empirical token entropy and more uniform vocabulary usage, indicating improved morphological coverage and reduced sparsity. Taken together, these findings suggest that morphology-aware tokenization can simultaneously improve training efficiency and the statistical properties of text representations for morphologically rich languages, with positive implications for downstream tasks.
The main scientific contributions of this work include:
- (i)
- The development of a CSE-guided framework for morphology-aware statistical tokenization and neural segmentation for Turkic languages (Kazakh case);
- (ii)
- The generation of a large-scale CSE-segmented Kazakh corpus;
- (iii)
- The training of Kazakh-specific SentencePiece tokenizers on this corpus;
- (iv)
- The extension of the FEMSeg-CRF model with phonological harmony constraints;
- (v)
- A comprehensive intrinsic, extrinsic, and external evaluation of the resulting models.
As future work, we plan to extend this framework to other Turkic languages, including Azerbaijani, Kyrgyz, Turkmen, Turkish, and Uzbek, to construct a larger manually annotated gold-standard corpus for Kazakh morphological segmentation and to investigate how different segmentation strategies and dataset characteristics influence neural model performance.
Author Contributions
U.T.: conceptualization, methodology—problem statement, development of the idea of a morphology-aware tokenizer and comparative experiment; data curation—corpora preparation, morphological segmentation, creation of correspondence tables for the tokenizer; software—tokenizer implementation, writing scripts for analysis; validation—conducting experiments (intrinsic evaluation), checking the correctness of calculations; writing—review and editing, supervision, project administration, funding acquisition. B.R.: data curation—corpora preparation and morphological segmentation for the extrinsic experiment; software—for the extrinsic experiment, training a QA model with the tokenized data, writing scripts for analysis; validation—conducting experiments (extrinsic evaluation), checking the correctness of calculations; writing—review and editing (part of the extrinsic experiment). All authors have read and agreed to the published version of the manuscript.
Funding
This research was supported by the Ministry of Science and Higher Education of the Republic of Kazakhstan, within the framework of the scientific project “Study of neural models for the formation of transcripts of speech and minutes of meetings in Turkic languages” (Grant No. AP23487816).
Institutional Review Board Statement
Not applicable.
Data Availability Statement
The software, preprocessing scripts, and reproducibility instructions are publicly available at https://github.com/walsher46/Kazakh-tokenizer/ (accessed on 25 January 2026). Due to GitHub (version, commit e15afca) storage limitations, only partial versions of the corpora are included in the repository. The complete datasets are not redistributed directly. However, detailed data access instructions, data format specifications, and representative samples are provided in the repository. The software is released under the MIT License, and the datasets are distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- Abudouwaili, G.; Ruzmamat, S.; Abiderexiti, K.; Wu, B.; Wumaier, A.A. Benchmark for Morphological Segmentation in Uyghur and Kazakh. Appl. Sci. 2024, 14, 5369. [Google Scholar] [CrossRef] [Scilit]
- Sennrich, R.; Haddow, B.; Birch, A. Neural Machine Translation of Rare Words with Subword Units. In Proceedings of the 54th Annual Meeting of The Association for Computational Linguistics, Berlin, Germany, 7–12 August 2016; pp. 1715–1725. [Google Scholar]
- Wu, Y.; Schuster, M.; Chen, Z.; Le, Q.V.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; et al. Google’s Neural Machine Translation System: Bridging the Gap Between Human and Machine Translation. arXiv 2016, arXiv:1609.08144. [Google Scholar]
- Kudo, T.; Richardson, J. SentencePiece: A Simple and Language-Independent Subword Tokenizer and Detokenizer for Neural Text Processing. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Brussels, Belgium, 31 October–4 November 2018; pp. 66–71. [Google Scholar]
- Xue, L.; Constant, N.; Roberts, A.; Kale, M.; Al-Rfou, R.; Siddhant, A.; Barua, A.; Raffel, C. mT5: A Massively Multilingual Pre-trained Text-to-Text Transformer. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Online, 6–11 June 2021. [Google Scholar]
- Wang, W.; Wei, F.; Dong, L. MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-trained Transformers. In NeurIPS; Neural Information Processing Systems Foundation, Inc.: San Diego, CA, USA, 2020. [Google Scholar]
- Hu, J.F. Tokenization Strategies for Low-Resource Agglutinative Languages in Word2Vec: Case Study on Turkish and Finnish. arXiv 2025, arXiv:2509.14238. [Google Scholar] [CrossRef] [Scilit]
- Pan, Y.; Li, X.; Yang, Y.; Dong, R. Morphological Word Segmentation on Agglutinative Languages for Neural Machine Translation. arXiv 2020, arXiv:2001.01589. [Google Scholar] [CrossRef] [Scilit]
- Arnett, C.; Bergen, B. Why do language models perform worse for morphologically complex languages? In Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, United Arab Emirates, 19–24 January 2025; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025; pp. 6607–6623. Available online: https://aclanthology.org/2025.coling-main.441/ (accessed on 8 December 2025).
- Tukeyev, U.A. A New Computational Model for Turkic Languages’ Morphology and Processing. J. Probl. Comput. Sci. Inf. Technol. 2023, 1. [Google Scholar] [CrossRef] [Scilit]
- Tukeyev, U.; Karibayeva, A.; Zhumanov, Z.H. Morphological Segmentation Method for Turkic Language Neural Machine Translation. Cogent Eng. 2020, 7, 1856500. [Google Scholar] [CrossRef] [Scilit]
- Tukeyev, U. Computational models of morphology and vowel/consonant harmony of the Kazakh language and their use for neural models. In Proceedings of the Materials of the International Scientific and Theoretical Conference Artificial Intelligence in the Space of Contemporary Art: Problems and Prospects, Baku, Azerbaijan, 11 April 2025; pp. 12–19. [Google Scholar] [CrossRef] [Scilit]
- Asgari, E.; El Kheir, Y.; Sadraei Javaheri, M.A. MorphBPE: A Morpho-Aware Tokenizer Bridging Linguistic Complexity for Efficient LLM Training Across Morphologies. arXiv 2025, arXiv:2502.00894. [Google Scholar] [CrossRef] [Scilit]
- Raj, B.; Suri, G.; Dewangan, V.; Sonavane, R. When Every Token Counts: Optimal Segmentation for Low-Resource Language Models. In Proceedings of the First Workshop on Language Models for Low-Resource Languages, Abu Dhabi, United Arab Emirates, 31 December 2025; Association for Computational Linguistics: Stroudsburg, PA, USA, 2025. [Google Scholar]
- Sorokin, A. Convolutional neural networks for low-resource morpheme segmentation: Baseline or state-of-the-art? In Proceedings of the 16th Workshop on Computational Research in Phonetics, Phonology, and Morphology, Florence, Italy, 2 August 2019; Association for Computational Linguistics: Stroudsburg, PA, USA, 2019; pp. 154–159. [Google Scholar]
- Tukeyev, U.; Karibayeva, A.; Turganbayeva, A.; Amirova, D. Universal Programs for Stemming, Segmentation, Morphological Analysis of Turkic Words. In Computational Collective Intelligence, Proceedings of the 13th International Conference, ICCCI 2021, Rhodes, Greece, 29 September–1 October 2021; Nguyen, N.T., Iliadis, L., Maglogiannis, I., Trawiński, B., Eds.; Lecture Notes in Computer Science; Springer: Cham, Switzerland, 2021; Volume 12876. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
