Context-Oriented Method for Resolving Lexical Ambiguities in Speech Synthesis for a Low-Resource Language
Abstract
1. Introduction
- -
- A CheWSData sentence corpus wasa created, containing homonyms in various contexts;
- -
- The corpus was marked up by linguists according to meanings, since homographs in the Chechen language differ in the length of vowels or in the presence of diphthongs;
- -
- Three parametric homonymy recognition algorithms were developed, and experimental studies of these algorithms were carried out;
- -
- A software module for recognizing homonyms in Chechen were created and integrated into the Chechen speech synthesis system.
2. Related Work
2.1. Speech Synthesis for Low-Resource Languages
2.2. Word Sense Disambiguation for Low-Resource Languages
- -
- does not rely on pre-trained multilingual models (e.g., XLM-R and mBERT), which often underperform for extremely low-resource languages like Chechen;
- -
- uses a purely context-driven positional weighting scheme that requires only a few hundred labeled sentences per homonym;
- -
- is explicitly designed for integration into a speech synthesis front-end, where the disambiguation result directly selects pronunciation variants.
- The relevance of developing and applying parametric methods for analyzing and resolving lexical ambiguity in low-resource languages, as well as for designing automatic speech synthesis systems, is substantiated using the Chechen language as a case study.
- A corpus of Chechen texts, CheWSData, was compiled, containing 15,035 manually selected sentences derived from the analysis of 5 million annotated words. This corpus reflects the natural frequency of polysemy across various grammatical categories of the Chechen language, exemplified by 100 identified homonyms.
- A method and a set of algorithms for estimating the positional occurrence of unique context words were developed. These enable the unambiguous determination of homonym meanings for each processed word exhibiting lexical ambiguity within a sentence.
- The software implementation of the proposed method made it possible to estimate weight coefficients and the significant range of words within the context of processed homonyms. It also facilitated a comparative analysis of the results against existing methods for word sense disambiguation in low-resource languages.
3. Materials and Methods
3.1. Method for Assessing the Positional Occurrence of Polysemantic Words for Resolving Lexical Ambiguities
3.2. Creation of the CheWSData Dataset
- (1)
- Preparation of the initial data in the form of an array of texts of various genres totaling 5 million words.
- (2)
- Compilation of dictionaries containing lists of homonyms and their possible meanings.
- (3)
- Homonym extraction: a list of homonyms requiring further analysis due to their polysemy was selected from an array of texts. A list of 100 homonyms was prepared.
- (4)
- Indexing and annotation: each homonym from the selected set underwent indexing and annotation within the general text bank to facilitate the subsequent extraction of all usage contexts.
- (5)
- Creation of context databases: for each homonym, a separate sentence database containing all context in which this homonym appears was formed.
- (6)
- Sense classification: The final stage is the classification of sentences within each database B(oj, wi). Each sentence is classified according to the specific meaning in which the homonym is used in that context.
3.3. Three Developed Algorithms for Homonym Sense Recognition
- -
- Test sentence: a sentence potentially containing a homonym whose meaning needs to be determined.
- -
- Sentence databases: a set of reference databases, each containing sentences that include one of the possible meanings of the homonym.
- -
- Word databases: a set of dictionaries or word databases associated with the corresponding sentence databases and used to form context vectors.
- AWEN method (based on Euclidean distance in vector space). This method involves the extraction and vectorization of word tags for each base B(oj, wi), followed by the construction of a numeric vector for the current test sentence TP(wi). Distances are then computed between the vector of TP(wi) and the sentence vectors from all bases B(oj, wi). The final step is selecting the homonym sense associated with the minimum distance roj_wi.
- AWA method (based on weighted matching with indexed tags). This method comprises the preparation of indexed word tags for each base B(oj, wi), vectorization of the test sentence TP(wi) and selection of the homonym sense based on the ratio of the weights Q.
- AWN method, which resolves homonyms based on the number of words neighboring the homonym and the word’s position relative to the homonym, involves the following steps: extracting word tags for each base B(oj, wi), computing the weights of word tags for each base B(oj, wi), and selecting the homonym sense according to the ratio between the weights Q(n_cur, p_cur, oj, wi).
4. Results
4.1. Results of Testing Homonymy Recognition Algorithms
4.2. Perceptual Evaluation
- Ткъа, хӀун де [die~] ас, нана, иштта дакъа [da~k@] кхаьчна-кх сoьга?—элира кӀанта. (It means in English: Well, what should I do, Mom, since such a share has fallen?—Said the boy).
- Цундела азаллера абаде кхаччалц йoлчу хана и хьайн дакъа [dak@] кху лахьти чoхь делхo деза хьан, цуьнга хьайна къинтӀерадалар дoьхуш, тӀаккха бен [bian] цӀанлур дац хьo, я ялсамани дахийта а, я хьайн дегӀана юха хьуo чудиллина иза къoн а деш, хьoьх кхин цкъа а адам дан (сделать) а. (It means in English: Therefore, for a time from eternity to eternity, you love your part in this lump, asking for its forgiveness, and only then will you be able to purify yourself, or enter paradise, or rejuvenate your body, making it human again).
- Вайн искусствехь—муьлхха дакъа [da~k@] ахь схьаэцча а, бен-башха а дац: музыка, театр, суртдиллар, литература—бoхь баьккхина-баьрччехь хетарш, мoттарш. (It means in English: In our art, no matter what field you choose, it doesn’t matter: music, theater, painting, literature—the peak in it is the peak of opinions, illusions).
- Цхьа хӀума дара тӀамна кӀелхьара адмашший, кхин йoлу [yol] садoлу хӀумнашший вoвшех къастoш: oйланаш яр, кхетам дала [da~l] адмашний бен [bian] белла ца хилар. (It means in English: There was something during the war that distinguished humans from animals: that thinking and reason were given by God only to humans).
- Иза а Бек Сараевс цхьа кoг кoша а бахана йoлу [yol] дела дӀахецнера, мел тиша белахь а, "кара ца вoгӀу oбарг" чу леста и бен [bie~] кхузахь латтийта дагахь. (It means in English: She was also released by Bek Saraev because “she had one foot in the grave,” and no matter how old she was, she was only going to stand here so as not to surrender to abrek).
5. Discussion
6. Conclusions and Future Work
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Appendix A
| Designation | Name |
|---|---|
| WN | Number of words being processed |
| wi | Number of the currently processed word |
| w (wi) | Currently processed word |
| W = {w(1), w(2), …, w (wi), …, w(WN)} | Set of words being processed |
| ON(wi) | Number of homonyms for the word w(wi) |
| oj(wi) | Number of the currently processed homonym for the word wi |
| o(oj,wi) | Currently processed homonym oj for the word w(wi) |
| O(wi) = {o(1,wi), o(2,wi), …, o(oj,wi), …, o(ON,wi)} | Set of homonyms being processed for the word wi |
| oj(wi) | Number of the currently processed database of sentences with the homonym o(oj,wi)for the word w(wi) |
| B(oj,wi) | Currently processed database of sentences with the homonym o(oj,wi). |
| B(wi) = {B(1,wi), B(2,wi), …, B(oj,wi), …, B(ON,wi)} } | Set of processed databases of sentences with the homonym o(oj,wi). |
| PN(oj,wi) | Number of sentences being processed with the homonym o(oj,wi). |
| p(oj,wi) | Number of the currently processed sentence with the homonym o(oj,wi) for the word w(wi). |
| P(p,oj,wi) | Currently processed sentence with the homonym o(oj,wi) for the word w(wi). |
| B(oj,wi) = {P(1,oj,wi), P(2,oj,wi), …, P(p,oj,wi), …, P(PN,oj,wi)} | Currently processed database of sentences P(p,oj,wi) with the homonym o(oj,wi). |
| N | Number of tag words closest to the homonym o(oj,wi) in the sentence P(p,oj,wi), which will be used to identify the meaning of the homonym. This value is fixed for all processed words in one version of the speech synthesis system. |
| n(p, oj, wi) | Number of the currently processed tag word in the sentence P(p,oj,wi) (1…N). |
| SP(n, p, oj, wi) | Currently processed tag word at position n in the sentence P(p,oj,wi) in the database B(oj,wi). |
| P(p,oj,wi) = {SP(1,p,oj,wi), SP(2,p,oj,wi), …, SP(n,p,oj,wi), …, SP(N,p,oj,wi)} | Sentence as a set of N tag words, establishing the context of the homonym o(oj,wi). |
| TN(oj, wi) | Number of unique tag words contained in sentences from the database B(oj,wi). |
| t(oj, wi) | Currently processed tag word from the sentence database B(oj,wi). |
| TV(g, t, p, oj, wi) | Positional occurrence score of the currently processed tag word t(oj, wi) at position g(1..N) in sentence P(p,oj,wi). |
| TV(t,p,oj,wi) = (TV(1,t,p,oj,wi), TV(2,t,p,oj,wi), …, TV(g,t,p,oj,wi), …, TV(N,t,p,oj,wi) | Numeric vector of positional occurrence scores of the currently processed tag word t(oj, wi) at position g(1..N) in sentence P(p,oj,wi). |
| STV(oj, wi) | Total numerical vector of positional occurrence scores aggregated over all tag words t = 1..N(oj,wi) in the context of homonym o(oj,wi) (vector of length N). |
| TP(wi) | Test input sentence containing the polysemantic word w(wi). |
| nt(wi) | Index of the currently processed tag word in sentence TP(wi) (1..N(wi)). |
| TSP(nt_cur, wi) | Currently processed tag word in sentence TP(wi). |
| TP(wi) ={TSP(1,wi), TSP(2,wi), …, TSP(nt_cur,wi), …, TSP(N,wi)} | Set of tag words in the input test sentence TP(wi). |
| STPV(oj,wi) | Total numerical vector of positional occurrence scores of tag words {TSP(nt_cur, wi)} in the contextual environment of the polysemantic word wi (vector of length N). |
| Vector of positional occurrence score weights of tag words. | |
| Calculated index of the homonym value for the word w(wi) in the input test sentence TP(wi). |
References
- Zen, H.; Sak, H. Unidirectional long short-term memory recurrent neural network with recurrent output layer for low-latency speech synthesis. In Proceedings of the 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brisbane, QLD, Australia, 19–24 April 2015; IEEE: New York, NY, USA, 2015; pp. 4470–4474. [Google Scholar] [CrossRef] [Scilit]
- Saito, Y.; Takamichi, S.; Saruwatari, H. Statistical parametric speech synthesis incorporating generative adversarial networks. IEEE/ACM Trans. Audio Speech Lang. Process. 2018, 26, 84–96. [Google Scholar] [CrossRef] [Scilit]
- Okamoto, T.; Toda, T.; Shiga, Y.; Kawai, H. Real-Time Neural Text-to-Speech with Sequence-to-Sequence Acoustic Model and WaveGlow or Single Gaussian WaveRNN Vocoders. In Proceedings of the INTERSPEECH, Graz, Austria, 15–19 September 2019; ISCA: White Oak, MN, USA, 2019; pp. 1308–1312. [Google Scholar]
- Wang, Y.; Chen, T.; Zhou, S.; Zhang, F.; Zou, R.; Hu, Q. An improved Wavenet network for multi-step-ahead wind energy forecasting. Energy Convers. Manag. 2023, 278, 116709. [Google Scholar] [CrossRef] [Scilit]
- Pascual, S.; Bhattacharya, G.; Yeh, C.; Pons, J.; Serrà, J. Full-band general audio synthesis with score-based diffusion. In Proceeding of the ICASSP 2023—2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Rhodes Island, Greece, 4–10 June 2023; IEEE: New York, NY, USA, 2023; pp. 1–5. [Google Scholar]
- Weiss, R.J.; Skerry-Ryan, R.J.; Battenberg, E.; Mariooryad, S.; Kingma, D.P. Wave-tacotron: Spectrogram-free end-to-end text-to-speech synthesis. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, 6–11 June 2021; IEEE: New York, NY, USA, 2021; pp. 5679–5683. [Google Scholar]
- Arık, S.Ö.; Chrzanowski, M.; Coates, A.; Diamos, G.; Gibiansky, A.; Kang, Y.; Li, X.; Miller, J.; Ng, A.; Raiman, J.; et al. Deep voice: Real-time neural text-to-speech. In Proceedings of the International Conference on Machine Learning, PMLR 2017, Sydney, NSW, Australia, 6–11 August 2017; JMLR: Norfolk, MA, USA, 2017; pp. 195–204. [Google Scholar]
- Pratap, V.; Tjandra, A.; Shi, B.; Tomasello, P.; Babu, A.; Kundu, S.; Elkahky, A.; Ni, Z.; Vyas, A.; Fazel-Zarandi, M.; et al. Scaling speech technology to 1,000+ languages. J. Mach. Learn. Res. 2024, 25, 97. [Google Scholar]
- Kipyatkova, I.; Kagirov, I.; Dolgushin, M. Use of Pre-Trained Multilingual Models for Karelian Speech Recognition. Inform. Autom. 2025, 24, 604–630. [Google Scholar] [CrossRef] [Scilit]
- Janardana Naidu, G.; Seshashayee, M. Sentiment Analysis Framework for Telugu Text Based on Novel Contrived Passive Aggressive with Fuzzy Weighting Classifier (CPSC-FWC). Inform. Autom. 2024, 23, 39–64. [Google Scholar] [CrossRef] [Scilit]
- Ren, Q.; Bo, Q.; Zhou, C.; Ji, Y.; Wu, N. SRC-IT2: Speech Rate-Controllable Mongolian Emotional Speech Synthesis Based on Improved Tacotron2. Electronics 2025, 14, 3835. [Google Scholar] [CrossRef] [Scilit]
- Ahmad, H.A.; Rashid, T.A. Central Kurdish Text-to-Speech Synthesis with Novel End-to-End Transformer Training. Algorithms 2024, 17, 292. [Google Scholar] [CrossRef] [Scilit]
- Mishev, K.; Karovska Ristovska, A.; Trajanov, D.; Eftimov, T.; Simjanoska, M. MAKEDONKA: Applied deep learning model for text-to-speech synthesis in Macedonian language. Appl. Sci. 2020, 10, 6882. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Wushouer, M.; Tuerhong, G.; Wang, H. Semi-supervised learning for robust emotional speech synthesis with limited data. Appl. Sci. 2023, 13, 5724. [Google Scholar] [CrossRef] [Scilit]
- Ren, Q.D.E.J.; Wang, L.; Zhang, W.; Li, L. Research on a Mongolian Text to Speech Model Based on Ghost and ILPCnet. Appl. Sci. 2024, 14, 625. [Google Scholar] [CrossRef] [Scilit]
- Liu, M.; Jiang, R.; Yang, H. Using Transfer Learning to Realize Low Resource Dungan Language Speech Synthesis. Appl. Sci. 2024, 14, 6336. [Google Scholar] [CrossRef] [Scilit]
- Karibayeva, A.; Karyukin, V.; Tukeyev, U.; Abduali, B.; Amirova, D.; Rakhimova, D.; Aliyev, R.; Shormakova, A. The Development and Experimental Evaluation of a Multilingual Speech Corpus for Low-Resource Turkic Languages. Appl. Sci. 2025, 15, 12880. [Google Scholar] [CrossRef] [Scilit]
- Karibayeva, A.; Karyukin, V.; Abduali, B.; Amirova, D. Speech Recognition and Synthesis Models and Platforms for the Kazakh Language. Information 2025, 16, 879. [Google Scholar] [CrossRef] [Scilit]
- Mir, T.A.; Lawaye, A.A. Word sense disambiguation corpus for Kashmiri. Nat. Lang. Process. 2025, 31, 631–654. [Google Scholar] [CrossRef] [Scilit]
- Aminu, H.; Saidu, I.R.; Odion, P.O. Curation of a polysemous word dataset for word sense disambiguation in Hausa language. J. Stat. Sci. Comput. Intell. 2025, 1, 175–186. [Google Scholar] [CrossRef] [Scilit]
- Sarmah, J.; Kumar Barman, A.; Kumar Sarma, S. Little Wins: Collecting, Preparing, and Publishing Resources for Assamese Word Sense Disambiguation. Comput. Sist. 2025, 29, 1317–1327. [Google Scholar] [CrossRef] [Scilit]
- Bibi, S.; Asghar, S.; Zubair, M. Breaking Barriers in URDU WSD: The Transfer Learning Enriched MAKS Framework. ACM Trans. Asian Low-Resour. Lang. Inf. Process. 2025, 24, 88. [Google Scholar] [CrossRef] [Scilit]
- Patil, C.S.; Patil, V.B. A Multilingual Exploration of Word Sense Disambiguation Using Transformer Models: “Dravidian and Devanagari Languages”. In Recent Advances in Computing Sciences; CRC Press: Boca Raton, FL, USA, 2025; pp. 158–160. [Google Scholar]
- Torunoğlu Selamet, D.; Şentaş, A.; Eryiğit, G. Gamified Crowd-sourcing for Word Sense Disambiguation of Turkish. ACM Trans. Asian Low-Resour. Lang. Inf. Process. 2025, 24, 130. [Google Scholar] [CrossRef] [Scilit]
- Barai, A.; Das, M.; Bhowmick, P.; Dey, U.; Dey, P.; Chowdhury, S. A Comprehensive Analysis of Word Sense Disambiguation in a Regional Language”. In Proceedings of the 2025 International Conference on Inventive Computation Technologies (ICICT), Kirtipur, Nepal, 23–25 April 2025; IEEE: New York, NY, USA, 2025; pp. 1260–1265. [Google Scholar] [CrossRef] [Scilit]
- Huynh, K.T.; Nguyen, D.H.; Nguyen, B.T. ViConBERT: Context-Gloss Aligned Vietnamese Word Embedding for Polysemous and Sense-Aware Representations. arXiv 2025, arXiv:2511.12249. [Google Scholar] [CrossRef] [Scilit]


| Reference | Language | Dataset Properties | Model | Quality | Notes |
|---|---|---|---|---|---|
| [12] | Kurdish | Kurdish speech corpus: 6078 utterances/13.63 h | KTTS | MOS 3.94 | Pre-training with a variational autoencoder (VAE) |
| [18] | Kazakh | Audio–text pair (271 h) | Kazakh TTS2 | DNSMOS 8.79–8.96 | Other models: Whisper, GPT-4 Transcribe, ElevenLabs, OpenAI TTS, Voiser, and TurkicTTS |
| [11] | Mongolian | Mongolian emotional speech corpus (2.25 h/2100 utterances) | SRC-IT2 | MOS 3.7 | Seven emotional categories |
| [13] | Macedonian | 20 h Macedonian high-quality speech audio dataset | TTS MACEDONKA | MOS 3.93 | Deep learning-based method—Deep Voice 3 |
| [14] | English | ESD dataset (350 utterances for each seven emotions) | SMAL-ET2 | MOS 3.77–4.15 EMOS 3.75–4.03 | The acoustic model Tacotron2 and the HiFi-GAN vocoder |
| [15] | Mongolian | The Mongolian speech text dataset (NMLR-Mon2Chs ST) 21,478 audio files/25 h | Ghost-ILPCnet | MOS 4.48 | With the Bang phoneme pre-training model |
| [16] | Dungan | Dungan corpus (4615 sentences/6 h) | MDSD-Tacotron2 + WaveRNN | MOS 4.17 DMOS 4.16 | The TTS framework of Tacotron2 + WaveRNN-based model |
| Reference | Language | Dataset Properties | Model | Quality Metrics | Notes |
|---|---|---|---|---|---|
| [19] | Kashmiri | WSD corpus for Kashmiri (19,854 sentences) | J48 IBk Naive Bayes Dl4jMlpClassifier SVM | F1: 0.64 F1: 0.70 F1: 0.66 F1: 0.70 F1: 0.70 | Corpus for 124 polysemantic words |
| [20] | Hausa | Hausa polysemous WSD dataset (2000 sentence) | XLM-R | F1: 0.79 | Zero-shot and fine-tuned XLM-R models |
| [21] | Assamese | Training dataset for SeAnDa (2000 sentence) | Naive Bayes Classifier | F1: 0.71 | List of 100 polysemantic words ASI |
| [22] | Urdu | Extended Urdu Corpus (EU): 25,000 words | XLM-RoBERTa SVM-RF XLM-RF XLM-SVM | F1: 0.63 F1: 0.68 F1: 0.69 F1: 0.68 | Multi-module software system for WSD MAKS |
| [23] | Dravidian Devanagari | BERT, CTRL | No data | No data | |
| [25] | Bengali | No data | No data | ||
| [26] | Vietnamese | ViConWSD (100,160 words) | ViConBERT | F1: 0.87 | With framework for contrastive learning |
| Statistics | Value |
|---|---|
| Number of synsets | 15,035 |
| Number of words | 152,204 |
| Average words per synset | 3 |
| Average sentences per word | 44.3 |
| Total number of polysemous and homonyms | 6918 (4.5%) |
| Part of Speech | Approximate | Count (%) |
|---|---|---|
| Noun | ~2560 | ~37 |
| Verb | ~3590 | ~52 |
| Adjective | ~550 | ~8 |
| Adverb | ~70 | ~1 |
| Pronoun | ~40 | ~0.6 |
| Others | ~100 | ~1.4 |
| The Homonym “Bala” | |||||
|---|---|---|---|---|---|
| Number of Sentences in the Database | 481 | ||||
| C | N | L | Q | F1 | Accuracy |
| 100 | 4 | 2 | [0.9, 1, 1, 0.9] | 80.33 | 80.05 |
| 6 | 3 | [0.8, 0.9, 1, 1, 0.9, 0.8] | 85.49 | 85.24 | |
| 6 | 2 | [0.9, 1, 1, 0.9, 0.8, 0.7] | 74.21 | 75.23 | |
| 6 | 4 | [0.7, 0.8, 0.9, 1, 1, 0.9] | 83.51 | 82.77 | |
| 8 | 4 | [0.7, 0.8, 0.9, 1, 1, 0.9, 0.8, 0.7] | 79.42 | 79.87 | |
| 8 | 5 | [0.6, 0.7, 0.8, 0.9, 1, 1, 0.9, 0.8] | 72.95 | 72.17 | |
| 8 | 3 | [0.8, 0.9, 1, 1, 0.9, 0.8, 0.7, 0.6] | 82.89 | 82.74 | |
| C | N | L | Q | F1 | Accuracy |
| 300 | 4 | 2 | [0.9, 1, 1, 0.9] | 81.63 | 81.63 |
| 6 | 3 | [0.8, 0.9, 1, 1, 0.9, 0.8] | 86.69 | 86.69 | |
| 6 | 2 | [0.9, 1, 1, 0.9, 0.8, 0.7] | 75.85 | 76.09 | |
| 6 | 4 | [0.7, 0.8, 0.9, 1, 1, 0.9] | 83.55 | 83.56 | |
| 8 | 4 | [0.7, 0.8, 0.9, 1, 1, 0.9, 0.8, 0.7] | 80.4 | 80.62 | |
| 8 | 5 | [0.6, 0.7, 0.8, 0.9, 1, 1, 0.9, 0.8] | 74.91 | 74.91 | |
| 8 | 3 | [0.8, 0.9, 1, 1, 0.9, 0.8, 0.7, 0.6] | 84.52 | 84.54 | |
| C | N | L | Q | F1 | Accuracy |
| 1500 | 4 | 2 | [0.9, 1, 1, 0.9] | 92.96 | 92.97 |
| 6 | 3 | [0.8, 0.9, 1, 1, 0.9, 0.8] | 91.07 | 91.09 | |
| 6 | 2 | [0.9, 1, 1, 0.9, 0.8, 0.7] | 91.56 | 91.60 | |
| 6 | 4 | [0.7, 0.8, 0.9, 1, 1, 0.9] | 90.34 | 90.42 | |
| 8 | 4 | [0.7, 0.8, 0.9, 1, 1, 0.9, 0.8, 0.7] | 89.4 | 89.42 | |
| 8 | 5 | [0.6, 0.7, 0.8, 0.9, 1, 1, 0.9, 0.8] | 90.11 | 90.22 | |
| 8 | 3 | [0.8, 0.9, 1, 1, 0.9, 0.8, 0.7, 0.6] | 91.45 | 91.45 | |
| Homographs | S | K | P | Km | Pm | Pm-P |
|---|---|---|---|---|---|---|
| Бала–give/die | 112 | 29 | 74% | 16 | 86% | 12% |
| елира–gave/died | 75 | 16 | 79% | 7 | 91% | 12% |
| Бен–only/nest | 52 | 20 | 62% | 8 | 85% | 23% |
| Дакъа–part/corpse | 85 | 26 | 70% | 18 | 79% | 9% |
| Дала–God/give | 118 | 38 | 68% | 24 | 80% | 12% |
| Дан–to make/lose | 42 | 15 | 64% | 7 | 83% | 19% |
| Де–a day/kill | 31 | 9 | 71% | 5 | 84% | 13% |
| Йаха–to go/live | 105 | 44 | 58% | 37 | 65% | 7% |
| Йoлу–having/a growing | 113 | 40 | 65% | 27 | 76% | 11% |
| Лар–trail/endure | 21 | 6 | 72% | 4 | 81% | 9% |
| Шун–your/tray | 112 | 21 | 81% | 12 | 89% | 8% |
| Method | Test Dataset | Different POS of Test Datasets | |||
|---|---|---|---|---|---|
| Noun | Verb | Adverb | Pronoun | ||
| Naive Bayes | 0.67 | 0.69 | 0.66 | 0.49 | 0.59 |
| Logistic Regression | 0.55 | 0.56 | 0.60 | 0.52 | 0.54 |
| SVM | 0.70 | 0.56 | 0.58 | 0.52 | 0.54 |
| Random Forest | 0.61 | 0.63 | 0.63 | 0.50 | 0.52 |
| AWEN (Ours) | 0.40 | 0.36 | 0.50 | 0.30 | 0.29 |
| AWA (Ours) | 0.74 | 0.76 | 0.76 | 0.82 | 0.73 |
| AWN (Ours) | 0.78 | 0.82 | 0.83 | 0.85 | 0.71 |
| Language | Dataset Properties | Model | F1 | Accuracy |
|---|---|---|---|---|
| Kashmiri | WSD corpus for Kashmiri (19,854 sentences) | J48 | 0.64 | 0.65 |
| IBk | 0.70 | 0.71 | ||
| Naive Bayes | 0.66 | 0.68 | ||
| SVM | 0.70 | 0.72 | ||
| Dl4jMlpClassifier | 0.70 | 0.72 | ||
| Hausa | Hausa Polysemous WSD dataset (2000 sentence) | XLM-R | 0.79 | 0.83 |
| Assamese | SeAnDa (2000 sentence) | Naive Bayes Classifier | 0.71 | 0.72 |
| Urdu | Corpus of Urdu texts (EU) 25,000 words | XLM-RoBERTa | 0.63 | 0.71 |
| SVM-RF | 0.68 | 0.72 | ||
| XLM-RF | 0.69 | 0.77 | ||
| XLM-SVM | 0.68 | 0.78 | ||
| Vietnamese | ViConWSD (100,160 words) | ViConBERT | 0.87 | 0.88 |
| Chechen | CheWSData (15,035 sentences) | Naive Bayes | 0.67 | 0.69 |
| Logistic Regression | 0.55 | 0.56 | ||
| SVM | 0.70 | 0.70 | ||
| Random Forest | 0.61 | 0.63 | ||
| XLM-R | 0.73 | 0.75 | ||
| mBERT | 0.72 | 0.73 | ||
| AWEN (Ours) | 0.40 | 0.43 | ||
| AWA (Ours) | 0.74 | 0.74 | ||
| AWN (Ours) | 0.78 | 0.80 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Izrailova, E.; Ronzhin, A.; Umarkhadzhiev, S.; Astemirov, A.; Figurek, A.; Sultanov, Z. Context-Oriented Method for Resolving Lexical Ambiguities in Speech Synthesis for a Low-Resource Language. Big Data Cogn. Comput. 2026, 10, 181. https://doi.org/10.3390/bdcc10060181
Izrailova E, Ronzhin A, Umarkhadzhiev S, Astemirov A, Figurek A, Sultanov Z. Context-Oriented Method for Resolving Lexical Ambiguities in Speech Synthesis for a Low-Resource Language. Big Data and Cognitive Computing. 2026; 10(6):181. https://doi.org/10.3390/bdcc10060181
Chicago/Turabian StyleIzrailova, Elisa, Andrey Ronzhin, Salaudin Umarkhadzhiev, Arslanbek Astemirov, Aleksandra Figurek, and Zelimkhan Sultanov. 2026. "Context-Oriented Method for Resolving Lexical Ambiguities in Speech Synthesis for a Low-Resource Language" Big Data and Cognitive Computing 10, no. 6: 181. https://doi.org/10.3390/bdcc10060181
APA StyleIzrailova, E., Ronzhin, A., Umarkhadzhiev, S., Astemirov, A., Figurek, A., & Sultanov, Z. (2026). Context-Oriented Method for Resolving Lexical Ambiguities in Speech Synthesis for a Low-Resource Language. Big Data and Cognitive Computing, 10(6), 181. https://doi.org/10.3390/bdcc10060181

