Incorporating Linguistic Normalization in Croatian NLP: Evaluating the Impact of Lemmatization on Disinformation Detection Performance
Abstract
1. Introduction
2. Related Work
2.1. Disinformation Detection and Linguistic Processing
2.2. Morphological Normalization and Lemmatization in NLP
2.3. Lemmatization for Croatian and South Slavic Languages
3. Methodology
- A description of the datasets used for these experiments and their linguistic characteristics;
- Preprocessing and lemmatization procedures;
- Modeling setup, including both traditional and transformer-based architectures;
- The evaluation protocol and implementation environment.
3.1. Datasets
3.2. Preprocessing
3.3. Model Setup
3.3.1. Support Vector Machines Model Configuration
3.3.2. Random Forrest Model Configuration
3.3.3. Feedforward Neural Network (NN) Model Configuration
3.3.4. Transformer Model Configuration
3.4. Computational Considerations
3.5. Evaluation Protocol
4. Results and Discussion
4.1. Experimental Setup
4.2. Quantitative Results
4.3. Discussion
5. Conclusions and Future Work
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Ljubi, I.; Grgić, Z.; Vuković, M.; Gledec, G. Detecting Disinformation in Croatian Social Media Comments. Future Internet 2025, 17, 178. [Google Scholar] [CrossRef] [Scilit]
- Vosoughi, S.; Roy, D.; Aral, S. The Spread of True and False News Online. Science 2018, 359, 1146–1151. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Toporkov, O.; Agerri, R. On the Role of Morphological Information for Contextual Lemmatization. Comput. Linguist. 2024, 50, 157–191. [Google Scholar] [CrossRef] [Scilit]
- Singh, G. The Role of Morphology in Natural Language Processing: A Comparative Study of Agglutinative and Fusional Languages. Int. J. Online Humanit. 2024, 10, 31–40. [Google Scholar] [CrossRef] [Scilit]
- Pramana, R.; Debora; Subroto, J.J.; Gunawan, A.A.S.; Anderies. Systematic Literature Review of Stemming and Lemmatization Performance for Sentence Similarity. In Proceedings of the 2022 IEEE 7th International Conference on Information Technology and Digital Applications (ICITDA), Yogyakarta, Indonesia, 4–5 November 2022; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Štefanec, V.; Farkaš, D.; Thakkar, G.; Tadić, M. Building a Large Language Model for Croatian. Proc. Conf. New Trends Transl. Technol. 2024, 2024, 204–209. [Google Scholar] [CrossRef] [Scilit]
- Emil, R.Ș.; Remus, B. A Review of Automatic Fake News Detection: From Traditional Methods to Large Language Models. Future Internet 2025, 17, 435. [Google Scholar] [CrossRef] [Scilit]
- Fields, J.; Chovanec, K.; Madiraju, P. A Survey of Text Classification With Transformers: How Wide? How Large? How Long? How Accurate? How Expensive? How Safe? IEEE Access 2024, 12, 6518–6531. [Google Scholar] [CrossRef] [Scilit]
- Pakray, P.; Gelbukh, A.; Brandypadhyay, S. Natural language processing applications for low-resource languages. Nat. Lang. Process. 2025, 31, 183–197. [Google Scholar] [CrossRef] [Scilit]
- Acs, J.; Hamerlik, E.; Schwartz, R.; Smith, N.A.; Kornai, A. Morphosyntactic probing of multilingual BERT models. Nat. Lang. Eng. 2024, 4, 753–792. [Google Scholar] [CrossRef] [Scilit]
- Tolegen, G.; Toleu, A.; Mussabayev, R. Contrastive Learning for Morphological Disambiguation Using Large Language Models in Low-Resource Settings. Appl. Sci. 2024, 14, 9992. [Google Scholar] [CrossRef] [Scilit]
- Drzik, D.; Kapusta, J. The importance of morphology-aware subword tokenization for NLP tasks in Slovak language modeling. Expert Syst. Appl. 2026, 312, 131492. [Google Scholar] [CrossRef] [Scilit]
- Supriya, M.; Acharya Udupi, D.; Nayak, A.; Srirangapatna Raghavendra, A. Developing a Hybrid Morphological Analyzer for Low-Resource Languages. Appl. Sci. 2025, 15, 5682. [Google Scholar] [CrossRef] [Scilit]
- Nivre, J.; de Marneffe, M.-C.; Ginter, F.; Hajic, J.; Manning, C.D.; Pyysalo, S.; Tyers, F.; Zeman, D. Universal Dependencies v2: An Evergrowing Multilingual Treebank Collection. In Proceedings of the 12th Language Resources and Evaluation Conference, Marseille, France, 11–16 May 2020; European Language Resources Association: Paris, France, 2020; pp. 4034–4043. [Google Scholar]
- Tiedemann, J.; Thottingal, S. The OPUS-MT project: Building open translation services for the world. In Proceedings of the 22nd Annual Conference of the European Association for Machine Translation, Lisbon, Portugal, 3–5 November 2020. [Google Scholar]
- Erjavec, T. MULTEXT-East: Morphosyntactic resources for Central and Eastern European languages. Lang. Resour. Eval. 2012, 46, 131–142. [Google Scholar] [CrossRef] [Scilit]
- Mitrofan, M.; Irimia, E.; Păiş, V. Multimodal Romanian language resources and tools: Challenges and perspectives. Discov. Data 2025, 3, 26. [Google Scholar] [CrossRef] [Scilit]
- Da San Martino, G.; Barron-Cedeno, A.; Wachsmuth, H.; Petrov, R.; Natkov, P.; Herbelot, A.; Zhu, X.; Palmer, A.; Schneider, N.; May, J.; et al. SemEval-2020 task 11: Detection of propaganda techniques in news articles. In Proceedings of the Fourteenth Workshop on Semantic Evaluation 2020, Barcelona, Spain, 12–13 December 2020; pp. 1377–1414. [Google Scholar] [CrossRef] [Scilit]
- Shu, K.; Silva, A.; Wang, S.; Liu, H. Fake News Detection on Social Media: A Data Mining Perspective. SIGKDD Explor. Newsl. 2017, 19, 22–36. [Google Scholar] [CrossRef]
- Shu, K.; Wang, S.; Lee, D.; Liu, H. Mining Disinformation and Fake News: Concepts, Methods, and Recent Advancements. In Disinformation, Misinformation, and Fake News in Social Media. Lecture Notes in Social Networks; Shu, K., Wang, S., Lee, D., Liu, H., Eds.; Springer: Cham, Switzerland, 2020. [Google Scholar] [CrossRef] [Scilit]
- Zhou, X.; Zafarani, R. A Survey of Fake News: Fundamental Theories, Detection Methods, and Opportunities. ACM Comput. Surv. 2020, 53, 109:1–109:40. [Google Scholar] [CrossRef] [Scilit]
- Ljubešić, N.; Lauc, D. BERTić -the transformer language model for Bosnian, Croatian, Montenegrin and Serbian. In Proceedings of the 8th Workshop on Balto-Slavic Natural Language Processing, Online, 20 April 2021; pp. 37–42. [Google Scholar]
- Liang, S.; Mawkanuli, T.; Levow, G.-A. Hybrid Neural-{LLM} Pipeline for Morphological Glossing in Endangered Language Documentation: A Case Study of Jungar Tuvan. In Proceedings of the Fifth Workshop on {NLP} Applications to Field Linguistics, Rabat, Morocco, 29 March 2026. [Google Scholar] [CrossRef] [Scilit]
- Baitenova, L.; Mukhamejanova, G.; Munaitbas, G.; Mambetov, S.; Mukanova, Z. A Multi-Branch Transformer-Enhanced Neural Framework for Joint Morphological Representation Learning. Comput. Mater. Contin. 2026, 88, 64. [Google Scholar] [CrossRef] [Scilit]
- Garcia, A.T.; Przybyla, P.; Wanner, L.; Christodoulopoulos, C.; Chakraborty, T.; Rose, C.; Peng, V. Exploring morphology-aware tokenization: A case study on Spanish language modeling. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, Suzhou, China, 4–9 November 2025; pp. 30505–30518. [Google Scholar] [CrossRef] [Scilit]
- Kondratyuk, D.; Gavenčiak, T.; Straka, M. LemmaTag: Jointly tagging and lemmatizing for morphologically-rich languages with BRNNs. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, 31 October–4 November 2018; pp. 4921–4928. [Google Scholar] [CrossRef] [Scilit]
- Nakov, P.; Martino, G.D.S.; Elsayed, T.; Barrón-Cedeño, A.; Míguez, R.; Shaar, S.; Alam, F.; Haouari, F.; Hasanain, M.; Babulkov, N.; et al. The CLEF-2021 CheckThat! Lab on Detecting Check-Worthy Claims, Previously Fact-Checked Claims, and Fake News. In Advances in Information Retrieval. ECIR 2021. Lecture Notes in Computer Science; Hiemstra, D., Moens, M.F., Mothe, J., Perego, R., Potthast, M., Sebastiani, F., Eds.; Springer: Cham, Switzerland, 2021; Volume 12657. [Google Scholar] [CrossRef] [Scilit]
- Agić, Ž.; Ljubešić, N.; Merkler, D. Lemmatization and Morphosyntactic Tagging of Croatian and Serbian. In Proceedings of the 4th Biennial International Workshop on Balto-Slavic Natural Language Processing BSNLP@ACL, Sofia, Bulgaria, 8–9 August 2013; pp. 48–57. [Google Scholar]
- Ljubešić, N.; Dobrovoljc, K. What does Neural Bring? Analysing Improvements in Morphosyntactic Annotation and Lemmatisation of Slovenian, Croatian and Serbian. In Proceedings of the 7th Workshop on Balto-Slavic Natural Language Processing, Florence, Italy, 2 August 2019. [Google Scholar] [CrossRef] [Scilit]
- Terčon, L.; Ljubešić, N.; Dobrovoljc, K. CLASSLA-Stanza: The Next Step for Linguistic Processing of South Slavic Languages. Contrib. Contemp. Hist. 2024, 65, 109–134. [Google Scholar] [CrossRef] [Scilit]
- Tadić, M.; Fulgosi, S. Building the Croatian Morphological Lexicon. In Proceedings of the EACL Workshop on Morphological Processing of Slavic Languages, Budapest, Hungary, 13 April 2003. [Google Scholar]
- Jongejan, B.; Dalianis, H. Automatic training of lemmatization rules that handle morphological changes in pre-, in- and suffixes alike. In Proceedings of the Joint Conference of the 47th Annual Meeting of the ACL and the 4th International Joint Conference on Natural Language Processing of the AFNLP, Singapore, 2–7 August 2009. [Google Scholar] [CrossRef] [Scilit]
- Sharoff, S.; Umanskaya, E.; Wilson, J. A Frequency Dictionary of Russian: Core Vocabulary for Learners; Routledge Publishing: London, UK, 2013. [Google Scholar] [CrossRef] [Scilit]
- Ljubešić, N.; Fišer, D.; Erjavec, T.; Šulc, A. Offensive Language Dataset of Croatian, English and Slovenian Comments FRENK 1.1, Clarin, 31 October 2025. Available online: https://www.clarin.si/repository/xmlui/handle/11356/1462 (accessed on 23 June 2026).
- Smith, J. Lying in print: The linguistic patterns of deception in the fabricated journalism of Stephen Glass. J. Corpora Discourse Stud. 2026, 10, 31–60. [Google Scholar] [CrossRef] [Scilit]
- Whitty, M.T.; Doherty, S. Enhancing mis- and disinformation detection and understanding its influence: Leveraging communication accommodation theory and information manipulation theory. In Behaviour & Information Technology; Taylor & Francis: London, UK, 2026; pp. 1–20. [Google Scholar] [CrossRef] [Scilit]
- Žabokrtsky, Z.; Ševčikova, M.; Straka, M.; Vidra, J.; Limburská, A. Merging data resources for inflectional and derivational morphology in Czech. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), Portorož, Slovenia, 23–28 May 2016; pp. 1307–1314. [Google Scholar] [CrossRef] [Scilit]
- Forman, G. An extensive empirical study of feature selection metrics for text classification. J. Mach. Learn. Res. 2003, 3, 1289–1305. [Google Scholar] [CrossRef]
- Kibriya, A.M.; Frank, E.; Pfahringer, B.; Holmes, G. Multinomial Naive Bayes for Text Categorization Revisited. In AI 2004: Advances in Artificial Intelligence. AI 2004. Lecture Notes in Computer Science; Webb, G.I., Yu, X., Eds.; Springer: Berlin/Heidelberg, Germany, 2004; Volume 3339. [Google Scholar] [CrossRef] [Scilit]
- Kovačić, B.; Brnjaković, A.; Thakkar, G. Beyond Words: Sentiment Analysis of Croatian Language Attitudes. In Proceedings of the Central European Conference on Information and Intelligent Systems, Varaždin, Croatia, 18–20 September 2024; pp. 217–224. [Google Scholar]
- Bajčetić, L.; Batanović, V.; Samardžić, T. Lemmatizing Serbian and Croatian via String Edit Prediction. In Proceedings of the Conference on Language Technologies & Digital Humanities, Ljubljana, Slovenia, 19–20 September 2024; pp. 6–22. [Google Scholar] [CrossRef]
- Rabus, A.; Ermakov, A.; Ferrazzo, I. Messy data, low-resource languages, and LLMs: Narrative analysis of pre-modern Slavic Lives of Saints. Comput. Humanit. Res. 2026, 2, e13. [Google Scholar] [CrossRef] [Scilit]
- Qi, P.; Zhang, Y.; Zhang, Y.; Bolton, J.; Manning, C.D. Stanza: A Python Natural Language Processing Toolkit for Many Human Languages. In Proceedings of the Association for Computational Linguistics (ACL) System Demonstrations, Seattle, WA, USA, 5–10 July 2020; pp. 101–108. [Google Scholar] [CrossRef] [Scilit]
- Piskorski, J.; Marcinczuk, M.; Yangarber, R. Cross-lingual Named Entity Corpus for Slavic Languages. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, Torino, Italy, 20–25 May 2024; pp. 4143–4157. [Google Scholar] [CrossRef] [Scilit]
- Chawla, N.V.; Bowyer, K.W.; Hall, L.O.; Kegelmeyer, W.P. SMOTE: Synthetic Minority Over-sampling Technique. J. Artif. Intell. Res. 2020, 16, 321–357. [Google Scholar] [CrossRef] [Scilit]
- Haidarh, M.; Mu, C.; Liu, Y.; He, X. Exploring traditional, deep learning and hybrid methods for hyperspectral image classification: A review. J. Inf. Intell. 2025. [Google Scholar] [CrossRef] [Scilit]
- Xu, G.; Qian, M.; Meng, L. Misinformation dissemination on social media: Key research themes and evolutionary paths between 2013 and 2023. Humanit Soc. Sci. Commun. 2025, 12, 1775. [Google Scholar] [CrossRef] [Scilit]
- Ormeño-Arriagada, P.; Puraivan, E.; Kloss, S.; Cofré-Morales, C.; Rodriguez, M. Interpretable Fake News Detection Using Linguistic Indicators Under Imbalanced and Low-Resource Conditions. Appl. Sci. 2026, 16, 5080. [Google Scholar] [CrossRef] [Scilit]




| Study | Language(s) | Task | Morphology Strategy | Model(s) | Dataset |
|---|---|---|---|---|---|
| [3] | 6 languages incl. Czech | Contextual lemmatization | Explicit vs. implicit morphology | XLM-R, mBERT, Morpheus | UniMorph, UD |
| [26] | Czech, German, Arabic, English | Joint tagging + lemmatization | Joint morphology learning | LemmaTag | PDT, TIGER, PADT, EWT corpora |
| [28] | Croatian, Serbian | Lemmatization, MSD tagging | Statistical lemmatization | CST, HunPos | SETIMES.HR, Wikipedia |
| [29] | Slovenian, Croatian, Serbian | MSD tagging, lemmatization | Lexicon-assisted morphology | reldi-tagger, StanfordNLP | ssj500k, hr500k, SETimes |
| [30] | South Slavic languages | Linguistic processing | Lexicon-assisted morphology | CLASSLA-Stanza | Multiple corpora |
| [32] | 10 European languages incl. Slovene | Lemmatization | Affix-based rules | Rule-based | Lexical resources |
| [33] | Russian | Topic modeling | Lemmatization | LDA | Russian Wikipedia |
| This study | Croatian | Disinformation detection | Lemmatization | TF-IDF + SVM, RF, NN, croBERT | Croatian disinformation dataset |
| Model | Final Hyperparameters |
|---|---|
| Linear SVM | class_weight = “balanced”, C ∈ {0.01, 0.1, 1, 10} |
| RBF-SVM | C ∈ {0.5, 1, 5}, γ ∈ {“scale”, 0.01, 0.001} |
| Random Forest | n_estimators = 300, max_depth ∈ {None, 20, 40}, max_features ∈ {“sqrt”, 0.2, 0.5}, min_samples_leaf ∈ {1, 2, 4} |
| Feed-forward Neural Network | hidden_layer_sizes = (256, 128), alpha ∈ {10−5, 10−4, 10−3}, learning_rate_init ∈ {10−4, 10−3} |
| croBERT | optimizer = AdamW, learning rate = 2 × 10−5, batch size = 16, epochs = 3 |
| Data | Model | Accuracy | Precision | Recall | F1 | ROC-AUC |
|---|---|---|---|---|---|---|
| original | RandomForest | 0.9223 | 0.9120 | 0.9533 | 0.9322 | 0.9442 |
| original | SVM | 0.9152 | 0.8982 | 0.9572 | 0.9268 | 0.9315 |
| original | LinearSVM | 0.9089 | 0.8886 | 0.9576 | 0.9218 | 0.9382 |
| original | NeuralNet | 0.8961 | 0.8797 | 0.9436 | 0.9106 | 0.9348 |
| original | croBERT | 0.9205 | 0.9096 | 0.9529 | 0.9308 | 0.9399 |
| lemmatized | RandomForest | 0.9232 | 0.9164 | 0.9496 | 0.9327 | 0.9450 |
| lemmatized | SVM | 0.8570 | 0.8133 | 0.9669 | 0.8835 | 0.9273 |
| lemmatized | LinearSVM | 0.9115 | 0.8921 | 0.9580 | 0.9239 | 0.9422 |
| lemmatized | NeuralNet | 0.8510 | 0.8490 | 0.8931 | 0.8705 | 0.9119 |
| lemmatized | croBERT | 0.9239 | 0.9145 | 0.9533 | 0.9335 | 0.9443 |
| Model | Effect of Lemmatization | p-Value | Cohen’s d | Interpretaion |
|---|---|---|---|---|
| Linear SVM | Small F1 increase | p > 0.05 | d < 0.2 | Negligible practical effect |
| RBF-SVM | F1 decrease | p < 0.01 | d > 0.5 | Medium-to-large negative effect |
| Random Forest | Small F1 increase | p > 0.05 | d < 0.2 | Negligible practical effect |
| Feed-forward NN | F1 decrease | p < 0.01 | d > 0.5 | Medium-to-large negative effect |
| croBERT | Small F1 increase | p > 0.05 | d < 0.2 | Negligible practical effect |
| Representation | Vocabulary Size |
|---|---|
| Original | 27.432 |
| Lemmatized | 16.067 |
| Reduction | 41.43% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Ljubi, I.; Horvat, M.; Gledec, G.; Vukovic, M. Incorporating Linguistic Normalization in Croatian NLP: Evaluating the Impact of Lemmatization on Disinformation Detection Performance. Electronics 2026, 15, 3826. https://doi.org/10.3390/electronics15173826
Ljubi I, Horvat M, Gledec G, Vukovic M. Incorporating Linguistic Normalization in Croatian NLP: Evaluating the Impact of Lemmatization on Disinformation Detection Performance. Electronics. 2026; 15(17):3826. https://doi.org/10.3390/electronics15173826
Chicago/Turabian StyleLjubi, Igor, Marko Horvat, Gordan Gledec, and Marin Vukovic. 2026. "Incorporating Linguistic Normalization in Croatian NLP: Evaluating the Impact of Lemmatization on Disinformation Detection Performance" Electronics 15, no. 17: 3826. https://doi.org/10.3390/electronics15173826
APA StyleLjubi, I., Horvat, M., Gledec, G., & Vukovic, M. (2026). Incorporating Linguistic Normalization in Croatian NLP: Evaluating the Impact of Lemmatization on Disinformation Detection Performance. Electronics, 15(17), 3826. https://doi.org/10.3390/electronics15173826

