Next Article in Journal
Hybrid LLM-Assisted Fault Diagnosis Framework for 5G/6G Networks Using Real-World Logs
Previous Article in Journal
A Secure Blockchain-Based MFA Dynamic Mechanism
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Hybrid Methods for Automatic Collocation Extraction in Building a Learners’ Dictionary of Italian

by
Damiano Perri
1,*,
Osvaldo Gervasi
1,
Sergio Tasso
1,
Stefania Spina
2,
Irene Fioravanti
2,
Fabio Zanda
2 and
Luciana Forti
3
1
Department of Math and Computer Science, University of Perugia, Via Luigi Vanvitelli, 1, 06123 Perugia, Umbria, Italy
2
Department of Italian Language, Literature and Arts, University for Foreigners of Perugia, Piazza Fortebraccio, 4, 06123 Perugia, Umbria, Italy
3
Department of Languages, Literature and Modern Cultures, University of Chieti ‘G. d’Annunzio’, Via dei Vestini, 31, 66100 Chieti, Abruzzo, Italy
*
Author to whom correspondence should be addressed.
Computers 2025, 14(12), 552; https://doi.org/10.3390/computers14120552
Submission received: 12 November 2025 / Revised: 9 December 2025 / Accepted: 10 December 2025 / Published: 12 December 2025

Abstract

The automatic construction of learners’ dictionaries requires robust methods for identifying non-literal word combinations, or collocations, which represent a significant challenge for second-language (L2) learners. This paper addresses the critical initial step of accurately extracting collocation candidates from corpora to build a learner’s dictionary for Italian. The adopted method and the implemented application are significant for learning the Italian language. We present a comparative study of three methodologies for identifying these candidates within a 41.7-million-word Italian corpus: a Part-Of-Speech-based approach, a syntactic dependency-based approach, and a novel Hybrid method that integrates both. The analysis yielded 2,097,595 potential collocations. Results indicate that the Hybrid method achieves superior performance in terms of Recall and Benchmark Match, identifying the most significant portion of candidates, 42.35% of the total. We conducted an in-depth analysis to refine the extracted dataset, calculating multiple statistical metrics for each candidate, which are described in detail in the paper. Such analysis allows for the classification of collocations by relevance, difficulty, and frequency of use, forming the basis for the future development of a high-quality, web-based dictionary tailored to the proficiency levels of Italian learners.
Keywords: Artificial Intelligence; Natural Language Processing; language learning; collocations; part-of-speech; syntactic dependency; L2 learners Artificial Intelligence; Natural Language Processing; language learning; collocations; part-of-speech; syntactic dependency; L2 learners

Share and Cite

MDPI and ACS Style

Perri, D.; Gervasi, O.; Tasso, S.; Spina, S.; Fioravanti, I.; Zanda, F.; Forti, L. Hybrid Methods for Automatic Collocation Extraction in Building a Learners’ Dictionary of Italian. Computers 2025, 14, 552. https://doi.org/10.3390/computers14120552

AMA Style

Perri D, Gervasi O, Tasso S, Spina S, Fioravanti I, Zanda F, Forti L. Hybrid Methods for Automatic Collocation Extraction in Building a Learners’ Dictionary of Italian. Computers. 2025; 14(12):552. https://doi.org/10.3390/computers14120552

Chicago/Turabian Style

Perri, Damiano, Osvaldo Gervasi, Sergio Tasso, Stefania Spina, Irene Fioravanti, Fabio Zanda, and Luciana Forti. 2025. "Hybrid Methods for Automatic Collocation Extraction in Building a Learners’ Dictionary of Italian" Computers 14, no. 12: 552. https://doi.org/10.3390/computers14120552

APA Style

Perri, D., Gervasi, O., Tasso, S., Spina, S., Fioravanti, I., Zanda, F., & Forti, L. (2025). Hybrid Methods for Automatic Collocation Extraction in Building a Learners’ Dictionary of Italian. Computers, 14(12), 552. https://doi.org/10.3390/computers14120552

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop