Next Article in Journal
SIMADL: Simulated Activities of Daily Living Dataset
Previous Article in Journal
A Data Set of Portuguese Traditional Recipes Based on Published Cookery Books
Article Menu

Export Article

Open AccessArticle

Associative Root–Pattern Data and Distribution in Arabic Morphology

1
Department of Computer Science, University of Petra, 11196 Amman, Jordan
2
Arabic Textware, 11181 Amman, Jordan
3
Brown University, Providence, RI 02912, USA
*
Author to whom correspondence should be addressed.
Received: 28 January 2018 / Revised: 18 March 2018 / Accepted: 26 March 2018 / Published: 29 March 2018
Full-Text   |   PDF [1052 KB, uploaded 3 May 2018]   |  

Abstract

This paper intends to present a large-scale dataset for Arabic morphology from a cognitive point of view considering the uniqueness of the root–pattern phenomenon. The center of attention is focused on studying this singularity in terms of estimating associative relationships between roots as a higher level of abstraction for words meaning, and all their potential occurrences with multiple morpho-phonetic patterns. A major advantage of this approach resides in providing a novel balanced large-scale language resource, which can be viewed as an instantiated global root–pattern network consisting of roots, patterns, stems, and particles, estimated statistically for studying the morpho-phonetic level of cognition of Arabic. In this context, this paper asserts that balanced root-distribution is an additional significant key criterion for evaluating topic coverage in an Arabic corpus. Furthermore, some additional novel probabilistic morpho-phonetic measures and their distribution have been estimated in the form of root and pattern entropies besides bi-directional conditional probabilities of bi-grams of stems, roots, and particles. Around 29.2 million webpages of ClueWeb were extracted, filtered from non-Arabic texts, and converted into a large textual dataset containing around 11.5 billion word forms and 9.3 million associative relationships. As this dataset is predominantly considering the root–pattern phenomenon in Semitic languages, the acquired data might be significant support for researchers interested in studying phenomena of Arabic such as visual word cognition, morpho-phonetic perception, morphological analysis, and cognitively motivated query expansion, spell-checking, and information retrieval. Furthermore, based on data distribution and frequencies, constructing balanced corpora will be easier. View Full-Text
Keywords: cognitive linguistics; Arabic morphology; corpus linguistics; associative root–pattern network; bi-gram analysis cognitive linguistics; Arabic morphology; corpus linguistics; associative root–pattern network; bi-gram analysis
Figures

Figure 1

This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited. (CC BY 4.0).
SciFeed

Share & Cite This Article

MDPI and ACS Style

Haddad, B.; Awwad, A.; Hattab, M.; Hattab, A. Associative Root–Pattern Data and Distribution in Arabic Morphology. Data 2018, 3, 10.

Show more citation formats Show less citations formats

Note that from the first issue of 2016, MDPI journals use article numbers instead of page numbers. See further details here.

Article Metrics

Article Access Statistics

1

Comments

[Return to top]
Data EISSN 2306-5729 Published by MDPI AG, Basel, Switzerland RSS E-Mail Table of Contents Alert
Back to Top