ArabicEduCrawler: AI-Assisted Focused Crawling and Corpus Construction for Arabic Educational Web Content
Abstract
1. Introduction
- We develop ArabicEduCrawler, a novel framework that integrates domain-aware source selection, FastText-based Arabic language identification, and LLM-assisted XPath extraction for early-stage quality control of Arabic educational web content.
- We employ Scrapy-Playwright to crawl both static and dynamic web pages, focusing on relevant educational domains while reducing unnecessary noise.
- We develop a dual-threshold Arabic language filtering method based on FastText to maintain high corpus purity and effectively remove irrelevant, non-Arabic, and mixed-language pages during data collection.
- We propose a sentence-aware chunking method for Arabic educational text that preserves sentence boundaries while dividing documents into token-bounded segments suitable for transformer-based models. In contrast to fixed-window segmentation, this method maintains sentence coherence while keeping chunk sizes appropriate for downstream indexing, retrieval, and annotation tasks.
- We create an automatic linguistic annotation pipeline using GateNLP and Stanza, which includes tokenization, lemmatization, part-of-speech tagging, and named entity recognition.
- We empirically validate the proposed approach through the construction of an Arabic Educational Web Corpus (AraEdu-WC), which contains 101,770 documents and 286,025 chunks, with a harvest ratio of 95.25%.
- We evaluate sentence-aware chunking against Character Text Splitter and Recursive Character Text Splitter using a FAISS index and ranx metrics in a retrieval experiment. The experiment uses 7505 Arabic query–document pairs derived from 500 source documents randomly selected from AraEdu-WC, where the queries were generated using Llama. The proposed sentence-aware chunking strategy achieves the highest Hit Rate@1, @5, and @10; Precision@1, @5, and @10; and MRR@10 across all evaluated embedding models.
- AraEdu-WC is released as an open, reusable resource to advance research in Arabic NLP and educational technology.
2. Background
2.1. Web Crawling
2.2. Web Scraping
2.3. Challenges in the Arabic NLP
3. Related Work
3.1. Corpus Construction for Low-Resource Languages
3.2. Crawling and Scraping for Research and Educational Content Discovery
3.3. Web Corpora for Pre-Trained Language Models
3.4. Research Gaps
4. Materials and Methods
4.1. Focused Web Crawling and Data Acquisition Layer
4.1.1. Data Sources and Domain Scope
- Alukah Network [44] is a network composed of thirteen websites, showcasing contributions from a notable assembly of writers and intellectuals. It hosts thousands of published articles across Islamic studies, literature, science, and cultural topics. The platform provides regularly updated content. The network supports both Arabic and English and includes a wide variety of content in various world languages.
- Adab [45] is a large digital repository of Arabic literary works, containing more than 90,000 poems and prose texts by 7316 classical and modern authors. It also provides biographical records and well-organized literary collections. With approximately 1.4 billion total views, the platform represents a linguistically rich and semantically diverse source of Arabic text.
- Shamela Library [46] is one of the largest digitized Arabic text repositories, with over 8000 books and about 7 million pages written by more than 3000 authors. It covers disciplines such as Islamic studies, history, Arabic linguistics, and literature. Its large volume makes it a significant resource for extracting Arabic texts.
- Arabic Wikipedia [47] contains over 1.3 million articles across a wide range of domains, making it a substantial encyclopedic resource. It has approximately 2 million registered users. This encyclopedia is collaboratively curated and offers structured, domain-diverse, and continuously updated Arabic content that contributes modern terminology and cross-disciplinary coverage.
4.1.2. Improving XPath Extraction Using LLMs
4.1.3. Focused Web Crawling and Scraping Architecture
4.1.4. FastText for Arabic Language Identification
| Algorithm 1 Arabic Language Identification Using FastText |
| Input: Text ; FastText model ; ; Primary Threshold ; Secondary Threshold ; Top-k Predictions k. Output: , , , .
|
4.1.5. Metadata Extraction and Storage
4.2. Arabic Natural Language Processing Layer
4.2.1. Sampling Strategy Using Yamane’s Formula
4.2.2. Data Preprocessing
4.2.3. Text Chunking
- Primary Sentence Segmentation: The document is first segmented using strong punctuation marks, including the period, question mark, exclamation mark, and their Arabic equivalents.
- Secondary Sub-Sentence Segmentation: If a primary sentence exceeds the maximum token limit (), it is further divided using weaker punctuation such as the Arabic comma and semicolon.
- Word-Level Fallback: In case no suitable punctuation is found, and the sentence surpasses the token limit, the algorithm will divide the text at the word level to comply with the token constraints.
4.2.4. From Raw Text to Annotations
5. Experimental Setup
5.1. Hardware and Software Environment
5.2. Evaluation Protocol
5.3. Focused Arabic Crawling Experiment
5.4. Sentence-Aware Chunking Evaluation
6. Results and Discussion
6.1. Crawling Performance
6.2. Analysis of FastText Predictions
6.2.1. Distribution of FastText Top-3 Predictions
6.2.2. Confidence Score Analysis
6.2.3. FastText and Rule-Based Filtering Results
6.3. Analysis of the Web Corpus
6.3.1. Corpus Statistics
6.3.2. Sentence and Token Analysis Across Chunks
6.3.3. Automatic Linguistic Annotation Analysis
6.3.4. Results of the Sentence-Aware Chunking Strategy
7. Limitations and Future Work
7.1. Limitations
7.2. Future Work
8. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| NLP | Natural Language Processing |
| GATE | General Architecture for Text Engineering |
| LLMs | Large Language Models |
| PLM | Pre-trained Language Model |
| XPath | XML Path Language |
| HTML | HyperText Markup Language |
| MSA | Modern Standard Arabic |
Appendix A. Prompt for Generating XPath Expressions
| Component | Prompt |
|---|---|
| Role | You are a professional software engineer specialized in web scraping and HTML parsing. |
| Task | Your task is to analyze the HTML code provided by the user and generate robust XPath expressions that extract the requested target values. |
| Instructions | 1. Carefully read the provided HTML snippet. 2. Identify structural patterns surrounding the target value. 3. Generate XPath expressions that reliably locate the target element. 4. Avoid hardcoding exact text values. 5. Use contains(., ’value’) instead of text()=’value’. 6. Prefer contains(@class, ’…’) and contains(@id, ’…’) over strict equality. 7. Ensure the XPath is robust against: - Additional classes being added - Minor ID changes - Extra wrapper elements |
| Output Format | Output only the XPath expressions, clearly labeled by target field. |
| Input | HTML snippet provided by the user. |
Appendix B. Crawling Parameter Setup with Scrapy–Playwright
| Parameter | Value | Purpose |
|---|---|---|
| Browser Type | Chromium (headless) | Dynamic page rendering |
| ROBOTSTXT_OBEY | True | Crawl policy compliance |
| DOWNLOAD_DELAY | 2.0 s | Request delay rate |
| RANDOMIZE_DOWNLOAD_DELAY | True | Reduced timing regularity |
| CONCURRENT_REQUESTS | 16 | Total request limit |
| CONCURRENT_REQUESTS_PER_DOMAIN | 8 | Per-domain request limit |
| CONCURRENT_REQUESTS_PER_IP | 8 | Per-IP request limit |
| AUTOTHROTTLE_ENABLED | True | Automatically throttling crawling speed based on server response |
| AUTOTHROTTLE_TARGET_CONCURRENCY | 2.0 | Limits server load |
| RETRY_ENABLED | True | Failed request recovery |
| RETRY_TIMES | 3 | Retry limit per request |
| DOWNLOAD_TIMEOUT | 30 s | Response wait limit |
| HTTPCACHE_ENABLED | True | Response caching |
| HTTPCACHE_EXPIRATION_SECS | 86,400 s | Cache duration period (24 h) |
Appendix C. Prompt for Generating Synthetic Arabic Questions
| Component | Prompt |
|---|---|
| Role | You are an AI assistant tasked with generating exactly 6 realistic Arabic user questions based on educational documents. |
| Task | Generate questions strictly based on the provided document excerpt. |
| Instructions | 1. Questions must be answerable using ONLY the information in the excerpt. 2. Do not assume or use external knowledge. 3. Do not ask about information not explicitly stated in the document. 4. Generate at least 2 questions in Modern Standard Arabic (dialect: “msa”). 5. Generate at least 2 questions in Saudi dialect (dialect: “saudi”). 6. Include the exact snippet in “location” that answers each question. 7. Focus on concrete facts, not interpretation or opinion. |
| Output Format | Output must be ONLY a single valid JSON object with the following structure: { “questions”: [ { “question_id”: 1, “question”: “question text”, “location”: “Quote from excerpt that contains the answer”, “dialect”: “msa | saudi” } ] } |
| Input | Document excerpt provided by the user. |
Appendix D. Dataset Details for the Chunking Retrieval Experiment
| Statistic | Value |
|---|---|
| Query-document pairs | 7505 |
| Unique source documents | 500 |
| Query generation model | llama3.1:8b |
| MSA queries | 4350 |
| Saudi dialect queries | 3155 |
| Source | Total Pairs | MSA | Saudi Dialect |
|---|---|---|---|
| Alukah | 3731 | 2215 | 1516 |
| Shamela | 1236 | 694 | 542 |
| Arabic Wikipedia | 1535 | 892 | 643 |
| Adab | 1003 | 549 | 454 |
| Total | 7505 | 4350 | 3155 |
| Ex. | Dialect | Query | Retrieved Answer | Document Snippet |
|---|---|---|---|---|
| E1 | Saudi | ما عدد العينات التي تم جمعها من القمر خلال برنامج أبوللو؟ How many samples were collected from the moon during the Apollo program? | تم تجميع ٣٤٨ كجم من العينات من القمر A total of 348 kg of samples were collected from the moon. | وفي فترة برنامج الفضاء أبوللو تم تجميع ٣٤٨ كجم من العينات من القمر During the Apollo space program, 348 kg of samples were collected from the moon. |
| Saudi | ما هو اسم البرنامج الذي تم خلاله تجميع العينات من القمر؟ What is the name of the program during which the moon samples were collected? | في فترة برنامج الفضاء أبوللو تم تجميع ٣٤٨ كجم من العينات من القمر During the Apollo space program, 348 kg of samples were collected from the moon. | وفي فترة برنامج الفضاء أبوللو تم تجميع ٣٤٨ كجم من العينات من القمر During the Apollo space program, 348 kg of samples were collected from the moon. | |
| MSA | من هو عالم الفلك الذي اكتشف الأقمار الأربعة في المشتري؟ Who is the astronomer who discovered Jupiter’s four moons? | وفي هذا الإتجاه نجد أن عالم الفلك الكوكبي الأول هو جاليليو جاليلي In this context, the first planetary astronomer is Galileo Galilei. | عالم الفلك الكوكبي الأول هو جاليليو جاليلي الذي اكتشف الأقمار الأربعة في المشترى The first planetary astronomer is Galileo Galilei, who discovered Jupiter’s four moons. | |
| MSA | كيف يتم حل شفرة تاريخ سطح الكوكب؟ How is the chronology of a planet’s surface decoded? | يتم حل شفرة تاريخ سطح الكوكب عن طريق وضع خرائط جيولوجية للمعالم الموجودة عليه من القمة إلى القاع طبقا لتتابع الطبقات وترتيبها The chronology of a planet’s surface is decoded by creating geological maps of its features from top to bottom, according to the order and succession of layers. | يتم حل شفرة تاريخ سطح الكوكب عن طريق وضع خرائط جيولوجية للمعالم الموجودة عليه من القمة إلى القاع طبقا لتتابع الطبقات وترتيبها The chronology of a planet’s surface is decoded by creating geological maps of its features from top to bottom, according to the order and succession of layers. | |
| E2 | Saudi | متى طور برنامج فيورستين التعليمي الإغنائي؟ When was the Feuerstein Instrumental Enrichment program developed? | برنامج فيورستين التعليمي الإغنائي، ١٩٨٠ The Feuerstein Instrumental Enrichment program, 1980. | وبرنامج فيورستين التعليمي الإغنائي، ١٩٨٠ And the Feuerstein Instrumental Enrichment program, 1980. |
| Saudi | مين اللي بنزل برنامج الفلسفة للأطفال؟ Who introduced the Philosophy for Children program? | برنامج الفلسفة للأطفال Philosophy for Children program. | برنامج الفلسفة للأطفال للبمان Lipman’s Philosophy for Children program. | |
| MSA | كم عدد وحدات برنامج كورت؟ How many units are in the CoRT program? | يتكون برنامج كورت من ست وحدات تعليمية The CoRT program consists of six educational units. | يتكون برنامج كورت من ست وحدات تعليمية تعطي جوانب عديدة للتفكير The CoRT program consists of six educational units that address various aspects of thinking. | |
| MSA | من هو مؤلف برنامج البناء العقلي لجيلفورد؟ Who is the author of Guilford’s Structure of Intellect program? | برنامج البناء العقلي لجيلفورد، الذي طورته الباحثة ميكر، ١٩٦٩ Guilford’s Structure of Intellect program, which was developed by researcher Meeker in 1969. | ومن بين البرامج المعروفة التي تمثل اتجاه العمليات المعرفية برنامج البناء العقلي لجيلفورد، الذي طورته الباحثة ميكر، ١٩٦٩ Among the well-known programs representing the cognitive operations approach is Guilford’s Structure of Intellect program, which was developed by researcher Meeker in 1969. |
Appendix E. Distribution of Drop Reasons Across Spiders
| Source | Too Short | Duplicate URL | Non-Arabic Detected 1 | Non-Arabic Detected 2 | Total Dropped |
|---|---|---|---|---|---|
| Alukah | 4478 | 0 | 360 | 18 | 4695 |
| Arabic Wiki | 1 | 104 | 153 | 0 | 231 |
| Shamela | 1 | 0 | 47 | 35 | 83 |
| Adab | 0 | 0 | 14 | 1 | 15 |
| Total | 4480 | 104 | 574 | 54 | 5212 |
Appendix F. Corpus Automatic Annotation Challenges and Lexical Visualization
| No. | Title | Text Sample | Token | UPOS | Lemma | NER | Annotation Observation |
|---|---|---|---|---|---|---|---|
| 1 | الأدب بين نفس المروءة ولهاث الإثارة | لم يكن الأدب العربي بشقيه المنظوم والمنثور إلا مرآة لروح صاحبه، وصورة حية لخلجات نفسه… كما تساءل ابن قتيبة في أدب الكاتب | لم | PART | لَم | O | POS ambiguity in proper names. NER correctly detects ابن قتيبة as PER, but UPOS assigns X to both name tokens. This suggests difficulty with historical Arabic names in literary texts. |
| يكن | VERB | كَان | O | ||||
| الأدب | NOUN | أَدَب | O | ||||
| العربي | ADJ | عَرَبِيّ | O | ||||
| ابن | X | ابن | B-PER | ||||
| قتيبة | X | قتيبة | E-PER | ||||
| 2 | منهج القرآن الكريم في تنمية التفكير التأملي | الخاتمة: يبرهن القرآن الكريم على أن التفكير التأملي ليس ترفا ذهنيا، بل ضرورة تربوية لصناعة إنسان مسؤول وقادر على اتخاذ قرارات مستنيرة | الخاتمة | NOUN | خَاتِمَة | O | The religious phrase is treated as lexical material, not a named entity. This is not necessarily incorrect, but scripture and religious references may need clearer NER guidelines. |
| يبرهن | VERB | أَبرَه | O | ||||
| القرآن | NOUN | قُرآن | O | ||||
| الكريم | ADJ | كَرِيم | O | ||||
| التفكير | NOUN | تَفكِير | O | ||||
| 3 | فروع الكيمياء | ويدرس بنية وخواص وتفاعلات المركبات والمواد العضوية التي تحتوي على عنصر الكربون ووضع ميخائيل لومونوسوف كتابا في الكيمياء الفيزيائية | الكربون | NOUN | كَربُون | S-MISC | POS ambiguity in foreign names. ميخائيل لومونوسوف is detected as PER, but assigned X in UPOS. This suggests difficulty tagging foreign proper names in scientific text. |
| ميخائيل | X | مِيخَائِيل | B-PER | ||||
| لومونوسوف | X | لومونوسوف | E-PER | ||||
| الكيمياء | NOUN | كِيمِيَاء | B-MISC | ||||
| الفيزيائية | ADJ | فِيزِيَائِيّ | E-MISC | ||||
| 4 | فورميمينو-٥ ثلاثي هيدروفولات | فورميمينو ثلاثي هيدروفولات هو مركب وسطي ينتج من عملية الهدم للحمض الأميني هستيدين بوساطة الغلوتاميت فورماميدويلترانسفيريز | فورميمينو | X | فورميمينو | S-MISC | Biomedical terms are mostly tagged as X in UPOS and MISC in NER. Main issue: specialized chemical and biochemical vocabulary lacks fine-grained entity categories. |
| هيدروفولات | X | هيدروفولات | S-MISC | ||||
| هستيدين | X | هستيدين | S-MISC | ||||
| الغلوتاميت | X | الغلوتاميت | S-MISC | ||||
| سايكلوديامينيز | X | سايكلوديامينيز | S-MISC | ||||
| 5 | أثر التناقض اللفظي في المعنى | هو أسلوب أدبي يجمع بين كلمتين متناقضتين ظاهريا لتشكيل مفهوم مثير للتفكير كلمة أوكسيمورون مشتقة من أوكسيس وموروس | هو | PRON | هُوَ | O | POS variation in borrowed terms. Terms such as أوكسيمورون, أوكسيس, and موروس are labeled as MISC, while UPOS varies between X and NOUN. This suggests ambiguity in assigning POS tags to foreign-origin or borrowed technical terms in Arabic text. |
| أسلوب | NOUN | أُسلُوب | O | ||||
| أدبي | ADJ | دِبِّيّ | O | ||||
| يجمع | VERB | جَمَع | O | ||||
| أوكسيمورون | X | أوكسيمورون | S-MISC | ||||
| أوكسيس | X | أوكسيس | S-MISC | ||||
| موروس | NOUN | مَورُوس | S-MISC |

References
- Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T.B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; Amodei, D. Scaling laws for neural language models. arXiv 2020, arXiv:2001.08361. [Google Scholar] [CrossRef] [Scilit]
- Bommasani, R.; Hudson, D.A.; Adeli, E.; Altman, R.; Arora, S.; von Arx, S.; Bernstein, M.S.; Bohg, J.; Bosselut, A.; Brunskill, E.; et al. On the opportunities and risks of foundation models. arXiv 2021, arXiv:2108.07258. [Google Scholar] [CrossRef] [Scilit]
- Rae, J.W.; Borgeaud, S.; Cai, T.; Millican, K.; Hoffmann, J.; Song, F.; Aslanides, J.; Henderson, S.; Ring, R.; Young, S.; et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv 2021, arXiv:2112.11446. [Google Scholar]
- Common Crawl Foundation. Common Crawl. 2007. Available online: https://commoncrawl.org/ (accessed on 19 April 2026).
- Wenzek, G.; Lachaux, M.A.; Conneau, A.; Chaudhary, V.; Guzmán, F.; Joulin, A.; Grave, E. CCNet: Extracting high quality monolingual datasets from web crawl data. In Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France, 11–16 May 2020; pp. 4003–4012. [Google Scholar]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 2020, 21, 1–67. [Google Scholar]
- Suárez, P.J.O.; Sagot, B.; Romary, L. Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures. In Proceedings of the 7th Workshop on the Challenges in the Management of Large Corpora (CMLC-7); Leibniz-Institut für Deutsche Sprache: Mannheim, Germany, 2019. [Google Scholar]
- Penedo, G.; Kydlíček, H.; Ben allal, L.; Lozhkov, A.; Mitchell, M.; Raffel, C.; Von Werra, L.; Wolf, T. The fineweb datasets: Decanting the web for the finest text data at scale. Adv. Neural Inf. Process. Syst. 2024, 37, 30811–30849. [Google Scholar]
- Hawasly, M.; Mohiuddin, M.T.; Mubarak, H.; Boughorbel, S. ArabicWeb-Edu: Educational Quality Data for Arabic LLM Training. In Proceedings of the Third Arabic Natural Language Processing Conference, Suzhou, China, 8–9 November 2025; pp. 436–447. [Google Scholar]
- Roziewski, S.; Kozłowski, M. LanguageCrawl: A generic tool for building language models upon common Crawl. Lang. Resour. Eval. 2021, 55, 1047–1075. [Google Scholar] [CrossRef] [Scilit]
- Dodge, J.; Sap, M.; Marasović, A.; Agnew, W.; Ilharco, G.; Groeneveld, D.; Mitchell, M.; Gardner, M. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Punta Cana, Dominican Republic, 7–11 November 2021; pp. 1286–1305. [Google Scholar]
- Kreutzer, J.; Caswell, I.; Wang, L.; Wahab, A.; Van Esch, D.; Ulzii-Orshikh, N.; Tapo, A.; Subramani, N.; Sokolov, A.; Sikasote, C.; et al. Quality at a Glance: An Audit of Web-Crawled Multilingual Datasets. Trans. Assoc. Comput. Linguist. 2022, 10, 50–72. [Google Scholar] [CrossRef] [Scilit]
- Remus, S.; Biemann, C. Domain-Specific Corpus Expansion with Focused Webcrawling. In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC’16), Portorož, Slovenia, 23–28 May 2016; pp. 3607–3611. [Google Scholar]
- Antoun, W.; Baly, F.; Hajj, H. AraBERT: Transformer-based Model for Arabic Language Understanding. In Proceedings of the 4th Workshop on Open-Source Arabic Corpora and Processing Tools, with a Shared Task on Offensive Language Detection, Marseille, France, 11–16 May 2020; pp. 9–15. [Google Scholar]
- Abdallah, A.; Kasem, M.; Abdalla, M.; Mahmoud, M.; Elkasaby, M.; Elbendary, Y.; Jatowt, A. Arabicaqa: A comprehensive dataset for arabic question answering. In Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, Washington, DC, USA, 14–18 July 2024; pp. 2049–2059. [Google Scholar]
- Bhatia, G.; Nagoudi, E.M.B.; El Mekki, A.; Alwajih, F.; Abdul-Mageed, M. Swan and ArabicMTEB: Dialect-Aware, Arabic-Centric, Cross-Lingual, and Cross-Cultural Embedding Models and Benchmarks. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, NM, USA, 29 April–4 May 2025; pp. 4669–4685. [Google Scholar] [CrossRef] [Scilit]
- Olston, C.; Najork, M. Web Crawling. Found. Trends Inf. Retr. 2010, 4, 175–246. [Google Scholar] [CrossRef] [Scilit]
- Liu, B.; Menczer, F. Web Crawling. In Web Data Mining: Exploring Hyperlinks, Contents, and Usage Data; Springer: Berlin/Heidelberg, Germany, 2011; pp. 311–362. [Google Scholar] [CrossRef] [Scilit]
- Zhai, C.; Massung, S. Text Data Management and Analysis: A Practical Introduction to Information Retrieval and Text Mining; Association for Computing Machinery and Morgan & Claypool: New York, NY, USA, 2016. [Google Scholar] [CrossRef] [Scilit]
- Kumar, M.; Bhatia, R.; Rattan, D. A Survey of Web Crawlers for Information Retrieval. WIREs Data Min. Knowl. Discov. 2017, 7, e1218. [Google Scholar] [CrossRef] [Scilit]
- Manning, C.; Raghavan, P.; Schuetze, H. Introduction to Information Retrieval; Cambridge University Press: Cambridge, UK, 2009. [Google Scholar]
- Diouf, R.; Sarr, E.N.; Sall, O.; Birregah, B.; Bousso, M.; Mbaye, S.N. Web Scraping: State-of-the-Art and Areas of Application. In Proceedings of the 2019 IEEE International Conference on Big Data (Big Data), Los Angeles, CA, USA, 9–12 December 2019; pp. 6040–6042. [Google Scholar] [CrossRef] [Scilit]
- Ayuso, E.; Dumfeh Brogya, M.S.; Kumar Ahlawat, V.; Sain, M. From Manual to Machine: How AI Is Redefining Web Scraping for Superior Efficiency: A Literature Review. In Proceedings of the 2024 International Conference on Communication, Control, and Intelligent Systems (CCIS), Mathura, India, 6–7 December 2024; pp. 1–9. [Google Scholar] [CrossRef] [Scilit]
- Obeid, O.; Zalmout, N.; Khalifa, S.; Taji, D.; Oudah, M.; Alhafni, B.; Inoue, G.; Eryani, F.; Erdmann, A.; Habash, N. CAMeL Tools: An Open Source Python Toolkit for Arabic Natural Language Processing. In Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France, 13–15 May 2020; pp. 7022–7032. [Google Scholar]
- Habash, N.; Soudi, A.; Buckwalter, T. On arabic transliteration. In Arabic Computational Morphology: Knowledge-Based and Empirical Methods; Springer: Berlin/Heidelberg, Germany, 2007; pp. 15–22. [Google Scholar]
- Boudchiche, M.; Mazroui, A.; Bebah, M.O.A.O.; Lakhouaja, A.; Boudlal, A. AlKhalil Morpho Sys 2: A robust Arabic morpho-syntactic analyzer. J. King Saud Univ. Comput. Inf. Sci. 2017, 29, 141–146. [Google Scholar] [CrossRef] [Scilit]
- Jarrar, M.; Zaraket, F.; Asia, R.; Amayreh, H. Diacritic-based matching of arabic words. ACM Trans. Asian Low-Resour. Lang. Inf. Process. (TALLIP) 2018, 18, 10. [Google Scholar] [CrossRef] [Scilit]
- Alothman, A.; Alsalman, A. Arabic morphological analysis techniques. Int. J. Adv. Comput. Sci. Appl. 2020, 11, 214–222. [Google Scholar] [CrossRef] [Scilit]
- Hossain, M.R.; Hoque, M.M.; Siddique, N.; Dewan, M.A.A. AraCovTexFinder: Leveraging the transformer-based language model for Arabic COVID-19 text identification. Eng. Appl. Artif. Intell. 2024, 133, 107987. [Google Scholar] [CrossRef] [Scilit]
- Mahi, G.S.; Verma, A. Development of Focused Crawlers for Building Large Punjabi News Corpus. J. Ict Res. Appl. 2021, 15, 205–215. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Su, K.; Tian, Y.; Matsumoto, T. WCC-EC 2.0: Enhancing neural machine translation with a 1.6 M+ web-crawled English-Chinese parallel corpus. Electronics 2024, 13, 1381. [Google Scholar] [CrossRef] [Scilit]
- de Jesus, G.; Nunes, S.S. Data Collection Pipeline for Low-Resource Languages: A Case Study on Constructing a Tetun Text Corpus. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy, 20–25 May 2024; pp. 4368–4380. [Google Scholar]
- Tahir, B.; Mehmood, M.A. Corpulyzer: A novel framework for building low resource language corpora. IEEE Access 2021, 9, 8546–8563. [Google Scholar] [CrossRef] [Scilit]
- Marquard, C.; Suleman, H. Focused Crawling for Automated IsiXhosa Corpus Building. In Proceedings of the South African Institute of Computer Scientists and Information Technologists, Pretoria, South Africa, 17–19 July 2023; pp. 19–31. [Google Scholar]
- Mehmood, M.A.; Tahir, B. Humkinar: Construction of a large scale web repository and information system for low resource Urdu language. IEEE Access 2024, 12, 128404–128423. [Google Scholar] [CrossRef] [Scilit]
- Fahrudin, T.M.; Funabiki, N.; Brata, K.C.; Naing, I.; Aung, S.T.; Muhaimin, A.; Prasetya, D.A. An improved reference paper collection system using web scraping with three enhancements. Future Internet 2025, 17, 195. [Google Scholar] [CrossRef] [Scilit]
- Mutlu, M.A.; Ulku, E.E.; Yildiz, K. A web scraping app for smart literature search of the keywords. PeerJ Comput. Sci. 2024, 10, e2384. [Google Scholar] [CrossRef] [Scilit]
- Ampadi Ramachandran, R.; Tell, L.A.; Rai, S.; Millagaha Gedara, N.I.; Xu, X.; Riviere, J.E.; Jaberi-Douraki, M. An Automated Customizable Live Web Crawler for Curation of Comparative Pharmacokinetic Data: An Intelligent Compilation of Research-Based Comprehensive Article Repository. Pharmaceutics 2023, 15, 1384. [Google Scholar] [CrossRef] [Scilit]
- Barwary, M.J.; Jacksi, K.; Al-Zebari, A. Constructing a multilingual e-learning ontology through web crawling and scraping. Int. J. Commun. Netw. Inf. Secur. 2023, 15, 137–153. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y. Visual Analysis of Finance Courses on Chinese University MOOC Platform Based on Web Crawler. In Proceedings of the 6th International Conference on Digital Technology in Education, Hangzhou, China, 16–18 September 2022; pp. 310–315. [Google Scholar]
- Aljemazi, M.A.; Khder, M.A. Octobot-Web Scarping towards retrieving Google Scholar Data. In Proceedings of the 2022 ASU International Conference in Emerging Technologies for Sustainability and Intelligent Systems (ICETSIS), Virtual, 22–23 June 2022; pp. 477–482. [Google Scholar]
- Alrashed, S.; Khizbullin, D.; Pugh, D.R. Fineweb-Edu-Ar: Machine-translated corpus to support Arabic small language models. arXiv 2024, arXiv:2411.06402. [Google Scholar]
- Aloui, M.; Chouikhi, H.; Chaabane, G.; Kchaou, H.; Dhaouadi, C. 101 billion arabic words dataset. arXiv 2024, arXiv:2405.01590. [Google Scholar] [CrossRef] [Scilit]
- El-Hmed, S.B.A. Alukah Network. 2006. Available online: https://www.alukah.net (accessed on 27 February 2026).
- Adab Foundation. World Encyclopedia of Arabic Literature. 1999. Available online: https://www.adab.com (accessed on 27 February 2026).
- Al-Maktaba Al-Shamela Foundation. Al-Maktaba al-Shamela, Digital Library. 2005. Available online: https://shamela.ws (accessed on 27 February 2026).
- Wikipedia Contributors. Arabic Wikipedia. 2003. Available online: https://ar.wikipedia.org (accessed on 27 February 2026).
- World Wide Web Consortium (W3C). XML Path Language (XPath) Version 1.0. 1999. Available online: https://www.w3.org/TR/1999/REC-xpath-19991116/ (accessed on 19 April 2026).
- Darmawan, I.; Maulana, M.; Gunawan, R.; Widiyasono, N. Evaluating web scraping performance using XPath, CSS selector, regular expression, and HTML DOM with multiprocessing technical applications. JOIV Int. J. Inform. Vis. 2022, 6, 904–910. [Google Scholar] [CrossRef] [Scilit]
- Huang, J.; Song, J. Automatic XPath generation agents for vertical websites by LLMs. J. King Saud Univ. Comput. Inf. Sci. 2025, 37, 74. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Wang, B.; Luan, X. Xpath agent: An efficient xpath programming agent based on llm for web crawler. arXiv 2024, arXiv:2502.15688. [Google Scholar] [CrossRef] [Scilit]
- Scrapy Developers. Architecture Overview. Available online: https://docs.scrapy.org/en/latest/topics/architecture.html (accessed on 19 April 2026).
- Joulin, A.; Grave, E.; Bojanowski, P.; Douze, M.; Jégou, H.; Mikolov, T. FastText.zip: Compressing text classification models. arXiv 2016, arXiv:1612.03651. [Google Scholar]
- Joulin, A.; Grave, E.; Bojanowski, P.; Mikolov, T. fastText Language Identification, 2024. Available online: https://github.com/facebookresearch/fastText (accessed on 19 April 2026).
- Bojanowski, P.; Grave, E.; Joulin, A.; Mikolov, T. Enriching word vectors with subword information. Trans. Assoc. Comput. Linguist. 2017, 5, 135–146. [Google Scholar] [CrossRef] [Scilit]
- Penedo, G.; Malartic, Q.; Hesslow, D.; Cojocaru, R.; Cappelli, A.; Alobeidli, H.; Pannier, B.; Almazrouei, E.; Launay, J. The RefinedWeb dataset for Falcon LLM: Outperforming curated corpora with web data, and web data only. arXiv 2023, arXiv:2306.01116. [Google Scholar] [CrossRef] [Scilit]
- Soldaini, L.; Kinney, R.; Bhagia, A.; Schwenk, D.; Atkinson, D.; Authur, R.; Bogin, B.; Chandu, K.; Dumas, J.; Elazar, Y.; et al. Dolma: An Open Corpus of Three Trillion Tokens for Language Model Pretraining Research. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 11–16 August 2024; pp. 15725–15788. [Google Scholar] [CrossRef] [Scilit]
- Alsubhi, J.; Alahmadi, M.D.; Alhusayni, A.; Aldailami, I.; Hamdine, I.; Shabana, A.; Iskandar, Y.; Khayyat, S. Optimizing RAG Pipelines for Arabic: A Systematic Analysis of Core Components. arXiv 2025, arXiv:2506.06339. [Google Scholar] [CrossRef] [Scilit]
- Wang, Z.; Gao, C.; Xiao, C.; Huang, Y.; Si, S.; Luo, K.; Bai, Y.; Li, W.; Duan, T.; Lv, C.; et al. Document Segmentation Matters for Retrieval-Augmented Generation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025, Vienna, Austria, 27 July–1 August 2025; pp. 8063–8075. [Google Scholar]
- Qu, R.; Tu, R.; Bao, F.S. Is Semantic Chunking Worth the Computational Cost? In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, NM, USA, 29 April–4 May 2025; pp. 2155–2177. [Google Scholar] [CrossRef] [Scilit]
- Merola, C.; Singh, J. Reconstructing context: Evaluating advanced chunking strategies for retrieval-augmented generation. In Proceedings of the International Workshop on Knowledge-Enhanced Information Retrieval; Springer: Cham, Switzerland, 2025; pp. 3–18. [Google Scholar]
- Cunningham, H.; Maynard, D.; Bontcheva, K.; Tablan, V.; Aswani, N.; Roberts, I.; Gorrell, G.; Funk, A.; Roberts, A.; Damljanovic, D.; et al. Developing Language Processing Components with GATE: Version 9 (User Guide); The University of Sheffield, Department of Computer Science: Sheffield, UK, 2023. [Google Scholar]
- Qi, P.; Zhang, Y.; Zhang, Y.; Bolton, J.; Manning, C.D. Stanza: A Python natural language processing toolkit for many human languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Online, 5–10 July 2020; pp. 101–108. [Google Scholar]
- Ollama, Inc. Ollama: Open-Source Software Tool for Large Language Model. 2023. Available online: https://ollama.com (accessed on 20 January 2026).
- Wang, L.; Yang, N.; Huang, X.; Yang, L.; Majumder, R.; Wei, F. Multilingual E5 Text Embeddings: A Technical Report. arXiv 2024, arXiv:2402.05672. [Google Scholar] [CrossRef] [Scilit]
- Reimers, N.; Gurevych, I. Sentence-bert: Sentence embeddings using siamese bert-networks. arXiv 2019, arXiv:1908.10084. [Google Scholar] [CrossRef] [Scilit]
- Nacar, O.; Koubaa, A.; Sibaee, S.; Al-Habashi, Y.; Ammar, A.; Boulila, W. GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Matryoshka Representation Learning and Hybrid Loss Training. arXiv 2025, arXiv:2505.24581. [Google Scholar] [CrossRef] [Scilit]
- Nacar, O.; Koubaa, A. Enhancing Semantic Similarity Understanding in Arabic NLP with Nested Embedding Learning. arXiv 2024, arXiv:2407.21139. [Google Scholar] [CrossRef] [Scilit]
- Douze, M.; Guzhva, A.; Deng, C.; Johnson, J.; Szilvasy, G.; Mazaré, P.E.; Lomeli, M.; Hosseini, L.; Jégou, H. The faiss library. IEEE Trans. Big Data 2025, 12, 346–361. [Google Scholar] [CrossRef] [Scilit]
- Bassani, E. ranx: A Blazing-Fast Python Library for Ranking Evaluation and Comparison. In Proceedings of the ECIR (2); Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2022; Volume 13186, pp. 259–264. [Google Scholar] [CrossRef] [Scilit]
- Bhat, S.R.; Rudat, M.; Spiekermann, J.; Flores-Herr, N. Rethinking Chunk Size For Long-Document Retrieval: A Multi-Dataset Analysis. arXiv 2025, arXiv:2505.21700. [Google Scholar]













| Study | Domain | Lang. | Source | Strategy | QC | LID | Dedup. | Size | Limitation |
|---|---|---|---|---|---|---|---|---|---|
| FineWeb-Edu [8] | Edu. | EN | FineWeb/ Common Crawl | LLM scoring & classifier filtering | Post hoc | FastText | MinHash | 1.3 T tokens | Demographic and religious bias skew |
| ArabicWeb-Edu [9] | Edu. | AR | 1 M Arabic Common Crawl documents | Qwen labeling & mGTE filtering | Post hoc | NR | NR | 1 M docs | Costly filtering; English-based criteria may not suit Arabic educational contexts. |
| FineWeb-Edu-Ar [42] | Edu. | AR | Translated FineWeb-Edu | MT construction using NLLB | Post-translation | N/A | Inherited | 189.4 M passages; 202.4 B tokens | Translation inaccuracies; English-centric knowledge |
| 101 B Arabic Words Dataset [43] | General | AR | Common Crawl WET | WET extraction, cleaning, normalization | Post hoc | NR | URL & MinHash | 101 B words; 89.1 M pages | No downstream evaluation; regional imbalance |
| ArabicEduCrawler (This study) | Edu. | AR | Targeted Arabic educational websites | AI-assisted crawling, scraping, and processing | During crawling & post-processing | FastText | MD5 | 101 k docs; 286 k passages; 50 M tokens; 289 k sent. | Smaller scale than broad web archives; depends on source availability |
| Domain | Description | Topics | Type | Format | Access |
|---|---|---|---|---|---|
| Alukah Network [44] | Islamic and cultural network presenting articles on science, literature, and contemporary issues from an Islamic perspective. | Islamic studies, literature, culture | Digital Portal | Text, PDF | Open |
| Adab [45] | Repository of Arabic literature featuring a rich collection of poetry and prose by renowned Arab writers. | Arabic literature, poetry | Encyclopedia | Text | Open |
| Shamela Library [46] | Digital Arabic e-book library offering a wide range of Islamic and literary works across multiple disciplines. | Islamic studies, history, literature, linguistics | Digital Library | Text, PDF | Open |
| Arabic Wikipedia [47] | Arabic edition of Wikipedia providing comprehensive information on various scientific disciplines. | Physics, chemistry, biology, mathematics | Encyclopedia | Text | Open |
| Source Domain | Resource Metadata | Crawling Metadata * |
|---|---|---|
| alukah.net [44] | Title, Author, Publication Date, View Count. | Crawl Timestamp, HTTP Status, Response Headers, Encoding, Crawl Depth, Referer URL, Source URL, Content Length, HTML Length, MD5 Hash. |
| adab.com [45] | Title, Author, Profile URL, Post Date, Page Number, Poem Length, Excerpt. | |
| shamela.ws [46] | Title, Author, Publisher, Edition, Book ID, Total Pages, Page Number. | |
| ar.wikipedia.org [47] | Title, Author, Subcategory, Last Modified Date. |
| Configuration | Crawling | GateNLP | Chunking Experiment |
|---|---|---|---|
| Platform | macOS | Windows 11 | |
| Hardware | Apple Silicon, 8 CPU cores, 16 GB memory | i7-13620H, 24 GB RAM, RTX 5060 8 GB | |
| Runtime | Docker | Python 3.12, PyTorch, CUDA 13.1 | |
| Image | python:3.11-slim | jupyter/base-notebook:python-3.11.5 | Not applicable |
| Domain | Seed URLs | Main Topic Area |
|---|---|---|
| ar.wikipedia.org | https://ar.wikipedia.org/wiki/تصنيف:كيمياء | Chemistry |
| https://ar.wikipedia.org/wiki/تصنيف:علوم_فيزيائية | Physics | |
| https://ar.wikipedia.org/wiki/تصنيف:علم_الأحياء | Biology | |
| https://ar.wikipedia.org/wiki/تصنيف:رياضيات | Mathematics | |
| alukah.net | https://www.alukah.net/sharia/0/ | Sharia |
| https://www.alukah.net/literature_language/0/ | Literature & Language | |
| https://www.alukah.net/culture/0/ | Culture | |
| https://www.alukah.net/library/0/ | Library | |
| https://www.alukah.net/social/0/ | Social Sciences | |
| adab.com | https://www.adab.com | Arabic Literature and Poetry |
| shamela.ws | https://shamela.ws | Islamic Studies, Literature, Linguistics, History |
| Case | Source | Sample Text | FastText Labels | Scores | Filtering Threshold | Filtering Justification |
|---|---|---|---|---|---|---|
| S1 | Arabic Wikipedia | …حساب التفاضل والتكامل من الاختلافات العثور على القيم القصوى للعمليات مشابه …لإيجاد القيم العظمى والصغرى للمعادلات | __label__arb_Arab __label__yue_Hant __label__azb_Arab | 0.896 0.046 0.015 | Accepted by primary | Arabic is the top prediction and its confidence is above the primary threshold of 0.8. |
| S2 | Arabic Wikipedia | …طرائق ح لانهاية في نظرية التحكم هي أحد طرائق بناء المتحكمات والتي …يمكن تطبيقها على الأنظمة الخطية واللاخطية | __label__arb_Arab __label__yue_Hant __label__azb_Arab | 0.883 0.069 0.011 | Accepted by primary | Arabic is the top prediction and its confidence is above the primary threshold of 0.8. |
| S3 | Alukah Network | …المهارات والقدرات المستهدفة بتبني نهج التقصي نهج التقصي هو منهجية تعليمية تشجع المتعلمين …من خلال طرح الأسئلة وإجراء البحوث والتحريات | __label__arb_Arab __label__pes_Arab __label__azb_Arab | 0.999 0.00004 0.00004 | Accepted by primary | The page is clearly Arabic educational content, with very high Arabic confidence. |
| S4 | Arabic Wikipedia | فلوريد الفضة وهو مركب كيميائي …مكون من عنصري الفضة والفلور .mw-parser-output .dmbox{display:flex;…} | __label__arb_Arab __label__yue_Hant __label__azb_Arab | 0.537 0.146 0.107 | Accepted by secondary | Arabic is still the top prediction, but CSS noise reduces confidence below 0.8. Since the score remains above 0.5, the page is retained. |
| S5 | Arabic Wikipedia | …المُلْتَقَى أو المَقْرَن مكان التقاء مسطحين مائيين @font-face{font-family: “TemplateStyles-Calibri-Quran”; …} ٩٤… (وَتَرَى الْمُجْرِمِينَ يَوْمَئِذٍ مُقَرَّنِينَ فِي الْأَصْفَادِ) | __label__arb_Arab __label__azb_Arab __label__yue_Hant | 0.756 0.060 0.045 | Accepted by secondary | Arabic is still the top prediction, but CSS templates and citation noise reduce confidence below 0.8. Since the score remains above 0.5, the page is retained. |
| S6 | Adab | لا أيها الزمن، لن تستطيع المباهاة بأن التغير …يطرأ عليّ إن أهراماتك التي شُيِّدَتْ بعزيمة جديدة No, Time, thou shalt not boast that I do change: Thy pyramids built up with newer might… | __label__arb_Arab __label__eng_Latn __label__azb_Arab | 0.649 0.332 0.009 | Accepted by secondary | Arabic is still the top prediction, but the presence of both Arabic and English translations lowers confidence below 0.8. Since the score remains above 0.5, the page is retained. |
| S7 | Arabic Wikipedia | (Homochiral) …التجانس اليدواني …أو تجانس عدم التناظر المرآتي) | – | – | Rejected | Rejected as duplicate content. The page had already been identified during crawling and was not retained as a unique Arabic page. |
| S8 | Alukah Network | “Kəlimə-i səva (Haqq Söz)” Mesajı Sevimli Peyğəmbər Məhəmməd (s.ə.s.) sünnəsinin əhlindən şiə əhlinə | __label__azj_Latn __label__bak_Cyrl __label__knc_Latn | 1.000 0.00001 0.00001 | Rejected | Rejected because Arabic was not confidently detected. FastText predicted Azerbaijani Latin with 1.000 confidence, and Arabic was absent from the top three labels. This case was therefore excluded by the language filter. |
| S9 | Shamela Library | قصة الفيل | – | – | Rejected | Rejected because the text was too short for reliable language detection, even though the sample itself contains Arabic words. |
| S10 | Alukah Network | Empty extracted text | – | – | Rejected | Rejected as too short because the extraction produced no usable text for language identification. |
| Phase/Layer | Statistic | Value |
|---|---|---|
| Crawling | Number of crawled pages | 109,633 |
| In-crawl filtering | Number of accepted documents | 104,421 |
| Preprocessing | Number of retained documents | 101,770 |
| Sentence-aware chunking | Number of chunks | 286,025 |
| GateNLP | Number of tokens | 50,366,982 |
| GateNLP | Number of sentences | 289,778 |
| GateNLP | Vocabulary size | 684,221 |
| GateNLP | Number of unique lemmas | 603,535 |
| GateNLP | Number of named entities | 3,429,315 |
| Chunking Method | Chunks | Docs | Qrels | Avg. Tokens | Median | Min | Max |
|---|---|---|---|---|---|---|---|
| CharacterTextSplitter | 1301 | 494 | 7476 | 265.33 | 301 | 22 | 442 |
| RecursiveCharacterTextSplitter | 1044 | 494 | 7476 | 333.91 | 449 | 22 | 450 |
| Sentence-Aware (Proposed) | 1291 | 494 | 7476 | 263.50 | 290 | 21 | 450 |
| Embedding Model | Dim. | Chunking Method | Hit Rate | Precision | Recall | MRR@10 | MAP@10 | nDCG@10 | ||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| @1 | @5 | @10 | @1 | @5 | @10 | @1 | @5 | @10 | ||||||
| E5-large | 1024 | Character | 0.5091 | 0.6499 | 0.7041 | 0.5091 | 0.2694 | 0.1820 | 0.1974 | 0.3611 | 0.4282 | 0.5699 | 0.3470 | 0.4354 |
| Recursive | 0.4831 | 0.6311 | 0.6851 | 0.4831 | 0.2415 | 0.1578 | 0.2153 | 0.3901 | 0.4566 | 0.5463 | 0.3642 | 0.4379 | ||
| Sentence-aware | 0.5185 | 0.6540 | 0.7081 | 0.5185 | 0.2725 | 0.1859 | 0.1938 | 0.3534 | 0.4252 | 0.5771 | 0.3442 | 0.4358 | ||
| MiniLM | 384 | Character | 0.0736 | 0.1471 | 0.1907 | 0.0736 | 0.0368 | 0.0271 | 0.0420 | 0.0757 | 0.0947 | 0.1058 | 0.0606 | 0.0816 |
| Recursive | 0.0602 | 0.1170 | 0.1552 | 0.0602 | 0.0274 | 0.0201 | 0.0423 | 0.0757 | 0.0940 | 0.0850 | 0.0594 | 0.0752 | ||
| Sentence-aware | 0.0940 | 0.1663 | 0.2063 | 0.0940 | 0.0490 | 0.0360 | 0.0438 | 0.0820 | 0.1049 | 0.1253 | 0.0674 | 0.0946 | ||
| Ara- Matryoshka | 768 | Character | 0.4176 | 0.5700 | 0.6328 | 0.4176 | 0.2375 | 0.1672 | 0.1624 | 0.3188 | 0.3920 | 0.4825 | 0.3036 | 0.3829 |
| Recursive | 0.4073 | 0.5564 | 0.6264 | 0.4073 | 0.2209 | 0.1499 | 0.1811 | 0.3508 | 0.4276 | 0.4719 | 0.3270 | 0.3953 | ||
| Sentence-aware | 0.4347 | 0.5766 | 0.6422 | 0.4347 | 0.2434 | 0.1746 | 0.1620 | 0.3132 | 0.3942 | 0.4969 | 0.3050 | 0.3898 | ||
| Ara-MPNet | 768 | Character | 0.1807 | 0.3253 | 0.4012 | 0.1807 | 0.0891 | 0.0632 | 0.0808 | 0.1489 | 0.1872 | 0.2420 | 0.1212 | 0.1693 |
| Recursive | 0.1497 | 0.2925 | 0.3731 | 0.1497 | 0.0755 | 0.0535 | 0.0795 | 0.1531 | 0.1956 | 0.2115 | 0.1206 | 0.1628 | ||
| Sentence-aware | 0.1958 | 0.3456 | 0.4255 | 0.1958 | 0.0976 | 0.0691 | 0.0830 | 0.1540 | 0.1935 | 0.2608 | 0.1263 | 0.1787 | ||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Alkhamisi, A.A.; Bamashmoos, F.; Alsaggaf, W. ArabicEduCrawler: AI-Assisted Focused Crawling and Corpus Construction for Arabic Educational Web Content. Appl. Sci. 2026, 16, 5964. https://doi.org/10.3390/app16125964
Alkhamisi AA, Bamashmoos F, Alsaggaf W. ArabicEduCrawler: AI-Assisted Focused Crawling and Corpus Construction for Arabic Educational Web Content. Applied Sciences. 2026; 16(12):5964. https://doi.org/10.3390/app16125964
Chicago/Turabian StyleAlkhamisi, Afyaa Atyan, Fatmah Bamashmoos, and Wafaa Alsaggaf. 2026. "ArabicEduCrawler: AI-Assisted Focused Crawling and Corpus Construction for Arabic Educational Web Content" Applied Sciences 16, no. 12: 5964. https://doi.org/10.3390/app16125964
APA StyleAlkhamisi, A. A., Bamashmoos, F., & Alsaggaf, W. (2026). ArabicEduCrawler: AI-Assisted Focused Crawling and Corpus Construction for Arabic Educational Web Content. Applied Sciences, 16(12), 5964. https://doi.org/10.3390/app16125964

