Next Article in Journal
MDM-GANSA: A Multi-Distribution Generative Shilling Attack for Recommender Systems
Previous Article in Journal
Immersion as Convergence: How Storytelling, Interaction, and Sensory Design Co-Produce Museum Virtual Reality Experiences
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

SD-CVD Corpus: Towards Robust Detection of Fine-Grained Cyber-Violence Across Saudi Dialects in Online Platforms

by
Abrar Alsayed
*,
Salma Elhag
and
Sahar Badri
Department of Information Systems, Faculty of Computing and Information Technology, King Abdulaziz University, Jeddah 21589, Saudi Arabia
*
Author to whom correspondence should be addressed.
Information 2026, 17(1), 76; https://doi.org/10.3390/info17010076
Submission received: 4 November 2025 / Revised: 2 January 2026 / Accepted: 4 January 2026 / Published: 12 January 2026

Abstract

This paper introduces Saudi Dialects Cyber Violence Detection (SD-CVD) corpus, a large-scale, class-balanced Saudi-dialect corpus for fine-grained cyber violence detection on online platforms. The dataset contains 88,687 Saudi Arabic tweets annotated using a three-level hierarchical scheme that assigns each tweet to one of 11 mutually exclusive classes, covering benign sentiment (positive, neutral, negative), cyberbullying, and seven hate-speech subtypes (incitement to violence, gender, national, social class, tribal, religious, and regional discrimination). To mitigate the class imbalance common in Arabic cyber violence datasets, data augmentation was applied to achieve a near-uniform class distribution. Annotation quality was ensured through multi-stage review, yielding excellent inter-annotator agreement (Fleiss’ κ > 0.89). We evaluate three modeling paradigms: traditional machine learning with TF–IDF and n-gram features (SVM, logistic regression, random forest), deep learning models trained on fixed sentence embeddings (LSTM, RNN, MLP, CNN), and fine-tuned transformer models (AraBERTv02-Twitter, CAMeLBERT-MSA). Experimental results show that transformers perform best, with AraBERTv02-Twitter achieving the highest weighted F1-score (0.882) followed by CAMeLBERT-MSA (0.869). Among non-transformer baselines, SVM is most competitive (0.853), while CNN performs worst (0.561). Overall, SD-CVD provides a high-quality benchmark and strong baselines to support future research on robust and interpretable Arabic cyber-violence detection.

1. Introduction

Social media has rapidly transformed communication patterns across the Arab world, reshaping how people express emotions, share opinions, and engage in events [1]. Among these platforms, Twitter stands out, with over 27 million daily tweets, offering users a space for interaction but also exposing them to cyber violence, including cyberbullying and online hate speech [2,3]. Despite the existence of regional and international legislative frameworks, violent digital content continues to rise in Saudi Arabia, where social media is widely used for both positive and harmful exchanges [4]. Prior research emphasizes the urgent need for effective online mechanisms to mitigate abusive behaviors and protect vulnerable users [1,3]. Studies confirm that toxic content has severe psychological and social effects and is closely linked to hate crimes and rights violations [5]. Therefore, developing automated systems to detect and limit such behavior is critical to safeguarding society and supporting national digital safety initiatives [6].
In this study, cyber violence refers to the intentional use of digital technologies to cause psychological, social, or financial harm to individuals, organizations, or communities. It encompasses two primary categories [7]:
Cyberbullying: Involves harassment, threats, defamation, or ridicule carried out through online or mobile communication channels. It particularly affects youth, leading to severe emotional and psychological distress [8].
Online Hate Speech: Refers to derogatory or discriminatory language directed at individuals or groups based on characteristics such as nationality, gender, religion, social class, or region [9]. In the context of this study, regional hate speech specifically addresses intra-Saudi variations across regions such as Najdi, Hijazi, and Qassimi.
The growth of cyber violence emerges from multiple social elements that interact with technological systems. People can create harmful content through internet anonymity, as they face no accountability, and social media platforms create echo chambers that strengthen extremist beliefs, while their algorithms boost sensational content for increased views [3]. The current digital environment of Arabic social media platforms experiences increased digital aggression because of these social and technological elements.
The Arabic language presents special linguistic challenges because it exists in three forms: Classical Arabic (CA), Modern Standard Arabic (MSA), and Arabic dialects (AD) [1]. Dialects play a significant role in informal communication, particularly on social media platforms like Twitter. Saudi Arabia ranks tenth in the world in terms of Twitter interaction and leads Arabic-speaking countries with 15.67 million active users [10]. This has been confirmed by a study conducted by Mubarak and Darwish, which shows that most dialect-classified Twitter content is in Saudi dialects, demonstrating the need for dialect-based resources [11]. Saudi Vision 2030 [12] emphasizes digital transformation and improving quality of life, which underscores the importance of maintaining a trusted and safe digital environment for communication [13].
Although notable progress has been made in English cyber violence detection, Arabic language resources remain limited [14]. Moreover, much of the existing work still frames the task as binary classification (e.g., hate vs. non-hate) [15] and often focuses on narrowly defined hate domains, such as religion- or gender-based hate speech. In addition, many publicly available corpora suffer from severe class imbalance, where a few dominant categories overshadow rarer forms of cyber violence content; this can bias learning and limit generalizability across contexts [15,16]. The problem is further complicated by the frequent use of emojis, hashtags, and coded expressions that mask hostility or shift meaning [3]. At the operational level, platforms that rely heavily on manual moderation face clear scalability constraints in high-volume environments [14]. Collectively, these limitations underscore the need for a balanced, multi-dimensional Saudi dialect corpus that captures both cyberbullying and diverse hate-speech categories, enabling more robust, fair, and reliable model performance across real-world Saudi social media settings.
This study aims to address this challenge by developing a comprehensive Arabic corpus for cyber violence detection, with a particular emphasis on Saudi dialects. The corpus is designed to support the development of advanced AI models for digital safety, enabling proactive monitoring, early warning of cyberbullying and online hate speech trends, and more effective protection of individuals and communities in online spaces. Accordingly, the study demonstrates both technical feasibility and a scalable strategic foundation for national-level digital protection.
This study makes the following key contributions to Arabic cyber violence research:
Digital Security Enhancement: We introduce the first large-scale Saudi-dialect corpus for cyber violence detection (88,687 tweets), enabling the development and benchmarking of effective models for identifying cyber violence in Saudi social media, and supporting decision-makers with evidence-driven insights to strengthen online safety and moderation.
Hierarchical Annotation: Introduces a multi-level framework distinguishing between benign, cyberbullying, and hate speech tweets, with detailed hate categories.
Balanced Dataset Design: Ensures category parity to mitigate imbalance found in prior corpora.
Comprehensive Model Evaluation: We systematically evaluate three major NLP approaches (machine learning-, deep learning-, and transformer-based models) and demonstrate their comparative performance on Saudi dialectal content.
Exploratory Data Analysis (EDA): We provide a comprehensive statistical characterization of the SD-CVD corpus, reporting label distributions, subtype frequencies, and word- and character-level length profiles. This analysis offers actionable insights that support corpus reuse, reproducibility, and future benchmarking.
Complementary Features: Incorporates sentiment and emoji-derived features to improve model performance
The remainder of this paper is structured as follows. Section 2 reviews related work on Arabic sentiment analysis and cyber violence detection. Section 3 presents the adopted methodology, including data collection, preprocessing, annotation framework, exploratory data analysis, and corpus optimization. Section 4 outlines the experimental setup and evaluation metrics. Section 5 presents and discusses results, and Section 6 concludes with recommendations for future work.

2. Related Work

Research on Arabic abusive and hateful content has progressed along two main lines: resource creation and model development, with an increasing focus on the Saudi dialect. Early sentiment corpora such as AraSenTi-Tweet (17,573 Saudi tweets) [11] and AraCust (20,000 telecom tweets) [17] demonstrated the importance of domain-specific corpora and highlighted the overlap between sentiment and offensive expressions.
A shift toward explicit offensiveness emerged with Alshalan and Al-Khalifa (2020), who introduced a Saudi hate speech corpus (9316 tweets), where CNN achieved F1 = 0.79 [18]. Similarly, Mubarak et al. (2023) collected large-scale emoji-anchored data covering vulgarity, violence, and hate, demonstrating how transformer-based models capture cultural nuances [3]. Multi-label approaches like Zaghouani et al. (2024) expanded scope through So Hateful!, a dataset annotated for hate, offensiveness, humor, and spam, with AraBERT outperforming other models [19].
Subsequent work introduced broader Arabic corpora. Ahmad et al. (2024) compiled a 403,688-tweet Jordanian dataset and compared TF-IDF, Word2Vec, and AraBERT representations, evaluating their performance across seven classifiers [5]. Charfi et al. (2024) advanced this work with the balanced, multi-dialect ADHAR corpus, addressing race, religion, and ethnicity [15]. Within Saudi Arabia, Asiri and Saleh (2024) introduced SOD, the most comprehensive Saudi corpus (24,000 tweets), featuring hierarchical labels, with data augmentation performed best, reaching and achieving F1 ≈ 0.91 using AraBERTv0.2-Twitter [1].
Alghamdi et al. (2021) proposed a two-level machine-learning framework for detecting violent content in Saudi Arabic tweets [20]. The first stage classifies tweets as violent or non-violent, and the second stage further distinguishes violent tweets as either cyberbullying or threatening. Using SVM and Naive Bayes with different feature extraction and preprocessing settings, SVM achieved stronger results, reporting an F1-score of 76.06% for Level 1 and 89.18% for Level 2 [20], while newer transformer models, such as Arabic BERT-Mini [21] and SaudiBERT [22], achieved high performance with reduced computational cost. Studies integrating sentiment and emoji features (e.g., Refs. [23,24]) or evolutionary optimization [6] demonstrated performance improvements, with surveys confirming that data balance and label design remain critical [25]. Complementary emotion-aware modeling that leverages affective cues has also improved hate-speech detection in multi-target settings [26].
Fine-grained multi-label detection gained attention through works like Alghamdi et al. (2021) (two-level violent tweet classifier, F1 ≈ 0.89) [20] and Al Anezi (2022), who introduced a seven-class DRNN model achieving 84% accuracy [27]. Collectively, these studies confirm that dialect-specific, balanced corpora, and hybrid architectures are vital for improving Arabic cyber violence and hate speech detection. Table 1 provides an overview of representative Arabic and Saudi-dialect corpora for sentiment and hate-speech detection. For each dataset, we report the main task, best-model performance, and documented limitations, enabling a concise comparison of coverage and research gaps.
The reviewed studies directly informed the design choices and research directions adopted in this work. While earlier efforts contributed valuable resources and modeling insights, a careful analysis of the related literature revealed several recurring limitations. Saudi dialect–specific resources remain comparatively limited [1,18], particularly for cyber violence and fine-grained harmful-content modeling. Moreover, many prior datasets and benchmarks continue to rely on binary classification, which can yield unreliable predictions for complex cyber violence phenomena [26], and on highly imbalanced label distributions that bias learning toward majority classes and reduce generalizability [15]. These gaps motivated the development of SD-CVD: a large-scale, balanced Saudi-dialect corpus with hierarchical annotations targeting cyberbullying and fine-grained hate speech subtypes across multiple Saudi dialects.
Prior research has also shown that linear models with explicit lexical features can remain highly competitive for text classification tasks [17], which guided our inclusion of strong TF–IDF-based baselines. At the same time, recent advances in Arabic transformer models motivated the incorporation of BERT-based benchmarks to ensure alignment with contemporary NLP practices. Furthermore, our annotation guidelines, error analysis, and evaluation protocol were shaped by observations reported in earlier studies regarding annotation ambiguity, dialectal bias, and sentiment overlap. Overall, the proposed corpus and experimental framework build on existing empirical findings while explicitly addressing open challenges identified in the literature.

3. Adopted Methodology

This section describes the pipeline for building the Saudi Dialect Cyber Violence Detection Corpus (SD-CVD). It covers data collection, preprocessing, the annotation framework, exploratory data analysis, and corpus optimization.

3.1. Data Collection

Tweets were collected between 2020 and 2025, covering major social and regional events that shaped online discourse during this period. The dataset includes Najdi, Hijazi, Eastern, and Qassimi dialects, ensuring balanced representation of Saudi linguistic diversity. Data harvesting was conducted through the Twitter API using a seed list of keywords, emojis, and hashtags related to cyber violence, including cyberbullying and hate speech. The list contained explicit insults and indirect derogatory expressions common in Saudi online contexts, as well as sentiment-related terms. Prior studies demonstrated a strong correlation between negative sentiment and cyber violence [24,28,29].
Access to the Twitter API was obtained via the Data 356 platform after completing registration, and Python scripts were developed in Google Colab to collect and store the tweets.
A lexicon-driven approach guided data collection, consisting of three elements:
Keywords: Related to cyberbullying, gender, tribal, nationality, religious, class, and regional discrimination, in addition to positive and negative words. Neutral expressions were not predefined; tweets without clear sentiment or cyber violence (e.g., factual news or questions) were labeled as Neutral.
Emojis: Emojis often substitute for offensive words [3] and were included to capture both cyber violence and sentiment cues.
Hashtags: Selected to reflect cyber violence themes aligned with the above keyword categories.
The lexicons were designed to capture diverse expressions of cyber violence and sentiment across Saudi dialects. Table 2 presents examples of the keywords, emojis, and hashtags included in the collection process.

3.2. Data Preprocessing

Data preprocessing enhanced classifier performance by cleaning and structuring tweets [30]. Scripts were written in Python 3.12.12 using the NLTK 3.9.1 library [17]. The main steps included the following:
Normalization: Handling elongated or repeated characters.
Tokenization: Splitting sentences into tokens.
Noise Reduction: Removing @mentions, URLs, numbers, special symbols, duplicates, retweets, and non-Arabic content.
Content Filtering: Excluding advertisements, irrelevant hashtags, and short tweets.
Semantic Duplicate Detection: Eliminating tweets with ≥94% similarity (TF-IDF + cosine similarity).

3.3. Annotation Framework

A three-level hierarchical annotation strategy was adopted to assign each tweet to one of the 11 target classes defined in this study. Figure 1 illustrates the hierarchical annotation framework.
Although the annotation schema in Figure 1 is hierarchical, the task is formulated as a hierarchical multiclass classification problem rather than a multi-label one. Each tweet follows a single decision path across the hierarchy and is assigned exactly one mutually exclusive final label among the 11 defined classes. In other words, a tweet annotated as cyber violent cannot simultaneously belong to the benign category, and a tweet labeled as hate speech is further assigned to only one specific hate subtype. This design choice simplifies model training and evaluation, ensures label consistency, and aligns with prior Arabic cyber violence and hate-speech corpora that adopt hierarchical but single-label annotation schemes [1,20].

3.3.1. Annotation Setup

To streamline the manual annotation process, we developed a custom, secure web-based annotation platform featuring an interactive, game-like interface to enable controlled access, progress tracking, and consistent workflow management. A total of 14 native Saudi annotators were selected from 26 applicants based on predefined qualification criteria to conduct the labeling task. Figure 2 illustrates the annotator recruitment and selection process.

3.3.2. Saudi Dialects Annotation Guidelines

To ensure robust coverage of Saudi dialectal variation, we recruited native Saudi annotators from multiple regions to capture diverse dialect usage. We developed dialect-aware annotation guidelines tailored to Saudi Arabic that define class boundaries, highlight key linguistic cues, and codify explicit decision rules [5,19]. Before launching full-scale labeling, annotators completed a training program that introduced the schema and operationalized the guidelines through supervised practice and iterative feedback, aligning judgments and improving labeling consistency.
Core Annotation Principles
Identify Saudi dialect tweets; mark non-Saudi tweets as Delete.
Read entire tweets to interpret tone and context.
Label insults, profanity, threats, or incitement as Cyber Violence; otherwise, classify as Benign.
Identify whether the target is an individual or group.
When uncertain, assign the tweet to Indeterminate Classification.
Treat emojis as part of text; they may express or intensify cyber violence intent.
Table 3 presents the categories considered in this study and their corresponding definitions. Providing clear definitions for each category ensured consistent understanding of class boundaries and supported systematic, high-quality annotation.

3.3.3. Annotation Process

After the training program, annotators accessed the labeling platform individually through secure sessions. Tweets were displayed one at a time for manual classification (Figure 3). For coordination and clarification, a WhatsApp group allowed annotators to communicate directly with supervisors. Consistent communication and software-controlled workflows improved annotation reliability.

3.3.4. Quality Assurance and Inter-Annotator Agreement

To ensure reliability, each labeled tweet underwent a secondary review by another annotator. Disagreements were resolved through a third annotator’s decision. Statistical reliability was evaluated using Fleiss’ Kappa, which exceeded 0.89, indicating excellent agreement among annotators [32]. This high score reflects the effectiveness of the training program and follow-up supervision.

3.3.5. Annotation Challenges

Several challenges emerged during annotation:
Supplications: Difficult to classify due to unclear sentiment direction. Following [11], such tweets were classified as Delete.
Quotes: Often positive but not opinion-based; classified as Neutral [11].
Target Ambiguity: Tweets lacking explicit targets were labeled Undetermined to prevent speculation.
News Content: News-style tweets were assigned to the Neutral category, consistent with prior Arabic sentiment studies [11].
Spelling Errors: Tweets containing extensive spelling errors were labeled Delete to improve data clarity and annotation consistency.

3.4. Exploratory Data Analysis

The exploratory data analysis (EDA) examined class distribution, tweet lengths, and word statistics to reveal linguistic and structural patterns of Saudi cyber violence. Figure 4 and Table 4 summarize the results.
In addition to the visual distribution illustrated in Figure 4, Table 4 presents detailed exploratory statistics for each class, including tweet length and character count, highlighting structural differences between benign and cyber violent categories.

3.5. Corpus Optimization

Imbalanced class distribution is a persistent issue in Arabic cyber violence datasets [15,33]. To mitigate bias toward majority classes [15,16], we applied targeted data augmentation to achieve near-equal representation across labels. The Positive class originally contained 8045 tweets; all other classes were incrementally augmented to comparable counts (~8000 each), prioritizing preservation of label semantics and dialectal cues. Table 5 presents the distribution of original and augmented tweets for each class. Consistent with prior work on augmentation for abusive-language detection, this strategy balances categories without additional data collection and improves robustness during training [1,34,35].
Our data augmentation strategy is based on LLM-driven text augmentation to expand the corpus. The Gemini (gemini-2.5-pro) model generates multiple paraphrased variants for each original tweet that preserve the original meaning and remain consistent with the assigned label. The augmentation prompt contains a complete label definition which requires three conditions for enforcement: (i) linguistic diversity via rephrasing, synonym substitution, word reordering, and minor contextual variations; (ii) strict label preservation; and (iii) natural expression in Saudi dialectal Arabic. The model produces different output results because of its non-deterministic decoding process (temperature = 0.7, top-p = 0.8, top-k = 40) which to increase diversity while maintaining coherence. The generated samples need further processing to remove quotation marks and short responses while keeping the number of generated outputs equal to the number of tweets. To ensure reproducibility and prevent duplication, processed tweet IDs are tracked using a JSON progress file, and all augmented samples are stored in a CSV file with appropriate metadata. Parallel generation is implemented using a thread pool to improve efficiency.
Throughout the augmentation process, dialect-specific lexical markers and regional linguistic cues were explicitly preserved, and no synthetic samples were generated across dialect boundaries, ensuring that the original dialectal characteristics of the corpus were maintained.
Table 5. Class-wise distribution of original and augmented tweets in the SD-CVD corpus.
Table 5. Class-wise distribution of original and augmented tweets in the SD-CVD corpus.
LabelNumber of Original
Tweets
Number of Augmented
Tweets
Positive8045-
Negative59032107
Neutral7061944
Cyberbullying70421046
Incitement to Violence7255805
Gender Discrimination7219825
National Discrimination55852428
Social Class Discrimination44153622
Tribal Discrimination40953832
Religious Discrimination55592480
Regional Discrimination42554163
Following augmentation, the final corpus comprised 88,687 tweets, ensuring a near-uniform distribution across all classes, reducing classifier bias toward the largest class, and enabling more reliable, comparable evaluation metrics.

4. Results

The following subsections describe the experimental setting used in this study, along with the results obtained from the conducted experiments. The results are organized into three approaches: machine learning, deep learning, and transformer-based models. This structure enables a detailed analysis of individual classifiers within each approach, while the final subsection provides an overall comparative analysis across the three approaches based on their experimental results.

4.1. Machine Learning (ML) Models

The study applies supervised learning because it suits sentiment analysis and cyber violence detection [36]. We evaluate three machine learning algorithms: SVM, LR, and RF as baseline models, as prior work shows they perform well for these tasks [37]. Since our task is text classification and requires numerical inputs, feature extraction is applied to transform raw text into features that capture word patterns, contextual relationships, and semantic content. Two feature extraction methods are used, combining n-grams with TF-IDF. The SD-CVD corpus is split 80/20 for training and testing across all experiments, consistent with prior works [1,16]. Evaluation metrics include accuracy, precision, recall, and F1-score to comprehensively assess model performance. To ensure interpretability, we report the results algorithm by algorithm. In addition, for the best-performing ML model, we conduct a deeper analysis using a confusion matrix.

4.1.1. Support Vector Machine (SVM)

Support Vector Machine (SVM) is a margin-based linear classifier well-suited for high-dimensional TF–IDF feature spaces. In this study, SVM achieved the best overall performance among all machine learning models. Table 6 summarizes the weighted performance metrics for the Support Vector Machine (SVM) model, highlighting its strong generalization across multiple text categories.
As illustrated in Table 6, the SVM model achieved weighted-average performance of accuracy = 0.854, precision = 0.852, recall = 0.854, and F1-score = 0.853. At the subclass level, the strongest results were obtained for regional discrimination (F1 ≈ 0.954), whereas the weakest performance occurred within the benign cluster, particularly for benign_neutral (F1 ≈ 0.651). Among the ML algorithms evaluated, SVM achieved the highest overall performance; therefore, its confusion matrix is presented and analyzed in detail to illustrate the underlying classification behavior. The confusion matrix in Figure 5 highlights a strong diagonal, with errors concentrated among the benign categories.
As shown in Figure 5, the confusion matrix of the SVM model illustrates a strong diagonal, confirming its high accuracy across most categories. The classifier demonstrates particularly high recognition for regional discrimination (1610), social class discrimination (1506), and tribal discrimination (1494), with similarly strong results for incitement to violence (1457) and gender discrimination (1492). In contrast, the benign categories show notable overlap: benign_negative (1182 correct) is frequently misclassified as benign_neutral or benign_positive, and benign_neutral (1042 correct) suffers from confusion with other benign labels. A moderate degree of misclassification also occurs between cyberbullying and gender discrimination, reflecting the semantic closeness of these categories. Further error analysis indicates that some cyberbullying tweets are misclassified as specific hate-speech subtypes, particularly gender-based hate speech. This confusion arises largely because personal attacks in Arabic often employ gender-specific insults or stereotypes that simultaneously target individuals and invoke group-based discriminatory language. Consequently, the linguistic overlap between cyberbullying and gender-related hate blurs the distinction between these categories, especially in short and context-limited tweets. This observation suggests that incorporating explicit target modeling or richer contextual information could help future models more effectively differentiate between individual-directed harassment and group-based discrimination. Overall, the figure confirms that SVM delivers robust classification for hate-speech subtypes, while most errors arise from semantic ambiguity within benign expressions.

4.1.2. Logistic Regression (LR)

Logistic Regression (LR) is a linear probabilistic model that estimates class membership based on maximum likelihood. In this study, LR achieved performance very close to SVM, confirming its strength in text classification tasks with sparse TF–IDF features. Table 7 presents the weighted performance metrics for the LR model, providing a linear baseline for comparison with SVM.
As indicated in Table 7, the LR model achieved weighted average performance of accuracy = 0.850, precision = 0.848, recall = 0.850, and F1-score = 0.849. At the subclass level, LR showed strong performance for regional discrimination (F1 ≈ 0.950). However, its weakest performance was observed for benign_neutral (recall ≈ 0.653), reflecting challenges in distinguishing subtle differences between neutral and negative expressions.

4.1.3. Random Forest (RF)

Random Forest (RF) is an ensemble-based learning algorithm that combines multiple decision trees through bagging to capture non-linear relationships. Despite its robustness in many domains, RF underperformed compared to the linear baselines in this study. Table 8 summarizes the weighted performance metrics of the RF algorithm, including accuracy, precision, recall, and F1-score.
According to the results reported in Table 8, the RF model achieved weighted average performance of accuracy = 0.769, precision = 0.768, recall = 0.769, and F1-score = 0.767. At the subclass level, RF performed best on social class discrimination (F1 ≈ 0.869). However, it showed a clear weakness on benign_neutral (F1 ≈ 0.574), highlighting its limitations when modeling high-dimensional, sparse TF–IDF feature spaces.

4.2. Deep Learning (DL) Models

In this study, four deep learning (DL) algorithms were evaluated using sentence-level embedding vectors for feature extraction. Embeddings are predictive, distributed representations of text that map words into dense numerical vectors, allowing models to learn semantic features during training. Sentence embeddings with fixed dimensionality (d = 768) were used as input for all deep learning models to ensure consistency and a fair comparison across architectures.
For a fair comparison across deep learning models, we used a unified training setup: AdamW (learning rate = 0.001, weight decay = 0.01) with cross-entropy loss. We applied ReduceLROnPlateau (factor = 0.7, patience = 3), monitoring validation macro-F1, along with early stopping with a patience of 5 epochs. Models were trained for up to 15 epochs, using a batch size of 32 for training and 64 for evaluation.
The SD-CVD corpus was split using the same ratio as in the ML experiments (80/20) for training and testing across all experiments. We report accuracy, precision, recall, and F1-score to evaluate performance. Results are presented model by model, and the best DL model is further analyzed using a confusion matrix.

4.2.1. Long Short-Term Memory (LSTM)

Long Short-Term Memory (LSTM) is a gated recurrent neural network architecture designed to capture long-range dependencies in sequential data. In this study, the LSTM-based classifier takes a single fixed-size sentence embedding per instance (sequence length = 1) and passes it through a two-layer LSTM encoder with a hidden size of 256, followed by a fully connected classification head that maps the final hidden representation to the target classes. Table 9 reports the weighted performance metrics of the LSTM model.
As summarized in Table 9, the model achieved weighted average performance of accuracy = 0.813, precision = 0.813, recall = 0.813, and F1-score = 0.812. At the subclass level, it performed strongly on regional discrimination (F1 ≈ 0.916). The weakest scores were observed for cyberbullying (F1 ≈ 0.666), suggesting difficulties in capturing subtle semantic variations within this class.

4.2.2. Recurrent Neural Network (RNN)

The Recurrent Neural Network (RNN) functions as a basic recurrent model that serves as a simpler alternative to LSTM. In our experiment, we employed an improved RNN classifier that processes sentence embeddings through LayerNorm normalization before using the recurrent block for processing. The embeddings undergo a single-step sequence transformation, resulting in a sequence length of 1 before they enter a 3-layer network. The RNN encoder contains 256 hidden units, uses ReLU activation, and applies dropout at a rate of 0.4. The model also uses dropout to regularize its final hidden state before passing it through a fully connected classification layer that generates predictions for the target classes. The RNN model’s performance metrics are reported in Table 10, which contains weighted values.
As shown in Table 10, the model achieved weighted average performance of accuracy = 0.812, precision = 0.814, recall = 0.812, and F1-score = 0.812. At the subclass level, the RNN performed strongly on regional discrimination (F1 ≈ 0.921). The weakest performance was observed for benign_neutral (F1 ≈ 0.642), reflecting challenges in distinguishing subtle neutral expressions from other benign subcategories.

4.2.3. Multilayer Perceptron (MLP)

The Multilayer Perceptron (MLP) operates as a feed-forward neural network which performs non-linear transformations on dense feature representations. The research results showed that MLP produced the highest overall performance among deep learning models which proved its capability to handle complex non-linear patterns in text information. The Enhanced MLP architecture uses sentence-level embedding vectors as input data which receives LayerNorm normalization at its first stage. The network uses a sequence of fully connected layers with hidden dimensions [64,128,256,512] to process the embeddings. The network contains four blocks which apply Linear → BatchNorm → ReLU → Dropout operations with decreasing dropout rates from 0.5 × 0.8i. The last hidden representation becomes transformed into target classes through a final linear layer. The weighted performance metrics of the MLP model appear in Table 11.
As illustrated in Table 11, the model achieved weighted averages of Accuracy = 0.831, Precision = 0.830, Recall = 0.831, and F1-score = 0.830, with its best subclass performance in social class discrimination (F1 ≈ 0.916) and its weakest results in the benign cluster, particularly benign_negative (F1 ≈ 0.718), indicating persistent difficulty in distinguishing benign subcategories. Since MLP was the highest performing deep learning model in this study, Figure 6 presents its confusion matrix, which shows a clear diagonal reflecting strong overall recognition, while most remaining errors are concentrated among benign labels due to their semantic overlap and ambiguous boundaries.
As shown in Figure 6, the confusion matrix of the Enhanced MLP model reveals a strong diagonal structure, indicating reliable classification performance across most hate-speech subtypes. The model demonstrates particularly high true positive rates for regional, social class, and tribal discrimination, confirming its effectiveness in capturing non-linear semantic patterns in these categories. In contrast, most misclassifications occur within the benign cluster, especially between benign_negative and benign_neutral, reflecting the inherent semantic overlap and subjectivity of sentiment boundaries in short Arabic tweets. A moderate level of confusion is also observed between cyberbullying and gender-based hate speech, which can be attributed to the frequent use of gendered insults in personal attacks, blurring the distinction between individual-directed harassment and group-based discrimination. Overall, the confusion matrix analysis confirms that the MLP model is robust for fine-grained hate-speech detection, while highlighting sentiment ambiguity and target overlap as key challenges for future model improvements.

4.2.4. Convolutional Neural Network (CNN)

The Convolutional Neural Network (CNN) uses convolutional filters to capture local patterns from the input representations. While CNNs are powerful for image processing and in text tasks where local cues are informative, they are generally less suited to modeling long-range contextual dependencies, which are often present in Arabic social media content. The CNN takes sentence embeddings as input, which are first normalized using LayerNorm and then reshaped into a 1D single-channel representation. The input is processed by three parallel Conv1D branches with 100 filters each and kernel sizes of 3, 5, and 7 (with padding). Each branch applies BatchNorm and ReLU, followed by global average pooling (AdaptiveAvgPool1d(1)). The pooled features are concatenated and regularized using Dropout (0.5), then passed through a 256-dimensional fully connected layer, followed by Dropout (0.3), and a final linear layer that outputs predictions for the target classes. The CNN performance metrics are reported in Table 12.
As indicated in Table 12, the CNN model achieved weighted average performance of accuracy = 0.571, precision = 0.560, recall = 0.571, and F1-score = 0.561, making it the weakest among all evaluated models. At the subclass level, it achieved relatively better performance on social class discrimination (F1 ≈ 0.756). In contrast, it performed poorly across the benign cluster, particularly benign_negative (F1 ≈ 0.315). Overall, these results suggest that CNN’s emphasis on local patterns is insufficient for distinguishing subtle sentiment boundaries and capturing context-dependent cues that characterize cyber violence content in Arabic social media.

4.3. Transformer-Based Models

Bidirectional Encoder Representations from Transformers (BERT) is a pretrained language model based on the Transformer architecture that uses only the encoder stack. Two BERT-based models were fine-tuned on the SD-CVD corpus: the first model, AraBERTv02-Twitter, and the second, CAMeLBERT-MSA. The SD-CVD corpus was randomly partitioned into train–dev–test splits of 70:15:15, and all reported results are based on the test set. For the hyperparameter settings, we adopted a unified configuration for both models. In particular, each model was fine-tuned with a maximum input sequence length of 512 tokens, a per-device batch size of 8 with gradient accumulation of 4 steps (effective batch size of 32), an initial learning rate of 5 × 10−5, and 15 training epochs. Training used the HuggingFace Trainer with weight decay of 0.01, cosine_with_restarts learning-rate scheduling with a warm-up ratio of 0.15, and label smoothing of 0.1. We employed class-weighted focal loss computed from the training split. Mixed-precision training (fp16) and early stopping (patience = 5) were enabled. Experiments were conducted in a GPU-accelerated Google Colab environment using Python and open-source libraries (HuggingFace Transformers, PyTorch, and scikit-learn).

4.3.1. AraBERTv02-Twitter

AraBERTv02-Twitter is a variant of the AraBERT family specifically adapted for Arabic social media text. Unlike the standard AraBERTv02 model, which is primarily trained on Modern Standard Arabic (MSA), AraBERTv02-Twitter was designed to better capture the linguistic characteristics of Arabic dialects commonly used on Twitter. The model was further pretrained using masked language modeling on ~60 M Arabic tweets (filtered from 100 M), improving coverage of dialectal and informal expressions. This pretraining was conducted on 77 GB of data with a 64 K vocabulary, and the model supports both MSA and Arabic dialects (AD) [38,39]. Table 13 illustrates the fine-tuning results of the AraBERTv02-Twitter model on the SD-CVD corpus.
Table 13 presents the results obtained using the AraBERTv02-Twitter model. The model achieved weighted average performance of accuracy = 0.883, precision = 0.883, recall = 0.883, and F1-score = 0.882. At the subclass level, the strongest performance was observed for regional discrimination (F1 ≈ 0.963), whereas the weakest performance occurred for benign_neutral (F1 ≈ 0.718). As AraBERTv02-Twitter achieved higher performance than CAMeLBERT-MSA, we conducted a more detailed analysis and present its confusion matrix in Figure 7. The matrix exhibits a pronounced diagonal, indicating strong overall classification, while the remaining errors are mainly concentrated among the benign classes due to semantic overlap and blurred boundaries between closely related benign expressions.
As shown in Figure 7, the AraBERTv02 Twitter confusion matrix is strongly diagonal, indicating that the model correctly assigns the majority of samples to their true classes and achieves consistently high recognition across the label set. Performance is particularly strong for the fine-grained hate speech subtypes, most notably regional, tribal, and social class discrimination, which suggests that the model effectively exploits contextual cues and dialectal usage patterns learned from large-scale Arabic Twitter data.
However, the residual errors exhibit clear structure rather than noise. Misclassifications are concentrated within the benign cluster, where benign neutral is frequently confused with benign negative, highlighting the intrinsic difficulty of separating subtle, mixed, or weakly expressed sentiment in short, informal social media posts. In addition, a moderate overlap is observed between cyberbullying and gender-based hate speech, which is a plausible outcome given that gendered expressions often function as insult markers and can blur the boundary between targeted hate and interpersonal harassment. Overall, this error analysis reinforces the robustness of AraBERTv02 Twitter for SD-CVD corpus, while also showing that sentiment ambiguity and target or intent overlap remain the primary sources of confusion, even under transformer-based representations.

4.3.2. CAMeLBERT-MSA

CAMeLBERT-MSA is a monolingual Arabic BERT-style model trained on 107 GB of Modern Standard Arabic from news, Wikipedia, OSCAR, and other large MSA corpora. It uses a 30 k WordPiece vocabulary, whole-word masking, and is pretrained for 1 M steps (first with a maximum sequence length of 128, then 512) using the original BERT pretraining setup. Prior work indicates that CAMeLBERT MSA performs well on several Arabic downstream tasks, such as text classification and sentiment analysis [40]. Table 14 reports the weighted performance metrics of the CAMeLBERT-MSA model on the SD-CVD corpus.
As shown in Table 14, the CAMeLBERT-MSA model achieves weighted averages of Accuracy = 0.871, Precision = 0.869, Recall = 0.871, and F1-score = 0.869. At the subclass level, the strongest results were obtained for regional discrimination (F1 ≈ 0.959), while the weakest performance occurred in benign_neutral (F1 ≈ 0.695).

4.4. Comparative Analysis

To consolidate the findings across individual models, this subsection presents a comparative summary of the F1-score across the three evaluated approaches: machine learning (ML)-, deep learning (DL)-, and transformer-based models (BERT models). As shown in Figure 8, the BERT-based models, particularly AraBERTv02 Twitter, followed by CAMeLBERT MSA, consistently outperformed both traditional ML and DL baselines. This result underscores the advantage of pretrained contextual representations for capturing complex linguistic patterns, implicit meaning, and dialect-specific nuances that are characteristic of Saudi Arabic.
Across the non-transformer baselines, the results indicate that ML models generally outperform DL models. In particular, linear classifiers such as SVM and LR achieved the highest accuracy and weighted metrics, surpassing the performance of recurrent and convolutional architectures. This trend suggests that, when using sparse, high-dimensional feature representations such as TF-IDF and n-gram features in settings characterized by class overlap and dialectal variability, linear decision boundaries can remain highly competitive and often more stable. In contrast, the CNN model yielded the weakest performance among all experiments, suggesting limited effectiveness under the current feature representation and training configuration.

5. Discussion

Arabic’s rich morphology and wide dialect variation make cyber violence detection especially difficult in informal tweets. The SD-CVD corpus addresses key gaps in prior Arabic resources by focusing on multiple Saudi dialects, enabling more robust and culturally informed detection. Its balanced distribution across eleven classes supports fair training and reliable evaluation, reducing bias caused by class imbalance. In addition, the corpus follows a rigorous collection and annotation process with multiple native Arabic annotators and reported inter-annotator agreement, strengthening the reliability and overall quality of the dataset.
Extensive experiments across ML-, DL-, and transformer-based models consistently confirm that SD-CVD provides an effective benchmark for accurate cyber violence classification. The results confirm the advantage of BERT-based models for Saudi dialect cyber violence detection, where AraBERTv02-Twitter achieved the top weighted F1-score (0.882), followed by CAMeLBERT-MSA (0.869). The gap is expected given AraBERTv02-Twitter’s closer pretraining domain (Arabic Twitter), which better reflects informal spelling, dialect markers, and emoji-heavy discourse typical of SD-CVD. These findings reinforce prior evidence that domain- and dialect-matched pretraining is crucial for robust Arabic harmful-content detection [1,3].
These outcomes are consistent with the findings of Ahmad et al. [5], who reported that SVM and LR achieved competitive results on a large-scale Jordanian dialect corpus, while fine-tuned transformer models provided superior contextual understanding when sufficient training data were available. Similarly to their study, our results confirm that carefully engineered textual features remain highly effective as baselines, but pretrained transformers provide additional gains when handling fine-grained and context-dependent cyber violence categories.
In contrast, Anezi [27] proposed a deep recurrent neural network and reported high accuracy on smaller, balanced datasets. However, as observed in our experiments, deep learning architectures based on fixed sentence embeddings did not surpass either SVM or BERT-based models, particularly when applied to a large, multi-dialect corpus with subtle interclass boundaries. This suggests that deep architectures may require either end-to-end contextual learning or significantly richer supervision to outperform linear and transformer-based approaches.
Similarly, Mubarak et al. [3] demonstrated that transformer-based models such as AraBERT outperform classical classifiers for Arabic offensive and hate speech detection, especially when emojis and dialectal cues are involved. Our findings reinforce this observation, as the best results were achieved by AraBERTv02-Twitter, which is explicitly pretrained on Arabic Twitter data and dialectal content.
Overall, the comparative analysis confirms that transformer-based models represent the state of the art for Saudi dialect cyber violence detection, while traditional ML models such as SVM remain strong, interpretable, and computationally efficient baselines. Deep learning models without contextual pretraining showed comparatively weaker performance, supporting previous conclusions that dataset scale, representation choice, and annotation granularity critically influence model effectiveness [3,27].
Beyond the supervised models evaluated in this study, recent work on Large Language Models (LLMs) has introduced alternative directions for social media moderation that go beyond traditional supervised classifiers. In contrast to task-specific fine-tuned models, LLM-based approaches often rely on instruction prompting, reasoning mechanisms, or logic-augmented inference to detect cyber violence content. One representative example is FOLAR and related reasoning aware co-design frameworks, in which LLM representations are combined with logical constraints to achieve more robust and consistent stance detection and harmful-content detection on social media platforms [41]. Similarly, LogiMDF, proposed by [42], integrates LLM predictions through logical multi-decision fusion, resulting in improved moderation reliability in complex online discussions and a reduction in spurious predictions.
Although recent studies have explored Large Language Models (LLMs) as an alternative direction for social media moderation, their direct application to Arabic dialectal cyber violence detection remains challenging and was therefore not evaluated in this study. Despite their strong generalization capabilities, LLMs often suffer from dialectal and cultural bias, as most are predominantly trained on English and Modern Standard Arabic, leading to reduced sensitivity to Saudi dialects, idiomatic expressions, and culturally grounded insults [43]. In addition, hallucination and over-generalization can introduce false positives, which is particularly problematic for benign or sarcastic content and for fine-grained cyber-violence detection tasks [44]. Furthermore, deploying LLMs for large-scale monitoring of user-generated content raises ethical and privacy concerns, including unintended disclosure of sensitive information and the reinforcement of social biases [45].

6. Conclusions

This study introduced SD-CVD, a large-scale and balanced Saudi dialect corpus for fine-grained cyber violence detection. The corpus covers diverse Saudi dialects, adopts a three-level hierarchical annotation scheme encompassing eleven classes across cyber violence and benign content, and achieves excellent inter-annotator agreement (Fleiss’ κ > 0.89), ensuring high annotation reliability and consistency. Extensive experiments demonstrate that fine-tuned transformer-based models deliver the strongest performance on SD-CVD, with AraBERTv02-Twitter achieving the highest weighted F1-score (0.882), followed by CAMeLBERT-MSA (0.869), highlighting the importance of contextual pretraining and domain alignment for capturing dialectal variation, implicit meaning, and culturally grounded expressions in Saudi social media content. Although transformer models outperform all other approaches, traditional machine learning baselines remain strong and competitive; in particular, SVM achieved the best ML performance (F1 = 0.853), confirming that linear classifiers combined with TF–IDF and n-gram features continue to provide interpretable, computationally efficient, and robust solutions for Arabic cyber violence detection. Deep learning models based on fixed sentence embeddings (e.g., LSTM, RNN, and MLP) achieved moderate results but did not surpass either transformer-based or linear ML approaches, underscoring the value of contextual representations and large-scale pretraining. Error analysis further indicates that most residual misclassifications arise from semantic overlap within benign categories, especially between neutral and negative tweets, and linguistic overlap between cyberbullying and certain hate-speech subtypes such as gender-based discrimination. Despite its strengths, SD-CVD has limitations: it does not explicitly model sarcasm or more implicit, context-dependent forms of cyber violence, and its static nature means periodic updates are necessary as online language and abuse patterns evolve. Nevertheless, SD-CVD fills a critical gap in Arabic NLP by offering a high-quality Saudi-dialect benchmark with balanced coverage and reproducible baselines, enabling more robust and inclusive cyber-violence detection research. Future work will explore hybrid and ensemble frameworks that combine transformer embeddings with lightweight linear classifiers and will extend the corpus to better capture sarcasm and implicit cyber violence, improving robustness, interpretability, and fairness in Saudi-dialect cyber-violence detection systems.

Author Contributions

Conceptualization, S.E.; methodology, S.E. and S.B.; software, A.A.; validation, A.A.; formal analysis, A.A.; investigation, A.A.; resources, A.A.; data curation, A.A.; writing—original draft, A.A.; writing—review and editing, S.E. and S.B.; visualization, A.A.; supervision, S.E.; project administration, S.E. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded by the General Authority for Defense Development (GADD) under Project No. 262.

Institutional Review Board Statement

In this study, all collected and processed data were fully public and did not contain any personal identifiers that could reveal or trace the identity of any individual. Therefore, in accordance with the Saudi Personal Data Protection Law (PDPL), prior ethical approval or IRB clearance was not required.

Informed Consent Statement

In this study, all collected and processed data were fully public and did not contain any personal identifiers that could reveal or trace the identity of any individual. Therefore, in accordance with the Saudi Personal Data Protection Law (PDPL), informed consent was not required.

Data Availability Statement

The dataset used in this study consists of Arabic social media posts that may contain offensive or harmful language. Due to ethical restrictions, the dataset cannot be made publicly available. However, it can be obtained from the corresponding author upon reasonable request.

Acknowledgments

This research was supported by the General Authority for Defense Development (GADD), under Pro-ject No. 262. The authors gratefully acknowledges this financial support, which covered the full re-search and development costs of the SD-CVD project.

Conflicts of Interest

The authors declare no conflict of interest.

References

  1. Asiri, A.; Saleh, M. Sod: A corpus for saudi offensive language detection classification. Computers 2024, 13, 211. [Google Scholar] [CrossRef]
  2. Kaur, S.; Singh, S.; Kaushal, S. Deep learning-based approaches for abusive content detection and classification for multi-class online user-generated data. Int. J. Cogn. Comput. Eng. 2024, 5, 104–122. [Google Scholar] [CrossRef]
  3. Mubarak, H.; Hassan, S.; Chowdhury, S.A. Emojis as anchors to detect arabic offensive language and hate speech. Nat. Lang. Eng. 2023, 29, 1436–1457. [Google Scholar] [CrossRef]
  4. Mousa, A.; Shahin, I.; Nassif, A.B.; Elnagar, A. Detection of arabic offensive language in social media using machine learning models. Intell. Syst. Appl. 2024, 22, 200376. [Google Scholar] [CrossRef]
  5. Ahmad, A.; Azzeh, M.; Alnagi, E.; Al-Haija, Q.A.; Halabi, D.; Aref, A.; AbuHour, Y. Hate speech detection in the arabic language: Corpus design, construction, and evaluation. Front. Artif. Intell. 2024, 7, 1345445. [Google Scholar] [CrossRef] [PubMed]
  6. Shannaq, F.; Hammo, B.; Faris, H.; Castillo-Valdivieso, P.A. Offensive language detection in arabic social networks using evolutionary-based classifiers learned from fine-tuned embeddings. IEEE Access 2022, 10, 75018–75039. [Google Scholar] [CrossRef]
  7. Council of Europe. Cyberviolence. Council of Europe Website, 2025. Available online: https://www.coe.int/en/web/cyberviolence (accessed on 18 December 2025).
  8. Ministry of Health. Available online: https://www.moh.gov.sa/HealthAwareness/EducationalContent/BabyHealth/Pages/Bullying.aspx (accessed on 25 March 2025).
  9. KAICIID. Stop Hate Speech. 2019. Available online: https://www.kaiciid.org/resources/publications/quick-guide-hate-speech-prevention (accessed on 3 November 2025).
  10. Statista. Number of Active Twitter Users in Selected Countries. Available online: https://www.statista.com/statistics/242606/number-of-active-twitter-users-in-selected-countries/ (accessed on 17 October 2025).
  11. Al-Twairesh, N.; Al-Khalifa, H.; Al-Salman, A.; Al-Ohali, Y. Arasenti-tweet: A corpus for arabic sentiment analysis of saudi tweets. Procedia Comput. Sci. 2017, 117, 63–72. [Google Scholar] [CrossRef]
  12. Saudi Vision 2030. Available online: https://www.vision2030.gov.sa/en (accessed on 19 December 2025).
  13. Dhiaa, M.; Atta, R.; Mohammed Abbas, A.; Kamel, M.; Mohammed, M.; Ali, H.; Ali, J.; Fayez, G.; Issa, M.; Dalal, A.; et al. A machine learning approach to cyberbullying detection in arabic tweets. Comput. Mater. Contin. 2024, 80, 1033. [Google Scholar] [CrossRef]
  14. Duwairi, R.; Hayajneh, A.; Quwaider, M. A deep learning framework for automatic detection of hate speech embedded in arabic tweets. Arab. J. Sci. Eng. 2021, 46, 4001–4014. [Google Scholar] [CrossRef]
  15. Charfi, A.; Besghaier, M.; Akasheh, R.; Atalla, A.; Zaghouani, W. Hate speech detection with adhar: A multi-dialectal hate speech corpus in arabic. Front. Artif. Intell. 2024, 7, 1391472. [Google Scholar] [CrossRef] [PubMed]
  16. Alzaqebah, M.; Jaradat, G.M.; Nassan, D.; Alnasser, R.; Alsmadi, M.K.; Almarashdeh, I.; Jawarneh, S.; Alwohaibi, M.; Al-Mulla, N.A.; Alshehab, N.; et al. Cyberbullying detection framework for short and imbalanced arabic datasets. J. King Saud Univ.-Comput. Inf. Sci. 2023, 35, 101652. [Google Scholar] [CrossRef]
  17. Almuqren, L.; Cristea, A. Aracust: A saudi telecom tweets corpus for sentiment analysis. PeerJ Comput. Sci. 2021, 7, e510. [Google Scholar] [CrossRef]
  18. Alshaalan, R.; Al-Khalifa, H. Hate Speech Detection in Saudi Twittersphere: A Deep Learning Approach. In Proceedings of the Fifth Arabic Natural Language Processing Workshop, Barcelona, Spain, 12 December 2020; pp. 12–23. [Google Scholar]
  19. Zaghouani, W.; Mubarak, H.; Biswas, M.R. So hateful! building a multi-label hate speech annotated arabic dataset. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italy, 20–25 May 2024; pp. 15044–15055. [Google Scholar]
  20. Alghamdi, D.; Al-Motery, R.; Alma’Abdi, R.; Alzamzami, O.; Babour, A. Automatic detection of cyberbullying and threatening in saudi tweets using machine learning. Int. J. Adv. Appl. Sci. 2021, 8, 17–25. [Google Scholar] [CrossRef]
  21. Almaliki, M.; Almars, A.M.; Gad, I.; Atlam, E.-S. Abmm: Arabic bert-mini model for hate-speech detection on social media. Electronics 2023, 12, 1048. [Google Scholar] [CrossRef]
  22. Qarah, F. Saudibert: A large language model pretrained on saudi dialect corpora. arXiv 2024, arXiv:2405.06239. [Google Scholar] [CrossRef]
  23. Alfreihat, M.; Almousa, O.S.; Tashtoush, Y.; Al-Sobeh, A.; Mansour, K.; Migdady, H. Emo-sl framework: Emoji sentiment lexicon using text-based features and machine learning for sentiment analysis. IEEE Access 2024, 12, 81793–81812. [Google Scholar] [CrossRef]
  24. Althobaiti, M.J. Bert-based approach to arabic hate speech and offensive language detection in twitter: Exploiting emojis and sentiment analysis. Int. J. Adv. Comput. Sci. Appl. 2022, 13, 972–980. [Google Scholar] [CrossRef]
  25. Subramanian, M.; Sathiskumar, V.E.; Deepalakshmi, G.; Cho, J.; Manikandan, G. A survey on hate speech detection and sentiment analysis using machine learning and deep learning models. Alex. Eng. J. 2023, 80, 110–121. [Google Scholar] [CrossRef]
  26. Chiril, P.; Pamungkas, E.W.; Benamara, F.; Moriceau, V.; Patti, V. Emotionally informed hate speech detection: A multi-target perspective. Cogn. Comput. 2022, 14, 322–352. [Google Scholar] [CrossRef]
  27. Al Anezi, F.Y. Arabic hate speech detection using deep recurrent neural networks. Appl. Sci. 2022, 12, 6010. [Google Scholar] [CrossRef]
  28. Aldjanabi, W.; Dahou, A.; Al-Qaness, M.A.A.; Elaziz, M.A.; Helmi, A.M.; Damaševičius, R. Arabic offensive and hate speech detection using a cross-corpora multi-task learning model. Informatics 2021, 8, 69. [Google Scholar] [CrossRef]
  29. Plaza-Del-Arco, F.M.; Molina-González, M.D.; Ureña-López, L.A.; Martín-Valdivia, M.T. A multi-task learning approach to hate speech detection leveraging sentiment analysis. IEEE Access 2021, 9, 112478–112489. [Google Scholar] [CrossRef]
  30. Hashmi, E.; Yayilgan, S.Y.; Hameed, I.A.; Yamin, M.M.; Ullah, M.; Abomhara, M. Enhancing multilingual hate speech detection: From language-specific insights to cross-linguistic integration. IEEE Access 2024, 12, 121507–121537. [Google Scholar] [CrossRef]
  31. Alhazmi, A.; Mahmud, R.; Idris, N.; Abo, M.E.M.; Eke, C. A systematic literature review of hate speech identification on Arabic Twitter data: Research challenges and future directions. PeerJ Comput. Sci. 2024, 10, e1966. [Google Scholar] [CrossRef]
  32. Moons, F.; Vandervieren, E. Measuring agreement among several raters classifying subjects into one or more (hierarchical) categories: A generalization of fleiss’ kappa. Behav. Res. Methods 2025, 57, 287. [Google Scholar] [CrossRef]
  33. Mubarak, H.; Rashed, A.; Darwish, K.; Samih, Y.; Abdelali, A. Arabic offensive language on twitter: Analysis and experiments. arXiv 2020, arXiv:2004.02192. [Google Scholar]
  34. Alrashidi, B.; Jamal, A.; Alkhathlan, A. Abusive content detection in arabic tweets using multi-task learning and transformer-based models. Appl. Sci. 2023, 13, 5825. [Google Scholar] [CrossRef]
  35. Cohen, S.; Presil, D.; Katz, O.; Arbili, O.; Messica, S.; Rokach, L. Enhancing social network hate detection using back translation and gpt-3 augmentations during training and test-time. Inf. Fusion 2023, 99, 101887. [Google Scholar] [CrossRef]
  36. Jovel, J.; Greiner, R. An introduction to machine learning approaches for biomedical research. Front. Med. 2021, 8, 771607. [Google Scholar] [CrossRef] [PubMed]
  37. Nahar, K.M.O.; Alauthman, M.; Yonbawi, S.; Almomani, A. Cyberbullying detection and recognition with type determination based on machine learning. Comput. Mater. Contin. 2023, 75, 5308–5319. [Google Scholar] [CrossRef]
  38. El Koshiry, A.M.; Eliwa, E.H.I.; El-Hafeez, T.A.; Omar, A. Arabic toxic tweet classification: Leveraging the AraBERT model. Big Data Cogn. Comput. 2023, 7, 170. [Google Scholar] [CrossRef]
  39. Qarah, F. EgyBERT: A large language model pretrained on Egyptian dialect corpora. arXiv 2024, arXiv:2408.03524. [Google Scholar] [CrossRef]
  40. Inoue, G.; Alhafni, B.; Baimukan, N.; Bouamor, H.; Habash, N. The interplay of variant, size, and task type in Arabic pre-trained language models. arXiv 2021, arXiv:2103.06678. [Google Scholar] [CrossRef]
  41. Dai, G.; Liao, J.; Zhao, S.; Fu, X.; Peng, X.; Huang, H.; Zhang, B. Large language model enhanced logic tensor network for stance detection. Neural Netw. 2024, 183, 106956. [Google Scholar] [CrossRef] [PubMed]
  42. Zhang, B.; Ma, J.; Fu, X.; Dai, G. Logic-augmented multi-decision fusion framework for stance detection on social media. Inf. Fusion 2025, 122, 103214. [Google Scholar] [CrossRef]
  43. Naous, T.; Ryan, M.; Xu, W.; Ritter, A.; Van Durme, B. Having Beer After Prayer? Measuring Cultural Bias in Large Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Bangkok, Thailand, 11–16 August 2024. [Google Scholar]
  44. Huang, T. Content moderation by large language models: From accuracy to legitimacy. Artif. Intell. Rev. 2025, 58, 320. [Google Scholar] [CrossRef]
  45. Gallegos, I.O.; Rossi, R.A.; Barrow, J.; Tanjim, M.M.; Kim, S.; Dernoncourt, F.; Yu, T.; Zhang, R.; Ahmed, N.K. Bias and fairness in large language models: A survey. Comput. Linguist. 2024, 50, 1098–1179. [Google Scholar] [CrossRef]
Figure 1. Three-level hierarchical annotation framework: Level 1 distinguishes Cyber Violent from Benign tweets; Level 2 labels Cyber Violent tweets as Cyberbullying or Online Hate Speech and Benign tweets as Positive, Negative, or Neutral; and Level 3 further decomposes Online Hate Speech into seven fine-grained discrimination classes.
Figure 1. Three-level hierarchical annotation framework: Level 1 distinguishes Cyber Violent from Benign tweets; Level 2 labels Cyber Violent tweets as Cyberbullying or Online Hate Speech and Benign tweets as Positive, Negative, or Neutral; and Level 3 further decomposes Online Hate Speech into seven fine-grained discrimination classes.
Information 17 00076 g001
Figure 2. Annotation team selection process.
Figure 2. Annotation team selection process.
Information 17 00076 g002
Figure 3. Screenshots of the tweet annotation interface used to label the SD-CVD corpus. The interface is presented in Arabic; the heading translates to “Tweet Classification System.” The screenshots illustrate the workflow: (right) annotator login, (left) displaying a tweet and indicating whether it contains cyber-violence, and (middle) assigning the primary and secondary class labels, with options to save the annotation (green), undetermined classification (orange), or delete the tweet (red).
Figure 3. Screenshots of the tweet annotation interface used to label the SD-CVD corpus. The interface is presented in Arabic; the heading translates to “Tweet Classification System.” The screenshots illustrate the workflow: (right) annotator login, (left) displaying a tweet and indicating whether it contains cyber-violence, and (middle) assigning the primary and secondary class labels, with options to save the annotation (green), undetermined classification (orange), or delete the tweet (red).
Information 17 00076 g003
Figure 4. Class distribution in the SD-CVD corpus.
Figure 4. Class distribution in the SD-CVD corpus.
Information 17 00076 g004
Figure 5. Confusion matrix of the Support Vector Machine (SVM) model.
Figure 5. Confusion matrix of the Support Vector Machine (SVM) model.
Information 17 00076 g005
Figure 6. Confusion matrix of the Multilayer Perceptron (MLP)model.
Figure 6. Confusion matrix of the Multilayer Perceptron (MLP)model.
Information 17 00076 g006
Figure 7. Confusion matrix of the AraBERTv02-Twitter model.
Figure 7. Confusion matrix of the AraBERTv02-Twitter model.
Information 17 00076 g007
Figure 8. Comparative F1-score performance of machine learning (ML)-, deep learning (DL)-, and BERT-based models.
Figure 8. Comparative F1-score performance of machine learning (ML)-, deep learning (DL)-, and BERT-based models.
Information 17 00076 g008
Table 1. Summary of Saudi dialect dataset studies.
Table 1. Summary of Saudi dialect dataset studies.
CitationDataset SizeMain TaskModel PerformanceLimitation
[1] Asiri and Saleh (2024)24,000 Saudi tweetsOffensive and hate speech detectionAraBERT with data augmentation achieved up to 0.91 F1-score (binary) and 0.83 (multiclass).Only AraBERTv0.2-Twitter was used; no comparison with other transformer-based models.
[19] Zaghouani et al. (2024)15,965 tweetsArabic text classification tasks (sentiment, emotion, offensive language, hate target/type identification)AraBERT outperformed classical baselines with an F1-score of 0.66.Annotator/regional bias, subjectivity in implicit hate, under-represented dialects, and class imbalance.
[5] Ahmad et al. (2024)403,688 tweetsMulticlass hate detectionFine-tuned AraBERT-based models reached an aggregate F1-score of approximately 0.60.Focused on Jordanian Arabic; labels are relatively coarse-grained and sentiment based.
[18] Alshalan and Al-Khalifa (2020)9316 Saudi tweetsHate speech detectionCNN achieved the best performance with an F1-score of 0.79.Binary classification limits fine-grained hate categorization.
[3] Mubarak et al. (2023)12,698 annotated Arabic tweetsOffensive language and hate speech detectionQARiB achieved F1 ≈ 82.3 for offensive language detection; AraBERT achieved F1 ≈ 80.1 for hate speech detection.Emoji-based sampling may over-represent certain offensive styles while under-representing others, reducing dataset representativeness.
[20] Alghamdi et al. (2021)2000 Saudi tweetsCyberbullying and threat detectionSVM achieved 76.06% F1 at Level 1 (violent vs. non-violent) and 89.18% F1 at Level 2 (cyberbullying vs. threats).Small dataset; keyword-based collection may bias toward explicit violence; no comparison with transformer-based architectures.
[11] Al-Twairesh et al. (2017)17,573 Saudi tweetsArabic sentiment analysisSVM with term-based features achieved F1: 62.27% (two-way), 58.17% (three-way), and 54.69% (four-way).Moderate inter-annotator agreement (κ = 0.60); performance drops in multiclass settings; limited robustness for implicit/ambiguous sentiment.
[17] Almuqren and Cristea (2021)20,000 telecom tweetsArabic sentiment analysis (telecom customer satisfaction)SVM achieved 91% accuracy.Binary sentiment (positive/negative) limits granularity, moderate agreement (κ ≈ 0.60), and domain-specific (telecom) may reduce generalizability.
[27] Al Anezi (2022)4203 commentsArabic hate speech detection (multiclass)Deep RNN achieved 99.73% accuracy (binary), 95.38% (three-class), and 84.14% (seven-class).Small dataset for deep learning; potential overfitting; and lacks comparison with transformer-based models.
[15] Charfi et al. (2024)4240 tweetsArabic hate speech detectionAraBERT reached 94% accuracy/F1 for hate vs. non-hate and 95% accuracy/F1 for hate category classification.Relatively small/static dataset; limited cross-dataset generalization; evolving hate expressions.
Table 2. Examples of lexicon components for data collection.
Table 2. Examples of lexicon components for data collection.
LexiconsListTranslation
Keywordsغبي، فاشل، جاهل، تفو عليك، حمير، همجي بقر، عالة على المجتمع، وجهك يسد النفس، خنزير، مغفل، سافل، مجانين، حقودين، ثور، نجس، عيال الشوارع، خنازير، كلب، يلعن، بلا كرامةStupid, loser, ignorant, shame on you, donkeys, savage, cows, burden on society, your face is repulsive, pig, fool, despicable, lunatics, spiteful, bull, filthy, street kids, pigs, dog, curse, dishonorable.
Emojis🦍 👊 💩 🔪 💣 🤮 👟 💀 🐸 🐶 🐻 ❄️ 🐷 🙈 🦉 🐖 🐭Gorilla, fist/punch, pile of poo, kitchen knife, bomb, face vomiting, sneaker, skull, frog, dog face, polar bear, pig face, see-no-evil monkey, owl, pig, mouse.
Hashtags# النساء_ناقصة_عقل ، #فاترحل_بنغلادش ، #الموت_ لإسرائيل ، # بنت_شوارع ، #مشاهير_الفلس ، #احرقوهم ، #سجن_صيدنايا#Women_Are_Lacking_Mind, #Bangladesh_Go_Away, #Death_To_Israel, #Street_Girl, #Bankrupt_Celebrities, #Burn_Them, #Saydnaya_Prison
Table 3. Definition of cyber violence categories and subcategories.
Table 3. Definition of cyber violence categories and subcategories.
CategoriesSubcategoriesDefinitions
Cyber ViolenceCyberbullyingHarassment, threats, or defamation through online or mobile platforms, exposing individuals—particularly youth—to digital violence [8].
Cyber ViolenceOnline Hate SpeechDiscriminatory or derogatory language directed at individuals or groups based on characteristics such as nationality, gender, religion, or region [9].
Online Hate SpeechIncitement to ViolencePosts legitimizing or calling for violence or threats [21,25,27].
Online Hate SpeechGender DiscriminationDerogatory remarks reinforcing gender stereotypes or bias [31].
Online Hate SpeechNational DiscriminationRidiculing or targeting others based on nationality or cultural prejudice [2,24].
Online Hate SpeechSocial Class DiscriminationVerbal attacks or mockery of individuals based on economic status [4].
Online Hate SpeechTribal DiscriminationOffensive remarks toward individuals or groups from specific tribes.
Online Hate SpeechReligious DiscriminationHostile comments elevating or demeaning individuals based on religion [6,19,30].
Online Hate SpeechRegional DiscriminationDiscriminatory language directed toward regional identities in Saudi Arabia such as Najdi, Hijazi, Eastern, or Qassimi.
Table 4. Exploratory data analysis for the SD-CVD Corpus.
Table 4. Exploratory data analysis for the SD-CVD Corpus.
StatisticBenignCyber Violence
NegativeNeutralPositiveCyber
Bullying
GenderIncitement
to Violence
NationalRegionalReligiousSocial
Class
Tribal
Count59037061804570427219725555854255555944154095
Word length
Mean20.5522.4616.0618.2823.6313.3620.4517.2820.4120.5015.02
Median1618121419101717172015
Min33233233333
Max91322299128184116128641518776
Character length
Mean112.55129.2498.24102.76134.1980.02119.9893.58115.20113.1480.11
Median8910671771095995929410777
Min129971181114111013
Max4751964160485610221338810376919505705
Table 6. Weighted performance metrics for SVM.
Table 6. Weighted performance metrics for SVM.
AlgorithmAccuracyPrecisionRecallF1-Score
SVM0.8540.8520.8540.853
Table 7. Weighted performance metrics for LR.
Table 7. Weighted performance metrics for LR.
AlgorithmAccuracyPrecisionRecallF1-Score
LR0.8500.8480.8500.849
Table 8. Weighted performance metrics for RF.
Table 8. Weighted performance metrics for RF.
AlgorithmAccuracyPrecisionRecallF1-Score
RF0.7690.7680.7690.767
Table 9. Weighted performance metrics for LSTM.
Table 9. Weighted performance metrics for LSTM.
AlgorithmAccuracyPrecisionRecallF1-Score
LSTM0.8130.813LSTM0.813
Table 10. Weighted performance metrics for RNN.
Table 10. Weighted performance metrics for RNN.
AlgorithmAccuracyPrecisionRecallF1-Score
RNN0.8120.8140.8120.812
Table 11. Weighted performance metrics for MLP.
Table 11. Weighted performance metrics for MLP.
AlgorithmAccuracyPrecisionRecallF1-Score
MLP0.8310.8300.8310.830
Table 12. Weighted performance metrics for CNN.
Table 12. Weighted performance metrics for CNN.
AlgorithmAccuracyPrecisionRecallF1-Score
CNN0.5710.5600.5710.561
Table 13. Weighted performance metrics for AraBERTv02-Twitter.
Table 13. Weighted performance metrics for AraBERTv02-Twitter.
AlgorithmAccuracyPrecisionRecallF1-Score
AraBERTv02-Twitter0.8830.8830.8830.882
Table 14. Weighted performance metrics for CAMeLBERT-MSA.
Table 14. Weighted performance metrics for CAMeLBERT-MSA.
AlgorithmAccuracyPrecisionRecallF1-Score
CAMeLBERT-MSA0.8710.8690.8710.869
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Alsayed, A.; Elhag, S.; Badri, S. SD-CVD Corpus: Towards Robust Detection of Fine-Grained Cyber-Violence Across Saudi Dialects in Online Platforms. Information 2026, 17, 76. https://doi.org/10.3390/info17010076

AMA Style

Alsayed A, Elhag S, Badri S. SD-CVD Corpus: Towards Robust Detection of Fine-Grained Cyber-Violence Across Saudi Dialects in Online Platforms. Information. 2026; 17(1):76. https://doi.org/10.3390/info17010076

Chicago/Turabian Style

Alsayed, Abrar, Salma Elhag, and Sahar Badri. 2026. "SD-CVD Corpus: Towards Robust Detection of Fine-Grained Cyber-Violence Across Saudi Dialects in Online Platforms" Information 17, no. 1: 76. https://doi.org/10.3390/info17010076

APA Style

Alsayed, A., Elhag, S., & Badri, S. (2026). SD-CVD Corpus: Towards Robust Detection of Fine-Grained Cyber-Violence Across Saudi Dialects in Online Platforms. Information, 17(1), 76. https://doi.org/10.3390/info17010076

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop