Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

Search Results (25)

Search Parameters:
Keywords = Arabic Speech-To-Text

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
24 pages, 463 KB  
Article
A Corpus-Based Pragmatic Study of the Formulation of Definitions and Legal Rulings in Sharia and Law College Curricula at Saudi Universities
by Fouad Ahmed Atallah
Languages 2026, 11(8), 158; https://doi.org/10.3390/languages11080158 - 30 Jul 2026
Viewed by 326
Abstract
Sharia and Law colleges in Saudi universities provide a distinctive educational setting in which classical Islamic jurisprudential discourse and contemporary statutory legal discourse coexist within the same curriculum. Despite this shared institutional context, the pragmatic characteristics of these coexisting normative genres remain largely [...] Read more.
Sharia and Law colleges in Saudi universities provide a distinctive educational setting in which classical Islamic jurisprudential discourse and contemporary statutory legal discourse coexist within the same curriculum. Despite this shared institutional context, the pragmatic characteristics of these coexisting normative genres remain largely unexplored from a corpus-based perspective. This study examines how genre differences influence the linguistic formulation of legal definitions and religious rulings across four officially prescribed texts: two classical Hanbali works—Rawḍat al-Nāẓir by Ibn Qudāma and Al-Rawḍ al-Murbiʿ by al-Buhūtī—and two contemporary Saudi statutes—the Civil Transactions Law and the Law of Criminal Procedure. A purpose-built corpus of approximately 480,000 words was analysed using a corpus-assisted discourse analysis design integrating quantitative frequency and keyword analysis with systematic manual pragmatic coding within a triangulated theoretical framework. The findings show that differences in disciplinary genre are systematically reflected in distinct pragmatic profiles. Rawḍat al-Nāẓir is characterised by assertive-definitional speech acts and methodological deontic expressions, whereas Al-Rawḍ al-Murbiʿ is dominated by directive speech acts associated with applied legal rulings. The statutory texts employ standardised legislative constructions, including negative-exceptive formulations, formal prohibition markers, and institution-specific obligation structures. The Civil Transactions Law further exhibits a hybrid pragmatic register through the incorporation of classical jurisprudential maxims into enacted statutory provisions. Based on the systematic literature review undertaken for this study, the findings provide what is, to the best of the authors’ knowledge, the first corpus-based pragmatic comparison of these coexisting curricular genres. The study thereby contributes to Arabic legal linguistics, corpus pragmatics, and the linguistic analysis of legal education, while offering empirically grounded insights for curriculum development in Sharia and Law programmes. Full article
(This article belongs to the Special Issue Corpus Pragmatics: Investigating Language Use in Context)
29 pages, 2632 KB  
Article
AI-Based Framework for Arabic Language Proficiency Assessment: A Deep Learning ASR Model with Enhanced Similarity Measures
by Sufian A. Badawi, Maen Takruri, Khouloud Salameh, Mohammad Al-Badawi, Nowar Alani, Isam ElBadawi, Aws Al-Qaisi and Ghaleb Aldoboni
Future Internet 2026, 18(5), 251; https://doi.org/10.3390/fi18050251 - 9 May 2026
Viewed by 905
Abstract
This work presents an innovative approach to test the Arabic language proficiency assessment via Automatic Speech Recognition (ASR) by enhancing the proficiency of the Whisper model in transcribing Arabic speech. The core of our research involved fine-tuning the Whisper model using a substantial, [...] Read more.
This work presents an innovative approach to test the Arabic language proficiency assessment via Automatic Speech Recognition (ASR) by enhancing the proficiency of the Whisper model in transcribing Arabic speech. The core of our research involved fine-tuning the Whisper model using a substantial, large-scale Arabic speech corpus, with a specific focus on Modern Standard Arabic. This process used a 2000-h Arabic-labeled speech corpus, the QASR dataset, and improved the model’s Word Error Rate (WER). After optimization, the fine-tuned Whisper model’s WER was reduced from 35% to 7% on the QASR dataset, corresponding to an absolute reduction of 28 percentage points (approximately 80% relative reduction). These results demonstrate the strong generalization ability of the fine-tuned model across multiple Arabic ASR benchmarks. A key component of our methodology was the development of a sophisticated scoring system. This system integrates various similarity metrics, such as cosine similarity, the Jaccard index, and the Levenshtein distance, with a machine learning regression model. This multifaceted system provides a comprehensive assessment of reading proficiency, proposing a practical automated assessment method that contributes to the field of AI language transcription and to its application in the assessment of students’ reading. Our research also introduces the ICONET dataset, an augmented Arabic speech corpus comprising 3160 h of diverse and tailored audio–text pairs designed for fine-tuning ASR models. This study demonstrates the potential of fine-tuning pretrained models for specific linguistic contexts (Arabic), establishing a foundation for future research in ASR and language technology. Full article
(This article belongs to the Topic Learning to Live with Gen-AI)
Show Figures

Figure 1

15 pages, 1962 KB  
Article
Design and Performance Evaluation of a Low-Cost High-SNR EOG Sensing System for Arabic Locked-In Syndrome Communication
by Saleh I. Alzahrani, Najat Alomari, Sarah Alkilani, Lama Alghamdi and Bushra Melhem
Sensors 2026, 26(8), 2425; https://doi.org/10.3390/s26082425 - 15 Apr 2026
Viewed by 595
Abstract
Locked-in Syndrome (LIS) is a neurological condition in which individuals remain conscious but experience complete paralysis of voluntary muscles, except for eye movements—highlighting the need for reliable assistive communication technologies. This study presents the design and evaluation of an Arabic electrooculogram (EOG)-based communication [...] Read more.
Locked-in Syndrome (LIS) is a neurological condition in which individuals remain conscious but experience complete paralysis of voluntary muscles, except for eye movements—highlighting the need for reliable assistive communication technologies. This study presents the design and evaluation of an Arabic electrooculogram (EOG)-based communication system with adaptive classification capabilities for LIS applications. A custom-designed EOG acquisition circuit incorporating filtering and amplification stages was implemented and compared with the OpenBCI Cyton board. The system employed a hybrid classification approach combining amplitude, temporal, and statistical features to distinguish between blinks and voluntary vertical eye movements. Testing with ten healthy subjects yielded a mean classification accuracy of 83.96% ± 4.59% and an information transfer rate of 10.43 letters per minute, corresponding to a 30.38% improvement over conventional approaches. The custom-designed circuit achieved a signal-to-noise ratio of 25.21 dB, outperforming the OpenBCI Cyton board by 8% while reducing system cost by 62%. The integration with a Morse code-based interface enabled Arabic letter composition, while the system incorporated auto-completion and text-to-speech functionalities to further enhance communication efficiency. This cost-effective solution addresses a critical gap in assistive technologies for Arabic-speaking individuals with LIS and shows strong potential for enhancing their communication abilities and overall quality of life. Full article
(This article belongs to the Special Issue Advanced Sensor Technologies for Neuroimaging and Neurorehabilitation)
Show Figures

Figure 1

22 pages, 3288 KB  
Article
An Intelligent Real-Time System for Sentence-Level Recognition of Continuous Saudi Sign Language Using Landmark-Based Temporal Modeling
by Adel BenAbdennour, Mohammed Mukhtar, Osama Almolike, Bilal A. Khawaja and Abdulmajeed M. Alenezi
Sensors 2026, 26(5), 1652; https://doi.org/10.3390/s26051652 - 5 Mar 2026
Viewed by 1051
Abstract
A persistent challenge for Deaf and Hard-of-Hearing individuals is the communication gap between sign language users and the hearing community, particularly in regions with limited automated translation resources. In Saudi Arabia, this gap is amplified by the reliance on Saudi Sign Language (SSL) [...] Read more.
A persistent challenge for Deaf and Hard-of-Hearing individuals is the communication gap between sign language users and the hearing community, particularly in regions with limited automated translation resources. In Saudi Arabia, this gap is amplified by the reliance on Saudi Sign Language (SSL) and the scarcity of real-time, sentence-level translation systems. This paper presents a real-time system for sentence-level recognition of continuous SSL and direct mapping to natural spoken Arabic. The proposed system operates end-to-end on live video streams or pre-recorded content, extracting spatio-temporal landmark features using the MediaPipe Holistic framework. For classification, the input feature vector consists of 225 features derived from hand and body pose landmarks. These features are processed by a Bidirectional Long Short-Term Memory (BiLSTM) network trained on the ArabSign (ArSL) dataset to perform direct sentence-level classification over a vocabulary of 50 continuous Arabic sign language sentences, supported by an idle-based segmentation mechanism that enables natural, uninterrupted signing. Experimental evaluation demonstrates robust generalization: under a Leave-One-Signer-Out (LOSO) cross-validation protocol, the model attains a mean sentence-level accuracy of 94.2%, outperforming the fixed signer-independent split baseline of 92.07%, while maintaining real-time performance suitable for interactive use. To enhance linguistic fluency, an optional post-recognition refinement stage is incorporated using a large language model (LLM), followed by text-to-speech synthesis to produce audible Arabic output; this refinement operates strictly as post-processing and is not included in the reported recognition accuracy metrics. The results demonstrate that direct sentence-level modeling, combined with landmark-based feature extraction and real-time segmentation, provides an effective and practical solution for continuous SSL sentence recognition in real-time. Full article
(This article belongs to the Special Issue Sensor Systems for Gesture Recognition (3rd Edition))
Show Figures

Figure 1

30 pages, 6201 KB  
Article
AFAD-MSA: Dataset and Models for Arabic Fake Audio Detection
by Elsayed Issa
Computation 2026, 14(1), 20; https://doi.org/10.3390/computation14010020 - 14 Jan 2026
Cited by 2 | Viewed by 2136
Abstract
As generative speech synthesis produces near-human synthetic voices and reliance on online media grows, robust audio-deepfake detection is essential to fight misuse and misinformation. In this study, we introduce the Arabic Fake Audio Dataset for Modern Standard Arabic (AFAD-MSA), a curated corpus of [...] Read more.
As generative speech synthesis produces near-human synthetic voices and reliance on online media grows, robust audio-deepfake detection is essential to fight misuse and misinformation. In this study, we introduce the Arabic Fake Audio Dataset for Modern Standard Arabic (AFAD-MSA), a curated corpus of authentic and synthetic Arabic speech designed to advance research on Arabic deepfake and spoofed-speech detection. The synthetic subset is generated with four state-of-the-art proprietary text-to-speech and voice-conversion models. Rich metadata—covering speaker attributes and generation information—is provided to support reproducibility and benchmarking. To establish reference performance, we trained three AASIST models and compared their performance to two baseline transformer detectors (Wav2Vec 2.0 and Whisper). On the AFAD-MSA test split, AASIST-2 achieved perfect accuracy, surpassing the baseline models. However, its performance declined under cross-dataset evaluation. These results underscore the importance of data construction. Detectors generalize best when exposed to diverse attack types. In addition, continual or contrastive training that interleaves bona fide speech with large, heterogeneous spoofed corpora will further improve detectors’ robustness. Full article
Show Figures

Figure 1

32 pages, 1254 KB  
Review
Arabic Natural Language Processing (NLP): A Comprehensive Review of Challenges, Techniques, and Emerging Trends
by Abdulaziz M. Alayba
Computers 2025, 14(11), 497; https://doi.org/10.3390/computers14110497 - 15 Nov 2025
Cited by 26 | Viewed by 11300
Abstract
Arabic natural language processing (NLP) has garnered significant attention in recent years due to the growing demand for automated text and Arabic-based intelligent systems, in addition to digital transformation in the Arab world. However, the unique linguistic characteristics of Arabic, including its rich [...] Read more.
Arabic natural language processing (NLP) has garnered significant attention in recent years due to the growing demand for automated text and Arabic-based intelligent systems, in addition to digital transformation in the Arab world. However, the unique linguistic characteristics of Arabic, including its rich morphology, diverse dialects, and complex syntax, pose significant challenges to NLP researchers. This paper provides a comprehensive review of the main linguistic challenges inherent in Arabic NLP, such as morphological complexity, diacritics and orthography issues, ambiguity, and dataset limitations. Furthermore, it surveys the major computational techniques employed in tokenisation and normalisation, named entity recognition, part-of-speech tagging, sentiment analysis, text classification, summarisation, question answering, and machine translation. In addition, it discusses the rapid rise of large language models and their transformative impact on Arabic NLP. Full article
Show Figures

Figure 1

37 pages, 3329 KB  
Article
Deobfuscating Iraqi Arabic Leetspeak for Hate Speech Detection Using AraBERT and Hierarchical Attention Network (HAN)
by Dheyauldeen Marzoog and Hasan Çakir
Electronics 2025, 14(21), 4318; https://doi.org/10.3390/electronics14214318 - 3 Nov 2025
Cited by 1 | Viewed by 2231
Abstract
The widespread use of leetspeak and dialectal Arabic on social media poses a critical challenge to automated hate speech detection systems. Existing Arabic NLP models, largely trained on Modern Standard Arabic (MSA), struggle with obfuscated, noisy, and dialect-specific text, leading to poor generalization [...] Read more.
The widespread use of leetspeak and dialectal Arabic on social media poses a critical challenge to automated hate speech detection systems. Existing Arabic NLP models, largely trained on Modern Standard Arabic (MSA), struggle with obfuscated, noisy, and dialect-specific text, leading to poor generalization in real-world scenarios. This study introduces a Hybrid AraBERT–Hierarchical Attention Network (HAN) framework for deobfuscating Iraqi Arabic leetspeak and accurately classifying hate speech. The proposed model employs a custom normalization pipeline that converts digits, symbols, and Latin-script substitutions (e.g., "3يب" → "عيب") into canonical Arabic forms, thereby enhancing tokenization and embedding quality. AraBERT provides deep contextualized representations optimized for Arabic morphology, while HAN hierarchically aggregates and attends to critical words and sentences to improve interpretability and semantic focus. Experimental evaluation on an Iraqi Arabic social media dataset demonstrates that the proposed model achieves 97% accuracy, 96% precision, 96% recall, 96% F1-score, and 0.98 ROC–AUC, outperforming standalone AraBERT and HAN models by up to 6% in F1-score and 4% in AUC. Ablation studies confirm the important role of the normalization stage (F1 = 0.91 without it) and the contribution of hierarchical attention in balancing precision and recall. Robustness testing under controlled perturbations (including character substitutions, symbol obfuscations, typographical noise, and class imbalance) shows performance retention above 91% F1, validating the framework’s noise tolerance and generalization capability. Comparative analysis with state-of-the-art approaches such as DRNNs, arHateDetector, and ensemble BERT systems further highlights the hybrid model’s effectiveness in handling noisy, dialectal, and adversarial text. Full article
Show Figures

Figure 1

17 pages, 2127 KB  
Article
Leveraging Large Language Models for Real-Time UAV Control
by Kheireddine Choutri, Samiha Fadloun, Ayoub Khettabi, Mohand Lagha, Souham Meshoul and Raouf Fareh
Electronics 2025, 14(21), 4312; https://doi.org/10.3390/electronics14214312 - 2 Nov 2025
Cited by 6 | Viewed by 4458
Abstract
As drones become increasingly integrated into civilian and industrial domains, the demand for natural and accessible control interfaces continues to grow. Conventional manual controllers require technical expertise and impose cognitive overhead, limiting their usability in dynamic and time-critical scenarios. To address these limitations, [...] Read more.
As drones become increasingly integrated into civilian and industrial domains, the demand for natural and accessible control interfaces continues to grow. Conventional manual controllers require technical expertise and impose cognitive overhead, limiting their usability in dynamic and time-critical scenarios. To address these limitations, this paper presents a multilingual voice-driven control framework for quadrotor drones, enabling real-time operation in both English and Arabic. The proposed architecture combines offline Speech-to-Text (STT) processing with large language models (LLMs) to interpret spoken commands and translate them into executable control code. Specifically, Vosk is employed for bilingual STT, while Google Gemini provides semantic disambiguation, contextual inference, and code generation. The system is designed for continuous, low-latency operation within an edge–cloud hybrid configuration, offering an intuitive and robust human–drone interface. While speech recognition and safety validation are processed entirely offline, high-level reasoning and code generation currently rely on cloud-based LLM inference. Experimental evaluation demonstrates an average speech recognition accuracy of 95% and end-to-end command execution latency between 300 and 500 ms, validating the feasibility of reliable, multilingual, voice-based UAV control. This research advances multimodal human–robot interaction by showcasing the integration of offline speech recognition and LLMs for adaptive, safe, and scalable aerial autonomy. Full article
Show Figures

Figure 1

11 pages, 1076 KB  
Proceeding Paper
Moroccan Institutional Chatbots: A Hybrid Approach with LLMs, Semantic Matching, and Dialect Adaptation for DARIJA
by Oumaima Ennasri, Brahim El Bhiri and Yann Ben Maissa
Eng. Proc. 2025, 112(1), 41; https://doi.org/10.3390/engproc2025112041 - 20 Oct 2025
Viewed by 2270
Abstract
With the rapid growth of LLM-based chatbots and their applications in fields such as health, education, and entertainment, there is a growing interest in developing systems capable of mimicking human behavior through conversation and natural language interaction. These chatbots are available in several [...] Read more.
With the rapid growth of LLM-based chatbots and their applications in fields such as health, education, and entertainment, there is a growing interest in developing systems capable of mimicking human behavior through conversation and natural language interaction. These chatbots are available in several languages, such as English, French, and Spanish. Unfortunately, Arabic chatbots—especially those that understand Arabic dialects—are still very limited. In this paper, we develop a chatbot for the Moroccan Arabic dialect, specifically designed for the public sector, such as the fiscal domain and government administration. These institutions require tools to reduce communication loads, limit human assistance, and minimize the time needed to find documents or complete payment procedures. Our optimized chatbot combines recent technologies like LLMs and semantic similarity. It supports Moroccan citizens by providing responses in the Moroccan dialect (Darija), both in text and speech, without requiring extensive resources. It also supports other citizens in French, Spanish, and English. Our chatbot was tested in a real use case in the tax domain, and the results were satisfactory, especially considering the general complexity of the Arabic language and the particular challenges of the Moroccan dialect. Full article
Show Figures

Figure 1

30 pages, 1673 KB  
Article
Adversarially Robust Multitask Learning for Offensive and Hate Speech Detection in Arabic Text Using Transformer-Based Models and RNN Architectures
by Eman S. Alshahrani and Mehmet S. Aksoy
Appl. Sci. 2025, 15(17), 9602; https://doi.org/10.3390/app15179602 - 31 Aug 2025
Cited by 6 | Viewed by 2824
Abstract
Offensive language and hate speech have a detrimental effect on victims and have become a significant problem on social media platforms. Recent research has developed automated techniques for detecting Arabic offensive language and hate speech but remains limited, and further research is required [...] Read more.
Offensive language and hate speech have a detrimental effect on victims and have become a significant problem on social media platforms. Recent research has developed automated techniques for detecting Arabic offensive language and hate speech but remains limited, and further research is required compared to the research on high-resource languages such as English due to limited resources, annotated corpora, and morphological analysis. Most social media users who use profanities attempt to modify their text while maintaining the same meaning, thereby deceiving detection methods that forbid offending phrases. Therefore, this study proposes an adversarially robust multitask learning framework for detection of Arabic offensive and hate speech. For this purpose, this study used the OSACT2020 dataset, augmented with additional posts collected from the X social media platform. To improve contextual understanding, classification models based on various configurations were constructed using four pre-trained Arabic language models integrated with various sequential layers that were trained and evaluated in three different settings: single-task learning with the original dataset, single-task learning with the augmented dataset, and multitask learning with the augmented dataset. The multitask MARBERTv2+BiGRU model achieved the best results, with an 88% macro-F1 for hate speech and 93% for offensive language on clean data. To improve the model’s robustness, adversarial samples were generated using attacks on both the character and sentence levels. These attacks subtly change the text to mislead the model while maintaining the overall appearance and meaning. The clean model’s performance dropped significantly under attack, especially for hate speech, to a 74% macro-F1; however, adversarial training, which re-trains the model using both clean and adversarial data, improved the results to a 78% macro-F1 for hate speech. Further improvements were achieved with input transformation techniques, boosting the macro-F1 to 81%. Notably, the adversarially trained model maintained high performance on clean data, demonstrating both robustness and generalization. Full article
(This article belongs to the Special Issue Machine Learning Approaches in Natural Language Processing)
Show Figures

Figure 1

20 pages, 3244 KB  
Article
SOUTY: A Voice Identity-Preserving Mobile Application for Arabic-Speaking Amyotrophic Lateral Sclerosis Patients Using Eye-Tracking and Speech Synthesis
by Hessah A. Alsalamah, Leena Alhabrdi, May Alsebayel, Aljawhara Almisned, Deema Alhadlaq, Loody S. Albadrani, Seetah M. Alsalamah and Shada AlSalamah
Electronics 2025, 14(16), 3235; https://doi.org/10.3390/electronics14163235 - 14 Aug 2025
Viewed by 1746
Abstract
Amyotrophic Lateral Sclerosis (ALS) is a neurodegenerative disorder that progressively impairs motor and communication abilities. Globally, the prevalence of ALS was estimated at approximately 222,800 cases in 2015 and is projected to increase by nearly 70% to 376,700 cases by 2040, primarily driven [...] Read more.
Amyotrophic Lateral Sclerosis (ALS) is a neurodegenerative disorder that progressively impairs motor and communication abilities. Globally, the prevalence of ALS was estimated at approximately 222,800 cases in 2015 and is projected to increase by nearly 70% to 376,700 cases by 2040, primarily driven by demographic shifts in aging populations, and the lifetime risk of developing ALS is 1 in 350–420. Despite international advancements in assistive technologies, a recent national survey in Saudi Arabia revealed that 100% of ALS care providers lack access to eye-tracking communication tools, and 92% reported communication aids as inconsistently available. While assistive technologies such as speech-generating devices and gaze-based control systems have made strides in recent decades, they primarily support English speakers, leaving Arabic-speaking ALS patients underserved. This paper presents SOUTY, a cost-effective, mobile-based application that empowers ALS patients to communicate using gaze-controlled interfaces combined with a text-to-speech (TTS) feature in Arabic language, which is one of the five most widely spoken languages in the world. SOUTY (i.e., “my voice”) utilizes a personalized, pre-recorded voice bank of the ALS patient and integrated eye-tracking technology to support the formation and vocalization of custom phrases in Arabic. This study describes the full development life cycle of SOUTY from conceptualization and requirements gathering to system architecture, implementation, evaluation, and refinement. Validation included expert interviews with Human–Computer Interaction (HCI) expertise and speech pathology specialty, as well as a public survey assessing awareness and technological readiness. The results support SOUTY as a culturally and linguistically relevant innovation that enhances autonomy and quality of life for Arabic-speaking ALS patients. This approach may serve as a replicable model for developing inclusive Augmentative and Alternative Communication (AAC) tools in other underrepresented languages. The system achieved 100% task completion during internal walkthroughs, with mean phrase selection times under 5 s and audio playback latency below 0.3 s. Full article
Show Figures

Figure 1

24 pages, 2410 KB  
Article
UA-HSD-2025: Multi-Lingual Hate Speech Detection from Tweets Using Pre-Trained Transformers
by Muhammad Ahmad, Muhammad Waqas, Ameer Hamza, Sardar Usman, Ildar Batyrshin and Grigori Sidorov
Computers 2025, 14(6), 239; https://doi.org/10.3390/computers14060239 - 18 Jun 2025
Cited by 14 | Viewed by 7205
Abstract
The rise in social media has improved communication but also amplified the spread of hate speech, creating serious societal risks. Automated detection remains difficult due to subjectivity, linguistic diversity, and implicit language. While prior research focuses on high-resource languages, this study addresses the [...] Read more.
The rise in social media has improved communication but also amplified the spread of hate speech, creating serious societal risks. Automated detection remains difficult due to subjectivity, linguistic diversity, and implicit language. While prior research focuses on high-resource languages, this study addresses the underexplored multilingual challenges of Arabic and Urdu hate speech through a comprehensive approach. To achieve this objective, this study makes four different key contributions. First, we have created a unique multi-lingual, manually annotated binary and multi-class dataset (UA-HSD-2025) sourced from X, which contains the five most important multi-class categories of hate speech. Secondly, we created detailed annotation guidelines to make a robust and perfect hate speech dataset. Third, we explore two strategies to address the challenges of multilingual data: a joint multilingual and translation-based approach. The translation-based approach involves converting all input text into a single target language before applying a classifier. In contrast, the joint multilingual approach employs a unified model trained to handle multiple languages simultaneously, enabling it to classify text across different languages without translation. Finally, we have employed state-of-the-art 54 different experiments using different machine learning using TF-IDF, deep learning using advanced pre-trained word embeddings such as FastText and Glove, and pre-trained language-based models using advanced contextual embeddings. Based on the analysis of the results, our language-based model (XLM-R) outperformed traditional supervised learning approaches, achieving 0.99 accuracy in binary classification for Arabic, Urdu, and joint-multilingual datasets, and 0.95, 0.94, and 0.94 accuracy in multi-class classification for joint-multilingual, Arabic, and Urdu datasets, respectively. Full article
(This article belongs to the Special Issue Recent Advances in Social Networks and Social Media)
Show Figures

Figure 1

18 pages, 373 KB  
Article
Machine Learning- and Deep Learning-Based Multi-Model System for Hate Speech Detection on Facebook
by Amna Naseeb, Muhammad Zain, Nisar Hussain, Amna Qasim, Fiaz Ahmad, Grigori Sidorov and Alexander Gelbukh
Algorithms 2025, 18(6), 331; https://doi.org/10.3390/a18060331 - 1 Jun 2025
Cited by 13 | Viewed by 3309
Abstract
Hate speech is a complex topic that transcends language, culture, and even social spheres. Recently, the spread of hate speech on social media sites like Facebook has added a new layer of complexity to the issue of online safety and content moderation. This [...] Read more.
Hate speech is a complex topic that transcends language, culture, and even social spheres. Recently, the spread of hate speech on social media sites like Facebook has added a new layer of complexity to the issue of online safety and content moderation. This study seeks to minimize this problem by developing an Arabic script-based tool for automatically detecting hate speech in Roman Urdu, an informal script used most commonly for South Asian digital communications. Roman Urdu is relatively complex as there are no standardized spellings, leading to syntactic variations, which increases the difficulty of hate speech detection. To tackle this problem, we adopt a holistic strategy using a combination of six machine learning (ML) and four Deep Learning (DL) models, a dataset from Facebook comments, which was preprocessed (tokenization, stopwords removal, etc.), and text vectorization (TF-IDF, word embeddings). The ML algorithms used in this study are LR, SVM, RF, NB, KNN, and GBM. We also use deep learning architectures like CNN, RNN, LSTM, and GRU to increase the accuracy of the classification further. It is proven by the experimental results that deep learning models outperform the traditional ML approaches by a significant margin, with CNN and LSTM achieving accuracies of 95.1% and 96.2%, respectively. As far as we are aware, this is the first work that investigates QLoRA for fine-tuning large models for the task of offensive language detection in Roman Urdu. Full article
(This article belongs to the Special Issue Linguistic and Cognitive Approaches to Dialog Agents)
Show Figures

Figure 1

18 pages, 585 KB  
Article
Improving Diacritical Arabic Speech Recognition: Transformer-Based Models with Transfer Learning and Hybrid Data Augmentation
by Haifa Alaqel and Khalil El Hindi
Information 2025, 16(3), 161; https://doi.org/10.3390/info16030161 - 20 Feb 2025
Cited by 7 | Viewed by 6142
Abstract
Diacritical Arabic (DA) refers to Arabic text with diacritical marks that guide pronunciation and clarify meanings, making their recognition crucial for accurate linguistic interpretation. These diacritical marks (short vowels) significantly influence meaning and pronunciation, and their accurate recognition is vital for the effectiveness [...] Read more.
Diacritical Arabic (DA) refers to Arabic text with diacritical marks that guide pronunciation and clarify meanings, making their recognition crucial for accurate linguistic interpretation. These diacritical marks (short vowels) significantly influence meaning and pronunciation, and their accurate recognition is vital for the effectiveness of automatic speech recognition (ASR) systems, particularly in applications requiring high semantic precision, such as voice-enabled translation services. Despite its importance, leveraging advanced machine learning techniques to enhance ASR for diacritical Arabic has remained underexplored. A key challenge in developing DA ASR is the limited availability of training data. This study introduces a transformer-based approach leveraging transfer learning and data augmentation to address these challenges. Using a cross-lingual speech representation (XLSR) model pretrained on 53 languages, we fine-tune it on DA and integrate connectionist temporal classification (CTC) with transformers for improved performance. Data augmentation techniques, including volume adjustment, pitch shift, speed alteration, and hybrid strategies, further mitigate data limitations, significantly reducing word error rates (WER). Our methods achieve a WER of 12.17%, outperforming traditional ASR systems and setting a new benchmark for DA ASR. These findings demonstrate the potential of advanced machine learning to address longstanding challenges in DA ASR and enhance its accuracy. Full article
Show Figures

Figure 1

20 pages, 1420 KB  
Article
A Survey of Grapheme-to-Phoneme Conversion Methods
by Shiyang Cheng, Pengcheng Zhu, Jueting Liu and Zehua Wang
Appl. Sci. 2024, 14(24), 11790; https://doi.org/10.3390/app142411790 - 17 Dec 2024
Cited by 10 | Viewed by 11372
Abstract
Grapheme-to-phoneme conversion (G2P) is the task of converting letters (grapheme sequences) into their pronunciations (phoneme sequences). It plays a crucial role in natural language processing, text-to-speech synthesis, and automatic speech recognition systems. This paper provides a systematical overview of the G2P conversion from [...] Read more.
Grapheme-to-phoneme conversion (G2P) is the task of converting letters (grapheme sequences) into their pronunciations (phoneme sequences). It plays a crucial role in natural language processing, text-to-speech synthesis, and automatic speech recognition systems. This paper provides a systematical overview of the G2P conversion from different perspectives. The conversion methods are first presented in the paper; detailed discussions are conducted on methods based on deep learning technology. For each method, the key ideas, advantages, disadvantages, and representative models are summarized. This paper then mentioned the learning strategies and multilingual G2P conversions. Finally, this paper summarized the commonly used monolingual and multilingual datasets, including Mandarin, Japanese, Arabic, etc. Two tables illustrated the performance of various methods with relative datasets. After making a general overall of G2P conversion, this paper concluded with the current issues and the future directions of deep learning-based G2P conversion. Full article
(This article belongs to the Collection Trends and Prospects in Multimedia)
Show Figures

Figure 1

Back to TopTop