Sign in to use this feature.

Years

Between: -

Subjects

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Journals

Article Types

Countries / Regions

remove_circle_outline
remove_circle_outline
remove_circle_outline
remove_circle_outline

Search Results (305)

Search Parameters:
Keywords = fake news detection

Order results
Result details
Results per page
Select all
Export citation of selected articles as:
36 pages, 626 KB  
Article
Comparative Performance of AI-Generated Fake News Detection Pipelines on Romanian News Content
by Claudiu Coman, Costel Marian Dalban, Vlad Bătrânu-Pințea, Georgiana Aron and Lucian Marina
Information 2026, 17(7), 698; https://doi.org/10.3390/info17070698 - 18 Jul 2026
Viewed by 354
Abstract
Fake news detection has become a major research topic at the intersection of artificial intelligence, data mining, and information security. In this paper, we evaluate the performance of English-trained algorithms on English translations of Romanian-sourced news articles, using a translation-mediated cross-domain evaluation design. [...] Read more.
Fake news detection has become a major research topic at the intersection of artificial intelligence, data mining, and information security. In this paper, we evaluate the performance of English-trained algorithms on English translations of Romanian-sourced news articles, using a translation-mediated cross-domain evaluation design. The study is based on source code generated with the assistance of artificial intelligence systems for a set of machine learning and transformer-based models. The code was subsequently implemented in Google Colab. 2026, trained on international benchmark datasets, and tested on Romanian news content. This design allowed the rapid prototyping of multiple detection pipelines and the systematic observation of their behavior in a media environment different from that represented in the training corpora. The models were evaluated comparatively using standard classification metrics, including accuracy, precision, recall, and F1-score, complemented by additional indicators relevant to model robustness and practical usability. The experimental results revealed significant differences in performance across algorithms when applied to English translations of Romanian-language news content after training on international datasets. However, this study does not provide a direct comparison between model performance on the international benchmark datasets and the Romanian test corpus; therefore, the gap between the international training corpus and the Romanian-sourced test corpus is interpreted as an exploratory limitation and as a direction for future research. Based on these findings, we propose an empirical classification of the tested models according to their predictive effectiveness, their contextual robustness across linguistic environments, and their operational relevance as filtering tools for institutional monitoring. The results show that AI-assisted coding workflows can provide a viable starting point for reproducible misinformation research, but they also underline the limitations of directly transferring models trained on non-Romanian data to local media ecosystems. The study offers both a replicable evaluation framework and practical insights for institutions involved in strategic communication, public security, and the monitoring of information threats. Full article
Show Figures

Graphical abstract

32 pages, 3378 KB  
Article
H-FuseNet: A Hybrid Multi-Representation Fusion Framework for Robust Misinformation Detection
by Abdullah, Muhammad Ateeb Ather, Kinza Sardar, Zulaikha Fatima, Grigori Sidorov, Carlos Guzmán Sánchez-Mejorada, Rolando Quintero Téllez and Miguel Jesús Torres Ruiz
Mach. Learn. Knowl. Extr. 2026, 8(7), 205; https://doi.org/10.3390/make8070205 - 13 Jul 2026
Viewed by 320
Abstract
This study investigates automated fake news detection as a reliability-oriented text classification problem in dynamic digital information environments. We propose H-FuseNet, a hybrid multi-representation fusion framework that combines pretrained transformer representations with deception-oriented handcrafted linguistic, stylistic, and semantic features. Using the WELFake dataset, [...] Read more.
This study investigates automated fake news detection as a reliability-oriented text classification problem in dynamic digital information environments. We propose H-FuseNet, a hybrid multi-representation fusion framework that combines pretrained transformer representations with deception-oriented handcrafted linguistic, stylistic, and semantic features. Using the WELFake dataset, we benchmark 15 baseline models, including classical classifiers, ensemble methods, recurrent and convolutional networks, and transformer fine-tuning models, under stratified 10-fold cross-validation with nested hyperparameter optimization. To examine generalization beyond a single benchmark, we train exclusively on WELFake and evaluate cross-dataset performance on three held-out external datasets: FakeNewsNet, CoAID, and LLM-generated misinformation. H-FuseNet integrates transformer document embeddings with a lightweight feature-processing MLP, optional contextual feature streams when metadata are available, and auxiliary supervision through pseudo-labeled headline body stance and clickbait signals. The proposed model achieves 98.9% mean accuracy and 0.998 ROC–AUC, while maintaining strong calibration, with a Brier score of 0.012 and Expected Calibration Error of 0.009, and low variance across folds. Cross-dataset evaluation yields accuracies of 87.34% on FakeNewsNet, 83.56% on CoAID, and 91.22% on LLM-generated misinformation, demonstrating robust generalization under distribution shift. Ablation analyses show that handcrafted features, auxiliary tasks, and learned fusion each contribute to performance, while Wilcoxon and McNemar tests indicate statistically significant differences against selected strong baselines. Error analysis shows that remaining failures mainly occur in professionally written misinformation that imitates neutral journalistic style. Overall, the results suggest that calibrated multi-representation fusion can improve the reliability of automated fake news detection systems. Full article
Show Figures

Figure 1

23 pages, 31865 KB  
Article
TERN: Type-Aware Evidence Reasoning for Multimodal Fake News Detection
by Mingshu Zhang, Hongyu Jin, Yuechuan Zhang, Bin Wei and Yaxuan Wang
Appl. Sci. 2026, 16(13), 6759; https://doi.org/10.3390/app16136759 - 6 Jul 2026
Viewed by 234
Abstract
Multimodal fake news detection remains challenging because deceptive posts exhibit heterogeneous manipulation patterns, while most existing methods still rely on a unified fusion strategy. This mismatch limits their ability to adapt to different evidence preferences across samples, encourages entanglement between deception cues and [...] Read more.
Multimodal fake news detection remains challenging because deceptive posts exhibit heterogeneous manipulation patterns, while most existing methods still rely on a unified fusion strategy. This mismatch limits their ability to adapt to different evidence preferences across samples, encourages entanglement between deception cues and topical semantics, and weakens decision making when textual, visual, and cross-modal signals conflict. To address these issues, we propose TERN, a type-aware evidence reasoning network for multimodal fake news detection. TERN induces latent deception types from image-side multimodal features through prototype-based clustering, uses the induced assignments as structural priors for downstream veracity prediction, disentangles type-discriminative factors from semantic content, and performs type-conditioned hierarchical reasoning over text semantics, image authenticity, and cross-modal consistency. Experiments on MR2-Chinese, MR2-English, Weibo, and PHEME show that TERN achieves an average accuracy of 93.21% and an average F1 score of 91.09% while also improving Matthews correlation coefficient over representative multimodal baselines. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

35 pages, 4618 KB  
Article
Design of an Iterative Cross-Modal and Context-Aware Deep Analytical Framework for Hate Speech and Fake Post Detection on Social Media Sets
by Rakesh Bharati, Jyoti Bharti and Vasudev Dehalwar
Appl. Sci. 2026, 16(13), 6419; https://doi.org/10.3390/app16136419 - 26 Jun 2026
Viewed by 353
Abstract
There is an enormous rise in the amount of user-generated content on social media. That makes it easier for hateful and fake messages to spread, and threatens both societal stability and public trust in institutions. Most of current solutions have fundamental limitations due [...] Read more.
There is an enormous rise in the amount of user-generated content on social media. That makes it easier for hateful and fake messages to spread, and threatens both societal stability and public trust in institutions. Most of current solutions have fundamental limitations due to modal limitations (i.e., each solution only uses one type of data at a time), lack of user context integration, poor synchronization across different types of data, and poor resilience to manipulation by adversaries. As a result, most solutions are subject to compound loss in terms of their ability to generalize well, classify correctly, or remain reliable when deployed in real-world environments. To address all of the above challenges, we propose a comprehensive and modular analytical framework consisting of five interconnected components that integrate contextual representation learning, multimodal semantic alignment, graph-based propagation modeling, adaptive inference, and consistency validation for hate speech and fake post detection. First is our Context-Driven Social Vector Extraction methodology, which provides enriched contextual embeddings by extracting and combining text-based metadata, image-based metadata, temporal metadata, and behavioral metadata. We use those embeddings in our second module, Multimodal Label Fusion via Mutual Co-Attention (CMF-MCA). Our CMF-MCA module incorporates two transformers with co-attention mechanisms that can mutually annotate text and images. In our third methodology, Semantic Propagation Graph for Hate and Fake Correlation (SPG-HFC), we implement a relational graph attention mechanism that captures both the influence of semantics and how communities propagate information about hate and fake posts. The fourth module, Adaptive Modality Routing via Reinforcement (AMR-R), routes based on the modality of the input and whether the input is simple enough to be classified using machine learning or complex enough to require deep learning. Finally, our Counterfactual Consistency Validation Engine (CCVE) is used after prediction to validate that the model’s predictions are consistent with the output data by creating counterfactuals and validating them. Therefore, in addition to improving the overall accuracy of hate speech and fake post detections, our proposed framework also improves its scalability and inference reliability. Additionally, because our framework allows multimodal classifications that include both context and behavior, it enables the scalable and trustworthy development of content moderation systems. Full article
Show Figures

Figure 1

24 pages, 2253 KB  
Article
Quantum-Inspired Semantic Encoding and Temporal Transformer Fusion (QuST-TF) for Misinformation Detection
by Krishna Kumar and Akila Venkatesan
Appl. Sci. 2026, 16(13), 6338; https://doi.org/10.3390/app16136338 - 24 Jun 2026
Viewed by 286
Abstract
Misinformation propagates more rapidly than factual content on social media, presenting significant challenges for automated misinformation detection. Existing approaches often focus solely on textual features without incorporating temporal information, treat timing and propagation as separate factors, or apply quantum-inspired methods primarily to multimodal [...] Read more.
Misinformation propagates more rapidly than factual content on social media, presenting significant challenges for automated misinformation detection. Existing approaches often focus solely on textual features without incorporating temporal information, treat timing and propagation as separate factors, or apply quantum-inspired methods primarily to multimodal data rather than text-centric misinformation. This study introduces QuST-TF (Quantum-inspired Semantic encoding and Temporal Transformer Fusion), a unified model designed to detect misinformation in tweets and news articles. QuST-TF integrates quantum-inspired (classical approximation) amplitude encoding, time-aware Transformer fusion, and propagation graph attention based on engagement data, without reliance on images, audio, or quantum hardware. Performance gains are achieved through quantum-inspired (classical approximation) nonlinear angular modulation (cosine and sine rotations) implemented via classical computation, rather than genuine quantum computing. All computations utilize classical Dense layers, Rectified Linear Unit (ReLU) activations, and cosine/sine functions on CPUs or GPUs; quantum hardware is not required. The quantum-inspired (classical approximation) layer applies classical rotation-based transformations to enrich the semantic representation of BERT (Bidirectional Encoder Representations and Transformer) embeddings. Temporal information is captured by a dual-attention Transformer encoder, while propagation graph attention monitors the spread of claims. Evaluation on FakeNewsNet and PHEME datasets demonstrates 91.4% and 95.5% accuracy, respectively, with 34% fewer trainable parameters compared to standard Transformers. Ablation studies indicate that quantum encoding is the most influential component (+3.0% versus without quantum encoding), surpassing the contributions of graph attention (+2.6%) and temporal attention (+2.2%). The integration of all three components yields a 1.3% synergistic improvement, confirming effective inter-module collaboration. Attention visualization enhances interpretability, supporting the utility of QuST-TF for fact-checking applications. Full article
Show Figures

Figure 1

22 pages, 923 KB  
Article
Early Detection of Fake News via Structured Social Interaction Simulation and Hierarchical Cross-Modal Fusion
by Ruihua Qi, Shuqin Chen, Weilong Li, Chenwei Zhang, Jiatai Lei, Haobo Lv and Yunhao Sun
Appl. Sci. 2026, 16(12), 6001; https://doi.org/10.3390/app16126001 - 13 Jun 2026
Viewed by 304
Abstract
The widespread dissemination and societal impact of fake news underscore the critical need for effective detection. Existing methods remain limited, as they often fail to learn joint representations from multi-modal data and rely heavily on complete social interaction signals. Such information is frequently [...] Read more.
The widespread dissemination and societal impact of fake news underscore the critical need for effective detection. Existing methods remain limited, as they often fail to learn joint representations from multi-modal data and rely heavily on complete social interaction signals. Such information is frequently unavailable in practice, especially during the early propagation stages. To address early fake news detection in social media, this paper proposes a hierarchical cross-modal fusion framework with structured LLM-simulated social interaction (HCF-LSIM). The framework employs a progressive cross-modal attention mechanism to systematically align semantic representations across multiple levels, integrating textual, thematic, and visual features. Additionally, HCF-LSIM designs an LLM-powered social interaction simulator that generates structured triplets from adapted user profiles, effectively compensating for missing real-time interaction data. Experiments on public benchmarks demonstrate strong performance, with accuracies of 93.5% on Weibo and 87.2% on X (formerly Twitter), ranking first on Weibo and second on Twitter. Full article
(This article belongs to the Section Computing and Artificial Intelligence)
Show Figures

Figure 1

17 pages, 10311 KB  
Article
DeepFakeX: A Comprehensive Multimodal Deepfake Dataset for Research and Analysis
by Sonia Salman, Jawwad Ahmed Shamsi and Rizwan Qureshi
Data 2026, 11(6), 141; https://doi.org/10.3390/data11060141 - 11 Jun 2026
Viewed by 1131
Abstract
The expanding capabilities of deep learning-based media synthesis have intensified concerns regarding the authenticity of digital content and the reliability of forensic analysis tools. In response to these challenges, this work introduces DeepFakeX, a collection of 800 synthetically generated videos available under controlled [...] Read more.
The expanding capabilities of deep learning-based media synthesis have intensified concerns regarding the authenticity of digital content and the reliability of forensic analysis tools. In response to these challenges, this work introduces DeepFakeX, a collection of 800 synthetically generated videos available under controlled access for research purposes. The dataset encompasses four distinct categories of AI-driven synthesis: facial identity replacement, audio track substitution, neural voice cloning, and combined audiovisual alteration. Unlike existing deepfake datasets that predominantly focus on facial synthesis, DeepFakeX covers a broader range of manipulation modalities, reflecting the diversity of synthetic media encountered in real-world settings. All deepfakes were generated using state-of-the-art, publicly available tools. Standardized post-processing procedures were applied to each video to ensure uniformity in terms of quality, duration and encoding format. DeepFakeX also emphasizes diversity in gender, age, ethnicity, and language. Video contexts span speeches, informational videos, movie clips, news broadcasts, and interviews that reflect content scenarios commonly encountered in real-world online environments. The dataset includes videos in both English and Urdu. The dataset’s quality and structural variability were assessed through visual and audio analyses using the Structural Similarity Index Measure (SSIM), Mel-Frequency Cepstral Coefficients (MFCCs), and Principal Component Analysis (PCA). The evaluation results revealed substantial variability within each manipulation category, along with clearly distinguishable patterns specific to each modality. DeepFakeX has been developed to facilitate rigorous and transparent research in deepfake detection, cross-modal forensic analysis, and AI-driven media forensics. It is hosted on Zenodo under controlled access for research use. Full article
Show Figures

Figure 1

16 pages, 4105 KB  
Article
SIFTNet: Structure-Guided Iterative Fusion with a Transformer Network for Fake News Detection
by Xuekun Zhang, Weijian Fan, Chi Zhang, Guowei Chen and Pengzhou Zhang
Electronics 2026, 15(12), 2582; https://doi.org/10.3390/electronics15122582 - 11 Jun 2026
Viewed by 248
Abstract
Fake news detection has become critical for safeguarding social media users and maintaining a reliable news ecosystem. However, existing methods rely mainly on context information and propagation structure and do not consider the news structure from framing theory. As a highly structured genre, [...] Read more.
Fake news detection has become critical for safeguarding social media users and maintaining a reliable news ecosystem. However, existing methods rely mainly on context information and propagation structure and do not consider the news structure from framing theory. As a highly structured genre, news implies writing intention and organizational logic in its discourse frame, which provides vital clues for authenticity verification. In this paper, we propose structure-guided iterative fusion with a transformer network for fake news detection (SIFTNet), which contains four modules: a structural label generator, an information architecture representation module, a structure-enhanced representation module, and a structure-guided iterative fusion module. Guided by framing theory, SIFTNet captures the semantics at both the local sentence level and global structure level. Extensive experiments demonstrate that our model achieves state-of-the-art performance on both Chinese and English datasets, exhibiting superior effectiveness and robustness. These findings validate the efficacy of applying framing theory to improve fake information detection. Full article
Show Figures

Figure 1

25 pages, 3608 KB  
Article
GC2MFND: Multi-Granularity Conflict and Domain-Guided Calibration for Multimodal Fake News Detection
by Yanming Sun, Mingyue Zhang and Fujun Zhang
Entropy 2026, 28(6), 672; https://doi.org/10.3390/e28060672 - 11 Jun 2026
Viewed by 397
Abstract
On current social media platforms, multimodal fake news has permeated various fields. Multi-domain fake news detection has garnered significant attention in the academic community. Existing multi-domain methods primarily employ feature fusion techniques based on text–image alignment, neglecting the extraction of conflicting information across [...] Read more.
On current social media platforms, multimodal fake news has permeated various fields. Multi-domain fake news detection has garnered significant attention in the academic community. Existing multi-domain methods primarily employ feature fusion techniques based on text–image alignment, neglecting the extraction of conflicting information across modalities and failing to address the domain-dependent nature of cross-modal feature conflicts. To address this, we propose a Multi-Granularity Conflict and Domain-Guided Calibration for Multimodal Fake News Detection model (GC2MFND). This model captures conflicting features through the domain-aware multi-granularity conflict extraction module and mitigates feature suppression using the domain-guided multimodal feature calibration module. Finally, it combines domain-adaptive aggregation with multi-view evidence integration to achieve robust decision-making under supervised contrastive learning constraints. Under known domain conditions, the experimental results demonstrate that GC2MFND outperforms existing multi-domain baseline methods, achieving accuracy rates of 95.3%, 95.7%, and 81.2% on the Weibo, Weibo21, and FineFake datasets, respectively, representing improvements of 1.1%, 1.2%, and 1.4% over the corresponding multi-domain baselines. Full article
(This article belongs to the Section Multidisciplinary Applications)
Show Figures

Figure 1

27 pages, 7120 KB  
Article
Systematic Fine-Tuning of Transformer Models for Domain-Specific Misinformation Detection in Spanish Social Media Text
by Gabriel Hurtado Avilés, José A. Reyes-Ortiz, Román A. Mora-Gutiérrez, Josué Padilla Cuevas and Óscar Herrera Alcántara
Informatics 2026, 13(6), 83; https://doi.org/10.3390/informatics13060083 - 9 Jun 2026
Viewed by 512
Abstract
While social media platforms are primary vectors for misinformation, automated detection systems remain largely confined to English. This paper presents a transferable, three-stage framework for fine-tuning transformer models to detect domain-specific deceptive content in Spanish. The pipeline comprises: (1) corpus unification, merging fragmented [...] Read more.
While social media platforms are primary vectors for misinformation, automated detection systems remain largely confined to English. This paper presents a transferable, three-stage framework for fine-tuning transformer models to detect domain-specific deceptive content in Spanish. The pipeline comprises: (1) corpus unification, merging fragmented datasets into a 61,674-article resource mapped into three classes (Real, Fake, Satire) to prevent stylistic confounding; (2) systematic model optimization, extensively benchmarking classical metaheuristics against eight transformer architectures (including mBERT, XLM-RoBERTa, and BETO) using strong regularization to mitigate overfitting; and (3) production deployment, encapsulating the optimized model as a containerized web application for real-time inference. Through rigorous experimentation, the Spanish-specific BETO encoder emerged as the strongest model for this task, achieving 89.18% overall accuracy. The model attains a near-perfect in-source F1-score on the satire class; however, a strict source-held-out test reveals that this performance is highly source-dependent—recall on satire from an unseen outlet drops to 0.08—indicating that single-source class construction leads the model to recognize the source rather than a generalizable category. We report this finding as a central methodological result: corpus design, and in particular the source diversity of each class, is the primary determinant of whether the framework generalizes. Adversarial robustness tests using named-entity masking and typo injection provide complementary evidence on the model’s reliance on semantic versus surface cues. The methodology is designed to be adaptable across domains: by substituting the training corpus, the same framework may in principle be retargeted to other digital threats, such as investment scams and phishing, provided that suitable labeled corpora are constructed and validated for each new domain. The complete framework, dataset, and application are released as open-source resources to support reproducible research and practical countermeasures against online misinformation. Full article
(This article belongs to the Special Issue Machine Learning in Social Media Analysis)
Show Figures

Figure 1

34 pages, 3502 KB  
Article
Complex-Time Framework for Authenticity and Identity in Personalized AI
by Gerardo Iovane, Giovanni Iovane, Antonio De Rosa and Francesco Barbato
Algorithms 2026, 19(6), 458; https://doi.org/10.3390/a19060458 - 5 Jun 2026
Viewed by 403
Abstract
The proliferation of AI-generated content and personalized AI systems has sharpened two fundamental and related computational problems: the progressive erosion of authentic identity in AI-mediated representations, and the growing difficulty of distinguishing human-originated from AI-generated behavioral and textual streams. This paper proposes a [...] Read more.
The proliferation of AI-generated content and personalized AI systems has sharpened two fundamental and related computational problems: the progressive erosion of authentic identity in AI-mediated representations, and the growing difficulty of distinguishing human-originated from AI-generated behavioral and textual streams. This paper proposes a rigorous computational framework in which digital identity is formalized as a holomorphic function of complex time T = (a + ib) ∈ ℂ, where the real component Re(T) encodes chronological progression and the imaginary component Im(T) spans a continuum from episodic memory (Im(T) < 0) through the present moment (Im(T) = 0) to prospective imagination (Im(T) > 0). We argue that holomorphicity—enforced via Cauchy–Riemann regularization during CTNN learning (Proposition 1)—provides a theoretically grounded encoding of identity coherence, and discuss its advantages over alternative mathematical choices, including Lipschitz continuity, C smoothness, piecewise analytic functions, and stochastic models. Under four explicit Assumptions 1–4 covering the Markovian structure and fixed context window of current LLM architectures, we establish via Lemmas 1 and 2 and Theorem 1 that AI-generated behavioral trajectories exhibit structural limitations in satisfying the Cauchy–Riemann conditions at temporal depths characteristic of human biographical memory—limitations that do not arise for human trajectories learned under CTNN regularization. Building on this result, we introduce the Human–AI Authenticity Discriminant (HAAD), a theoretically grounded classifier with a fully specified calibration algorithm and sensitivity analysis (κ ΔAUROC ≤ 0.04 over ±30% perturbation). Five metrics—TCS, ISI, PAS, GAS, and HAAD—are derived analytically from the holomorphic structure. The algorithmic framework is instantiated on four real-world datasets: MovieLens 25M, the Pushshift Reddit corpus, the Stack Overflow Data Dump, and the LIAR dataset. On the LIAR benchmark, TDT-HAAD achieves AUROC = 0.82 (95% CI: [0.79, 0.85]), exceeding a RoBERTa-based LLM detector baseline (AUROC = 0.75, DeLong p < 0.01); an ablation study supports the structural contribution of each component. A credibility harvesting signature is detectable 45.3 ± 12.1 days before standard temporal models reach statistical significance. Full article
Show Figures

Figure 1

12 pages, 1608 KB  
Article
Deep Neural Network Architectures for Fake News and Misinformation Detection
by Mariam Ibrahim and Ruba Elhafiz
J. Cybersecur. Priv. 2026, 6(3), 97; https://doi.org/10.3390/jcp6030097 - 5 Jun 2026
Viewed by 389
Abstract
The prompt spread of misleading information through recent information and communication technologies (ICT) admonishes social convention and credence. Developing trustworthy algorithms that can automatically identify fake content becomes increasingly difficult. We investigate a hybrid artificial intelligence (AI) strategy that integrates machine learning (ML) [...] Read more.
The prompt spread of misleading information through recent information and communication technologies (ICT) admonishes social convention and credence. Developing trustworthy algorithms that can automatically identify fake content becomes increasingly difficult. We investigate a hybrid artificial intelligence (AI) strategy that integrates machine learning (ML) and deep learning (DL) to enhance fake news detection. The model’s deep learning entity evaluates confined text arrangements and inclusive text values using a Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory (BiLSTM) with an attention layer. Conventional machine learning classifiers, mostly Support Vector Machine (SVM), Random Forest (RF), and Logistic Regression (LR), are trained synchronously employing Term Frequency–Inverse Document Frequency (TF-IDF). A simple ensemble averaging strategy is used on both machine learning and deep learning predictions. The model demonstrates strong generalization across various text types when evaluated on the LIAR dataset and a Kaggle-style fake news dataset. The combined system performs noticeably better than each of the separate models in terms of accuracy, precision, recall, F1, and AUC. Full article
(This article belongs to the Section Security Engineering & Applications)
Show Figures

Figure 1

34 pages, 2657 KB  
Article
KazFakeCorpus: A Bilingual Corpus with Multi-Level Semantic Annotation for Fake News Detection
by Zhanar Lamasheva, Anargul Nekessova, Mansiya Kantureyeva, Madina Sambetbayeva, Mira Kaldarova and Aksaule Nazymkhan
Big Data Cogn. Comput. 2026, 10(6), 183; https://doi.org/10.3390/bdcc10060183 - 1 Jun 2026
Viewed by 571
Abstract
This paper addresses the lack of bilingual annotated resources for automatic fake news detection in the Kazakh–Russian media space, as well as the limitations of binary annotation, which does not always allow disinformation to be represented as a complex and interpretable phenomenon. The [...] Read more.
This paper addresses the lack of bilingual annotated resources for automatic fake news detection in the Kazakh–Russian media space, as well as the limitations of binary annotation, which does not always allow disinformation to be represented as a complex and interpretable phenomenon. The aim of the study is to develop KazFakeCorpus and propose a multi-level annotation scheme that captures not only the final veracity of a message, but also the type of fake content, the disinformation technique, communicative intent, modality, and the characteristics of the source and evidence base. The corpus was constructed on the basis of official news materials published on the Gov.kz portal for the REAL class and synthetically generated messages for the FAKE class, complemented by an external validation set of authentic fake news from independent fact-checking sources to assess generalization. After data collection, the texts underwent cleaning, normalization, balancing, and sampling. The final resource includes 4276 texts in Kazakh and Russian, with an average length of approximately 200 words and a balanced distribution across languages and classes. Annotation was carried out in the Label Studio environment by two independent experts: a linguist and a fact-checking specialist. Before the main annotation phase, a pilot study was conducted on a subsample of 120 texts, the results of which were used to refine the categories and prepare the annotation guidelines. Krippendorff’s alpha was used to assess inter-annotator agreement; the obtained values, ranging from 0.79 to 0.88, indicate sufficient stability of the annotation across the key categories. The corpus analysis showed that misattribution (32.5%) is the most frequent disinformation technique, followed by clickbait (23.0%) and emotional pressure (16.4%). The results show that the proposed scheme makes it possible to treat fake news not only as a binary class but also as a multi-level semantic object that includes mechanisms of information distortion and features of content presentation. The practical contribution of the study lies in the creation of a bilingual corpus and annotation protocol that can be used in disinformation research, interpretable text analysis, and cross-lingual studies. Full article
(This article belongs to the Section Data Mining and Machine Learning)
Show Figures

Figure 1

26 pages, 3339 KB  
Article
Fake News Detection Using Text-Based Graph Convolutional Networks
by Faisal A. Alshuwaier and Fawaz A. Alsulaiman
Computers 2026, 15(6), 352; https://doi.org/10.3390/computers15060352 - 30 May 2026
Viewed by 734
Abstract
Detecting fake news is a challenging task and an important area of research for social media researchers. This task also involves clarifying accountability mechanisms that demonstrate the credibility of quotable sources, such as networks that document the spread of misinformation. Deep learning techniques, [...] Read more.
Detecting fake news is a challenging task and an important area of research for social media researchers. This task also involves clarifying accountability mechanisms that demonstrate the credibility of quotable sources, such as networks that document the spread of misinformation. Deep learning techniques, particularly neural networks that rely on popular graph representation techniques such as graph convolutional networks (GCNs), are increasingly being utilized to detect fake news, fake accounts, and rumors spreading through social media. In this paper, features were extracted using TF-IDF, Bag-of-Words, and bigrams. The evaluation was conducted using the standard Kaggle/ISOT and GossipCop datasets, which include news headlines and published models. Using the extracted features, the proposed GCN-based model/classifier achieved a high detection accuracy of 95% by combining TF-IDF and Bag-of-Words representations. The results demonstrate that the extracted features improve the efficiency of the detection model. Full article
(This article belongs to the Special Issue Advances in Semantic Multimedia and Personalized Digital Content)
Show Figures

Figure 1

29 pages, 3277 KB  
Article
MiniLM-CNN-LSTM: A Lightweight Hybrid Transformer Model for Malicious URL Detection
by Emad-ul-Haq Qazi, Muhammad Hamza Faheem and Abdulrazaq Almorjan
Technologies 2026, 14(6), 316; https://doi.org/10.3390/technologies14060316 - 24 May 2026
Viewed by 817
Abstract
Phishing and malicious websites are a serious threat on the internet. Attackers use fake links to trick users and steal their private information. Detecting these links is difficult because attackers change their tricks often. Many old methods cannot detect new or hidden threats. [...] Read more.
Phishing and malicious websites are a serious threat on the internet. Attackers use fake links to trick users and steal their private information. Detecting these links is difficult because attackers change their tricks often. Many old methods cannot detect new or hidden threats. Some recent models use deep learning (DL), but they are large, slow, and hard to use in real-time systems. In this paper, we present a lightweight and accurate model called MiniLM-CNNLSTM. It combines a small transformer model (MiniLM) with a hybrid DL network using Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) layers. The transformer learns the meaning of URLs. The CNN finds important patterns. The LSTM captures the order of characters. We also add handcrafted features that help the model detect tricky URLs. We test our method on two public datasets: the Phishing Site URLs dataset and the Malicious URLs dataset from Kaggle. We use 3-fold cross-validation and early stopping to ensure fair and stable results. The MiniLM-CNN-LSTM model outperformed previous benchmarks by achieving an average three-fold cross-validation accuracy of 98.98%, a precision of 98.63%, a recall of 98.29%, an F1-score of 98.46%, and a false positive rate of 0.68%. The proposed model has a higher accuracy, precision, recall, F1-score and a lower false positive rate, which enhances the accuracy by 1.88, precision by 3.77, recall by 4.17 and decreases the false positive rate by 61.58% compared with the strongest baseline (Distil BERT + CNN-LSTM), showing significant practical improvements. The results show that our approach is fast, small, and highly effective. It can detect phishing and malicious links with high accuracy. This makes it a good choice for real-time security systems like browsers, email filters, or firewalls. Full article
(This article belongs to the Special Issue Research on Security and Privacy of Data and Networks)
Show Figures

Figure 1

Back to TopTop