1. Introduction
As information and communication services become more digitized, fake news phenomena grow. According to a European Commission study (2018), fake news is “all forms of false, inaccurate, or misleading information designed, presented, and promoted to cause public harm intentionally or for profit”. The propagation of false news on the internet is a major issue in today’s society [
1]. With the expansion of online communication, a rising number of individuals use digital platforms to get news and exchange information.
The scientific community has demonstrated an increasing interest in fake news identification in recent years. Consequently, a number of fake news detection systems have been created in an effort to automatically determine whether or not news is phony [
2].
The majority of these methods formulate the problem of automatically identifying false information as a supervised binary classification problem. In order to obtain good performance, the news is divided into two classes (false and legitimate news), and the classifier is trained and validated using a dataset of labeled data [
3]. However, as datasets grow larger and more complex, the capacity of deep learning to automatically build hierarchical representations from raw data has drawn significant attention in the field of false news detection [
4].
Semantic linkages, temporal dependencies, and contextual information in textual data have been captured using methods including Recurrent Neural Networks (RNNs), CNNs, and attention mechanisms [
5]. Moreover, automated large-scale text data categorization has become a popular area of study. Mapping textual input to a collection of predetermined labels is the main goal of text classification, a basic aspect of natural language processing (NLP). It has been widely used in the fields of public opinion monitoring, spam detection, and sentiment categorization [
6].
Building on these developments, this research introduces a hybrid algorithm that integrates conventional machine learning and deep learning in an ensemble configuration to detect false news. The BiLSTM network, which is the foundation of the architecture, has an adaptive attention mechanism that records contextual dependencies. In addition, the deep learning model integrates the TF–IDF vectorizer with classical classifiers. By combining these components, the framework leverages the distinct benefits of the two learning approaches. Classical Machine Learning (ML) can identify better decision boundaries in high-dimensional sparse feature sets, whereas deep learning can capture sequence patterns and semantic context. Additionally, a variety of datasets, such as the LIAR dataset and the Kaggle-style fake news corpus, were employed to assess the suggested methodology. The following aspects summarize our paper’s contributions.
Dual architecture: The proposed deep learning architecture combines traditional TF-IDF-based machine learning models with BiLSTM with a specific attention layer.
Multi-dataset assessment: We show that the suggested algorithm works well on a variety of datasets.
Pipeline prepared for deployment: We provide a whole pipeline that may be expanded or repurposed in useful false news detection systems, from data pretreatment and training to model saving and ensemble inference.
The remainder of the paper is presented as follows.
Section 2 covers some of the studies conducted in this topic.
Section 3 goes over the main preliminaries.
Section 4 displays the methodology, while
Section 5 discusses the assessment and results. Finally, in
Section 6, we discuss our inputs and future endeavors.
2. Related Work
One of the significant issues that may be resolved using DL is the propagation of fake news. An intelligent-detection-system-based ensemble voting classifier is suggested in [
7] to handle news classification. Naïve Bayes, k-nearest neighbors (k-NN), support vector machine (SVM), random forest (RF), artificial neural network (ANN), logistic regression, gradient boosting, Ada boosting, and other machine learning methods were used by the authors to detect forged news. Additionally, following cross-validation, the top three machine learning algorithms were included in the ensemble voting classifier, with a reported maximum accuracy of 94.5%.
Additionally, ref. [
8] proposes two deep learning models that are successful in addressing the challenge of detecting false news in online news material across several domains. On the FakeNews AMT and Celebrity datasets, the systems perform admirably, surpassing the existing handmade feature-engineering-based systems by a noteworthy margin of 3.08% and 9.3%, respectively. A ML model is created by [
9] to identify stance phrases that contain news headlines like discusses, unrelated, agrees, and disagrees. To find attitude sentences, a deep model using string similarity characteristics is used. The model includes efficient text representation, document categorization, and natural language inference. In particular, bi-directional recurrent neural networks (BiRNNs) and neural attention are used on the temporal dimension with a max-pooling layer.
The several features—syntactic, semantic, lexical, word, and morphological—obtained from the reviews are processed using deep RNN and SVM. An inquiry for the identification of false news is offered by [
10]. Semi-supervised learning is helpful when there is a significant number of unlabeled data and a small number of labeled data available for training and model assessment. A unique two-path deep semi-supervised learning framework is then created by [
11], with one path being for supervised learning and the other for unsupervised learning. While the unsupervised learning path may learn from a vast quantity of unlabeled data, the supervised learning path can only learn from a small amount of tagged data. Additionally, these two CNN-implemented pathways are jointly optimized to finish semi-supervised learning. Additionally, a common CNN is constructed to feed the low-level features into these two routes by extracting them from both labeled and unlabeled data. A Word-CNN-based semi-supervised learning model is developed and tested on the LIAR and PHEME datasets in order to validate this approach. In the conclusion, ref. [
12] investigates the Vehicular Ad-hoc Network (VANET) to carry out vehicle communication to exchange different messages for efficient travel meant for the passenger. However, there are instances when intruders spread misleading information about traffic jams, accidents, etc., which negatively impacts vehicle security.
In order to identify fake information, an entropy-based method is created in [
13]. A theory-driven methodology is presented for detecting false news. This approach examines news material at the lexicon, discourse, syntax, and semantic levels. Every level of the news is captured, and a supervised ML model is used to identify false news based on social relevance and forensic psychology theory. The authors investigated characteristics that can be used to identify false news, including linkages between fake news, patterns of fake news, and interpretability of fake news for feature engineering. Ref. [
14] gradually created several deep learning models to identify false information and classify it into predetermined, fine-grained categories. The input is represented by CNN and Bi-directional Long Short Term Memory (Bi-LSTM), and the combined output is fed into a Multi-layer Perceptron (MLP) model.
False news classification is carried out in [
15]. The authors used hybrid CNN and RNN models to detect false news on Twitter. The suggested framework has an accuracy of 82% in identifying and categorizing false news messages from tweets. Without relying on prior domain knowledge, the proposed method is able to automatically learn and extract relevant features associated with false news content.
A multi-modal cross-attention technique for AI-generated content (AIGC) video forgery detection is proposed in [
16]. A cross-attention module is utilized as a fusion and inconsistency diagnostic tool with an accuracy of 94.32%. Yet, the generalization effectiveness requires more investigation.
A fusion detection Transformer (F-DETR) is proposed in [
17] with a heterogeneous-scale multi-branch structure. The performance of DETRs primarily relies on the attention mechanism, which captures dependencies among different spatial positions in the image, providing new opportunities for the incorporation of many multi-scale features. Yet, their method did not yield the best outcomes on the COCO dataset and lacks major uniqueness.
Table 1 provides an overview of related work in comparison to our hybrid model.
3. Preliminaries
3.1. Datasets and Labeling Schemes
To evaluate the proposed framework, two text-based datasets are used: the LIAR dataset and the Kaggle Fake News dataset.
LIAR dataset: The LIAR dataset consists of 12,836 brief utterances tagged for truthfulness, subject, context/venue, speaker, state, party, and history. LIAR is an order of magnitude greater in size than similar resources now accessible. In addition, unlike crowdsourced datasets, LIAR examples are collected in a more natural setting, such as political discussion, TV advertising, Facebook posts, tweets, interviews, press releases, and so on. In each case, the labeler gives a detailed analytical report to back up each conclusion, as well as links to any supporting papers [
18].
Kaggle Fake News dataset: The fake news element of this dataset is built from Kaggle’s fake news dataset, which includes material from the 2016 US election campaign. The true news segment is gathered from media outlets such as the New York Times, Wall Street Journal, Bloomberg, National Public Radio, and the Guardian during 2015 or 2016. The dataset’s repository contains around 6.3 k news items, with an equal distribution of false and true news, with political news accounting for half of the total corpus [
19].
3.2. Feature Extraction and Deep Learning
TF-IDF: The TF-IDF algorithm works by comparing the relative frequency of words in a single text to the inverse percentage of that word over the whole document corpus. Intuitively, this algorithm assesses the relevance of a certain word in a specific manuscript. Terms that occur often in a single or small set of texts have higher TF-IDF values than common terms like articles and prepositions [
20,
21].
CNN Architecture: Starting with a tokenized sentence, it is transformed to a sentence matrix with rows containing word vector representations of each token. We designate the dimensionality of the word vectors as
. The sentence matrix has a dimensionality of
, where
s is the sentence length. Consider a filter matrix w with a region size of
. It will have
parameters that need to be calculated [
22]. The sentence matrix is denoted as
and
represents the sub-matrix from row
to row
. The convolution operator produces the output sequence
by repeatedly applying the filter to sub-matrices of
[
23]. To create the feature map
for this filter, we add a bias term (
) and an activation function (
) to each
[
24]. Each feature map is thus subjected to a pooling function in order to generate a fixed-length vector [
24]. The final classification may be produced by concatenating the outputs from each filter map into a fixed-length, “top-level” feature vector and feeding it through a softmax function. One can use “dropout” [
25] as a regularization technique at this softmax layer. The categorical cross-entropy loss is the training goal that has to be reduced. The bias term in the activation function, the weight vector of the softmax function, and the weight vector of the filter are among the parameters that need to be calculated. Stochastic Gradient Descent (SGD) and back-propagation are used for optimization [
26].
BiLSTM: Sequential data processing is the primary use of RNN. It forecasts the future output by storing the pertinent portions of the input data. Memory cells in the RNN are used to store the most pertinent data from previous inputs. Furthermore, it is unable to manage long-term dependence. Consequently, a unique kind of RNN known as a long short-term memory network (LSTM) is created to address the long-term dependencies [
27].
Three gating concepts—input (IG), output (OG), and forget gates (FG)—are used to regulate the candidate hidden state (CHS), current state (CS), and hidden sequence (HS) in addition to the information flow (read, write, and reset) across the gradient [
28]. In particular, information is processed unidirectionally by the LSTM network, either from left to right or from right to left [
29].
Attention mechanism: The attention layer generates a context vector for the learnt input vectors. It has a significant impact on machine translation, text summarization, text categorization, and question-answering systems. In this study, the attention layer is created on top of the BiLSTM network to update the weights. Specifically, the attention layer applies greater weights to the most relevant and essential words in the input sequence. The advantage of this method is that it preserves lengthier input sequences [
30]. The incorporation of the attention layer allows the model to concentrate on the most informative terms in the input text, which enhances interpretability and classification performance. The attention mechanism improves the BiLSTM model’s capacity to identify forged news patterns by highlighting significant textual contents.
3.3. Evaluation Measures
The number of correctly identified class examples (true positives,
), correctly identified examples that do not belong to the class (true negatives,
), and examples that are either incorrectly assigned to the class (false positives,
) or not recognized as class examples (false negatives,
) can all be used to assess how accurate a classification is. For the binary classification example, these four counts make up the confusion matrix [
31].
Table 2 presents the measures used most often for binary classification based on the values of the confusion matrix.
4. Methodology
Figure 1 illustrates the overall workflow of the proposed hybrid fake news detection pipeline. The flowchart provides a high-level visualization of the data preprocessing steps, the parallel deep learning and classical machine learning branches, the ensemble averaging mechanism, and the final evaluation stage.
Every deep learning model starts by tokenization. For example, each piece of news gets transformed into a series of digits. In a limited vocabulary, each word corresponds to an integer value that indicates its position. Every sequence of tokens is capped at a length of L = 200. For each sequence, the tokens are embedded into a dense vector of size 128. Two stacks of 1D-convolutional networks hold the embedding matrix, with each one being followed by a max-pooling and a dropout layer. From the embedding, the n-grams feature representations, built by the convolutional layers, capture and retain the stylistic and syntactic structures used in distinguishing between fake and true news.
Now, the output from the CNN goes to a BiLSTM that processes the text in both directions to capture long-range dependencies. To increase such interpretability and focus on the relevant pieces of text, an Attention Layer is added to the output of the BiLSTM. The attention weights indicate which tokens contributed the most to making a particular classification decision. Lastly, predicting from deep learning output is achieved by sending the attention vector through multiple fully connected layers. To complement deep learning, the algorithm employs three TF–IDF-based ML models: Support Vector Machine (linear kernel), Random Forest, and Logistic Regression.
The ensemble technique comprises the averaging of both traditional ML outputs and deep model predictions. By balancing the shortcomings of each model class, this ensemble approach generates a forecast that is more reliable and safer. Before aggregation, all classifier outputs are probability-calibrated using Platt scaling to ensure comparable output ranges across models. The final ensemble prediction is then obtained using uniform weighted averaging over the calibrated probabilities.
Algorithm 1 demonstrates how the proposed system uses ensemble learning combined with deep and classical ML to identify fake content. This includes preparing datasets, cleaning documents, generating features, training deep neural networks, classical ML with TF-IDF, and ensemble learning for the final prediction. The architecture focuses on global contextual information and the extraction of local text structures. Furthermore, it uses the strong decision functions and interpretability of classical ML classifiers. The key parameter values are: sequence length (200 tokens), embedding dimension (128), convolutional filters (up to 128 filters), BiLSTM units (64 per direction), dropout rates between 0.3 and 0.4, and TF–IDF feature sizes between 10,000 and 15,000 terms (depending on the classifier).
Table 3 shows the hyperparameters used for fine-tuning the models based on commonly used settings in recent fake news detection studies.
| Algorithm 1: Hybrid Fake News Detection Pipeline |
Output: Trained models, ensemble predictions, evaluation metrics BEGIN 1: Input fetch LIAR and Kaggle-style datasets; add source labels. 2: Combine: merge all data; drop empty rows; keep text, label. 3: Split: 80/20 train–test with stratified sampling. 4: Tokenizer: fit on training text; set vocab size. 5: Sequences: convert text to integer sequences; pad to fixed max length. 6: Embedding: map tokens to dense vectors; apply dropout. 7: Model. CNN: apply Conv1D layers for n-gram features; max pooling + dropout. 8: Model. BiLSTM: process CNN output bidirectionally (return sequences). 9: Attention: compute attention weights; obtain context vector. 10: Dense layers: apply two fully connected layers + dropout. 11: DL output: final sigmoid probability. 12: Train DL model: train with class weights; track accuracy, precision, recall, AUC. 13: TF–IDF: extract features. 14: Train ML models: SVM, Random Forest, Logistic Regression. 15: Average: combine DL + ML probabilities: 16: Evaluate: compute confusion matrix; accuracy, precision, recall, F1-score, AUC. END |
5. Results and Discussion
The comparative accuracy of the models is shown in
Figure 2. The BiLSTM + Attention model achieved an accuracy of approximately 0.94, outperforming the investigated classical TF–IDF-based models. The ensemble model, however, yielded the highest accuracy of 0.96, confirming the benefit of combining heterogeneous classifiers.
Similarly,
Figure 3 presents the F1-score comparison. The BiLSTM + Attention model demonstrated a strong balance between precision and recall (F1-score ≈ 0.94). The ensemble model further improved this performance (F1-score ≈ 0.945), indicating more consistent classification across both real and fake news instances.
The AUC comparison in
Figure 4 shows that the ensemble model provided the best overall separability between the two classes, with an AUC close to 0.99. The BiLSTM + Attention model also performed strongly (AUC = 0.9832), while classical models followed with slightly lower but still competitive AUC values of 0.97.
The experimental results show that the CNN–BiLSTM + Attention architecture is highly effective for detecting fake news, primarily because it captures both local linguistic patterns (via convolutional layers) and long-range contextual dependencies (via bidirectional LSTMs). The attention mechanism further enhances the model by focusing on the most informative segments of text, contributing to its high F1-score and AUC.
Classical machine learning techniques are less sophisticated, but also provide valuable contributions, as they are able to set high linear and tree-based decision boundaries across the high-dimension TF–IDF literature. Their consistent performance across various metrics validates the effectiveness of integrating modern and classical approaches. In our experiments, the proposed ensemble approach outperformed all individual model components, as it achieved the highest results across all three metrics: accuracy, F1-score and AUC. This is probably because the individual components are lexically classical and therefore achieve robust decision boundaries, and also semantically deep learning, which is one of the unique factors of the ensemble. This contributes to improved reliability and generalization.
Qualitative assessment based on sample prediction has more than enough evidence to support this. Well-written, factual literature is predicted to be real and very convincing, while all are confirmed to be fake without any evidence. Most remaining prediction errors occur in borderline cases—typically factual content expressed with a sensational tone or text mimicking journalistic style without clear evidence. Incorporating multimodal features may help address such instances. Well-written factual articles were consistently classified as real.
Recent studies have tackled the use of advanced deep learning as well as hybrid models in the detection of fake news. However, these are unfortunately still lower compared to the models introduced in this paper. For example, study [
32] utilized hierarchical attention architectures, but only produced 0.89 total accuracy with an F1-score of 0.88. Study [
33] utilizes transformers, which are considered to be advanced in the field, but still only yielded 0.92 total accuracy with a 0.91 F1 score.
Recently, the field has seen the use of ensemble learning models. Ref. [
34] employed ensemble graph neural networks (EGNN) for the detection of fake news with integrated text features, but still used graph-based relations to achieve similar, if not better results. However, with our proposed hybrid ensemble, we achieved better accuracy, F1-score, and AUC. The result shows the effectiveness of the models where convolution of features, contextual modeling of BiLSTM, attention mechanism and TF-IDF were employed in creating an ensemble model that performed well.
6. Conclusions
In this paper, we proposed a hybrid framework for fake news detection, incorporating deep learning and classical machine learning. We have shown how both these approaches can work together and improve predictive performance. Using CNN for feature extraction and bidirectional LSTM (BiLSTM) for long-dependency feature learning, together with an attention layer for the most relevant parts of the text, provided deep learning with better capability of learning the language and the context of the misinformation and, hence, better capability of detecting it. In the meantime, the TF-IDF-based Support Vector Machine, Random Forest, and Logistic Regression were able to offer a good, firm decision boundary and added useful robustness to the model due to their ability to handle sparse feature sets. The ensemble models achieved good performance across all evaluated metrics (accuracy 0.96, F1 0.945, and AUC 0.99). The findings substantiate the claim that heterogeneous models outperform homogenous models due to their ability to capture composite misinformation at the stylistic, semantic, and structural levels. The validated analytical metrics in combination with the confusion matrix further established that the hybrid systems differentiate more than the individual models, thus improving the overall system performance. The proposed pipeline is a practical, scalable, and adaptable solution for fake news detection overall. Thanks to a modular architecture, it can be flexibly augmented to accommodate more datasets, features, or neural models, including new transform-based models and multimodal systems. This work relied on a single train–test split and dataset heterogeneity. Future work could be more complex, incorporating user behavioral patterns, network propagation signals, or cross-lingual representations. It could also analyze domain adaptation for rapidly changing misinformation campaigns.
Author Contributions
Conceptualization, M.I. and R.E.; methodology, M.I. and R.E.; software, M.I. and R.E.; validation, M.I. and R.E.; formal analysis, M.I. and R.E.; investigation, M.I. and R.E.; resources, M.I. and R.E.; data curation, M.I. and R.E.; writing—original draft preparation, M.I. and R.E.; writing—review and editing, M.I. and R.E.; visualization, M.I. and R.E.; supervision, M.I.; project administration, M.I.; funding acquisition, M.I. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Data Availability Statement
The raw data supporting the conclusions of this article will be made available by the authors on request.
Conflicts of Interest
The authors declare no conflicts of interest.
References
- D’Ulizia, A.; Caschera, M.C.; Ferri, F.; Grifoni, P. Fake news detection: A survey of evaluation datasets. PeerJ Comput. Sci. 2021, 7, e518. [Google Scholar] [CrossRef] [Scilit]
- Bondielli, A.; Marcelloni, F. A survey on fake news and rumour detection techniques. Inf. Sci. 2019, 497, 38–55. [Google Scholar] [CrossRef] [Scilit]
- Dwivedi, S.M.; Wankhade, S.B. Survey on fake news detection techniques. In International Conference on Image Processing and Capsule Networks; Springer International Publishing: Cham, Switzerland, 2020; pp. 342–348. [Google Scholar]
- Alnabhan, M.Q.; Branco, P. Fake news detection using deep learning: A systematic literature review. IEEE Access 2024, 12, 114435–114459. [Google Scholar] [CrossRef] [Scilit]
- Zhu, Y.; Wang, G.; Li, S.; Huang, X. A novel rumor detection method based on non-consecutive semantic features and comment stance. IEEE Access 2023, 11, 58016–58024. [Google Scholar] [CrossRef] [Scilit]
- Wang, K.; Ding, Y.; Han, S.C. Graph neural networks for text classification: A survey. Artif. Intell. Rev. 2024, 57, 190. [Google Scholar] [CrossRef] [Scilit]
- Mahabub, A. A robust technique of fake news detection using Ensemble Voting Classifier and comparison with other classifiers. SN Appl. Sci. 2020, 2, 525. [Google Scholar] [CrossRef] [Scilit]
- Saikh, T.; De, A.; Ekbal, A.; Bhattacharyya, P. A deep learning approach for automatic detection of fake news. arXiv 2020, arXiv:2005.04938. [Google Scholar] [CrossRef] [Scilit]
- Borges, L.; Martins, B.; Calado, P. Combining similarity features and deep representation learning for stance detection in the context of checking fake news. J. Data Inf. Qual. (JDIQ) 2019, 11, 1–26. [Google Scholar] [CrossRef] [Scilit]
- Mansouri, R.; Naderan-Tahan, M.; Rashti, M.J. A semi-supervised learning method for fake news detection in social media. In 2020 28th Iranian Conference on Electrical Engineering (ICEE); IEEE: New York, NY, USA, 2020; pp. 1–5. [Google Scholar]
- Dong, X.; Victor, U.; Qian, L. Two-path deep semisupervised learning for timely fake news detection. IEEE Trans. Comput. Soc. Syst. 2020, 7, 1386–1398. [Google Scholar] [CrossRef] [Scilit]
- Gaurav, A.; Gupta, B.B.; Castiglione, A.; Psannis, K.; Choi, C. A novel approach for fake news detection in vehicular ad-hoc network (VANET). In International Conference on Computational Data and Social Networks; Springer International Publishing: Cham, Switzerland, 2020; pp. 386–397. [Google Scholar]
- Zhou, X.; Jain, A.; Phoha, V.V.; Zafarani, R. Fake news early detection: A theory-driven model. Digit. Threat. Res. Pract. 2020, 1, 1–25. [Google Scholar] [CrossRef] [Scilit]
- Roy, A.; Basak, K.; Ekbal, A.; Bhattacharyya, P. A deep ensemble framework for fake news detection and classification. arXiv 2018, arXiv:1811.04670. [Google Scholar] [CrossRef] [Scilit]
- Ajao, O.; Bhowmik, D.; Zargari, S. Fake news identification on twitter with hybrid cnn and rnn models. In Proceedings of the 9th International Conference on Social Media and Society; Association for Computing Machinery: New York, NY, USA, 2018; pp. 226–230. [Google Scholar]
- Sheng, H.; Wang, X.; Zhang, C.; Wang, J.; Duan, P.; Wang, Y. AIGC video detection based on the fusion of spatial-frequency-optical flow multimodal features. J. Syst. Eng. Electron. 2026, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Liu, F.; Zheng, Q.; Tian, X.; Shu, F.; Jiang, W.; Wang, M.; Elhanashi, A.; Saponara, S. Rethinking the multi-scale feature hierarchy in object detection transformer (DETR). Appl. Soft Comput. 2025, 175, 113081. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.Y. “Liar, liar pants on fire”: A new benchmark dataset for fake news detection. arXiv 2017, arXiv:1705.00648. [Google Scholar] [CrossRef] [Scilit]
- Khan, J.Y.; Khondaker, M.T.I.; Afroz, S.; Uddin, G.; Iqbal, A. A benchmark study of machine learning models for online fake news detection. Mach. Learn. Appl. 2021, 4, 100032. [Google Scholar] [CrossRef] [Scilit]
- Ramos, J. Using tf-idf to determine word relevance in document queries. In Proceedings of the First Instructional Conference on Machine Learning, Washington, DC, USA, 21–24 August 2003; pp. 29–48. [Google Scholar]
- Salton, G.; Buckley, C. Term-weighting approaches in automatic text retrieval. Inf. Process. Manag. 1988, 24, 513–523. [Google Scholar] [CrossRef] [Scilit]
- Mikolov, T.; Sutskever, I.; Chen, K.; Corrado, G.S.; Dean, J. Distributed representations of words and phrases and their compositionality. In Advances in Neural Information Processing Systems; NeurIPS: San Diego, CA, USA, 2013; Volume 26. [Google Scholar]
- Zhang, Y.; Wallace, B.C. A sensitivity analysis of (and practitioners’ guide to) convolutional neural networks for sentence classification. In Proceedings of the Eighth International Joint Conference on Natural Language Processing, Taipei, Taiwan, 27 November–1 December 2017; pp. 253–263. [Google Scholar]
- Boureau, Y.L.; Ponce, J.; LeCun, Y. A theoretical analysis of feature pooling in visual recognition. In Proceedings of the 27th International Conference on Machine Learning (ICML-10); Omnipress: Madison, WI, USA, 2010; pp. 111–118. [Google Scholar]
- Hinton, G.E.; Srivastava, N.; Krizhevsky, A.; Sutskever, I.; Salakhutdinov, R.R. Improving neural networks by preventing co-adaptation of feature detectors. arXiv 2012, arXiv:1207.0580. [Google Scholar] [CrossRef] [Scilit]
- Rumelhart, D.E.; Hinton, G.E.; Williams, R.J. Learning representations by back-propagating errors. Nature 1986, 323, 533–536. [Google Scholar] [CrossRef] [Scilit]
- Hochreiter, S.; Schmidhuber, J. Long short-term memory. Neural Comput. 1997, 9, 1735–1780. [Google Scholar] [CrossRef] [Scilit]
- Kim, Y. Convolutional neural networks for sentence classification. arXiv 2014, arXiv:1408.5882. [Google Scholar] [CrossRef] [Scilit]
- Trueman, T.E.; Kumar, A. Attention-based C-BiLSTM for fake news detection. Appl. Soft Comput. 2021, 110, 107600. [Google Scholar] [CrossRef] [Scilit]
- Yang, Z.; Yang, D.; Dyer, C.; He, X.; Smola, A.; Hovy, E. Hierarchical attention networks for document classification. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Association for Computational Linguistics: Stroudsburg, PA, USA, 2016; pp. 1480–1489. [Google Scholar]
- Sokolova, M.; Lapalme, G. A systematic analysis of performance measures for classification tasks. Inf. Process. Manag. 2009, 45, 427–437. [Google Scholar] [CrossRef] [Scilit]
- Ying, L.; Yu, H.; Wang, J.; Ji, Y.; Qian, S. Multi-level multi-modal cross-attention network for fake news detection. IEEE Access 2021, 9, 132363–132373. [Google Scholar] [CrossRef] [Scilit]
- Raza, N.; Abdulkadir, S.J.; Abid, Y.A.; Albouq, S.S.; Alwadain, A.; Rehman, A.U.; Sumiea, E.H.; Farhan, M. Enhancing fake news detection with transformer-based deep learning: A multidisciplinary approach. PLoS ONE 2025, 20, e0330954. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Malik, A.; Behera, D.K.; Hota, J.; Swain, A.R. Ensemble graph neural networks for fake news detection using user engagement and text features. Results Eng. 2024, 24, 103081. [Google Scholar] [CrossRef] [Scilit]
| Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |