Next Article in Journal
Short-Term Wind Power Forecasting via Multimodal Adaptive Graph Neural Networks with Credibility-Modulated Aggregation
Previous Article in Journal
Constructing MIDA5: A Design Science Approach for a User-Centered Data Analytics Methodology for Business Process Improvement
Previous Article in Special Issue
EC-MFR: A Hierarchical Edge–Cloud Collaborative Framework for Multimodal Fact-Checking
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Comparative Performance of AI-Generated Fake News Detection Pipelines on Romanian News Content

by
Claudiu Coman
1,2,*,
Costel Marian Dalban
1,3,
Vlad Bătrânu-Pințea
1,
Georgiana Aron
1 and
Lucian Marina
4
1
Department of Social Sciences and Communication, Transilvania University of Brașov, 500036 Brașov, Romania
2
The Academy of Romanian Scientists Bucharest, 050044 Bucharest, Romania
3
Department of Sociology and Social Work, Alexandru Ioan Cuza University of Iași, 700506 Iași, Romania
4
Department of Social Sciences, “1 Decembrie 1918” University of Alba Iulia, 510009 Alba Iulia, Romania
*
Author to whom correspondence should be addressed.
Information 2026, 17(7), 698; https://doi.org/10.3390/info17070698
Submission received: 15 May 2026 / Revised: 10 July 2026 / Accepted: 16 July 2026 / Published: 18 July 2026

Abstract

Fake news detection has become a major research topic at the intersection of artificial intelligence, data mining, and information security. In this paper, we evaluate the performance of English-trained algorithms on English translations of Romanian-sourced news articles, using a translation-mediated cross-domain evaluation design. The study is based on source code generated with the assistance of artificial intelligence systems for a set of machine learning and transformer-based models. The code was subsequently implemented in Google Colab. 2026, trained on international benchmark datasets, and tested on Romanian news content. This design allowed the rapid prototyping of multiple detection pipelines and the systematic observation of their behavior in a media environment different from that represented in the training corpora. The models were evaluated comparatively using standard classification metrics, including accuracy, precision, recall, and F1-score, complemented by additional indicators relevant to model robustness and practical usability. The experimental results revealed significant differences in performance across algorithms when applied to English translations of Romanian-language news content after training on international datasets. However, this study does not provide a direct comparison between model performance on the international benchmark datasets and the Romanian test corpus; therefore, the gap between the international training corpus and the Romanian-sourced test corpus is interpreted as an exploratory limitation and as a direction for future research. Based on these findings, we propose an empirical classification of the tested models according to their predictive effectiveness, their contextual robustness across linguistic environments, and their operational relevance as filtering tools for institutional monitoring. The results show that AI-assisted coding workflows can provide a viable starting point for reproducible misinformation research, but they also underline the limitations of directly transferring models trained on non-Romanian data to local media ecosystems. The study offers both a replicable evaluation framework and practical insights for institutions involved in strategic communication, public security, and the monitoring of information threats.

Graphical Abstract

1. Introduction

The post-truth era has recently redefined the global information landscape, continuously transforming the public digital space from a place of democratic deliberation into a modern battlefield for operations whose ultimate goal is to exert influence. As information consumption slowly shifts toward social media platforms and feeds based on different algorithms, the scale, swiftness, and especially the diversity and sophistication of misinformation have outpaced traditional human-controlled verification methods. Fake news, defined as deliberate false information disguised and promoted as traditional journalism, is no longer merely a nuisance; it is a systemic threat to security, especially to the cognitive security of modern societies. The fundamental goal of this research is to bridge the gap between advanced computational linguistics and the specific needs of non-English media ecosystems. By proposing an AI-assisted framework, this study seeks to evaluate how detection models perform when directly applied to Romanian media and cultural contexts. This specific research is driven by the overall necessity to move beyond existing barriers and models towards more practical and refined solutions that can protect local information environments.
The accelerated digitalization of public communication is changing the way information is created, disseminated, and consumed by society. Social networks and online platforms have come to replace traditional mass media as the main sources of information, thereby creating an environment that enables the rapid spread of different types of content. This has transformed the online information environment into one that allows the proliferation of false and misleading content, a phenomenon whose consequences extend beyond individual-level effects, threatening public institutions or even social cohesion. Although, on average, people consider fake news less plausible than true news [1], misinformation proliferates where truth loses its value and where public opinion is shaped by political beliefs and emotions rather than facts.

1.1. The Global Scale of Digital Misinformation

Globally, the World Economic Forum’s Global Risks Report ranked misinformation and disinformation as the most severe near-term risk facing humanity for two consecutive years in 2024 and 2025 [2,3]. The scale of this problem was also examined in research by scholars [4], who found that social media users who frequently share content become less sensitive to the truth value of what they share, while a 2025 national survey of social media users in Australia found that 75% of respondents had encountered some form of misinformation. These findings suggest that the spread of misinformation is a feature of life online, driven by routine and the reward architecture of platforms.
What has changed in recent years is not just the growing presence of misinformation online but the methods by which it has come to be produced, disseminated, and incorporated into articles that appear trustworthy at first glance—a practice increasingly adopted by online publications whenever possible. According to NewsGuard [5], by March 2026, more than 3000 news sites were identified that dealt solely with the dissemination of erroneous information generated with the help of artificial intelligence. These platforms are not isolated cases, as they appear in search engine results alongside credible sources; yet they plagiarize articles from top publications and modify them, distorting the factual accuracy of the events they report.

1.2. Vulnerabilities in the Romanian Media Ecosystem

At the national level, Romania also faces such problems in terms of the rapid spread of disinformation, both among the population and through fake news channels. According to Eurobarometer data from 2025, 55% of Romanians report being exposed “often” or “very often” to disinformation and fake news within a given week, well above the EU average of 36%. Another Eurobarometer survey from 2025 on the Digital Decade found that 78% of Romanian citizens consider it important for public authorities to become involved in combating the spread of fake news. These figures highlight Romania’s vulnerability to disinformation [6].
This unstable environment is very well-expressed by the case of the 2024 Romanian elections. According to the Reuters Institute’s Digital News Report from 2025 [7], an independent candidate managed to win through a campaign based on disinformation propaganda during the first round of the elections, relying solely on his online campaign. This candidate gained strategic traction across social platforms such as Facebook, YouTube, and TikTok with anti-NATO, anti-European Union, and pro-Russia messages. However, the elections were annulled after documents were made public by the intelligence services revealing over 25,000 fake TikTok accounts, more than 5000 Telegram channels, and documented links to state-affiliated networks [8]. This event highlighted the need for automatic detection tools capable of operating in the Romanian media environment. Therefore, this context underscores the purpose of the present research: the comparative evaluation of a set of classifier models based on machine learning and transformer architectures, selected to represent the main stages of algorithmic evolution in natural language processing (NLP), on Romanian news content, as well as to assess the transferability of systems trained on international reference datasets in a non-English linguistic environment.

2. Literature Review

2.1. Definition of Fake News

The definition of the term fake news cannot be reduced to a single sentence due to its complex nature. Conceptualizations of the term are generally organized around two criteria: factuality, the degree to which content corresponds to verifiable reality; and intentionality, the deliberate intention to deceive an audience [9]—these two dimensions generating the greatest disagreements [10]. A strict definition treats fake news exclusively as articles that are both intentionally and verifiably false [11], while a broader formulation describes them as fabricated information that mimics the form of legitimate journalism [12]. Both formulations, however, have been criticized as insufficient, as a definition focused exclusively on falsehood omits manipulation of content; therefore, the definition may be complemented with the idea of information that is not entirely false in reality, but only manipulated or taken out of context [13]. This imprecision in defining the concept has led to the proposal of the term “information disorder” as a much more precise analytical framework [14].
In order to correctly understand and classify a piece of information, it is necessary to analyze three fundamental elements: the degree of truth of the content, the intention behind its production, and the actor who produces it [15]. On the first axis, the American Psychological Association [16] distinguishes misinformation, i.e., false content spread without the intent to deceive, from disinformation, which is intentionally fabricated, while scholars [17] have expanded the taxonomy with malicious information (malinformation): authentic information used with malicious intentions. On the axis of intent, satirical content presents a low degree of literal truth but no intent to deceive [18], while the essential identifying characteristic of fake news is the deliberate presentation of false content as if it were true [19]. The actor dimension has become more complex with the advent of large language models (LLMs), which can generate high-quality deceptive content on a large scale with a more formal and syntactically coherent linguistic profile, complicating automatic detection [20]. Content designed to “closely resemble the truth” [21] may circumvent both human skepticism and algorithmic classification, while the structural diversity of actors, state-sponsored networks, clickbait operations, and malicious actors require differentiated detection strategies, which generalist classifiers rarely achieve in practice [21,22].
For the purposes of this study, fake news refers to textual content, including journalistic articles, digital publications, and news-platform posts, that is intentionally fabricated or manipulated, produced, and then disseminated online with the deliberate purpose of deceiving the public. This definition aligns with formulations emphasizing the high degree of intentionality of fake news as “incorrect information intentionally disseminated to serve certain interests, political or economic in nature” [23], while also acknowledging the broader framework of information disorder as its epistemic context.

2.2. Detecting Fake News and Switching to Automated Systems

Fake news is deliberately designed to mimic legitimate journalism, leading even educated readers and AI-based systems without human oversight to fail regularly in distinguishing fabricated content from authentic content [24]. The automation of detection is therefore not merely a matter of efficiency but an operational necessity.
Natural language processing (NLP), a subfield of artificial intelligence that translates plain text into structured data that machines can analyze and classify [25], provides the technical basis for this automation. Thus, it allows for the automatic classification of items into predefined categories strictly based on their content [26]. Numerous processing tasks are required for machines to decode and understand human language [27], as it is extremely complex and varied given that people express themselves in multiple ways [28]. In such situations, NLP tools can not only identify the contexts of discussions and emotions but also extract manipulative linguistic patterns and detect irony or alarmist tones, essential elements to differentiate an authentic news story from disinformation. Three paradigms organize NLP-based detection: content-based methods, social-context-based methods, and evidence-based methods [29]. The present study operates entirely within the content-based paradigm, the most manageable and reproducible approach as it relies solely on characteristics that can be extracted from the article itself [30].

2.2.1. Traditional Models: Naive Bayes and KNN

Early research has been based on identifying characteristics such as clickbait patterns, punctuation density, and polarity of feelings. Naive Bayes is commonly cited as a classic example in the literature [31], and “assumes that the presence (or absence) of a certain characteristic is not related to the presence (or absence) of any other characteristic” [32]. Empirically, this limitation is well-documented, as in a comparative study that evaluated five algorithms on balanced fake news datasets, Naive Bayes scored an F1-score of 0.88, the lowest of all the models tested, while KNN achieved 92% accuracy but remained sensitive to sparse, high-dimensional representations generated by large corpora of text, limiting its scalability [33].

2.2.2. Ensemble Models: Random Forest and XGBoost

To overcome the limitations of individual classifiers, ensemble methods combine multiple decision trees to improve robustness. Random Forest aggregates predictions from random subsets of data by majority vote [34], while XGBoost constructs trees sequentially, each correcting for residual errors of its predecessor [20]. To address this limitation, the present research design directly analyzes the structural stability of these classifiers when applied to a different linguistic environment. The performance gap compared to classical models is substantial and reproduced in a direct comparison on the ISOT dataset, where both XGBoost and Random Forest achieved 100% accuracy, surpassing even deep learning methods [35]. As demonstrated by Essa, Omar, and Alqahtani [36], who tested traditional machine learning models, including XGBoost, on multiple real-world fake news datasets with various embeddings, performance degrades substantially when the linguistic environment changes—a pattern directly anticipated in the interlingual design of the present study.

2.2.3. Advanced Models: BERT and GCN Text

Deep learning has introduced artificial neural network architectures capable of capturing the meaning of text rather than merely its statistical properties; these architectures “have gained significant traction due to their superior performance in capturing complex patterns” [37]. Within this generation, transformer-based architectures demonstrate high effectiveness in fake news detection tasks [38] due to their architecture that allows attention to be paid to the relationships between words. The most prominent representative, BERT [39], reads text in both directions simultaneously to understand contextual relationships: “it is distinguished by its ability to consider the context of a word, both to the left and right of it within the sentence, unlike previous models that analyzed text only in one direction” [32]. The performance advantage over ensemble-type methods is empirically established in cross-lingual settings; in a direct comparison of the models, XGBoost achieved an accuracy of 81.0%, while BERT reached 98.0% on the same dataset, indicating that BERT-type architectures consistently outperform classical and ensemble models in classification tasks [40]. TextGCN models the relationships between words and documents these in the form of graph structures, allowing evidence to be aggregated and structural patterns of disinformation to be detected from multiple sources [41].

2.3. Performance Evaluation: Confusion Matrix and Classification Metrics

The confusion matrix breaks down predictions into four categories: true positives, true negatives, false positives, and false negatives [42]. From this, four metrics are derived. Accuracy, which is “the percentage of instances that are correctly categorized” [43], is the most intuitive but also the most misleading indicator under conditions of class imbalance: a classifier that always predicts the majority class achieves high accuracy without detecting a single instance of fake news. Precision “measures how well the model can identify positive instances, reduce false positive outcomes” [43] and quantifies the cost of false alarms, which is critical in implementation contexts where incorrectly labeling legitimate journalism as fabricated entails legal and reputational consequences. Recall, which measures “how well a pattern can identify each positive occurrence, thereby minimizing false negative outcomes” [43], captures the opposite risk of failing to detect real misinformation, which is the error with the greatest public consequences.

2.4. Trends in Research

Given the multidimensional nature of fake news and the architectural evolution of NLP detection flows, it is essential to recognize that the cross-lingual transferability of these systems to a morphologically complex and resource-constrained environment such as Romania’s requires a rigorous approach.
Benchmarking of fake news detection algorithms plays a key role in understanding their detection mechanisms and performance. Consequently, studies to this end by Buzea, Trausan-Matu, and Rebedea [44] established comparative benchmarks for the automatic detection of fake news in the Romanian online press. Their work demonstrated that transformer architectures generally outperform traditional models but face difficulties with subtle misinformation and decontextualized content, findings that directly support the H1 and H4 hypotheses of the present study. Furthermore, Bucos and Drăgulescu [45] demonstrated that machine translation augmentation substantially improves detection performance in Romanian content, while Moisi et al. [18] compared classical machine learning algorithms with transformer models on Romanian news, confirming the superiority of advanced architectures and emphasizing the critical importance of preserving diacritics in preprocessing to prevent cross-lingual transfer from degrading model performance.
In the present paper, by training the models on established international datasets and testing them on a localized corpus of 100 Romanian news sources (50 authentic and 50 fabricated), the study does not measure performance under ideal conditions but the practical applicability of global detection tools within the Romanian media ecosystem. Evaluation by standard classification metrics, such as accuracy, precision, recall and F1-score, complemented by a qualitative analysis of confusion matrices, will allow the study to test the premise that advanced models capture contextual nuances superior to classical algorithms (H1), while also verifying whether traditional statistical classifiers can maintain competitive results when the boundary between authentic journalism and disinformation is based on distinct lexical frequencies rather than complex semantic structures (H3).

3. Methodology

The purpose of this research is to comparatively evaluate the performance of several automated fake news detection pipelines, developed with the support of artificial intelligence systems and applied to Romanian journalistic content. The study aims to analyze the extent to which algorithms trained on international datasets can be transferred and used effectively to classify news articles in the Romanian media space. Therefore, the research is not just limited to evaluating the performance of individual algorithms but also investigates a broader process: generating code with the help of artificial intelligence, implementing it in Google Colab, training models on international databases, and then testing them on a Romanian news corpus.
Specific objectives
The research specifically pursues the following objectives:
  • Development of an AI-assisted experimental framework for generating, deploying, and testing multiple fake news classification pipelines.
  • Implementation in Google Colab of classic machine learning models and advanced natural language processing models.
  • Training algorithms on international datasets used in disinformation detection research.
  • Testing models on a corpus of Romanian articles that is culturally and contextually different from the data used in the training stage.
  • Comparative evaluation of model performance through standard classification indicators: accuracy, precision, recall, and F1-score.
  • Analysis of the robustness of models in conditions where training and test data differ in geographic origin, journalistic convention, and political context, specifically by evaluating models trained on internationally sourced English datasets and on English translations of Romanian journalistic content.
  • Identification of the types of errors produced by each model by analyzing the real items classified as fake and fake items classified as real.
Research hypotheses
The research starts from the following hypotheses:
H1. 
Advanced natural language processing models, such as BERT and TextGCN, will outperform classical algorithms due to their ability to analyze contextual relationships between words.
H2. 
The performance of all models will be affected by the gap between the internationally sourced training data and the Romanian-sourced test content, as the training corpus does not fully reflect the thematic, political, and journalistic particularities of the Romanian media space, even when that content is evaluated in English translations.
H3. 
Classic machine learning models can achieve competitive results in situations where the difference between fake and real news is reflected mainly at the lexical level through the frequency and distribution of words.
H4. 
Most classification errors will occur in the case of articles in which legitimate journalistic language overlaps with the language of disinformation, especially in political, institutional, satirical, or alarmist texts.
Scope and methodological limitations
Although the study was designed to meet methodological and academic standards, the present study has several limitations that should be acknowledged. The first limitation concerns the gap between the training and test corpora. Although the Romanian test articles were translated into English prior to evaluation—meaning the models operated within the same language throughout—the internationally sourced training data do not reflect the journalistic style, political context, or thematic specificity of the Romanian media space. Additionally, the translation step itself may introduce semantic shifts that affect classification outcomes. The second limitation concerns the size of the Romanian test corpus. Although the set of 100 articles is balanced, it cannot cover the entire diversity of the Romanian media ecosystem. The third limitation concerns the dependence of the pipelines on AI-generated or AI-assisted code. Even though this code allows for rapid prototyping, it requires validation, adjustment, and methodological verification to avoid implementation or interpretation errors. Nevertheless, human validation remained essential throughout the use of AI-assisted forms. The fourth limitation is that the models evaluated do not verify the factual accuracy of the statements in the articles and only classify texts based on the patterns learned from the data. Therefore, an article can be classified as real, not because the information is verified, but because its language resembles that of authentic articles in the training data.
Research design and experimental framework
The study has an experimental, comparative, and mixed-methods design oriented toward evaluating the performance of several automated fake news detection pipelines. The models were developed with the support of the Google Gemini artificial intelligence systems, Gemini version 3.1 Pro, which assisted with the generation and adaptation of the source code. The resulting code was then implemented and run in Google Colab. 2026. This environment allowed rapid prototyping, standardization of experimental conditions, and comparison of algorithms under similar conditions. The comparative component consisted of testing several models, from classic machine learning algorithms to advanced transformer- and graph-based architectures, on the same corpus of Romanian articles. The mixed component combined quantitative evaluation, performed through metrics such as accuracy, precision, recall, and F1-score, with qualitative analysis of classification errors.
Six models were evaluated in the study that were selected to cover several algorithmic families:
  • Naive Bayes—A probabilistic model frequently used in the classification of texts, based on the distribution and frequency of terms.
  • K-Nearest Neighbors—KNN, a model based on the similarity between texts, which classifies documents according to their proximity to other labeled documents.
  • Random Forest—An assembly-type model, consisting of several decision trees, capable of combining multiple classification rules.
  • XGBoost—A boosting model based on successively built decision trees, with each stage seeking to correct the errors of the previous one.
  • BERT—A transformer model, capable of analyzing words according to the context in which they appear.
  • TextGCN—A model based on graphical networks, which analyzes the relationships between words and documents by means of a structure of nodes and connections.
Data used
The research used two levels of data: training data and test data. The training data consisted of international datasets used in research on the detection of fake news. Thus, three international databases were selected, totaling 195,701 news items, both real and fake. The sources used were as follows:
(a) WELFake Dataset—An aggregated dataset, constructed by combining four distinct news sources, containing articles labeled as either real or fake, and intended for textual classification tasks [46]. Available at: https://doi.org/10.5281/zenodo.4561253.
(b) ISOT Fake News Dataset—Developed by the University of Victoria, Canada, the dataset contains real articles taken from Reuters and fake articles identified by fact-checking organizations, and is frequently used as a reference in comparative auto-detection studies [47]. Available at: https://onlineacademiccommunity.uvic.ca/isot/2022/11/27/fake-news-detection-datasets/ (Accessed on 5 March 2026).
(c) Misinformation and Fake News Text Dataset 79K—A publicly available dataset on the Kaggle platform, comprising approximately 79,000 labeled text articles, which is used to train and test disinformation classification models [48]. Available at: https://www.kaggle.com/datasets/stevenpeutz/misinformation-fake-news-text-dataset-79k (Accessed on 6 March 2026).
To ensure optimal performance of the detection models, the three databases were subjected to a systematic preprocessing pipeline prior to training. The datasets were merged into a single centralized corpus, with article titles and full-body texts. A standardization and cleaning procedure was then applied, in which special characters, web links, and redundant whitespaces were removed. In the next step, deduplication and class balancing were performed; duplicate entries were identified and removed, and the corpus was balanced to an equal number of real and fake news articles, resulting in 124,580 unique articles. All preprocessing steps were implemented in Python and are provided in Appendix A.1—“Database processing code”. All three datasets are exclusively in English, and no translation was applied to the training corpus. The models were therefore trained entirely in English. Training and testing were performed on entirely separate datasets: the models were trained on the international corpus and tested on the independent Romanian corpus.
The test data were represented by a corpus of Romanian articles. The test set was balanced and included 100 items, of which 50 were labeled as false and 50 as true. The labels assigned to the fake articles in the corpus were drawn from fact-checking websites. Articles classified as authentic were selected from news sources with editorial credibility and were individually reviewed by the authors to verify their consistency with the factual context. Since the models were trained in English, it was necessary to translate the Romanian test texts into English prior to evaluation. The test articles were preprocessed using the same cleaning pipeline applied to the training corpus. The translation was performed using Google Translate, followed by a manual review by the authors to correct errors. The detailed description of the articles included in the test corpus, accompanied by the titles, sources and tags related to each article, is presented in Appendix A.2—“Training test database”.
The articles were selected from content published or circulated during the last three months prior to data collection and were included only if they were related to Romania or to topics directly relevant to the Romanian information space. Real articles were selected from Romanian mainstream online media sources, whereas fake articles were selected sequentially from fact-checking websites. Items that did not match the article-based design of the study, such as social media posts, video-only materials, or deepfake video cases, were excluded. The corpus was intended not to be statistically representative of the Romanian media ecosystem but to provide a pilot external test set for observing how models trained on international datasets behave when applied to Romanian-sourced news content.
Experimental procedure
The present study reports an external evaluation of the trained models on the Romanian test corpus. Although the models were trained using international benchmark datasets, the manuscript does not include a separate evaluation of model performance on those international datasets. Consequently, the study does not quantify a direct transfer-performance gap between international and Romanian data. Instead, it provides an exploratory assessment of how the tested pipelines behaved when applied to English translations of Romanian news content.
In the first stage, the models to be compared were selected. In the second stage, classification pipelines were generated or adapted with the support of artificial intelligence systems. The resulting code was organized and run in Google Colab, which provided approximately 12 GB of RAM and access to a shared T4 GPU. Standard CPU resources were used for classical and ensemble models, while the T4 GPU accelerator was used for BERT and TextGCN. The implementation was developed in Python 3.10, using the libraries: pandas, scikit-learn, XGBoost, transformers (Hugging Face), and torch (PyTorch). For BERT, the BERT-base-uncased variant was used, which is a model for English text. The complete hyperparameter configuration for all six models, including all relevant parameter values, is provided in Table A4.
Given the resource constraints of the cloud environment, the source code initially generated by the AI was reviewed and adapted before execution. The primary category of adaptations concerned memory management: for all classical and ensemble models (Naive Bayes, KNN, Random Forest, XGBoost), the maximum vocabulary size used by the TF-IDF vectorizer was explicitly limited, ranging from 10,000 to 50,000 of the most frequent terms, depending on the requirements of each algorithm. For transformer-based and neural models (BERT, TextGCN), additional constraints were applied to the tokenizer’s maximum sequence length and to the training batch size, with the number of training epochs reduced accordingly. For all models involving random processes, the random seed was set to 42. These adjustments were necessary to prevent session crashes due to RAM exhaustion and are provided in Table A3.
The code generated by the AI system served as a functional starting point for each pipeline. Where the generated code executed without errors on the target dataset and environment, it was accepted with only minor adaptations, specifically with file path references. When runtime errors occurred, most frequently due to memory exhaustion, the error output was diagnosed and the parameter was adjusted manually until successful execution was achieved. The prompt submitted to the AI for this stage was:
“Write complete Python code for Google Colab that: (1) mounts Google Drive; (2) loads a CSV training file with columns text and label; (3) cleans nulls from those columns; (4) vectorizes the text using TF-IDF; (5) creates and trains the classifier; (6) saves the trained model and vectorizer as .pkl files to Google Drive. Print a progress message before each step.”
In the second stage, the trained models were applied to the Romanian test corpus of 100 news items. For each item, the algorithm generated a binary prediction: “fake” or “real”. The prompt submitted to the AI system for this stage was:
“Write complete Python code for Google Colab that: (1) mounts Google Drive; (2) loads the saved model and vectorizer; (3) loads the CSV test file; (4) preprocesses the text using the same cleaning procedure applied during training; (5) generates a prediction for each article; (6) saves the results to a CSV file. Print a progress message before each step.”
In the third stage, the predictions were compared with the ground-truth labels. Based on this comparison, performance metrics were calculated and confusion matrices were constructed. Subsequently, the classification errors were qualitatively analyzed to identify the linguistic and thematic patterns that influenced the performance of the models. The prompt submitted to the AI system for this stage was:
“Write Python code that: (1) loads the ground truth file and the predictions file; (2) aligns the two columns for comparison; (3) computes a full classification report including accuracy, precision, recall, F1-score, macro-F1 and weighted-F1; (4) plots a confusion matrix heatmap with axis labels and a title specifying the model and the test set”.
The quantitative assessment was based on standard indicators used in the automatic classification:
  • Accuracy measures the total proportion of correct classifications.
  • Precision indicates the proportion of correct predictions out of the total predictions made for a given class.
  • Recall measures the model’s ability to identify all examples belonging to a class.
  • F1-score represents the harmonic mean between precision and recall, and is useful for assessing the balance between the two.
  • Macro-F1 computes the unweighted mean of per-class F1-scores, treating each class with equal importance.
  • Weighted-F1 computes the same mean weighted by the number of instances per class. Since the test set is perfectly balanced (50 fake and 50 real articles), both metrics produce identical values across all models.
In addition to these metrics, the analysis included the values in the confusion matrix: true positives, true negatives, false positives, and false negatives.
In this study, the term “AI-generated” refers strictly to the use of artificial intelligence tools as implementation support for drafting, adapting, and debugging the source code of experimental pipelines. It does not imply that the AI-assisted coding process itself is evaluated as an independent research variable, nor that the study compares AI-generated code with human-written code; rather, the scientific focus remains on the comparative performance of the resulting fake news detection models under the same experimental conditions.

4. Results

This section presents the results obtained by testing the six pipelines for the automatic detection of fake news on a balanced Romanian corpus, consisting of 100 articles: 50 articles labeled as fake, and 50 articles labeled as real. Results are reported using standard classifications of accuracy, precision, recall, and F1-score and by the distribution of errors in the confusion matrix. In order to avoid terminological ambiguities generated by the different use of TP, TN, FP, and FN indicators depending on which class is treated as positive, the confusion matrices are presented below in descriptive format and directly report the combinations of actual labels and model prediction.
The analysis pursues two complementary levels. The first level is comparative and highlights the relative performance of the six models. The second level is interpretative and follows the types of articles that have generated classification errors, especially situations in which the vocabulary of legitimate journalism overlaps with the vocabulary of disinformation, satire, political discourse, or institutional communication.

4.1. Overall Comparative Performance of Models

Table 1 summarizes the results obtained by the six models for the two classes analyzed, fake and real. Because the test set is balanced, accuracy provides a useful insight into overall performance; however, the interpretation of the results should be supplemented by an analysis of the recall and the F1-score for each class, as some models showed strong tendencies to favor one of the categories.
The results reported in this section refer exclusively to the Romanian test corpus. Therefore, the analysis should not be interpreted as a complete measurement of cross-lingual transfer performance but as an exploratory external test of models trained on international data and applied to English translations of Romanian-language news content.
Because the Romanian pilot test corpus contains only 100 articles, 95% confidence intervals were calculated for model accuracy using the Wilson score interval. Accuracy was treated as a binomial proportion, corresponding to the number of correct predictions out of the total number of tested articles. These intervals provide a simple estimate of uncertainty around the reported accuracy values and support a more cautious interpretation of the differences observed between models.
The results indicate important differences between the models. Random Forest achieved the highest accuracy at 77%, and the lowest total number of errors of 23. BERT and TextGCN performed similarly, with accuracies of 73% and 72%, respectively, while Naive Bayes achieved 71% accuracy. XGBoost and KNN performed significantly worse, with accuracies of 57% and 49%, respectively.
These results show that (Figure 1), under the conditions of the present experiment, advanced natural language processing models did not automatically surpass classical machine learning models. Although BERT and TextGCN have superior contextual or relational modeling capabilities, Random Forest produced the most balanced overall performance. This observation is important for the interpretation of cross-lingual transfer, because models trained on international datasets may lose their architectural advantage when applied to a Romanian corpus with different linguistic and thematic particularities.
The distribution of errors complements the picture provided by precision. Naive Bayes favored the real class, generating 65 real predictions and only 35 fake predictions, which explains the high number of fake articles classified as real. KNN had the opposite behavior, classifying 95 of the 100 articles as fake, which led to the almost complete non-recognition of authentic journalism. XGBoost also showed a sharp trend towards the fake class, with 81 fake predictions. Random Forest, BERT, and TextGCN had more balanced distributions, although all three retained a slight orientation toward identifying fake content.

4.2. Analysis of Results by Model

4.2.1. Naive Bayes

Naive Bayes achieved an accuracy of 71% (Table 2), misclassifying 29 of the 100 articles. Of the 50 fake items, the model correctly identified 28 and misclassified 22 as real. Of the 50 real items, the model correctly identified 43 and erroneously classified seven as fake.
The rating ratio indicates a high precision for the fake class of 0.80 but a lower recall of 0.56. This means that the model is relatively reliable when labeling an item as fake but fails to detect a considerable portion of the fake items in the test set. For the real class, the recall of 0.86 indicates a good ability to recognize authentic items, but the precision of 0.66 shows that a significant proportion of predictions labeled real include a substantial number of fake items.
The analysis of real articles classified as fake shows that Naive Bayes is sensitive to potentially sensationalist or controversial vocabulary. Articles about surveys, important financial gains, unusual events, or cultural and historical themes have been interpreted as having features associated with disinformation. This limitation can be explained by the way the model works: the text is treated as a collection of words, and the classification is based on lexical frequencies and probabilities, without analyzing context, journalistic intent, or source credibility.
In the case of fake articles classified as real, the main problem was the use of common or seemingly legitimate vocabulary. Satirical texts, articles built around distorted real events, and materials using the names of institutions, such as the Romanian Government or NATO, were harder to distinguish from authentic journalism. Technical and legislative vocabulary, such as ‘modular reactors’, ‘engineering and design studies’, ‘draft law’, or ‘decisive vote of MPs’, contributed to this confusion, as the model cannot verify whether the terms are used in a factual or manipulative context.
Therefore, Naive Bayes performs relatively well in situations in which the lexical differences between fake and real articles are clear, but it becomes vulnerable when the two categories use similar vocabulary. The model is too permissive with some of the fake content, as 44% of the fake items in the test set were classified as real.

4.2.2. K-Nearest Neighbors (KNN)

KNN achieved the lowest performance among the models analyzed, with an accuracy of 49% and 51 classification errors. The model correctly identified 47 of the 50 fake items, but it misclassified 48 of the 50 real items. Thus, although the recall for the fake class is very high, 0.94, the recall for the real class is only 0.04, which indicates an almost complete inability to recognize authentic journalism.
The model generated 95 fake predictions and only five real predictions. This distribution indicates an extreme tendency to associate almost any item with the fake class. Methodologically, the result suggests an incompatibility between the structure of the KNN algorithm and the classification task applied in this cross-lingual context. The model does not build its own interpretative rules but classifies the texts according to their proximity to the previously observed examples in a multidimensional space of lexical characteristics.
The error of classifying real articles as fake (Table 3) can be explained by the overlap in high-frequency vocabulary between the two classes. Terms such as “Romania”, “energy”, “government”, “crisis”, “prices”, “elections”, and “people” appear in both real and fake news. In such a lexical space, KNN fails to distinguish between the classes, especially if one of the groups is denser or better represented in the training data. Articles that were atypical, short, or focused on unusual events amplified this problem, as their specific terms could not be associated clearly enough with relevant neighbors.
The three fake articles classified as real share vocabulary that resembles legitimate political analysis. The presence of proper names, quotes, geopolitical references, and terms such as “hypothesis”, “economy”, “press”, or “geopolitics” have positioned these texts closer to the cluster of real news than that of conspiracy content. The result confirms that KNN did not learn an operational definition of disinformation but instead reacted to lexical proximities.
Overall, KNN has little relevance to the task under review. Its performance reflects not only low precision but also a severe imbalance of predictions, which makes it unsuitable for operational uses in detecting fake news.

4.2.3. Random Forest

Random Forest achieved the best overall performance in the experiment, with an accuracy of 77% and 23 errors. The model correctly identified 42 of the 50 fake items and 35 of the 50 real items. It produced eight fake → real errors and 15 real → fake errors, indicating a relatively balanced distribution compared to KNN and XGBoost.
The ranking report shows an F1-score of 0.79 for the fake class and 0.75 for the real class. The precision for the real class is the highest of all models with 0.81, which indicates that when Random Forest assigns the real label, the prediction is generally reliable. At the same time, the 0.84 recall for the fake class shows a good ability to detect fake content.
Some of the real articles misclassified by Random Forest coincide with those misclassified by Naive Bayes, especially articles about surveys, unusual events, situations with sensationalist features, or cultural–historical themes. This overlap suggests the existence of difficult lexical areas for several types of models and not just for a single algorithm. In the case of Random Forest, the difficulty arises when the decision trees encounter combinations of characteristics associated in the training data with disinformation, even though the analyzed text is authentic journalism.
Three confounding factors were particularly noticeable (Table 4). The first is the length of the article. Very short texts offer few features for stable splits, whereas very long texts increase the likelihood of signal words associated with fake news. The second factor is represented by cultural or historical topics, where the vocabulary about legends, curses, spirits, deposits, or narrative stories can resemble the rhetoric of conspiracy texts. The third factor is the articles on regulations and legal provisions, where the institutional and procedural vocabulary can be confused with the language of theories about abuse of power or social control.
In the case of fake articles classified as real, Random Forest was especially vulnerable to misinformation built on seemingly plausible legislative, political, or institutional events. The names of public figures, references to institutions, geopolitical mentions, and narrative structures close to political journalism created a lexical profile that the model associated with real articles. The manipulation takes place here at the semantic and pragmatic level, not just at the lexical level, and the model cannot factually verify the statements.
Despite these limitations, Random Forest offered the best balance between detecting fake content and recognizing authentic journalism. Compared to the other models, its result suggests that a classical assembly-type architecture can remain competitive under cross-lingual conditions, especially when the classification task depends on lexical distributions and recurrent combinations of features.

4.2.4. XGBoost

XGBoost achieved 57% accuracy, with 43 articles misclassified. The model correctly identified 44 of the 50 fake items, but it correctly classified only 13 of the 50 real items. It generated 81 fake predictions and only 19 real predictions, which indicates a strong trend of overlabeling articles as fake.
This distribution is reflected in the high recall for the fake class with 0.88 and the very low recall for the real class with 0.26. The precision for the fake class is 0.54, which means that almost half of the fake predictions are actually real articles mislabeled. For the real class, the precision of 0.68 suggests that the model may be relatively correct when assigning this label, but it does so too rarely to be considered balanced.
The 37 real items classified as fake indicate a major transfer and calibration problem. Among these errors are articles that have caused difficulties for the other models as well: surveys, articles with analytical language, materials with cultural or historical topics, texts with large volumes, or formulations with sensationalist potential. This suggests that XGBoost has internalized certain patterns associated with disinformation too strongly and applied them excessively to the Romanian corpus.
Other problematic cases involved articles (Table 5) with metaphors, legislative language, or alarmist tones, including expressions such as “radical law”, “phasing out”, “ban”, “crisis”, or “chaos”. In these situations, the model interpreted the political, economic, and institutional vocabulary as a signal of disinformation, even though some texts were real articles.
The fake articles classified as real used either names, institutions, too apparently precise procedures, or technical and journalistic language that masked the fabricated nature of the information. In addition, the model had difficulty identifying subtle satire, as phrases such as ‘prime minister’ or ‘hunger strike’ can function as legitimate signals if analyzed without access to the pragmatic level of the text.
XGBoost’s performance is mostly relevant as a negative result. Although the model generally performs well in many classification tasks, it did not bring an advantage over Random Forest in this experiment. The result suggests either insufficient parameter calibration or high sensitivity to the distribution of the drive data, which led to an excessive orientation towards the fake class. In the form tested, XGBoost is less useful in an operational context in which it is also necessary to protect authentic journalism from misclassification.

4.2.5. BERT

BERT achieved 73% accuracy with 27 classification errors. The model correctly identified 41 of the 50 fake items and 32 of the 50 real items. It generated 59 fake predictions and 41 real predictions, having a relatively balanced distribution compared to KNN and XGBoost.
The classification report indicates an F1-score of 0.75 for the fake class and 0.70 for the real class. The 0.82 recall for the fake class shows a good ability to identify fake items, while the 0.78 precision for the real class indicates high reliability when the model classifies an item as genuine. However, the model misclassified 18 real items as fake and nine fake items as real. In the case of real articles classified as fake, two types of limitations were apparent. The first concerns articles built around polls, statistics, or public opinion topics. For example, the article about the INSCOP survey on giving up the time change contains percentages, divided opinions, and a topic with political potential, elements that can appear both in authentic journalism and in manipulative content. The model cannot verify the legitimacy of the source or the validity of the cited data.
The second limitation concerns the length of the texts (Table 6). BERT typically processes a maximum of 512 tokens, which means that long items can be classified based on the first part of the text. In the case of the article about “gold fever” in Romania, the initial part is more narrative and closer to the folkloric register, while the analytical elements that support the journalistic character appear later. This structure can lead to misclassification, as the model does not have effective access to the entire discursive architecture of the article.
Fake articles classified as real highlight a major difficulty of the model in relation to satire, irony, and absurdity. Five of these articles are satirical or strongly political in nature. BERT can correctly interpret the local grammatical and semantic coherence of sentences, but it does not have a sufficient mechanism to assess the plausibility of the situations described. Examples such as biodegradable microphones planted in homes, vouchers worth thousands of lei for cherries, or influencing corn prices through the consumption of popcorn by politicians require real-world knowledge and a pragmatic interpretation of ironic intention.
Compared to models based strictly on frequencies or decision trees, BERT benefits from more advanced contextual analysis. However, its performance remains dependent on the diversity of training data and the adequate representation of local forms of satire, political opinion, and disinformation. The results indicate a relatively stable model, but not one robust enough to outperform Random Forest under the conditions of this experiment.

4.2.6. TextGCN

TextGCN achieved 72% accuracy, with 28 articles misclassified. The model correctly identified 42 of the 50 fake items and 30 of the 50 real items. It generated 62 fake predictions and 38 real predictions, which indicates a moderate trend of favoring the fake class.
The performance indicators show an F1-score of 0.75 for the fake class and 0.68 for the real class. The 0.84 recall for the fake class indicates good misinformation detection capability, but the 0.60 recall for the real class shows that 40% of authentic items were misclassified as fake. The precision for the real class of 0.79 is still high, which means that real predictions are relatively reliable, but they occur less often than would be necessary in a balanced system.
The main explanation for the errors comes from the architecture of the model. TextGCN constructs a graph in which documents and words are represented as nodes, and the links reflect co-occurrence or association relationships. The classification decision depends not only on the internal vocabulary of an article but also on how its terms are connected to other nodes in the network. Therefore, a real article that uses terms commonly found in disinformation may receive negative signals through the propagation of the information within the graph.
This logic explains the relatively high number of real articles classified as fake. Terms such as “crisis”, “European”, “Romania”, “war”, “government”, or “energy” can appear in both categories. If these words are strongly connected to fake articles, they can transfer a negative lexical signal to the real articles. Thus, the advantage of the model, its ability to capitalize on the relationships between documents and words also becomes, in this context, its main vulnerability.
In the case of fake articles classified as real, TextGCN struggled with satire and disinformation built on real political events. Absurd combinations, such as “biodegradable microphones”, very large vouchers for cherries, or global corn prices allegedly being influenced by popcorn orders, cannot be evaluated by simply analyzing the relationships between the nodes. Each word may have a coherent position in the graph, but the pragmatic combination of terms remains absurd. Detecting these nuances requires a real-world plausibility assessment that the TextGCN architecture does not provide.
The result of TextGCN is close to that of BERT, but the causes of errors are different. If BERT is limited by incomplete contextual processing of long texts and the difficulty of interpreting satire, TextGCN is mostly affected by lexical contamination by propagation of signals in the graph. This is relevant for the use of the model in media ecosystems where the same political or institutional terms circulate in both legitimate journalism and disinformation.

4.3. Comparative Synthesis and Typology of Errors

The comparative analysis shows that the results do not fully confirm the expectation that advanced deep learning models would automatically outperform classical models. Random Forest achieved the best overall performance, followed by BERT, TextGCN, and Naive Bayes. XGBoost and KNN produced unbalanced results, especially by over labeling articles as fake.
The results allow the identification of five recurrent types of errors. The first type consists of real articles with sensationalist or alarmist vocabulary. These include news about unusual events, very large values, controversial surveys, economic crises, or local situations with an emotional impact. Models tend to associate these elements with disinformation even though they can also appear in authentic journalism.
The second type also involves authentic articles misclassified as fake, but the triggering mechanism is different. These articles cover political events, electoral dynamics, party disputes, coalitions, or geopolitical affairs using the vocabulary of political journalism: party names, names of public figures, references to institutional crises, elections, or governmental decisions. This register overlaps substantially with the vocabulary of political disinformation, making the two categories difficult to distinguish.
The third type of error is represented by fake articles that use institutional, political, or technical vocabulary. References to governments, NATO, legislative procedures, energy infrastructure, elections, or public policies increased the likelihood that some false texts will be interpreted as real. This category is particularly important, because contemporary disinformation does not always appear in the form of obviously absurd texts and can mimic a legitimate journalistic style.
The fourth type consists of fake articles. To the human reader, the satirical character may be obvious, but models have difficulty interpreting non-literal intent. BERT and TextGCN, although more architecturally advanced, classified several satirical articles as real because they were grammatically or lexically coherent.
The fifth type of error is related to the effects of cross-lingual transfer. The models were trained on international datasets and tested on Romanian articles, which implies differences in language, journalistic style, political themes, and cultural registers. The results suggest that these differences affect both classical and advanced models but through different mechanisms: some models overreact to lexical frequencies, whereas others to graph structures or limited contextual fragments.
Table 7 reports, for each of the six models, the total number of misclassifications and the number of errors falling into each of the four categories, which are strictly related to the nature of the articles. Categories “1. Sensationalist vocabulary in authentic journalism” and “2. Political register overlap in authentic journalism” represents real news predicted as fake. Categories “3. Institutional vocabulary in fabricated content” and “4. Failure to identify satirical content” represent fake news predicted as real. The fifth type of error, related to cross-cultural transfer, is not represented as a separate column in the table because it constitutes a background condition rather than a classifiable feature of the articles themselves.
The four categories presented reflect the editorial classification that already existed in the source publications and fact-checking platforms from which the corpus was assembled.
For authentic news articles that were misclassified as fake, the categorization follows the topic of the article, as can be inferred directly from the title and source. Articles covering political events, electoral dynamics, party disputes, or institutional decisions were distinguished from articles covering non-political topics with sensationalist vocabulary. This distinction reflects the observable thematic content of each article.
The qualitative explanations of the misclassifications are interpretative and based on the observed linguistic and thematic characteristics of the misclassified articles. They should not be understood as definitive causal proof but as plausible error patterns that require further validation on a larger corpus and through additional statistical testing.

4.4. Practical Implications and Cautious Interpretation

In the context of this study, operational relevance is evaluated by analyzing how each model’s distribution of prediction errors—including false positives and false negatives—affects its practical utility for institutional deployment. From the perspective of the present exploratory experiment, Random Forest obtained the highest observed performance and the most balanced distribution of errors among the tested models. However, this result should not be interpreted as a recommendation for operational deployment given the limited size and non-representative nature of the Romanian test corpus.
BERT and TextGCN are promising in their ability to capture contextual or relational information, but the results show that these advantages are not sufficient in the absence of a better adaptation to the Romanian journalistic conventions, contextual political language, and local specificity of disinformation.
Naive Bayes may function as a baseline or reference model, but its major risk lies in classifying a large number of fake articles as real. KNN and XGBoost, in the tested form, are less useful for practical applications because they exhibit severe prediction imbalances. KNN classifies almost all articles as fake, while XGBoost excessively labels many real articles as fake.
For institutions involved in information threat monitoring, strategic communication, or public security, these findings should be interpreted only as preliminary evidence of the behavior of the tested models. The results suggest that automatic classifiers may be useful as experimental screening tools, but they cannot be considered reliable operational systems without validation on larger, more diverse, and representative Romanian news datasets.
Furthermore, when these operational insights are contextualized alongside existing Romanian fake news detection benchmarks, the implications of the empirical results become clearer. While Buzea et al. [44] demonstrated high performance using deep learning architectures, and Bucos and Drăgulescu [45] highlighted the necessity of data augmentation for transformer models handling smaller Romanian datasets, our findings introduce an important operational counterpoint. The findings confirm their baseline observations that transformer architectures (such as BERT) encounter significant bottlenecks when confronted with specific localized Romanian disinformation without extensive local fine-tuning. However, by demonstrating that traditional ensemble methods like Random Forest can achieve the highest observed predictive stability on a localized corpus without requiring complex data-augmentation loops, this study complements prior benchmarks, shifting the discussion from raw architectural optimization toward the practical feasibility of efficient institutional deployment.

5. Conclusions

Section 4 shows that the performance of fake news detection models depends not only on algorithmic complexity but also on the compatibility between the training data, the test corpus, and the type of content analyzed. Random Forest obtained the highest observed performance within the limited Romanian test corpus used in this study, but this result should be interpreted descriptively and cautiously and not as evidence of operational robustness.
Naive Bayes demonstrated moderate performance but with significant vulnerabilities in detecting sophisticated fake items. KNN and XGBoost highlighted important limits of calibration and transfer.
A comparison of the results with the research hypotheses shows that the hypotheses are confirmed to different degrees. H1 is only partially confirmed, as the advanced BERT and TextGCN models achieved competitive results but did not outperform Random Forest, which had the best overall performance. H2 is confirmed, as all models were affected by the gap between the internationally sourced training data and the Romanian-sourced test corpus, with performance limited by differences in journalistic convention, political context, and topic distribution rather than language itself. H3 is confirmed because classical models, especially Random Forest and Naive Bayes, achieved competitive results, demonstrating that lexical distributions can remain relevant in the classification of fake news. However, this confirmation needs to be qualified, given the poor performance of KNN and XGBoost. H4 is confirmed because the analysis of errors has shown that the main difficulties arise when legitimate journalistic language overlaps with the language of disinformation, especially in political, institutional, satirical, or alarmist articles.
Overall, the results support the idea that automatic detection of disinformation in the Romanian media space requires locally adapted models, representative training sets, and assessments that combine quantitative indicators with qualitative error analysis. Without these elements, models risk confusing authentic journalism with disinformation or allowing false articles to pass that effectively mimic the style of legitimate media.
In this context, a relevant direction for future research aims to develop and validate a “gold standard” data universe on a larger scale, designed exclusively for the Romanian media ecosystem. The expansion of the volume of data would also allow for the training of LLMs (large language models) that are optimized for the needs and specificity of local language and culture. Also, the proper integration of new automated verification mechanisms is an essential step in improving the accuracy of detecting hard-to-detect disinformation, which is increasingly taking ultra-sophisticated forms. Through the integration of real-time fact-checking features, future iterations could serve as foundational components for assisting fact-checkers in analyzing text structures rather than just automated tools for blocking content, and could assess not only how a message is written but also whether or not the central claims in that message are supported by verified, truthful, and factual data. The continued development of abusive disinformation systems presents a real danger, which requires increased attention and adaptability to keep up and prevent new models used for misinformation, contributing to the broader academic development of automated media literacy frameworks.
A further limitation concerns the absence of a direct comparison between the models’ performance on the international benchmark datasets and their performance on the English translation of the Romanian test corpus. As a result, the study cannot precisely quantify the transfer-performance gap between international and Romanian data. Future research should address this limitation by reporting model performance on both international and Romanian datasets, comparing the results statistically, and evaluating whether performance decreases when models are transferred to the Romanian media context.
The Romanian test corpus includes only 100 articles and cannot be considered representative of the full Romanian media ecosystem. Therefore, the observed performance of Random Forest, although the highest among the tested models, does not justify recommending the model for real-world operational deployment. Any operational use would require validation on substantially larger, more diverse, and continuously updated Romanian news datasets, including different media sources, topics, genres, and forms of misinformation.

Practical Recommendations

Based on the exploratory results, automatic fake news detection models should be used only as preliminary screening tools and not as autonomous decision-making systems. Even the best-performing model in this study, Random Forest, produced errors, while BERT and TextGCN also showed limitations in cases involving satire, political language, institutional discourse, and alarmist vocabulary.
For public security, strategic communication, and institutional monitoring, these models may help prioritize content for human review, especially during elections, crises, or periods of intensified disinformation. However, any practical use should involve expert verification, contextual analysis, and fact-checking. Before institutional deployment, the models must be validated on larger, more diverse, and representative Romanian datasets.

Author Contributions

Conceptualization, C.C. and C.M.D.; methodology, C.C., C.M.D. and V.B.-P.; software, V.B.-P. and G.A.; validation, C.C., C.M.D., V.B.-P., G.A. and L.M.; formal analysis, C.C., C.M.D. and V.B.-P.; investigation, C.C., C.M.D., V.B.-P. and G.A.; resources, C.C. and L.M.; data curation, V.B.-P. and G.A.; writing—original draft preparation, C.C., C.M.D., V.B.-P. and G.A.; writing—review and editing, C.C., C.M.D. and L.M.; visualization, V.B.-P. and G.A.; supervision, C.C. and L.M.; project administration, C.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The original contributions presented in this study are included in the article. Further inquiries can be directed to the corresponding author.

Acknowledgments

During the preparation of this manuscript, the authors used Google Gemini version 3.1 Pro for the purposes of generating, adapting, and debugging source code for the experimental pipelines. The authors have reviewed, validated, and edited the generated output and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
AIArtificial Intelligence
APCArticle Processing Charge
BERTBidirectional Encoder Representations from Transformers
CNANational Audiovisual Council
EUEuropean Union
F1-scoreHarmonic Mean of Precision and Recall
FNFalse Negative
FPFalse Positive
GCNGraph Convolutional Network
GPUGraphics Processing Unit
ISOTInformation Security and Object Technology
KNNK-Nearest Neighbors
LLMLarge Language Model
Macro-F1Unweighted Mean of F1-Scores Calculated for Each Class
NATONorth Atlantic Treaty Organization
NLPNatural Language Processing
RAMRandom Access Memory
SMRSmall Modular Reactor
TF-IDFTerm Frequency–Inverse Document Frequency
TextGCNText Graph Convolutional Network
TNTrue Negative
TPTrue Positive
URLUniform Resource Locator
Weighted-F1F1-Score Averaged Across Classes According to Number of Instances in Each Class
XGBoostExtreme Gradient Boosting

Appendix A

Appendix A.1. Database Processing Code

import pandas as pd
import re
import string
from google.colab import drive
drive.mount(‘/content/drive’)
path = “..............”
print(“1. Loading the 3 raw databases…”)
df_isot_true = pd.read_csv(path + ‘True.csv’)
df_isot_fake = pd.read_csv(path + ‘Fake.csv’)
df_isot_true[‘label’] = 1
df_isot_fake[‘label’] = 0
df_welfake = pd.read_csv(path + ‘WELFake.csv’)
df_misinfo = pd.read_csv(path + ‘Misinfo.csv’)
print(“2. Merging all articles…”)
df_total = pd.concat([df_isot_true, df_isot_fake, df_welfake, df_misinfo], ignore_index = True)
print(“3. Preparing full text…”)
df_total[‘title’] = df_total[‘title’].fillna(“”)
df_total[‘text’] = df_total[‘text’].fillna(“”)
df_total[‘text_complet’] = df_total[‘title’].astype(str) + “ ” + df_total[‘text’].astype(str)
print(“4. Cleaning text for artificial intelligence”)
def clean_text(text):
text = str(text).lower()
text = re.sub(r’https?://\S+|www\.\S+’, ‘’, text)
text = re.sub(r’<.*?>+’, ‘’, text)
text = re.sub(r’[%s]’ % re.escape(string.punctuation), ‘’, text)
text = re.sub(r’\n’, ‘ ‘, text)
text = re.sub(r’\w*\d\w*’, ‘’, text)
text = re.sub(r’\s+’, ‘ ‘, text).strip()
return text
df_total[‘fulltext’] = df_total[‘fulltext’].apply(cleantext)
print(“5. Remove duplicates.”)
initial_count = len(df_total)
df_total = df_total.drop_duplicates(subset = [‘fulltext’])
print(f“-> Found and deleted {initial_count—len(df_total)} duplicate items!”)
print(“6. Balance the database”)
min_count = df_total[‘label’].value_counts().min()
df_reale = df_total[df_total[‘label’] == 1].sample(n = min_count, random_state = 42)
df_false = df_total[df_total[‘label’] == 0].sample(n = min_count, random_state = 42)
df_final = pd.concat([df_reale, df_false]).sample(frac = 1, random_state = 42).reset_index(drop = True)
df_final = df_final[[‘text_complete’, ‘label’]]
df_final.columns = [‘text’, ‘label’]
print(“\n--- INVENTORY ---”)
print(df_final[‘label’].value_counts())
print(f“Total items for training: {len(df_final)}”)

Appendix A.2. Training Test Database

The Romanian articles were translated into English by the authors, and the original texts were not used. The full texts and translations of the original articles are not reproduced in this Appendix to ensure compliance with applicable copyright norms. Instead, Table A1—Fake news articles and Table A2—Real news articles provide the English translation of each article’s title, their source, and a URL linking to the original source.
Table A1. Fake news articles.
Table A1. Fake news articles.
No. Article Title (English Translation)SourceURL
1The pro-WAR candidate. Nicușor Dan’s advisor says the war with Russia will start from Romanian territory. Tolontan: Nicușor will act like Zelenskyyr3media.rohttps://r3media.ro/pro-razboi-consilier-nicusor-dan-razboi-rusia-teritoriul-romaniei-aliatii-ne-au-desemnat-ca-stat-pe-linia-frontului-tolontan-zelenski/ (Accessed on 28 April 2026)
2The referendums in occupied Ukraine are legitimate, and Russia is militarily superior to the Westveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-referendumurile-din-ucraina-ocupata-sunt-legitime-iar-rusia-e-superioara-militar-occidentului (Accessed on 28 April 2026)
3European Union states will abandon Ukraine on the brink of the cold season, and secret services will orchestrate protests to justify thisveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-serviciile-secrete-vor-orchestra-proteste-pentru-a-justifica-abandonarea-ucrainei-de-catre-ue (Accessed on 28 April 2026)
4Romanians and Europe will freeze this winterveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/dezinformare-romanii-si-europa-vor-ingheta-la-iarna-din-cauza-gazingrad-ului (Accessed on 28 April 2026)
5The ‘world occult’ plans to implant chips in childrenveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-oculta-mondiala-planuieste-sa-implanteze-cipuri-copiilor (Accessed on 28 April 2026)
6Romania will receive many Africans and Asians who have taken refuge from Russia’s invasion of Ukraine, and authorities are keeping their nationality secretveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-romania-va-fi-invadata-de-africani-si-asiatici-veniti-din-ucraina (Accessed on 28 April 2026)
7The United States will test small modular reactors (SMRs) in Romania, causing environmental problems and nuclear accident riskveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-sua-fac-experimente-nucleare-periculoase-in-romania-2 (Accessed on 28 April 2026)
8Ukraine’s military intelligence tried, together with the United States, to trick Romania into sending special forces to Kherson to be attacked by the Russiansveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-ucraina-si-sua-au-incercat-sa-atraga-romania-in-razboi-cu-rusia (Accessed on 29 April 2026)
9Donald Trump confirms almost identically what Călin Georgescu was saying about green energy, wind turbines and fuel pricesortodoxinfo.rohttps://ortodoxinfo.ro/2026/01/21/video-donald-trump-confirma-aproape-la-indigo-ce-spunea-calin-georgescu/ (Accessed on 29 April 2026)
10Father Zenovie from Nechit Monastery: ‘People who get vaccinated will become zombies’antena3.rohttps://www.antena3.ro/actualitate/parintele-zenovie-manastirea-nechit-neam-oameni-vaccin-zombie-603969.html (Accessed on 29 April 2026)
11The anti-COVID vaccine causes aortic dissection and sudden deathveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-vaccinul-anti-covid-provoaca-disectie-de-aorta-si-moarte-subita (Accessed on 29 April 2026)
12Romania is invaded by Asiansveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-romania-este-invadata-de-asiatici (Accessed on 29 April 2026)
13Minister of Defense, harsh statements at the NATO-Industry Forum 2025: It is a matter of time until we see warmediaflux.rohttps://mediaflux.ro/ministrul-apararii-declaratii-dure-la-nato-industry-forum-2025-este-o-chestiune-de-timp-pana-cand-vom-vedea-razboi/ (Accessed on 29 April 2026)
14Nicușor Dan has decided, together with the governing coalition, to send the entire future gas production from the Black Sea to Ukraineveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-romania-va-da-ucrainei-gazele-extrase-din-marea-neagra (Accessed on 29 April 2026)
15Moldova’s rapprochement with the EU actually means its absorptionveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-integrarea-europeana-va-duce-la-disparitia-moldovei (Accessed on 29 April 2026)
16Blow for millions of Romanians. Wood heating and stoves could be bannednewsweek.rohttps://newsweek.ro/actualitate/lovitura-pentru-milioane-de-romani-incalzirea-cu-lemne-si-sobele-ar-putea-fi-interzise (Accessed on 29 April 2026)
17The cardboard civil society in Romaniaortodoxinfo.rohttps://ortodoxinfo.ro/2026/04/24/societatea-civila-de-carton-din-romania/ (Accessed on 29 April 2026)
18The satans from the CNA banned Realitatea Plus. Dastardly attack by the system on the only television where Călin Georgescu and the AUR party appearortodoxinfo.rohttps://ortodoxinfo.ro/2026/04/07/satanele-de-la-cna-au-interzis-realitatea-plus-atac-marsav-al-sistemului-la-singura-televiziune-la-care-apare-calin-georgescu-si-partidul-aur/ (Accessed on 29 April 2026)
19The President of Romania, Călin Georgescu, message of AWAKENING: ‘The curse will be broken. The hour has come!’ortodoxinfo.rohttps://ortodoxinfo.ro/2025/12/16/video-presedintele-romaniei-calin-georgescu-mesaj-de-trezire-blestemul-se-va-rupe-a-venit-ceasul/ (Accessed on 29 April 2026)
20The spring planting campaign begins. The SRI will plant 200,000 microphones in Romanians’ homestimesnewroman.rohttps://www.timesnewroman.ro/life-death/incepe-campania-de-plantare-de-primavara-sri-va-planta-200-000-de-microfoane-in-casele-romanilor/ (Accessed on 29 April 2026)
21Romania is getting ready to demolish dams, hallucinating decisions. Our rulers are not talking about the EU directivesstiridinbucovina.rohttps://stiridinbucovina.ro/romania-se-pregateste-sa-demoleze-baraje-decizii-halucinante-guvernantii-nostri-nu-vorbesc-despre-directivele-ue/ (Accessed on 29 April 2026)
22The socialist government of Spain goes even further with censorshipactivenews.rohttps://www.activenews.ro/stiri/Guvernul-socialist-din-Spania-merge-si-mai-departe-cu-cenzura.-Instrument-guvernamental-de-supraveghere-a-discursului-instigator-la-ura-de-pe-Telegram-X-Facebook-si-TikTok-203594 (Accessed on 29 April 2026)
23Right-hand drive cars will no longer be able to circulate on public roads, the senators votedavocatnet.rohttps://www.avocatnet.ro/articol_59658/Ma%C8%99inile-cu-volan-pe-dreapta-nu-vor-mai-putea-circula-pe-drumurile-publice-au-votat-senatorii.html (Accessed on 29 April 2026)
24CASH money is disappearing! The order came from the top of Europerealitatea.nethttps://www.realitatea.net/stiri/economie/dispar-banii-cash-ordinul-venit-de-la-varful-europei_680770ff2bceea38c4622613 (Accessed on 29 April 2026)
25The tax on the sun, a new villainy. And not the last onebursa.rohttps://www.bursa.ro/taxa-pe-soare-o-noua-ticalosie-nu-si-ultima-60344843 (Accessed on 29 April 2026)
26Bolojan sold the Tarnița hydroelectric plant to France. It is good for the French and it is not good for the Romaniansbzi.rohttps://www.bzi.ro/bolojan-a-vandut-hidrocentrala-tarnita-frantei-e-buna-pentru-francezi-si-nu-e-buna-pentru-romani-5478173 (Accessed on 30 April 2026)
27The worrying message sent by Mădălin Ionescu, just one day after he was replaced by Natalia Mateuț at Antena Starsradiofxnet.rohttps://radiofxnet.ro/stirilaminut/mesajul-ingrijorator-transmis-de-madalin-ionescu-la-doar-o-zi-dupa-ce-a-fost-inlocuit-de-natalia-mateut-la-antena-stars-nici-acum-nu-mi-am-revenit/?fbclid=IwY2xjawHo1oBleHRuA2FlbQIxMQABHYJyb1hj3Q8GJgf9Gwh_ZJyBPNvxqLO_t1J6x0lzc2N7LhRiRMfqYuUbCQ_aem_ZsbK3_YMEXZjqpoVGqlxfQ (Accessed on 30 April 2026)
28‘Blackout’ in Europe! Huge power outage. The Army has already intervened: Prepare yourselvescapital.rohttps://www.capital.ro/blackout-in-europa-pana-de-curent-uriasa-armata-a-intervenit-deja-pregatiti-va.html (Accessed on 30 April 2026)
29Romania, forced to organize the second round of the presidential election, according to the 2024 election calendarveridica.rohttps://www.veridica.ro/en/fake-news-disinformation-propaganda/fake-news-romania-forced-to-organize-the-second-round-of-the-presidential-election-according-to-the-2024-election-calendar (Accessed on 30 April 2026)
30The USA are dragging Romania into a war with Iranveridica.rohttps://www.veridica.ro/en/fake-news-disinformation-propaganda/fake-news-the-usa-are-dragging-romania-into-a-war-with-iran (Accessed on 30 April 2026)
31Macron wants a war with Russia, but the army opposes itveridica.rohttps://www.veridica.ro/en/fake-news-disinformation-propaganda/fake-news-macron-wants-a-war-with-russia-but-the-army-opposes-it (Accessed on 30 April 2026)
32EU and Moldova say Russia is a threat to mask their failuresveridica.rohttps://www.veridica.ro/en/fake-news-disinformation-propaganda/fake-news-eu-and-moldova-say-russia-is-a-threat-to-mask-their-failures (Accessed on 30 April 2026)
33Moldovans pay less for electricity produced in Romania than Romanians themselvesveridica.rohttps://www.veridica.ro/en/fake-news-disinformation-propaganda/fake-news-moldovans-pay-less-for-electricity-produced-in-romania-than-romanians-themselves (Accessed on 30 April 2026)
34Romanian consumers will no longer be able to purchase energy from Hidroelectrica and Nuclearelectrica, which will supply electricity to Moldovaveridica.rohttps://www.veridica.ro/en/fake-news-disinformation-propaganda/fake-news-romanian-consumers-will-no-longer-be-able-to-purchase-energy-from-hidroelectrica-and-nuclearelectrica-which-will-supply-electricity-to-the-republic-of-moldova (Accessed on 30 April 2026)
35Romania is under EU tyrannyeuvsdisinfo.euhttps://euvsdisinfo.eu/report/romania-is-under-eu-tyranny/ (Accessed on 30 April 2026)
36Bomb: Russia violates the Sanctions with the complicity of NATO and the EU—Romania sinks into the swamps of Suicidal Fanatical Servilityactivenews.rohttps://www.activenews.ro/opinii/Bomba-Rusia-incalca-Sanctiunile-cu-complicitatea-NATO-si-UE-%E2%80%93-Romania-se-afunda-in-smarcurile-Slugarniciei-Fanatice-Sinucigase-192151 (Accessed on 30 April 2026)
37Alex Jones: In Romania a war related to Ukraine can explode to block Trump’s installation. NATO staged a coup d’état against the Romanian people by eliminating Georgescuactivenews.rohttps://www.activenews.ro/stiri/Bomba-lui-Alex-Jones-In-Romania-poate-sa-explodeze-un-razboi-legat-de-Ucraina-pentru-blocarea-instalarii-lui-Trump.-NATO-a-dat-o-lovitura-de-stat-impotriva-poporului-roman-prin-eliminarea-lui-Georgescu-si-instaureaza-o-dictatura-cu-marioneta-Iohann-194299 (Accessed on 30 April 2026)
38IT WAS NOT ROMANIA THAT CANCELED THE ELECTIONS IN ROMANIA—The shocking hypothesis of a Trump influencer and former US State Department official Mike Benzactivenews.rohttps://www.activenews.ro/stiri-mass-media/NU-ROMANIA-A-ANULAT-ALEGERILE-DIN-ROMANIA-Ipoteza-socanta-a-unui-influencer-de-al-lui-Donald-Trump-si-a-unui-fost-oficial-al-Departamentului-de-Stat-al-SUA-Mike-Benz.-Are-legatura-cu-Ucraina-VIDEO-TRADUS-193962 (Accessed on 30 April 2026)
39The bomb from Washington: ‘The silent coup d’état given by the EU in Romania through the installation of Nicușor Dan’activenews.rohttps://www.activenews.ro/stiri/Bomba-de-la-Washington-Lovitura-de-stat-silentioasa-data-de-UE-in-Romania-prin-instalarea-lui-Nicusor-Dan.-Intrebari-pe-care-nicio-democratie-serioasa-nu-si-poate-permite-sa-le-ignore-197368 (Accessed on 30 April 2026)
40Will NATO deploy NUCLEAR WEAPONS in Romania?activenews.rohttps://www.activenews.ro/stiri/Va-amplasa-NATO-in-Romania-ARME-NUCLEARE-170847 (Accessed on 30 April 2026)
41Năsui’s bomb. Why wasn’t Iohannis at the reception from the NATO Summitactivenews.rohttps://www.activenews.ro/opinii/Bomba-lui-Nasui.-De-ce-nu-a-fost-Iohannis-la-receptia-de-la-Summitul-NATO-190665 (Accessed on 30 April 2026)
42The Jews caused the floods in Romaniaveridica.rohttps://www.veridica.ro/fake-news-dezinformare-propaganda/fake-news-evreii-au-provocat-inundatiile-din-romania (Accessed on 30 April 2026)
43Shock—Devastating analysis in the Ukrainian press: How Zelensky became the owner of Ukraineinpolitics.rohttps://inpolitics.ro/soc-analiza-devastatoare-in-presa-ucraineana-cum-a-ajuns-zelensky-proprietarul-ucrainei_1860085865.html (Accessed on 30 April 2026)
44The desperation of a dinosaur, PSD with AUR teethcatavencii.rohttps://www.catavencii.ro/disperarea-unui-dinozaur-psd-cu-dinti-de-aur/ (Accessed on 30 April 2026)
45The Government will grant vouchers for the purchase of cherriescatavencii.rohttps://www.catavencii.ro/guvernul-va-acorda-vouchere-pentru-achizitia-de-cirese/ (Accessed on 30 April 2026)
46The corn quotation grows after Simion and AUR started ordering popcorncatavencii.rohttps://www.catavencii.ro/creste-cotatia-porumbului-dupa-ce-simion-si-aur-au-inceput-sa-comande-popcorn/ (Accessed on 30 April 2026)
47After successfully unclogging the WC on Artemis II, the Romanian plumber received a new work: the Strait of Hormuzcatavencii.rohttps://www.catavencii.ro/dupa-ce-a-desfundat-cu-succes-wc-ul-de-pe-artemis-ii-instalatorul-roman-primeste-o-noua-lucrare-stramtoarea-ormuz/ (Accessed on 30 April 2026)
48Bolojan went personally to the hunger strikers to savor a hamburgerneurococi.rohttps://neurococi.ro/posts/bolojan-s-a-dus-personal-la-grevistii-foamei-sa-savureze-un-hamburgher (Accessed on 30 April 2026)
49CFR acquired a BMW locomotive that chases the trains of the poor and flashes themneurococi.rohttps://neurococi.ro/posts/cfr-a-achizitionat-o-locomotiva-bmw-care-urmareste-trenurile-si-le-da-flashuri (Accessed on 30 April 2026)
5018 people with wigs and fake mustaches detained on the airport in The Hagueneurococi.rohttps://neurococi.ro/posts/18-persoane-cu-peruci-si-mustati-false-retinute-pe-aeroportul-din-haga (Accessed on 30 April 2026)
Table A2. Real news articles.
Table A2. Real news articles.
No.Article Title (English Translation)SourceURL
1Great Britain takes a historic step: younger generations will never be able to buy cigarettes againmediafax.rohttps://www.mediafax.ro/sanatate/marea-britanie-face-un-pas-istoric-generatiile-tinere-nu-vor-mai-putea-cumpara-niciodata-tigari-23724675 (Accessed on 29 April 2026)
2A man stole two scooters and an electric scooter in just a few hours from Sector 3mediafax.rohttps://www.mediafax.ro/social/un-barbat-a-furat-doua-scutere-si-o-trotineta-electrica-in-doar-cateva-ore-din-sectorul-3-23723580 (Accessed on 29 April 2026)
3Travelers at European airports wait up to three hours at border controls due to the new EU entry-exit system (EES)mediafax.rohttps://www.mediafax.ro/externe/noul-sistem-de-intrare-iesire-al-ue-genereaza-intarzieri-pasagerii-asteapta-pana-la-trei-ore-la-controalele-de-frontiera-23720703 (Accessed on 29 April 2026)
4Economedia: The number of foreign workers in Romania has reached almost 150,000. Why immigration will increase in the coming yearsmediafax.rohttps://www.mediafax.ro/economic/economedia-numarul-muncitorilor-straini-din-romania-a-ajuns-la-aproape-150-000-de-ce-va-creste-imigratia-in-anii-urmatori-23720136 (Accessed on 29 April 2026)
5INSCOP Survey: 46% of Romanians want to give up changing the timeobservatornews.rohttps://observatornews.ro/social/sondaj-inscop-46-dintre-romani-vor-renuntarea-la-schimbarea-orei-650437.html (Accessed on 29 April 2026)
6The strategy that empties our pockets. How the tiny difference of one ban manipulates our perceptionadevarul.rohttps://adevarul.ro/stiri-interne/societate/strategia-care-ne-goleste-buzunarele-cum-ne-2521149.html (Accessed on 29 April 2026)
7The end of a historic robbery. Why the theft in the Netherlands was felt as a personal loss by Romaniansadevarul.rohttps://adevarul.ro/stiri-interne/societate/finalul-unui-jaf-traumatic-de-ce-furtul-din-2524448.html (Accessed on 29 April 2026)
8A 2-m long snake was found in the house of a woman from Galați county. Frightened, the owner called 112adevarul.rohttps://adevarul.ro/stiri-interne/societate/un-sarpe-lung-de-2-metri-a-fost-gasit-in-casa-unei-2524034.html (Accessed on 29 April 2026)
9Romania, roamed by thousands of feral horses. The places where free herds can be metadevarul.rohttps://adevarul.ro/stiri-interne/societate/romania-cutreierata-de-mii-de-cai-salbaticiti-2520199.html (Accessed on 29 April 2026)
10The places in Romania gripped by the gold rush. The curse of Dacian treasures and the legends of the Apuseni vâlveadevarul.rohttps://adevarul.ro/stiri-interne/societate/locurile-din-romania-cuprinse-de-febra-aurului-2516182.html (Accessed on 29 April 2026)
11Explosion in the rental market: More and more Romanians delay buying homes and choose the temporary solution. “People are much more cautious”digi24.rohttps://www.digi24.ro/imobiliare/explozie-pe-piata-chiriilor-tot-mai-multi-romani-amana-sa-cumpere-locuinte-si-aleg-solutia-temporara-oamenii-sunt-mult-mai-precauti-3738229 (Accessed on 29 April 2026)
12What fines drivers who drive a dirty car risk and when the technical inspection (ITP) can be rejecteddigi24.rohttps://www.digi24.ro/stiri/actualitate/social/ce-amenzi-risca-soferii-care-circula-cu-masina-murdara-si-cand-poate-fi-respinsa-inspectia-tehnica-itp-3737833 (Accessed on 29 April 2026)
13Earth Day 2026: Why it is celebrated on April 22 and how the event became a global phenomenondigi24.rohttps://www.digi24.ro/stiri/actualitate/evenimente/ziua-planetei-pamant-2026-de-ce-este-sarbatorita-pe-22-aprilie-si-cum-a-devenit-evenimentul-un-fenomen-global-3737707 (Accessed on 29 April 2026)
14The great draw: The secret plan by which Trump and Iran want to avoid total war. “No one can afford this conflict anymore”adevarul.rohttps://adevarul.ro/stiri-interne/evenimente/marea-remiza-planul-secret-prin-care-trump-si-2524548.html (Accessed on 29 April 2026)
15The USA begins the reimbursement of the illegal taxes imposed by Trump. Billions of dollars are at stakedigi24.rohttps://m.digi24.ro/stiri/externe/sua-incep-rambursarea-taxelor-ilegale-impuse-de-trump-sunt-in-joc-miliarde-de-dolari-3736397 (Accessed on 29 April 2026)
16The IMF’s warning to Europe, amid the war: “Accelerate the ‘Green Deal’. Do not protect the population from fuel costs”digi24.rohttps://m.digi24.ro/stiri/externe/ue/avertismentul-fmi-catre-europa-pe-fondul-razboiului-accelerati-pactul-verde-nu-protejati-populatia-de-costurile-combustibililor-3735625 (Accessed on 1 May 2026)
17The FBI investigates possible links between the deaths or disappearances of 11 scientists. Trump: “It’s a pretty serious matter”digi24.rohttps://m.digi24.ro/stiri/externe/sua/fbi-investigheaza-posibile-legaturi-intre-decesele-sau-disparitiile-a-11-oameni-de-stiinta-trump-e-o-chestiune-destul-de-serioasa-3739093 (Accessed on 1 May 2026)
18The first elections organized in the Gaza Strip in the last 20 years. How the voting wenthotnews.rohttps://hotnews.ro/primele-alegeri-organizate-in-fasia-gaza-in-ultimii-20-de-ani-votul-s-a-incheiat-insa-cu-o-prezenta-scazuta-la-urne-2229267 (Accessed on 1 May 2026)
19A major airline makes tickets cheaper for passengers who want to take only very small luggage on boardhotnews.rohttps://hotnews.ro/o-companie-aeriana-majora-ieftineste-biletele-pentru-pasagerii-ce-vor-sa-ia-la-bord-doar-bagaje-foarte-mici-2227254 (Accessed on 1 May 2026)
20Europe risks chaos in airports: fuel crisis threatens the 2026 summer seasondigi24.rohttps://www.digi24.ro/magazin/timp-liber/vacante/europa-risca-haos-in-aeroporturi-criza-de-combustibil-ameninta-sezonul-estival-2026-3742829 (Accessed on 1 May 2026)
21The European Union will prepare a detailed draft on the mutual defense clausedigi24.rohttps://www.digi24.ro/stiri/externe/ue/uniunea-europeana-va-pregati-un-proiect-detaliat-privind-clauza-de-asistenta-reciproca-3742861 (Accessed on 1 May 2026)
22The moment when a Romanian scammer nicknamed ‘The Prince of Dubai’ was caught by the Polish police. How he got this nicknamedigi24.rohttps://www.digi24.ro/stiri/externe/ue/momentul-in-care-un-escroc-roman-supranumit-printul-din-dubai-a-fost-prins-de-politia-poloneza-cum-a-capatat-aceasta-porecla-3718595 (Accessed on 1 May 2026)
23A man won 44 million euros at the 6/49 Lottery and told his wife nothing for several daysadevarul.rohttps://adevarul.ro/stiri-externe/europa/un-barbat-a-castigat-44-de-milioane-de-euro-la-2525252.html (Accessed on 1 May 2026)
24Explosion on the real estate market in Europe: Hungary climbs over 20% in housing prices, Romania—above the EU averageadevarul.rohttps://adevarul.ro/stiri-externe/europa/explozie-pe-piata-imobiliara-din-europa-ungaria-2522048.html (Accessed on 1 May 2026)
25Devastating attack in the Ukrainian city of Nikopol. A Russian drone killed four people and injured 16 othersadevarul.rohttps://adevarul.ro/stiri-externe/europa/atac-devastator-in-orasul-ucrainean-nikopol-o-2521066.html (Accessed on 1 May 2026)
26The kerosene crisis. What Romanians’ holidays will look like this summerdigi24.rohttps://m.digi24.ro/stiri/actualitate/criza-kerosenului-cum-vor-arata-vacantele-romanilor-in-aceasta-vara-3732353 (Accessed on 1 May 2026)
27Counterfeit goods of over 3 million lei discovered in a truck registered in Romania. The transport was coming from Turkeydigi24.rohttps://www.digi24.ro/stiri/economie/marfa-contrafacuta-de-peste-3-milioane-de-lei-descoperita-intr-un-camion-inmatriculat-in-romania-transportul-venea-din-turcia-3718385 (Accessed on 1 May 2026)
28Romanian drones for Ukraine. Who is developing the ‘Shahed killer’, equipped with a turbo engine and artificial intelligencestartupcafe.rohttps://startupcafe.ro/drone-romanesti-pentru-ucraina-cine-dezvolta-killer-ul-de-shahed-dotat-cu-motor-turbo-si-inteligenta-artificiala-98049 (Accessed on 1 May 2026)
29Gold: central banks were buying it at a record pace. Why they are selling it nowhotnews.rohttps://hotnews.ro/aur-bancile-centrale-il-cumparau-intr-un-ritm-record-de-ce-il-vand-acum-2219324 (Accessed on 1 May 2026)
30The mayors’ revelry on public money continues. The mayors will squander 18 billion lei in 2026, although the Court of Accounts found countless irregularitieslibertatea.rohttps://www.libertatea.ro/stiri/buget-2026-sume-alocate-primari-romania-18-miliarde-lei-nereguli-curtea-conturi-5657634 (Accessed on 2 May 2026)
31ANRE has put into consultation the rules for energy communities. How Romanians will be able to produce and share energydigi24.rohttps://www.digi24.ro/stiri/economie/energie/anre-a-pus-in-consultare-regulile-pentru-comunitatile-de-energie-cum-vor-putea-romanii-sa-produca-si-sa-imparta-energie-3709023 (Accessed on 2 May 2026)
32What effects the Government’s crisis measures will have on the fuel market. Experts do not rule out rationingdigi24.rohttps://www.digi24.ro/stiri/economie/energie/ce-efecte-vor-avea-masurile-de-criza-ale-guvernului-pe-piata-carburantilor-expertii-nu-exclud-rationalizarea-3692425 (Accessed on 2 May 2026)
33How Romania’s meat production evolved in January 2026. Differences between the pork, poultry and beef sectorsdigi24.rohttps://www.digi24.ro/stiri/economie/cum-a-evoluat-productia-de-carne-a-romaniei-in-luna-ianuarie-2026-diferente-intre-sectorul-porcin-avicol-si-cel-bovin-3669063 (Accessed on 2 May 2026)
34The EU tries to unblock the free trade agreement with Mercosur. Italy is at the center of the last-minute negotiationsdigi24.rohttps://www.digi24.ro/stiri/economie/agricultura/ue-incearca-sa-deblocheze-acordul-de-liber-schimb-cu-mercosur-italia-este-in-centrul-negocierilor-de-ultim-moment-3573579 (Accessed on 2 May 2026)
35How much the Gulf crisis made airline tickets more expensive and which are the most affected flightsdigi24.rohttps://www.digi24.ro/magazin/timp-liber/vacante/cu-cat-a-scumpit-criza-din-golf-biletele-de-avion-si-care-sunt-cele-mai-afectate-zboruri-3740123 (Accessed on 2 May 2026)
36Romania’s budget deficit dropped to 7.9% of GDP in 2025. However, the level remains the highest in the European Uniondigi24.rohttps://www.digi24.ro/stiri/economie/deficitul-bugetar-al-romaniei-a-scazut-la-79-din-pib-in-2025-nivelul-ramane-insa-cel-mai-ridicat-din-uniunea-europeana-3739931 (Accessed on 2 May 2026)
37BNR: Inflation will rise above estimates until summer, amid more expensive energy and the war in the Middle Eastdigi24.rohttps://www.digi24.ro/stiri/economie/bnr-inflatia-va-creste-peste-estimari-pana-in-vara-pe-fondul-scumpirii-energiei-si-al-razboiului-din-orientul-mijlociu-3715153 (Accessed on 2 May 2026)
38Diana Șoșoacă could be investigated in Romania. The Committee in the European Parliament, unanimous vote to lift immunitydigi24.rohttps://www.digi24.ro/stiri/actualitate/politica/diana-sosoaca-ar-putea-fi-anchetata-in-romania-comisia-din-parlamentul-european-a-dat-vot-pentru-ridicarea-imunitatii-ce-urmeaza-3740831 (Accessed on 2 May 2026)
39The trade unions proposed, at the meeting with PSD, a government program for the negotiations to form a new coalitiondigi24.rohttps://www.digi24.ro/stiri/actualitate/social/sindicatele-au-propus-la-intalnirea-cu-psd-un-nou-program-de-guvernare-pentru-negocierile-de-formare-a-unei-noi-coalitii-3739819 (Accessed on 2 May 2026)
40The Government has opened registrations for the Official Internship Program addressed to young peoplemediafax.rohttps://www.mediafax.ro/social/guvernul-a-deschis-inscrierile-pentru-programului-oficial-de-internship-adresat-tinerilor-23726131 (Accessed on 2 May 2026)
41Nicușor Dan does not accept a minority government with the support of AUR neither from PSD, nor from PNLadevarul.rohttps://adevarul.ro/politica/nicusor-dan-nu-accepta-un-guvern-minoritar-cu-2525023.html (Accessed on 2 May 2026)
42How AUR got to play all the power games. There is a point where PSD and President Nicușor Dan meetadevarul.rohttps://adevarul.ro/blogurile-adevarul/cum-a-ajuns-aur-sa-faca-toate-jocurile-de-putere-2523943.html (Accessed on 2 May 2026)
43The crisis in the pro-European coalition in Romania weakens the ‘barrier’ against the far right, a foreign analysis showsadevarul.rohttps://adevarul.ro/politica/criza-din-coalitia-proeuropeana-din-romania-2524457.html (Accessed on 2 May 2026)
44PSD leaders voted unanimously to withdraw ministers from the Government, after consultations at Cotroceniadevarul.rohttps://adevarul.ro/politica/liderii-psd-au-intrat-in-sedinta-dupa-consultarile-2524470.html (Accessed on 2 May 2026)
45The Romanian Government launches an in vitro fertilization support program intended for infertile couplesmediafax.rohttps://www.mediafax.ro/social/guvernul-romaniei-lanseaza-program-de-sprijin-pentru-fertilizare-in-vitro-destinat-cuplurilor-infertile-23725087 (Accessed on 2 May 2026)
46New rules for hiring foreigners: no taxes for workers and stricter controlsmediafax.rohttps://www.mediafax.ro/politic/noi-reguli-pentru-angajarea-strainilor-fara-taxe-pentru-lucratori-si-controale-mai-stricte-23725090 (Accessed on 2 May 2026)
47What Bucharesters think about a possible resignation of Prime Minister Ilie Bolojan. Voting intentions at the local elections (poll)digi24.rohttps://www.digi24.ro/stiri/actualitate/politica/ce-cred-bucurestenii-despre-o-eventuala-demisie-a-premierului-ilie-bolojan-intentiile-de-vot-la-alegerile-locale-sondaj-3745319 (Accessed on 2 May 2026)
48How Russia is trying to influence the parliamentary elections in Hungary to support Viktor Orbánadevarul.rohttps://adevarul.ro/stiri-externe/europa/cum-incearca-rusia-sa-influenteze-alegerile-2513142.html (Accessed on 2 May 2026)
49What the last two electoral years have shown. Political scientist: “We have a clearer picture. And we still have to piece together bits of reality”adevarul.rohttps://adevarul.ro/politica/ce-au-aratat-ultimii-doi-ani-electorali-2496931.html (Accessed on 2 May 2026)
50Călin Georgescu received for his birthday a gift of another 60 days of judicial control. Supporters waited with cake and sang ‘Happy Birthday’libertatea.rohttps://www.libertatea.ro/stiri/calin-georgescu-primit-inca-60-zile-control-judiciar-ziua-lui-sustinatori-asteptat-tort-cantat-la-multi-ani-5676482 (Accessed on 2 May 2026)
Table A3. Experimental procedures—code adjustments.
Table A3. Experimental procedures—code adjustments.
Algorithm/ParameterAI-Generated ValueValue Used
Naive Bayes: TF-IDF max_featuresUnlimited50,000
KNN: TF-IDF max_featuresUnlimited10,000
Random Forest: TF-IDF max_featuresUnlimited20,000
XGBoost: TF-IDF max_featuresUnlimited20,000
BERT: max_length (tokenizer truncation)512 tokens (BERT’s architectural maximum)256 tokens
BERT: per_device_train_batch_size168
BERT: num_train_epochs31
TextGCN: TF-IDF max_featuresUnlimited20,000
TextGCN: batch_size512256
TextGCN: epochs105
Table A4. Experimental procedures—hyperparameter configuration.
Table A4. Experimental procedures—hyperparameter configuration.
AlgorithmParameterValue Used
Naive BayesModel variantMultinomialNB()
TF-IDF max_features50,000
Random_state42
KNNN_neighbors5
TF-IDF max_features10,000
Random ForestN_estimators100
N_jobs−1
TF-IDF max_features20,000
Random_state42
XGBoostN_estimators100
Learning_rate0.1
Max_depth6
Tree_methodhist
TF-IDF max_features20,000
Random_state42
BERTModel variantBert-base-uncased
TokenizerBertTokenizerFast
Max_length256 tokens
Paddingmax_length
Per_device_train_batch_size8
Num_train_epochs1
TextGCNTF-IDF max_features/input nodes20,000
Hidden layer128 units + ReLU
Dropout0.5
Output layer2 units
OptimizerAdam (lr = 0.001)
Loss functionCrossEntropyLoss
Epochs5
Batch_size256

References

  1. Acerbi, A.; Altay, S.; Mercier, H. Research note: Fighting misinformation or fighting for information? HKS Misinf. Rev. 2022, 3, 1–15. [Google Scholar] [CrossRef] [Scilit]
  2. World Economic Forum. The Global Risks Report 2024, 19th ed.; World Economic Forum: Geneva, Switzerland, 2024; Available online: https://www.weforum.org/publications/global-risks-report-2024/ (accessed on 9 June 2026).
  3. World Economic Forum. The Global Risks Report 2025, 20th ed.; World Economic Forum: Geneva, Switzerland, 2025; Available online: https://www.weforum.org/publications/global-risks-report-2025/ (accessed on 9 June 2026).
  4. Ceylan, G.; Anderson, I.A.; Wood, W. Sharing of misinformation is habitual, not just lazy or biased. Proc. Natl. Acad. Sci. USA 2023, 120, e2216614120. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. NewsGuard. AI Tracking Center: 3006 AI Content Farm Websites Identified. Available online: https://www.newsguardtech.com/special-reports/ai-tracking-center/ (accessed on 9 June 2026).
  6. European Parliament. Social Media Survey 2025. Eurobarometer Surveys. Available online: https://europa.eu/eurobarometer/surveys/detail/3592 (accessed on 9 June 2026).
  7. Newman, N.; Ross Arguedas, A.; Robertson, C.T.; Nielsen, R.K.; Fletcher, R. Reuters Institute Digital News Report 2025; Reuters Institute for the Study of Journalism: Oxford, UK, 2025. [Google Scholar] [CrossRef]
  8. Solomon, M.-A. Hybrid Warfare Through Disinformation: The Case of Romania’s Presidential Election. Friends of Europe, 10 July 2025. Available online: https://www.friendsofeurope.org/insights/critical-thinking-hybrid-warfare-through-disinformation-the-recent-case-of-romania/ (accessed on 9 June 2026).
  9. Tandoc, E.C.; Lim, Z.W.; Ling, R. Defining “fake news”: A typology of scholarly definitions. Digit. Journal. 2018, 6, 137–153. [Google Scholar] [CrossRef] [Scilit]
  10. Dima, A.; Ilis, E.; Florea, D.; Dascalu, M. Detection of fake news in Romanian: LLM-based approaches to COVID-19 misinformation. Information 2025, 16, 796. [Google Scholar] [CrossRef] [Scilit]
  11. Allcott, H.; Gentzkow, M. Social media and fake news in the 2016 election. J. Econ. Perspect. 2017, 31, 211–236. [Google Scholar] [CrossRef] [Scilit]
  12. Lazer, D.M.J.; Baum, M.A.; Benkler, Y.; Berinsky, A.J.; Greenhill, K.M.; Menczer, F.; Metzger, M.J.; Nyhan, B.; Pennycook, G.; Rothschild, D.; et al. The science of fake news. Science 2018, 359, 1094–1096. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. El Mikati, I.K.; Hoteit, R.; Harb, T.; El Zein, O.; Piggott, T.; Melki, J.; Mustafa, R.A.; Akl, E.A. Defining misinformation and related terms in health-related literature: Scoping review. J. Med. Internet Res. 2023, 25, e45731. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  14. Wardle, C.; Derakhshan, H. Information Disorder: Toward an Interdisciplinary Framework for Research and Policy Making; Council of Europe: Strasbourg, France, 2017; Available online: https://rm.coe.int/information-disorder-toward-an-interdisciplinary-framework-for-researc/168076277c (accessed on 9 June 2026).
  15. Erhardt, K.; Pentland, A. Disambiguating disinformation: Extending beyond the veracity of online content. arXiv 2022, arXiv:2206.12915. [Google Scholar] [CrossRef] [Scilit]
  16. American Psychological Association. Using Psychological Science to Understand and Fight Health Misinformation: An APA Consensus Statement. Available online: https://www.apa.org/pubs/reports/health-misinformation (accessed on 9 June 2026).
  17. Raponi, S.; Khalifa, Z.; Oligeri, G.; Di Pietro, R. Fake news propagation: A review of epidemic models, datasets, and insights. ACM Trans. Web 2022, 16, 12. [Google Scholar] [CrossRef] [Scilit]
  18. Moisi, E.V.; Mihalca, B.C.; Coman, S.M.; Pater, A.M.; Popescu, D.E. Romanian fake news detection using machine learning and transformer-based approaches. Appl. Sci. 2024, 14, 11825. [Google Scholar] [CrossRef] [Scilit]
  19. Alkhateri, S.M.A.B.H.; Devi, S.I.; Jano, Z.; Al-shami, S.A. Attitudes towards fake news: A systematic literature review. Webology 2021, 18, 368–376. [Google Scholar] [CrossRef] [Scilit]
  20. Nitu, M.; Dascalu, M. Beyond lexical boundaries: LLM-generated text detection for Romanian digital libraries. Future Internet 2024, 16, 41. [Google Scholar] [CrossRef] [Scilit]
  21. Aimeur, E.; Amri, S.; Brassard, G. Fake news, disinformation and misinformation in social media: A review. Soc. Netw. Anal. Min. 2023, 13, 30. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  22. Domenico, G.D.; Sit, J.; Ishizaka, A.; Nunan, D. Fake news, social media and marketing: A systematic review. J. Bus. Res. 2021, 124, 329–341. [Google Scholar] [CrossRef] [Scilit]
  23. Farhangian, F.; Cruz, R.M.O.; Cavalcanti, G.D.C. Fake news detection: Taxonomy and comparative study. Inf. Fusion 2024, 103, 102140. [Google Scholar] [CrossRef] [Scilit]
  24. Peter, N. The impact of digitalization on a company’s risk management. MAP Soc. Sci. 2023, 3, 41–50. [Google Scholar] [CrossRef] [Scilit]
  25. Busioc, C.; Ruseti, S.; Dascalu, M. A literature review of NLP approaches to fake news detection and their applicability to Romanian-language news analysis. Transilvania 2020, 10, 65–71. [Google Scholar] [CrossRef] [Scilit]
  26. Veziroğlu, M.; Veziroğlu, E.; Bucak, İ.Ö. Performance comparison between Naive Bayes and machine learning algorithms for news classification. In Bayesian Inference—Recent Trends; Bucak, İ.Ö., Ed.; IntechOpen: London, UK, 2024. [Google Scholar] [CrossRef] [Scilit]
  27. Nagarhalli, T.P.; Vaze, V.; Rana, N.K. Impact of machine learning in natural language processing: A review. In Proceedings of the 2021 Third International Conference on Intelligent Communication Technologies and Virtual Mobile Networks (ICICV), Tirunelveli, India, 4–6 February 2021; pp. 1529–1534. [Google Scholar] [CrossRef] [Scilit]
  28. Khensous, G.; Labed, K.; Labed, Z. Exploring the evolution and applications of natural language processing in education. Rom. J. Inf. Technol. Autom. Control 2023, 33, 61–74. [Google Scholar] [CrossRef] [Scilit]
  29. Sheng, Q.; Zhang, X.; Cao, J.; Zhong, L. Integrating pattern- and fact-based fake news detection via model preference learning. In Proceedings of the 30th ACM International Conference on Information & Knowledge Management, Gold Coast, QLD, Australia, 1–5 November 2021; pp. 1640–1650. [Google Scholar] [CrossRef] [Scilit]
  30. Vishnupriya, G.; Jeriel K, A.; RNS, A.; Ajay, G.; Giftson J, A. Combating fake news in the digital age: A review of AI-based approaches. In Proceedings of the 2024 IEEE 9th International Conference for Convergence in Technology (I2CT), Pune, India, 6–7 April 2024; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  31. Alameri, S.A.; Mohd, M. Comparison of fake news detection using machine learning and deep learning techniques. In Proceedings of the 2021 3rd International Cyber Resilience Conference (CRC), Langkawi, Malaysia, 29–31 January 2021; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  32. Forno, A.J.D.; Richetti, G.P.; Knaesel, V.H. Fake news detection algorithms—A systematic literature review. Data Knowl. Eng. 2025, 158, 102441. [Google Scholar] [CrossRef] [Scilit]
  33. Al-Alshaqi, M.; Rawat, D.B.; Liu, C. Ensemble techniques for robust fake news detection: Integrating transformers, natural language processing, and machine learning. Sensors 2024, 24, 6062. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Reddy, A.V.; Harika, A.; Deshmukh, A. Mitigating misinformation: A comparative analysis of machine learning models for fake news detection. In Proceedings of International Conference on Sustainable Science and Technology for Tomorrow (SciTech 2024); Thatikonda, S.K., Pandey, B., Shankar, D., Rodriguez, R.V., Eds.; Atlantis Press: Amsterdam, The Netherlands, 2025; Volume 2, pp. 151–162. [Google Scholar]
  35. Ilugbiyin, K.R. A comparison of machine learning classifiers for fake news identification using the ISOT dataset: XGBoost and Random Forest achieve 100% accuracy. Int. J. Res. Sci. Innov. 2025, 12, 1213CS005. [Google Scholar] [CrossRef] [Scilit]
  36. Essa, E.; Omar, K.; Alqahtani, A. Fake news detection based on a hybrid BERT and LightGBM models. Complex Intell. Syst. 2023, 9, 6581–6592. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  37. Bashaddadh, O.; Omar, N.; Mohd, M.; Nor Akmal Khalid, M. Machine learning and deep learning approaches for fake news detection: A systematic review of techniques, challenges, and advancements. IEEE Access 2025, 13, 90433–90466. [Google Scholar] [CrossRef] [Scilit]
  38. Islam, M.S.; Zhang, L. A review on BERT: Language understanding for different types of NLP task. Preprints 2024, 20240101857. [Google Scholar] [CrossRef] [Scilit]
  39. Devlin, J.; Chang, M.-W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv 2018, arXiv:1810.04805. [Google Scholar] [CrossRef] [Scilit]
  40. Raza, S.; Paulen-Patterson, D.; Ding, C. Fake news detection: Comparative evaluation of BERT-like models and large language models with generative AI-annotated data. arXiv 2024, arXiv:2412.14276. [Google Scholar] [CrossRef] [Scilit]
  41. Yao, L.; Mao, C.; Luo, Y. Graph convolutional networks for text classification. In Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; Volume 33, pp. 7370–7377. [Google Scholar] [CrossRef] [Scilit]
  42. Singh, P.; Singh, N.; Singh, K.K.; Singh, A. Diagnosing of disease using machine learning. In Machine Learning and the Internet of Medical Things in Healthcare; Singh, K.K., Elhoseny, M., Singh, A., Elngar, A.A., Eds.; Academic Press: Cambridge, MA, USA, 2021; pp. 89–111. [Google Scholar] [CrossRef] [Scilit]
  43. Geng, S. Analysis of the different statistical metrics in machine learning. Highlights Sci. Eng. Technol. 2024, 88, 350–356. [Google Scholar] [CrossRef] [Scilit]
  44. Buzea, M.C.; Trausan-Matu, S.; Rebedea, T. Automatic fake news detection for Romanian online news. Information 2022, 13, 151. [Google Scholar] [CrossRef] [Scilit]
  45. Bucos, M.; Drăgulescu, B. Enhancing fake news detection in Romanian using transformer-based back translation augmentation. Appl. Sci. 2023, 13, 13207. [Google Scholar] [CrossRef] [Scilit]
  46. Verma, P.K.; Agrawal, P.; Prodan, R. WELFake dataset for fake news detection in text data. IEEE Trans. Comput. Soc. Syst. 2021, 8, 1–13. [Google Scholar] [CrossRef]
  47. Ahmed, H.; Traore, I.; Saad, S. Detection of online fake news using n-gram analysis and machine learning techniques. In Intelligent, Secure, and Dependable Systems in Distributed and Cloud Environments; Traore, I., Woungang, I., Awad, A., Eds.; Springer: Cham, Switzerland, 2017; Volume 10618, pp. 127–138. [Google Scholar]
  48. Peutz, S. Misinformation & Fake News Text Dataset 79k. Kaggle, 2022. [Data Set]. Available online: https://www.kaggle.com/datasets/stevenpeutz/misinformation-fake-news-text-dataset-79k (accessed on 9 June 2026).
Figure 1. Distribution of predictions and errors across the confusion matrices.
Figure 1. Distribution of predictions and errors across the confusion matrices.
Information 17 00698 g001
Table 1. Comparative performance of the models on the Romanian test corpus.
Table 1. Comparative performance of the models on the Romanian test corpus.
ModelAccuracy95% CI for AccuracyFake PrecisionRecall FakeF1 FakeReal PrecisionRecall RealF1 RealMacro-F1/Weighted-F1Total Errors
Naive Bayes0.710.61–0.790.800.560.660.660.860.750.7029
KNN0.490.39–0.590.490.940.650.400.040.070.3651
Random Forest0.770.68–0.840.740.840.790.810.700.750.7723
XGBoost0.570.47–0.660.540.880.670.680.260.380.5243
BERT0.730.64–0.810.690.820.750.780.640.700.7327
TextGCN0.720.63–0.800.680.840.750.790.600.680.7228
Table 2. Real articles classified as fake by Naive Bayes.
Table 2. Real articles classified as fake by Naive Bayes.
Item No.Title of the ArticleReal LabelAI Prediction
6INSCOP Survey: 46% of Romanians want to give up changing the timeRealFake
7The strategy that empties our pockets. How the tiny difference of one ban manipulates our perceptionRealFake
8The end of a historic robbery. Why the theft in the Netherlands was felt as a personal loss by RomaniansRealFake
9A 2-m long snake was found in the house of a woman from Galati county.RealFake
11The places in Romania gripped by the gold rushRealFake
24A man won 44 million euros at the 6/49 Lottery and told his wife nothing for several daysRealFake
34How Romania’s meat production evolved in January 2026. Differences between pork, poultry and beef sectorsRealFake
Table 3. Fake articles classified as real by KNN.
Table 3. Fake articles classified as real by KNN.
Crt. No.Title of the ArticleReal LabelAI Prediction
60Donald Trump confirms almost identically what Călin Georgescu was saying (Davos speech, green energy)FakeReal
89IT WAS NOT ROMANIA THAT CANCELED THE ELECTIONS IN ROMANIA—The shocking hypothesis of a Donald Trump influencer and Mike BenzFakeReal
92Năsui’s bomb. Why wasn’t Iohannis at the reception from the NATO SummitFakeReal
Table 4. Fake articles classified as real by Random Forest.
Table 4. Fake articles classified as real by Random Forest.
Crt. No.Title of the ArticleReal LabelAI Prediction
55Romanians and Europe will freeze this winterFakeReal
73The socialist government of Spain goes even further with censorship.FakeReal
74Right-hand drive cars will no longer be able to circulate on public roads, the senators votedFakeReal
77Bolojan sold the Tarnița hydroelectric plant to France. It is good for the French and not good for RomaniansFakeReal
80Romania, forced to organize the second round of the presidential election, according to the 2024 election calendarFakeReal
88Alex Jones’s bomb: In Romania a war related to Ukraine can explode to block Trump. NATO staged a coup d’état against RomaniansFakeReal
91Will NATO deploy NUCLEAR WEAPONS in Romania?FakeReal
99Bolojan went personally to the hunger strikers to savor a hamburger.FakeReal
Table 5. Fake articles classified as real by XGBoost.
Table 5. Fake articles classified as real by XGBoost.
Crt. No.Title of the ArticleReal LabelAI Prediction
55Romanians and Europe will freeze this winterFakeReal
69The satans from the CNA banned Realitatea Plus. Dastardly attack by the system on the only TV where Georgescu appearsFakeReal
74The socialist government of Spain goes even further with censorship.FakeReal
74Right-hand drive cars will no longer be able to circulate on public roads, the senators votedFakeReal
80Romania, forced to organize the second round of the presidential election, according to the 2024 election calendarFakeReal
99Bolojan went personally to the hunger strikers to savor a hamburger.FakeReal
Table 6. Fake articles classified as real by BERT.
Table 6. Fake articles classified as real by BERT.
Crt. No.Title of the ArticleReal LabelAI Prediction
71The spring planting campaign begins. The SRI will plant 200,000 microphones in Romanians’ homesFakeReal
76The tax on the sun, a new villainy. And not the last oneFakeReal
95The desperation of a dinosaur, PSD with AUR teethFakeReal
96The Government will grant vouchers for the purchase of cherriesFakeReal
97The corn quotation grows after Simion and AUR started ordering popcornFakeReal
73The socialist government of Spain goes even further with censorship.FakeReal
74Right-hand drive cars will no longer be able to circulate on public roads, the senators votedFakeReal
80Romania, forced to organize the second round of the presidential election, according to the 2024 election calendarFakeReal
99Bolojan went personally to the hunger strikers to savor a hamburger.FakeReal
Table 7. Error categories.
Table 7. Error categories.
ModelTotal ErrorsSensationalist Vocabulary in Authentic JournalismPolitical Register Overlap in Authentic JournalismInstitutional Vocabulary in Fabricated ContentFailure to Identify Satirical Content
Naive Bayes2970184
KNN51272130
Random Forest2311471
XGBoost43241351
BERT2714445
TextGCN2813735
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Coman, C.; Dalban, C.M.; Bătrânu-Pințea, V.; Aron, G.; Marina, L. Comparative Performance of AI-Generated Fake News Detection Pipelines on Romanian News Content. Information 2026, 17, 698. https://doi.org/10.3390/info17070698

AMA Style

Coman C, Dalban CM, Bătrânu-Pințea V, Aron G, Marina L. Comparative Performance of AI-Generated Fake News Detection Pipelines on Romanian News Content. Information. 2026; 17(7):698. https://doi.org/10.3390/info17070698

Chicago/Turabian Style

Coman, Claudiu, Costel Marian Dalban, Vlad Bătrânu-Pințea, Georgiana Aron, and Lucian Marina. 2026. "Comparative Performance of AI-Generated Fake News Detection Pipelines on Romanian News Content" Information 17, no. 7: 698. https://doi.org/10.3390/info17070698

APA Style

Coman, C., Dalban, C. M., Bătrânu-Pințea, V., Aron, G., & Marina, L. (2026). Comparative Performance of AI-Generated Fake News Detection Pipelines on Romanian News Content. Information, 17(7), 698. https://doi.org/10.3390/info17070698

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop