Artificial Intelligence for Detecting Electoral Disinformation on Social Media: Models, Datasets, and Evaluation
Abstract
1. Introduction
2. Methodology
2.1. Information Sources, Search Fields, and Search Date
2.2. Search Strategy
- Group 1: AI/ML methods.(“artificial intelligence” OR “machine learning” OR “deep learning” OR “natural language processing” OR “text mining” OR transformer* OR BERT OR “large language model*” OR “generative ai” OR GPT* OR ChatGPT OR multimodal* OR “vision-language” OR “cross-modal” OR “graph neural network*” OR GNN OR “convolutional neural network*” OR CNN OR “support vector machine*” OR SVM OR “random forest” OR “gradient boosting” OR XGBoost)
- Group 2: Information disorder and manipulation.(disinformation OR misinformation OR “fake news” OR “false information” OR propaganda OR rumor* OR “information disorder” OR “computational propaganda” OR “influence operation*” OR “information manipulation” OR “coordinated inauthentic behavior” OR astroturf* OR deepfake* OR “manipulated media”)
- Group 3: Electoral context.(election* OR electoral OR campaign* OR vot* OR referendum* OR “electoral integrity” OR politic*)
- Group 4: Social media and platforms.(“social media” OR “social network*” OR microblog* OR “online platform*” OR Twitter OR X OR Facebook OR YouTube OR TikTok OR Instagram OR WhatsApp OR Telegram OR Reddit OR Weibo OR WeChat)
- Group 5: Detection, datasets, and evaluation.(detect* OR classif* OR identif* OR recogn* OR predict* OR “fact check*” OR “fact-check*” OR “claim verification” OR “stance detection” OR dataset* OR benchmark* OR corpus OR evaluation OR validation OR explainab* OR interpretab*)
2.3. Corpus Construction and Quality Control
2.4. Content Analysis of Highly Cited Papers
2.4.1. Selection Criteria
2.4.2. Extraction Template
2.4.3. Internal Consistency
2.4.4. Reporting in the Results
2.5. Bibliometric Indicators and Analytical Procedures
2.6. Keyword Strategy, Completeness Rationale, and Noise Reduction
Group 1: Generic lexical and document/population terms.deep, artificial, machine, big, natural, fake, intelligence, learning, information, communication, news, media, article, human, humans, naGroup 2: AI/ML umbrella terms, methodological labels, pipeline actions, evaluation terms, and data descriptors.model, models, algorithm, algorithms, analysis, processing, training, method, methods, methodology, framework, approach, approaches, technology, technologies, system, systems, dataset, datasets, data, detect, detection, classify, classification, identify, identification, recognize, recognition, predict, prediction, accuracy, performance, impact, evaluation, metric, metrics, artificial intelligence, ai, machine learning, machine-learning, machine learning (ml), deep learning, natural language processing, nlp, language processing (nlp), bibliometric analysis, surveys, data modelsGroup 3: Digital environment/content terms, phenomenon/context keywords, platforms, and topic-specific outlier.internet, online, platform, platforms, content, text, image, images, video, videos, network, networks, social, language, disinformation, misinformation, fake news, false information, election, elections, electoral, campaign, campaigns, voting, vote, referendum, referendums, political, politics, social media, social network, social networks, social networking, microblog, online platform, online platforms, twitter, x, facebook, youtube, tiktok, instagram, whatsapp, telegram, reddit, weibo, wechat, fake news detection, COVID-19, social networking (online)
3. Results
3.1. Annual Scientific Production
- Note that 2026 is treated as a partial-year count for annual scientific production, but as the cut-off year for citation-rate normalization (see Section 2.4.1).
3.2. Scientific Production by Country
3.3. Inter-Country Collaboration Patterns
3.4. Most Relevant Institutions
3.5. Most Relevant Sources
3.6. Most Relevant Articles
3.7. Word-Clouds and Keyword Co-Occurrence
3.8. Thematic Map
- In the Motor Themes quadrant, Figure 8 highlights a health and infodemic-adjacent cluster dominated by vaccine hesitancy, public health, pandemic, public opinion, coronavirus, and trust and accompanied by terms that emphasize population-scale information dynamics and perception, such as information dissemination, public perception, coronavirus disease 2019, and COVID-19 vaccination. In the same quadrant, a method-centric theme appears near the quadrant boundary, grouping widely used detection model families and learning paradigms—long short-term memory (LSTM), convolutional neural network, fake detection, support vector machine, and learning systems—together with Naïve Bayes, ensemble learning, and learning algorithms. This positioning indicates that these modelling choices are both well developed and structurally relevant within the keyword space of the corpus.
- The Basic Themes quadrant concentrates the most central, field-defining vocabulary of AI-enabled disinformation research on social media. A large foundational cluster is anchored by bots and feature extraction, and includes rumor detection, transformers, BERT, behavioral research, blogs, large language models, propaganda, and word embedding. A second basic cluster, positioned closer to the density axis, integrates content and discourse-oriented risk constructs with analytical tasks and techniques, including sentiment analysis, fact checking, polarization, generative ai, hate speech, topic modeling, health, social network analysis, text mining, and extremism. Notably, tf-idf appears near the central boundary, while logistic regression lies close to the quadrant intersection, indicating their recurrent but comparatively less consolidated thematic roles.
- In the Niche Themes quadrant, Figure 8 identifies a specialized, internally cohesive cluster around synthetic-media threats and integrity/security perspectives, grouping deepfake, deepfake detection, blockchain, social media platforms, cyber security, and digital forensics. A second niche cluster captures a compact set of concepts related to diffusion and influence dynamics—transfer learning, spread, and deception—suggesting a focused but less structurally central research pocket within the overall field.
- Finally, the Emerging or Declining Themes quadrant includes low-centrality and low-density themes that appear more isolated in the current conceptual landscape, notably arabic language and political communication. Their location indicates that, within this corpus, these topics are either nascent and not yet fully integrated into the broader thematic structure, or they represent lines of work whose relative prominence is decreasing.
4. Discussion
4.1. Main Growth Signals and Consolidation of the Publication Ecosystem
4.2. What “AI for Electoral Disinformation” Means in Practice: Heterogeneity of Roles
4.3. Conceptual Backbone from Keywords: Actors, Detection Pipelines, and Modern NLP
4.4. Socio-Political Harm Constructs as Central Detection Targets: Polarization, Hate Speech, and Extremism
4.5. Collateral but Structurally Important Themes: Pandemic-Driven Infodemic Research and Blockchain
4.6. Evaluation Practices and the Persistent Problem of External Validity
Recommended Minimum Reporting Checklist for Election-Relevant Evaluation
4.7. Implications for Future Research
- From detection to impact and exposure-aware evaluation.The corpus contains influential examples where the analytic target shifts from binary veracity to measurable societal impact conditional on exposure. Extending this orientation to electoral settings would require linking content-level predictions to platform-scale exposure proxies and to outcomes that operationalize socio-political harms, including polarization dynamics, hate speech prevalence, and extremism-related narratives, while maintaining transparent assumptions and privacy-preserving measurement. In this context, evaluation should move beyond aggregate performance and incorporate cost-sensitive validation, calibration, and Robustness to temporal and narrative shifts, given that election-time risks are highly context-dependent.
- Robustness under generative and multimodal threat regimes.Research on deepfakes and synthetic media detection constitutes a coherent niche that is likely to expand as multimodal generative systems mature and become more accessible. Research priorities include evaluating under distribution shifts, developing multimodal benchmarks, and developing provenance-aware mitigation strategies that remain valid as content-generation capabilities evolve.
- Broadening representativeness across regions, languages, and electoral contexts.The country and collaboration structure indicates uneven representation. Expanding multilingual and non-Western electoral datasets and enabling reproducible evaluation across heterogeneous contexts is essential to avoid a field whose empirical claims generalize primarily to data-rich environments. The latter is particularly important when harm constructs (such as hate speech or extremism) depend on local linguistic cues, legal definitions, and cultural context.
- Integrating governance constraints into technical evaluation.Highly cited conceptual work emphasizes that accountability, contestation, and trade-offs between error and other costs shape algorithmic moderation. Future detection research would benefit from evaluation protocols that explicitly reflect operational constraints (asymmetric costs of false negatives during election windows, appeal and auditing mechanisms, and transparency requirements) rather than relying only on aggregate performance metrics.
4.8. Limitations of This Review
5. Conclusions
Supplementary Materials
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- OECD. Facts Not Fakes: Tackling Disinformation, Strengthening Information Integrity; OECD Publishing: Paris, France, 2024. [Google Scholar] [CrossRef]
- Center for an Informed Public; Digital Forensic Research Lab; Graphika; Stanford Internet Observatory. The Long Fuse: Misinformation and the 2020 Election. In Stanford Digital Repository: Election Integrity Partnership; v1.3.0.; Stanford University: Stanford, CA, USA, 2021. [Google Scholar] [CrossRef]
- Pacheco, D.; Hui, P.-M.; Torres-Lugo, C.; Truong, B.T.; Flammini, A.; Menczer, F. Uncovering Coordinated Networks on Social Media: Methods and Case Studies. Proc. Int. AAAI Conf. Web Soc. Media 2021, 15, 455–466. [Google Scholar] [CrossRef]
- Caled, D.; Silva, M.J. Digital media and misinformation: An outlook on multidisciplinary strategies against manipulation. J. Comput. Soc. Sci. 2022, 5, 123–159. [Google Scholar] [CrossRef] [PubMed]
- Cinus, F.; Minici, M.; Luceri, L.; Ferrara, E. Exposing Cross-Platform Coordinated Inauthentic Activity in the Run-Up to the 2024 U.S. Election. In Proceedings of the ACM on Web Conference 2025 (WWW ’25); Association for Computing Machinery: New York, NY, USA, 2025; pp. 541–559. [Google Scholar] [CrossRef]
- Nielsen, R.K.; Fletcher, R. Democratic Creative Destruction? The Effect of a Changing Media Landscape on Democracy. In Social Media and Democracy: The State of the Field, Prospects for Reform; Persily, N., Tucker, J.A., Eds.; SSRC Anxieties of Democracy; Cambridge University Press: Cambridge, UK; New York, NY, USA, 2020; pp. 139–162. [Google Scholar] [CrossRef]
- Vasist, P.N.; Chatterjee, D.; Krishnan, S. The Polarizing Impact of Political Disinformation and Hate Speech: A Cross-country Configural Narrative. Inf. Syst. Front. 2024, 26, 663–688. [Google Scholar] [CrossRef] [PubMed]
- Margetts, H. Rethinking Democracy with Social Media. Political Q. 2019, 90, 107–123. [Google Scholar] [CrossRef]
- Augenstein, I.; Baldwin, T.; Cha, M.; Chakraborty, T.; Ciampaglia, G.L.; Corney, D.; DiResta, R.; Ferrara, E.; Hale, S.; Halevy, A.; et al. Factuality challenges in the era of large language models and opportunities for fact-checking. Nat. Mach. Intell. 2024, 6, 852–863. [Google Scholar] [CrossRef]
- Montoro-Montarroso, A.; Cantón-Correa, J.; Rosso, P.; Chulvi, B.; Panizo-Lledot, Á.; Huertas-Tato, J.; Calvo-Figueras, B.; Rementeria, M.J.; Gómez-Romero, J. Fighting disinformation with artificial intelligence: Fundamentals, advances and challenges. Prof. Inf. 2023, 32, e320322. [Google Scholar] [CrossRef]
- Hackenburg, K.; Margetts, H. Evaluating the persuasive influence of political microtargeting with large language models. Proc. Natl. Acad. Sci. USA 2024, 121, e2403116121. [Google Scholar] [CrossRef]
- Park, C.Y.; Mendelsohn, J.; Field, A.; Tsvetkov, Y. Challenges and Opportunities in Information Manipulation Detection: An Examination of Wartime Russian Media. In Findings of the Association for Computational Linguistics: EMNLP 2022; Association for Computational Linguistics: Abu Dhabi, United Arab Emirates, 2022; pp. 5209–5235. [Google Scholar] [CrossRef]
- Cordoba-Cabús, A.; Casero-Ripollés, A.; Alonso-Muñoz, L.; Tirado-García, A. Disinformation and emotions in the 2024 European Parliament and US presidential elections. Soc. Sci. Humanit. Open 2026, 13, 102502. [Google Scholar] [CrossRef]
- Bhoi, S.; Jha, V.K.; Kumar, R. Information Disorder in Democracy: A Thematic Analysis of Mis/Disinformation in the 2024 Indian General Election. Int. Inf. Libr. Rev. 2026, 1–17. [Google Scholar] [CrossRef]
- Syafhendry, S.; Ganaie, N.A.; Yama, A. Smart elections or rigged algorithms: The rise of artificial intelligence in electoral governance in Southeast Asia. Front. Political Sci. 2025, 7, 1672310. [Google Scholar] [CrossRef]
- Suherman, A.; Hidayatullah, M.; Sa’ban, L.O.M.A.; Mulia, F. Information Disorder and Electoral Integrity: The Impact of Disinformation on Elections in Developing Democracies. In Elections in the Global South and North: A Comprehensive Overview; Aluko, O.I., Ed.; IntechOpen: London, UK, 2026; Chapter 2. [Google Scholar] [CrossRef]
- Marsden, C.; Meyer, T.; Brown, I. Platform values and democratic elections: How can the law regulate digital disinformation? Comput. Law Secur. Rev. Int. J. Technol. Law Pract. 2019, 36, 105373. [Google Scholar] [CrossRef]
- López-López, P.C.; Barredo-Ibáñez, D.; Jaráiz-Gulías, E. Research on Digital Political Communication: Electoral Campaigns, Disinformation, and Artificial Intelligence. Societies 2023, 13, 126. [Google Scholar] [CrossRef]
- Islam, M.B.E.; Haseeb, M.; Batool, H.; Ahtasham, N.; Muhammad, Z. AI Threats to Politics, Elections, and Democracy: A Blockchain-Based Deepfake Authenticity Verification Framework. Blockchains 2024, 2, 458–481. [Google Scholar] [CrossRef]
- López-Borrull, A.; Lopezosa, C. Mapping the Impact of Generative AI on Disinformation: Insights from a Scoping Review. Publications 2025, 13, 33. [Google Scholar] [CrossRef]
- Ibrahim, N.T.; Attia, N.A. The impact of disinformation generated by AI on democracy case studies: The US presidential elections in 2016 & 2024. Rev. Econ. Political Sci. 2025; in press. [Google Scholar] [CrossRef]
- Lipińska, M. Research Methods of the Impact of AI on Elections–Systematic Review. In AI for People, Democratizing AI, Proceedings of the Second EAI International Conference, CAIP 2023, Bologna, Italy, 24–26 November 2023; Ziosi, M., Sartor, G., Cunha, J.M., Trotta, A., Wicke, P., Eds.; Lecture Notes of the Institute for Computer Sciences, Social Informatics and Telecommunications Engineering; Springer Nature AG: Cham, Switzerland, 2024; Volume 591, pp. 63–70. [Google Scholar] [CrossRef]
- Aria, M.; Cuccurullo, C. bibliometrix: An R-tool for comprehensive science mapping analysis. J. Inf. 2017, 11, 959–975. [Google Scholar] [CrossRef]
- Gorwa, R.; Binns, R.; Katzenbach, C. Algorithmic content moderation: Technical and political challenges in the automation of platform governance. Big Data Soc. 2020, 7, 2053951719897945. [Google Scholar] [CrossRef]
- Vaccari, C.; Chadwick, A. Deepfakes and Disinformation: Exploring the Impact of Synthetic Political Video on Deception, Uncertainty, and Trust in News. Soc. Media Soc. 2020, 6, 2056305120903408. [Google Scholar] [CrossRef]
- Ferrara, E. Disinformation and Social Bot Operations in the Run Up to the 2017 French Presidential Election. First Monday 2017, 22, 8. [Google Scholar] [CrossRef]
- Sahoo, S.R.; Gupta, B.B. Multiple features based approach for automatic fake news detection on social networks using deep learning. Appl. Soft Comput. J. 2021, 100, 106983. [Google Scholar] [CrossRef]
- Hussain, A.; Tahir, A.; Hussain, Z.; Sheikh, Z.; Gogate, M.; Dashtipour, K.; Ali, A.; Sheikh, A. Artificial Intelligence–Enabled Analysis of Public Attitudes on Facebook and Twitter Toward COVID-19 Vaccines in the United Kingdom and the United States: Observational Study. J. Med. Internet Res. 2021, 23, e26627. [Google Scholar] [CrossRef] [PubMed]
- Shao, C.; Hui, P.-M.; Wang, L.; Jiang, X.; Flammini, A.; Menczer, F.; Ciampaglia, G.L. Anatomy of an online misinformation network. PLoS ONE 2018, 13, e0196087. [Google Scholar] [CrossRef] [PubMed]
- ALDayel, A.; Magdy, W. Stance Detection on Social Media: State of the Art and Trends. Inf. Process. Manag. 2021, 58, 102597. [Google Scholar] [CrossRef]
- Karnouskos, S. Artificial Intelligence in Digital Media: The Era of Deepfakes. IEEE Trans. Technol. Soc. 2020, 1, 138–147. [Google Scholar] [CrossRef]
- Verma, P.K.; Agrawal, P.; Amorim, I.; Prodan, R. WELFake: Word Embedding Over Linguistic Features for Fake News Detection. IEEE Trans. Comput. Soc. Syst. 2021, 8, 881–893. [Google Scholar] [CrossRef]
- Umer, M.; Imtiaz, Z.; Ullah, S.; Mehmood, A.; Choi, G.S.; On, B.-W. Fake News Stance Detection Using Deep Learning Architecture (CNN-LSTM). IEEE Access 2020, 8, 156695–156706. [Google Scholar] [CrossRef]
- Choudhary, A.; Arora, A. Linguistic Feature Based Learning Model for Fake News Detection and Classification. Expert Syst. Appl. 2021, 169, 114171. [Google Scholar] [CrossRef]
- Guo, B.; Ding, Y.; Yao, L.; Liang, Y.; Yu, Z. The Future of False Information Detection on Social Media: New Perspectives and Trends. ACM Comput. Surv. 2020, 53, 68. [Google Scholar] [CrossRef]
- Ferrara, E. GenAI against humanity: Nefarious applications of generative artificial intelligence and large language models. J. Comput. Soc. Sci. 2024, 7, 549–569. [Google Scholar] [CrossRef]
- Allen, J.; Watts, D.J.; Rand, D.G. Quantifying the impact of misinformation and vaccine-skeptical content on Facebook. Science 2024, 384, eadk3451. [Google Scholar] [CrossRef]








| Affiliation(s) | Articles/Affiliation |
|---|---|
| King Saud University | 23 |
| University of Florida | 15 |
| Symbiosis Institute of Technology | 14 |
| Delhi Technological University, University of Southern California | 13 |
| Carnegie Mellon University, Universiti Kebangsaan Malaysia | 12 |
| Bucharest University of Economic Studies | 11 |
| Arizona State University, Northwestern Polytechnical University, The University of Texas at Austin, University of Electronic Science and Technology of China, University of Münster | 9 |
| National Institute of Technology Hamirpur, Northwestern University, Princess Nourah bint Abdulrahman University, University of Arkansas at Little Rock, University of California, San Diego | 8 |
| K L (Deemed to be University), Lovely Professional University, Simon Fraser University, Tsinghua University, University of California, Berkeley, University of North Carolina | 7 |
| Indiana University Bloomington, Institut Català de la Salut, Kangwon National University, Mansoura University, Monash University, Nanyang Technological University, Sejong University, University of Cádiz, University College Dublin, Complutense University of Madrid | 6 |
| Source(s) | Articles/Journal |
|---|---|
| IEEE Access (Q1) | 32 |
| Expert Systems with Applications (Q1), Social Network Analysis and Mining (Q1) | 14 |
| Multimedia Tools and Applications (Q1) | 13 |
| IEEE Transactions on Computational Social Systems (Q1) | 12 |
| Journal of Medical Internet Research (Q1), PLOS ONE (Q1) | 10 |
| Applied Sciences (Switzerland) (Q2), JMIR Infodemiology (Q2), Social Media + Society (Q1) | 8 |
| Scientific Reports (Q1) | 7 |
| Applied Soft Computing (Q1), Mathematics (Q2) | 6 |
| CMC-Computers, Materials & Continua (Q2), Journal of Computational Social Science (Q2) | 5 |
| ACM Transactions on Asian and Low-Resource Language Information Processing (Q2), Computational and Mathematical Organization Theory (Q2), EPJ Data Science (Q1), IEEE Transactions on Knowledge and Data Engineering (Q1), Information (Switzerland) (Q2), Information Processing & Management (Q1), International Journal of Advanced Computer Science and Applications (Q3), International Journal of Intelligent Systems and Applications in Engineering (Na), Multimedia Systems (Q1), Proceedings of the ACM on Human-Computer Interaction (Q1) | 4 |
| Title | Year | Citations | Citations/Year | Study |
|---|---|---|---|---|
| Algorithmic content moderation: Technical and political challenges in the automation of platform governance | 2020 | 481 | 68.71 | [24] |
| Deepfakes and Disinformation: Exploring the Impact of Synthetic Political Video on Deception, Uncertainty, and Trust in News | 2020 | 442 | 63.14 | [25] |
| Disinformation and Social Bot Operations in the Run Up to the 2017 French Presidential Election | 2017 | 288 | 28.80 | [26] |
| Multiple features based approach for automatic fake news detection on social networks using deep learning | 2021 | 249 | 41.50 | [27] |
| Artificial Intelligence–Enabled Analysis of Public Attitudes on Facebook and Twitter Toward COVID-19 Vaccines in the United Kingdom and the United States: Observational Study | 2021 | 228 | 38.00 | [28] |
| Anatomy of an online misinformation network | 2018 | 215 | 23.89 | [29] |
| Stance detection on social media: State of the art and trends | 2021 | 166 | 27.67 | [30] |
| Artificial Intelligence in Digital Media: The Era of Deepfakes | 2020 | 163 | 23.29 | [31] |
| WELFake: Word Embedding Over Linguistic Features for Fake News Detection | 2021 | 156 | 26.00 | [32] |
| Fake News Stance Detection Using Deep Learning Architecture (CNN-LSTM) | 2020 | 135 | 19.29 | [33] |
| Linguistic feature based learning model for fake news detection and classification | 2021 | 126 | 21.00 | [34] |
| The Future of False Information Detection on Social Media: New Perspectives and Trends | 2020 | 116 | 16.57 | [35] |
| GenAI against humanity: nefarious applications of generative artificial intelligence and large language models | 2024 | 106 | 35.33 | [36] |
| Quantifying the impact of misinformation and vaccine-skeptical content on Facebook | 2024 | 73 | 24.33 | [37] |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Díaz, F.; Cerna, N.; Liza, R.; Motta, B. Artificial Intelligence for Detecting Electoral Disinformation on Social Media: Models, Datasets, and Evaluation. Information 2026, 17, 292. https://doi.org/10.3390/info17030292
Díaz F, Cerna N, Liza R, Motta B. Artificial Intelligence for Detecting Electoral Disinformation on Social Media: Models, Datasets, and Evaluation. Information. 2026; 17(3):292. https://doi.org/10.3390/info17030292
Chicago/Turabian StyleDíaz, Félix, Nhell Cerna, Rafael Liza, and Bryan Motta. 2026. "Artificial Intelligence for Detecting Electoral Disinformation on Social Media: Models, Datasets, and Evaluation" Information 17, no. 3: 292. https://doi.org/10.3390/info17030292
APA StyleDíaz, F., Cerna, N., Liza, R., & Motta, B. (2026). Artificial Intelligence for Detecting Electoral Disinformation on Social Media: Models, Datasets, and Evaluation. Information, 17(3), 292. https://doi.org/10.3390/info17030292

