Attention-Level Causal Intervention Framework for Multimodal Fake News Detection
Abstract
1. Introduction
- We formulate multimodal fake news detection from an attention-level causal perspective, treating attention representations as mediators within attention-dependent prediction pathways. This formulation provides a basis for applying front-door intervention to mitigate the influence of partially observed and unobserved confounding factors;
- We propose an Attention-level Causal Intervention Framework (ACIM) with a Causal Attention Layer Module (CALM). CALM jointly models in-sample and cross-sample attention through ISA and CSA and is integrated into textual encoding, visual encoding, and cross-modal fusion to reduce sample-specific spurious correlations throughout multimodal representation learning;
- We instantiate ACIM using BERT and Swin Transformer encoders and conduct extensive experiments on the Twitter and PHEME datasets, including baseline comparisons, ablation studies, sensitivity analyses, and case studies. The experimental findings provide empirical support for the performance and stability of the proposed framework.
2. Related Works
3. Method
3.1. Overview
3.2. Causal Attention Layer Module (CALM)
- Selector: selects appropriate mediating variables from ;
- Classifier: utilizes to predict .
3.3. Causal Intervention in Transformer Encoders
3.3.1. Textual Feature Extraction
3.3.2. Visual Feature Extraction
3.3.3. Causal-Aware Multimodal Fusion
- Text feature sequence ;
- Image feature sequence .
- where B is the batch size, and denote the number of text tokens and image patches respectively, and is the hidden dimension.
3.4. Loss Function
4. Experiments and Results
4.1. Experimental Setup
4.1.1. Datasets
4.1.2. Experimental Settings
4.1.3. Evaluation Metrics
- TP (True Positive): the number of samples whose true label is positive and are correctly predicted as positive.
- FP (False Positive): the number of samples whose true label is negative but are incorrectly predicted as positive.
- FN (False Negative): the number of samples whose true label is positive but are incorrectly predicted as negative.
- TN (True Negative): the number of samples whose true label is negative and are correctly predicted as negative.
4.2. Comparison with Baselines
- GRU [18]: An RNN-based method that uses a multi-layer GRU network to model posts as variable-length time series for hidden representation learning in fake news detection.
- CNN [19]: Extracts feature representations via convolutional neural networks, converting related posts into fixed-length sequences for fake news identification and early detection.
- SAFE [26]: A similarity-aware fusion model that extracts independent text and image embeddings, computes cross-modal similarity features, and concatenates them for fake news prediction.
- EANN [25]: A multi-task adversarial network that learns shared text and image representations through an auxiliary event-adversarial objective.
- TextGCN [80]: Models the entire corpus as a heterogeneous graph using graph convolutional networks to jointly learn richer word and document embeddings for fake news detection.
- MVAE [6]: A multimodal variational autoencoder that learns a shared latent space and simultaneously reconstructs both text and image features for downstream classification tasks.
- SpotFake [81]: Extracts text features using a pretrained language model (BERT) and image features using VGG-19 pretrained on ImageNet for multimodal fake news detection.
- HMCAN [79]: Jointly models multimodal contextual information and hierarchical textual semantics using BERT and ResNet, and fuses them through a multimodal contextual attention network for fake news detection.
- CSFND [82]: A context-sensitive framework that incorporates contextual semantic information to resolve inconsistencies between semantic and decision spaces, enabling clearer boundary delineation under different news contexts.
- BMR [83]: A multi-view representation bootstrapping method that improves fake news detection performance through multi-view feature extraction, an enhanced fusion mechanism, and an adaptive weighting strategy.
- NLIN [84]: Unifies multimodal inputs into a text space and employs an encoder–decoder with a prompt-based reasoning architecture to address the modality disparity issue in multimodal fake news detection.
- MVML [85]: Simultaneously models fine-grained semantic relationships between text–scene and text–object pairs, leverages graph attention for intra-modal reasoning, applies cross-view attention for multi-perspective fusion, and utilizes a mutual learning mechanism to enhance classifier performance.
- CCD [14]: A causal debiasing framework that employs causal intervention to mitigate psycholinguistic confounding and uses counterfactual reasoning to reduce the image-only bias in multimodal fake news detection.
4.3. Ablation Study
- Base (w/o causal): The baseline model without any causal intervention, i.e., standard multimodal feature concatenation followed by a classifier.
- w/o V: Removes the CALM module from the visual branch, retaining causal modeling only in the text branch and fusion layer.
- w/o L: Removes the CALM module from the text branch, retaining causal modeling only in the visual branch and fusion layer.
- w/o F: Removes the CALM module from the fusion layer, retaining causal modeling only in the text and visual branches.
- ACIM: The full model with CALM modules integrated across all three stages: text, image, and fusion.
4.4. Case Study
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Fu, Q. Analysis of Industrial and Economic News Reporting. J. Media Stud. 2025, 2, 39–41. [Google Scholar]
- Statista. Most Popular Social Networks Worldwide as of January 2024, Ranked by Number of Monthly Active Users. Available online: https://www.statista.com/statistics/272014/global-social-networks-ranked-by-number-of-users/ (accessed on 1 September 2026).
- Zhou, X.; Zafarani, R. A Survey of Fake News: Fundamental Theories, Detection Methods, and Opportunities. ACM Comput. Surv. 2020, 53, 1–40. [Google Scholar]
- Zhang, J.; Wei, B.; Song, P.J. Social Media Fake News Detection Model under Cognitive Domain Operations. Command Control Simul. 2025, 47, 72–78. [Google Scholar] [CrossRef]
- Vosoughi, S.; Roy, D.; Aral, S. The Spread of True and False News Online. Science 2018, 359, 1146–1151. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Khattar, D.; Goud, J.S.; Gupta, M.; Varma, V. MVAE: Multimodal Variational Autoencoder for Fake News Detection. In Proceedings of the World Wide Web Conference (WWW), San Francisco, CA, USA, 13–17 May 2019; pp. 2915–2921. [Google Scholar]
- Pulido, C.M.; Villarejo-Carballido, B.; Redondo-Sama, G.; Gomez, A. COVID-19 Infodemic: More Retweets for Science-Based Information on Coronavirus than for False Information. Int. Sociol. 2020, 35, 377–392. [Google Scholar] [CrossRef] [Scilit]
- Geirhos, R.; Jacobsen, J.; Michaelis, C. Shortcut Learning in Deep Neural Networks. Nat. Mach. Intell. 2020, 2, 665–673. [Google Scholar] [CrossRef] [Scilit]
- Pearl, J. Causality: Models, Reasoning and Inference; Cambridge University Press: Cambridge, UK, 2009. [Google Scholar]
- Hu, L.; Wei, S.; Zhao, Z.; Wu, B. MMCAN: Multi-Modal Co-Attention Network for Fact Checking via Explainability. IEEE Trans. Knowl. Data Eng. 2024, 36, 337–350. [Google Scholar]
- Saioni, M.; Giannone, C. Multimodal Attention is all you need. In Proceedings of the Tenth Italian Conference on Computational Linguistics (CLiC-it 2024), Pisa, Italy, 4–6 December 2024; pp. 873–879. [Google Scholar]
- Wu, F.; Jin, H.; Hu, C.; Ji, Y.; Jing, X.-Y.; Jiang, G.-P. Efficient cross-modal prompt learning with semantic enhancement for domain-robust fake news detection. In Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, United Arab Emirates, 19–24 January 2025; pp. 4175–4185. [Google Scholar]
- Hu, L.; Chen, Z.; Zhao, Z.; Yin, J.; Nie, L. Causal inference for leveraging image-text matching bias in multi-modal fake news detection. IEEE Trans. Knowl. Data Eng. 2022, 35, 11141–11152. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Hu, L.; Li, W.; Shao, Y.; Nie, L. Causal Intervention and Counterfactual Reasoning for Multi-modal Fake News Detection. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL), Toronto, ON, Canada, 9–14 July 2023; pp. 627–638. [Google Scholar]
- Yu, J.; Wang, S.; Yin, H.; Sun, Z.; Xie, R.; Zhang, B.; Rao, Y. Multimodal clickbait detection by de-confounding biases using causal representation inference. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Miami, FL, USA, 12–16 November 2024; pp. 10300–10317. [Google Scholar]
- Liu, Q.; Wu, J.; Wu, S.; Wang, L. Out-of-Distribution Evidence-Aware Fake News Detection via Dual Adversarial Debiasing. IEEE Trans. Knowl. Data Eng. 2024, 36, 6801–6813. [Google Scholar] [CrossRef] [Scilit]
- Castillo, C.; Mendoza, M.; Poblete, B. Information Credibility on Twitter. In Proceedings of the 20th International Conference on World Wide Web, Hyderabad, India, 28 March–1 April 2011; pp. 675–684. [Google Scholar]
- Ma, J.; Gao, W.; Mitra, P.; Kwon, S.; Jansen, B.J.; Wong, K.F.; Cha, M. Detecting Rumors from Microblogs with Recurrent Neural Networks. In Proceedings of the 25th International Joint Conference on Artificial Intelligence (IJCAI), New York, NY, USA, 9–15 July 2016; pp. 3818–3824. [Google Scholar]
- Yu, F.; Liu, Q.; Wu, S.; Wang, L.; Tan, T. A Convolutional Approach for Misinformation Identification. In Proceedings of the 26th International Joint Conference on Artificial Intelligence (IJCAI), Melbourne, Australia, 19–25 August 2017; pp. 901–907. [Google Scholar]
- Kaliyar, R.K.; Goswami, A.; Narang, P.; Sinha, S. FakeBERT: Fake News Detection in Social Media with a BERT-based Deep Learning Approach. Multimed. Tools Appl. 2021, 80, 11765–11788. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Verdoliva, L. Media Forensics and DeepFakes: An Overview. IEEE J. Sel. Top. Signal Process. 2020, 14, 910–932. [Google Scholar] [CrossRef] [Scilit]
- Qi, P.; Cao, J.; Yang, T.; Guo, J.; Li, J. Exploiting Multi-Domain Visual Information for Fake News Detection. In Proceedings of the IEEE International Conference on Data Mining (ICDM), Beijing, China, 8–11 November 2019; pp. 518–527. [Google Scholar]
- Jin, Z.; Cao, J.; Zhang, Y.; Luo, J. News Verification by Exploiting Conflicting Social Viewpoints in Microblogs. In Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016; pp. 2972–2978. [Google Scholar]
- Jin, Z.; Cao, J.; Guo, H.; Zhang, Y.; Luo, J. Multimodal Fusion with Recurrent Neural Networks for Rumor Detection on Microblogs. In Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, CA, USA, 23–27 October 2017; pp. 795–816. [Google Scholar]
- Wang, Y.; Ma, F.; Jin, Z.; Yuan, Y.; Xun, G.; Jha, K.; Su, L.; Gao, J. EANN: Event Adversarial Neural Networks for Multi-Modal Fake News Detection. In Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, London, UK, 19–23 August 2018; pp. 849–857. [Google Scholar]
- Zhou, X.; Wu, J.; Zafarani, R. SAFE: Similarity-Aware Multi-Modal Fake News Detection. In Proceedings of the Pacific-Asia Conference on Knowledge Discovery and Data Mining (PAKDD), Singapore, 11–14 May 2020; pp. 354–367. [Google Scholar]
- Jing, J.; Wu, H.; Sun, J.; Fang, X.; Zhang, H. Multimodal Fake News Detection via Progressive Fusion Networks. Inf. Process. Manag. 2023, 60, 103120. [Google Scholar] [CrossRef] [Scilit]
- Zhang, H.; Fang, Q.; Qian, S.; Xu, C. Multi-Modal Knowledge-Aware Event Memory Network for Social Media Rumor Detection. In Proceedings of the 27th ACM International Conference on Multimedia, Nice, France, 21–25 October 2019; pp. 1942–1951. [Google Scholar]
- Monti, F.; Frasca, F.; Eynard, D.; Mannion, D.; Bronstein, M.M. Fake News Detection on Social Media using Geometric Deep Learning. In Proceedings of the ICLR Workshop, New Orleans, LA, USA, 6–9 May 2019. [Google Scholar]
- Alam, F.; Cresci, S.; Chakraborty, T.; Silvestri, F.; Dimitrov, D.; Da San Martino, G.; Shaar, S.; Firooz, H.; Nakov, P. A Survey on Multimodal Disinformation Detection. In Proceedings of the 29th International Conference on Computational Linguistics (COLING), Gyeongju, Republic of Korea, 12–17 October 2022; pp. 6625–6643. [Google Scholar]
- Comito, C.; Caroprese, L.; Zumpano, E. Multimodal Fake News Detection with Deep Learning: A Survey. Information 2023, 14, 98. [Google Scholar] [CrossRef] [Scilit]
- Nasser, M.; Arshad, N.I.; Ali, A.; Alhussian, H.; Saeed, F.; Da’u, A.; Nafea, I. A systematic review of multimodal fake news detection on social media using deep learning models. Results Eng. 2025, 26, 104752. [Google Scholar] [CrossRef] [Scilit]
- Kou, F.; Wang, B.; Li, H.; Zhu, C.; Shi, L.; Zhang, J.; Qi, L. Potential Features Fusion Network for Multimodal Fake News Detection. ACM Trans. Multimed. Comput. Commun. Appl. 2025, 21, 87. [Google Scholar] [CrossRef] [Scilit]
- Shen, L.; Long, Y.; Cai, X.; Razzak, I.; Chen, G.; Liu, K.; Jameel, S. GAMED: Knowledge Adaptive Multi-Experts Decoupling for Multimodal Fake News Detection. In Proceedings of the 18th ACM International Conference on Web Search and Data Mining (WSDM), Hannover, Germany, 10–14 March 2025. [Google Scholar]
- Ma, Z.; Luo, M.; Guo, H.; Zeng, Z.; Hao, Y.; Zhao, X. Event-radar: Event-driven multi-view learning for multimodal fake news detection. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Bangkok, Thailand, 11–16 August 2024; pp. 5809–5821. [Google Scholar]
- Zhang, Y.; Ma, J.; Jia, Y. MCAN: Multimodal cross-aware network for fake news detection by extracting semantic-physical feature consistency. J. Supercomput. 2025, 81, 299. [Google Scholar] [CrossRef] [Scilit]
- Qian, M.; Guo, S.; Chen, Y.; Wang, C.; Ye, Z. MFND-DCL: Multimodal fake news detection based on dual contrastive learning. Inf. Process. Manag. 2026, 63, 104812. [Google Scholar] [CrossRef] [Scilit]
- Radford, A.; Kim, J.W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al. Learning Transferable Visual Models From Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning (ICML), Online, 18–24 July 2021; pp. 8748–8763. [Google Scholar]
- Li, J.; Li, D.; Xiong, C.; Hoi, S. BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation. In Proceedings of the 39th International Conference on Machine Learning (ICML), Baltimore, MD, USA, 17–23 July 2022; pp. 12888–12900. [Google Scholar]
- Zhang, L.; Zhang, X.; Zhou, Z.; Zhang, X.; Yu, P.S.; Li, C. Knowledge-aware multimodal pre-training for fake news detection. Inf. Fusion 2025, 114, 102715. [Google Scholar] [CrossRef] [Scilit]
- Zheng, X.; Luo, M.; Wang, X. Unveiling fake news with adversarial arguments generated by multimodal large language models. In Proceedings of the 31st International Conference on Computational Linguistics, Abu Dhabi, United Arab Emirates, 19–24 January 2025; pp. 7862–7869. [Google Scholar]
- Hu, S.; Hu, J.; Zhang, H. Synergizing llms with global label propagation for multimodal fake news detection. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Vienna, Austria, 27 July–1 August 2025; pp. 1426–1440. [Google Scholar]
- Peters, J.; Bühlmann, P.; Meinshausen, N. Causal Inference by Using Invariant Prediction: Identification and Confidence Intervals. J. R. Stat. Soc. Ser. B 2016, 78, 947–1012. [Google Scholar] [CrossRef] [Scilit]
- Arjovsky, M.; Bottou, L.; Gulrajani, I.; Lopez-Paz, D. Invariant Risk Minimization. arXiv 2019, arXiv:1907.02893. [Google Scholar]
- Kaushik, D.; Hovy, E.; Lipton, Z. Learning the Difference that Makes a Difference with Counterfactually-Augmented Data. In Proceedings of the 8th International Conference on Learning Representations (ICLR), Online, 26 April–1 May 2020. [Google Scholar]
- Niu, Y.; Tang, K.; Zhang, H.; Lu, Z.; Hua, X.S.; Wen, J.R. Counterfactual VQA: A Cause-Effect Look at Language Bias. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Online, 19–25 June 2021; pp. 12700–12710. [Google Scholar]
- Tang, K.; Niu, Y.; Huang, J.; Shi, J.; Zhang, H. Unbiased Scene Graph Generation from Biased Training. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Online, 14–19 June 2020; pp. 3716–3725. [Google Scholar]
- Yang, X.; Zhang, H.; Cai, J. Deconfounded Image Captioning: A Causal Retrospect. IEEE Trans. Pattern Anal. Mach. Intell. 2023, 45, 14311–14326. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Cheng, L.; Guo, R.; Shu, K.; Liu, H. Causal Understanding of Fake News Dissemination on Social Media. In Proceedings of the 27th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, Online, 14–18 August 2021; pp. 148–157. [Google Scholar]
- Sun, M.; Zhang, X.; Ma, J.; Xie, S.; Liu, Y.; Yu, P.S. Inconsistent Matters: A Knowledge-Guided Dual-Consistency Network for Multi-Modal Rumor Detection. IEEE Trans. Knowl. Data Eng. 2023, 35, 12736–12749. [Google Scholar] [CrossRef] [Scilit]
- Anderson, P.; He, X.; Buehler, C.; Teney, D.; Johnson, M.; Gould, S.; Zhang, L. Bottom-Up and Top-Down Attention for Image Captioning and Visual Question Answering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 6077–6086. [Google Scholar]
- Bahdanau, D.; Cho, K.; Bengio, Y. Neural Machine Translation by Jointly Learning to Align and Translate. In Proceedings of the International Conference on Learning Representations (ICLR), San Diego, CA, USA, 7–9 May 2015. [Google Scholar]
- Herdade, S.; Kappeler, A.; Boakye, K.; Soares, J. Image Captioning: Transforming Objects into Words. Adv. Neural Inf. Process. Syst. 2019, 32, 11137–11147. [Google Scholar]
- Luo, R.; Price, B.; Cohen, S.; Shakhnarovich, G. Discriminability Objective for Training Descriptive Captions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 6964–6974. [Google Scholar]
- Carion, N.; Massa, F.; Synnaeve, G.; Usunier, N.; Kirillov, A.; Zagoruyko, S. End-to-end object detection with transformers. In Proceedings of the European Conference on Computer Vision, Online, 23–28 August 2020; pp. 213–229. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. Adv. Neural Inf. Process. Syst. 2017, 30, 5998–6008. [Google Scholar]
- Pearl, J. The Mediation Formula: A Guide to the Assessment of Causal Pathways in Nonlinear Models. Prev. Sci. 2012, 13, 426–436. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Kobayashi, G.; Kuribayashi, T.; Yokoi, S.; Inui, K. Attention is not only a weight: Analyzing transformers with vector norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online, 16–20 November 2020; pp. 7057–7075. [Google Scholar] [CrossRef] [Scilit]
- Reed, W.J. The Pareto, Zipf and Other Power Laws. Econ. Lett. 2001, 74, 15–19. [Google Scholar] [CrossRef] [Scilit]
- Hendricks, L.A.; Burns, K.; Saenko, K.; Darrell, T.; Rohrbach, A. Women Also Snowboard: Overcoming Bias in Captioning Models. In Proceedings of the European Conference on Computer Vision (ECCV), Munich, Germany, 8–14 September 2018; pp. 793–811. [Google Scholar]
- Wang, X.; Girshick, R.; Gupta, A.; He, K. Non-Local Neural Networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, UT, USA, 18–22 June 2018; pp. 7794–7803. [Google Scholar]
- Battaglia, P.W.; Hamrick, J.B.; Bapst, V.; Sanchez-Gonzalez, A.; Zambaldi, V.; Malinowski, M.; Tacchetti, A.; Raposo, D.; Santoro, A.; Faulkner, R. Relational inductive biases, deep learning, and graph networks. arXiv 2018, arXiv:1806.01261. [Google Scholar]
- Chen, M.; Radford, A.; Child, R.; Wu, J.; Jun, H.; Dhariwal, P.; Luan, D.; Sutskever, I. Generative Pretraining from Pixels. In Proceedings of the 37th International Conference on Machine Learning (ICML), Online, 12–18 July 2020; pp. 1–10. [Google Scholar]
- Lu, J.; Batra, D.; Parikh, D.; Lee, S. ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks. Adv. Neural Inf. Process. Syst. 2019, 32, 13–23. [Google Scholar]
- Pearl, J. Causal Diagrams for Empirical Research. Biometrika 1995, 82, 669–688. [Google Scholar] [CrossRef] [Scilit]
- Rubin, D.B. Causal Inference Using Potential Outcomes: Design, Modeling, Decisions. J. Am. Stat. Assoc. 2005, 100, 322–331. [Google Scholar]
- Peters, J.; Janzing, D.; Schölkopf, B. Elements of Causal Inference: Foundations and Learning Algorithms; MIT Press: Cambridge, MA, USA, 2017. [Google Scholar]
- Zhang, D.; Zhang, H.; Tang, J.; Hua, X.-S.; Sun, Q. Causal intervention for weakly-supervised semantic segmentation. Adv. Neural Inf. Process. Syst. 2020, 33, 655–666. [Google Scholar]
- Wang, T.; Huang, J.; Zhang, H.; Sun, Q. Visual Commonsense R-CNN. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Online, 14–19 June 2020; pp. 10760–10770. [Google Scholar]
- Antol, S.; Agrawal, A.; Lu, J.; Mitchell, M.; Batra, D.; Zitnick, C.L.; Parikh, D. VQA: Visual Question Answering. In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Santiago, Chile, 7–13 December 2015; pp. 2425–2433. [Google Scholar]
- Chen, Y.; Li, D.; Zhang, P.; Sui, J.; Lv, Q.; Tun, L.; Shang, L. Cross-Modal Ambiguity Learning for Multimodal Fake News Detection. In Proceedings of the ACM Web Conference (WWW), Lyon, France, 25–29 April 2022; pp. 2897–2905. [Google Scholar]
- Vinyals, O.; Toshev, A.; Bengio, S.; Erhan, D. Show and Tell: A Neural Image Caption Generator. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Boston, MA, USA, 7–12 June 2015; pp. 3156–3164. [Google Scholar]
- Xue, J.; Wang, Y.; Tian, Y.; Li, Y.; Shi, L.; Wei, L. Detecting Fake News by Exploring the Consistency of Multimodal Data. Inf. Process. Manag. 2021, 58, 102610. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Yang, H.; Zhang, J.; Zhang, L.; Liu, Y.; Wang, S. MRAN: Multimodal Relationship-Aware Attention Network for Fake News Detection. Comput. Stand. Interfaces 2024, 89, 103822. [Google Scholar] [CrossRef] [Scilit]
- Xu, K.; Ba, J.; Kiros, R.; Cho, K.; Courville, A.; Salakhudinov, R.; Zemel, R.; Bengio, Y. Show, Attend and Tell: Neural Image Caption Generation with Visual Attention. In Proceedings of the 32nd International Conference on Machine Learning (ICML), Lille, France, 6–11 July 2015; pp. 2048–2057. [Google Scholar]
- Boididou, C.; Andreadou, K.; Papadopoulos, S.; Dang-Nguyen, D.T.; Boato, G.; Riegler, M.; Kompatsiaris, Y. Verifying Multimedia Use at MediaEval 2015. MediaEval 2015, 3, 7. [Google Scholar]
- Zubiaga, A.; Liakata, M.; Procter, R. Exploiting Context for Rumour Detection in Social Media. In Proceedings of the International Conference on Social Informatics, Oxford, UK, 13–15 September 2017; pp. 109–123. [Google Scholar]
- Wu, Y.; Zhan, P.; Zhang, Y.; Wang, L.; Xu, Z. Multimodal Fusion with Co-Attention Networks for Fake News Detection. In Proceedings of the Association for Computational Linguistics—Joint Conference on Natural Language Processing (ACL-IJCNLP), Online, 1–6 August 2021; pp. 2560–2569. [Google Scholar] [CrossRef] [Scilit]
- Qian, S.; Wang, J.; Hu, J.; Fang, Q.; Xu, C. Hierarchical Multi-modal Contextual Attention Network for Fake News Detection. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval, Online, 11–15 July 2021; pp. 153–162. [Google Scholar]
- Yao, L.; Mao, C.; Luo, Y. Graph Convolutional Networks for Text Classification. In Proceedings of the AAAI Conference on Artificial Intelligence, Honolulu, HI, USA, 27 January–1 February 2019; pp. 7370–7377. [Google Scholar]
- Singhal, S.; Shah, R.R.; Chakraborty, T.; Kumaraguru, P.; Satoh, S. SpotFake: A Multi-Modal Framework for Fake News Detection. In Proceedings of the IEEE Fifth International Conference on Multimedia Big Data (BigMM), Singapore, 11–13 September 2019; pp. 39–47. [Google Scholar]
- Peng, L.; Jian, S.; Kan, Z.; Qiao, L.; Li, D. Not all fake news is semantically similar: Contextual semantic representation learning for multimodal fake news detection. Inf. Process. Manag. 2024, 61, 103564. [Google Scholar] [CrossRef] [Scilit]
- Ying, Q.; Hu, X.; Zhou, Y.; Qian, Z.; Zeng, D.; Ge, S. Bootstrapping Multi-View Representations for Fake News Detection. In Proceedings of the 37th AAAI Conference on Artificial Intelligence, Washington, DC, USA, 7–14 February 2023; pp. 5384–5392. [Google Scholar]
- Zhang, Q.; Liu, J.; Zhang, F.; Xie, J.; Zha, Z.J. Natural Language-Centered Inference Network for Multi-Modal Fake News Detection. In Proceedings of the 33rd International Joint Conference on Artificial Intelligence (IJCAI), Jeju, Republic of Korea, 3–9 August 2024; pp. 2542–2550. [Google Scholar]
- Cui, W.; Zhang, X.; Shang, M. Multi-View Mutual Learning Network for Multimodal Fake News Detection. Expert Syst. Appl. 2025, 279, 127407. [Google Scholar] [CrossRef] [Scilit]










| Dataset | # Fake News | # Real News | # Images |
|---|---|---|---|
| 6020 | 7898 | 514 | |
| PHEME | 1972 | 3830 | 3670 |
| Predicted Positive | Predicted Negative | |
|---|---|---|
| Actual Positive | TP | FN |
| Actual Negative | FP | TN |
| Dataset | Method | Accuracy | Precision | Recall | F1 |
|---|---|---|---|---|---|
| GRU [18] | 0.634 | 0.581 | 0.812 | 0.677 | |
| CNN [19] | 0.549 | 0.508 | 0.597 | 0.549 | |
| SAFE [26] | 0.766 | 0.777 | 0.795 | 0.786 | |
| EANN [25] | 0.648 | 0.810 | 0.498 | 0.617 | |
| TextGCN [80] | 0.703 | 0.808 | 0.365 | 0.503 | |
| MVAE [6] | 0.745 | 0.801 | 0.719 | 0.758 | |
| SpotFake [81] | 0.771 | 0.784 | 0.744 | 0.764 | |
| HMCAN [79] | 0.897 | 0.971 | 0.801 | 0.878 | |
| HMCAN + CCD [14] | 0.874 | 0.820 | 0.792 | 0.806 | |
| CSFND [82] | 0.873 | 0.899 | 0.799 | 0.846 | |
| MVML [85] | 0.882 | 0.915 | 0.801 | 0.893 | |
| ACIM (ours) | 0.906 ± 0.003 | 0.881 ± 0.001 | 0.844 ± 0.002 | 0.860 ± 0.001 | |
| PHEME | GRU [18] | 0.832 | 0.782 | 0.712 | 0.745 |
| CNN [19] | 0.779 | 0.732 | 0.606 | 0.663 | |
| SAFE [26] | 0.811 | 0.827 | 0.559 | 0.667 | |
| EANN [25] | 0.681 | 0.685 | 0.664 | 0.694 | |
| MVAE [6] | 0.852 | 0.806 | 0.719 | 0.760 | |
| SpotFake [81] | 0.823 | 0.743 | 0.745 | 0.744 | |
| HMCAN [79] | 0.881 | 0.830 | 0.838 | 0.834 | |
| HMCAN + CCD [14] | 0.859 | 0.764 | 0.689 | 0.724 | |
| BMR [83] | 0.884 | 0.872 | 0.840 | 0.855 | |
| NLIN [84] | 0.903 | 0.875 | 0.883 | 0.879 | |
| MVML [85] | 0.893 | 0.872 | 0.846 | 0.855 | |
| ACIM (ours) | 0.909± 0.001 | 0.905 ± 0.003 | 0.875 ± 0.003 | 0.890 ± 0.001 |
| Dataset | Macro-F1 | Balanced Accuracy | MCC |
|---|---|---|---|
| 0.895 ± 0.001 | 0.892 ± 0.004 | 0.791 ± 0.002 | |
| PHEME | 0.907 ± 0.001 | 0.905 ± 0.002 | 0.813 ± 0.003 |
| Dataset | Method | Accuracy | Fake News | Real News | ||||
|---|---|---|---|---|---|---|---|---|
| Precision | Recall | F1 | Precision | Recall | F1 | |||
| w/o causal | 0.862 ± 0.004 | 0.803 ± 0.005 | 0.798 ± 0.009 | 0.801 ± 0.006 | 0.893 ± 0.004 | 0.896 ± 0.003 | 0.894 ± 0.004 | |
| w/o ISA | 0.898 ± 0.004 | 0.864 ± 0.006 | 0.838 ± 0.011 | 0.851 ± 0.007 | 0.915 ± 0.004 | 0.930 ± 0.005 | 0.923 ± 0.004 | |
| w/o CSA | 0.894 ± 0.006 | 0.858 ± 0.008 | 0.832 ± 0.017 | 0.845 ± 0.011 | 0.912 ± 0.005 | 0.927 ± 0.006 | 0.919 ± 0.005 | |
| w/o Aux. Loss | 0.901 ± 0.003 | 0.868 ± 0.004 | 0.843 ± 0.007 | 0.855 ± 0.005 | 0.918 ± 0.003 | 0.932 ± 0.004 | 0.925 ± 0.003 | |
| w/o V | 0.890 ± 0.003 | 0.848 ± 0.004 | 0.832 ± 0.012 | 0.840 ± 0.005 | 0.911 ± 0.003 | 0.921 ± 0.004 | 0.916 ± 0.003 | |
| w/o L | 0.884 ± 0.005 | 0.830 ± 0.008 | 0.838 ± 0.018 | 0.834 ± 0.013 | 0.913 ± 0.004 | 0.908 ± 0.005 | 0.911 ± 0.004 | |
| w/o F | 0.895 ± 0.003 | 0.852 ± 0.004 | 0.845 ± 0.007 | 0.848 ± 0.005 | 0.918 ± 0.003 | 0.922 ± 0.003 | 0.920 ± 0.002 | |
| ACIM | 0.906 ± 0.003 | 0.881 ± 0.001 | 0.844 ± 0.002 | 0.860 ± 0.001 | 0.919 ± 0.003 | 0.939 ± 0.002 | 0.929 ± 0.002 | |
| PHEME | w/o causal | 0.850 ± 0.005 | 0.836 ± 0.004 | 0.800 ± 0.015 | 0.817 ± 0.008 | 0.860 ± 0.005 | 0.886 ± 0.004 | 0.873 ± 0.005 |
| w/o ISA | 0.900 ± 0.004 | 0.894 ± 0.006 | 0.864 ± 0.010 | 0.879 ± 0.007 | 0.904 ± 0.004 | 0.926 ± 0.005 | 0.915 ± 0.004 | |
| w/o CSA | 0.896 ± 0.006 | 0.891 ± 0.008 | 0.856 ± 0.016 | 0.873 ± 0.010 | 0.899 ± 0.005 | 0.925 ± 0.007 | 0.912 ± 0.005 | |
| w/o Aux. Loss | 0.904 ± 0.003 | 0.898 ± 0.004 | 0.870 ± 0.007 | 0.884 ± 0.005 | 0.908 ± 0.003 | 0.929 ± 0.004 | 0.918 ± 0.003 | |
| w/o V | 0.862 ± 0.004 | 0.850 ± 0.005 | 0.815 ± 0.010 | 0.832 ± 0.006 | 0.870 ± 0.004 | 0.896 ± 0.003 | 0.883 ± 0.004 | |
| w/o L | 0.883 ± 0.004 | 0.880 ± 0.004 | 0.835 ± 0.008 | 0.857 ± 0.005 | 0.885 ± 0.003 | 0.918 ± 0.004 | 0.901 ± 0.003 | |
| w/o F | 0.885 ± 0.003 | 0.880 ± 0.004 | 0.840 ± 0.006 | 0.860 ± 0.004 | 0.888 ± 0.003 | 0.918 ± 0.003 | 0.903 ± 0.003 | |
| ACIM | 0.909 ± 0.001 | 0.905 ± 0.003 | 0.875 ± 0.003 | 0.890 ± 0.001 | 0.912 ± 0.002 | 0.934 ± 0.002 | 0.923 ± 0.002 | |
| Parameters | Value | Accuracy | Fake News | Real News | ||||
|---|---|---|---|---|---|---|---|---|
| Precision | Recall | F1 | Precision | Recall | F1 | |||
| num_heads | 2 | 0.889 ± 0.005 | 0.839 ± 0.006 | 0.842 ± 0.012 | 0.841 ± 0.008 | 0.916 ± 0.004 | 0.914 ± 0.005 | 0.915 ± 0.004 |
| 4 | 0.898 ± 0.004 | 0.854 ± 0.005 | 0.852 ± 0.007 | 0.853 ± 0.005 | 0.921 ± 0.003 | 0.922 ± 0.004 | 0.922 ± 0.003 | |
| 8 | 0.906 ± 0.003 | 0.881 ± 0.001 | 0.844 ± 0.002 | 0.860 ± 0.001 | 0.919 ± 0.003 | 0.939 ± 0.002 | 0.929 ± 0.002 | |
| 16 | 0.901 ± 0.004 | 0.857 ± 0.006 | 0.858 ± 0.009 | 0.858 ± 0.007 | 0.924 ± 0.004 | 0.924 ± 0.005 | 0.924 ± 0.004 | |
| dropout | 0 | 0.896 ± 0.005 | 0.852 ± 0.006 | 0.848 ± 0.011 | 0.850 ± 0.007 | 0.919 ± 0.004 | 0.922 ± 0.005 | 0.920 ± 0.004 |
| 0.1 | 0.906 ± 0.003 | 0.881 ± 0.001 | 0.844 ± 0.002 | 0.860 ± 0.001 | 0.919 ± 0.003 | 0.939 ± 0.002 | 0.929 ± 0.002 | |
| 0.2 | 0.904 ± 0.004 | 0.863 ± 0.005 | 0.860 ± 0.006 | 0.862 ± 0.004 | 0.926 ± 0.003 | 0.927 ± 0.004 | 0.927 ± 0.003 | |
| 0.4 | 0.898 ± 0.006 | 0.854 ± 0.008 | 0.852 ± 0.014 | 0.853 ± 0.009 | 0.921 ± 0.005 | 0.922 ± 0.006 | 0.922 ± 0.005 | |
| K | 16 | 0.893 ± 0.006 | 0.853 ± 0.008 | 0.836 ± 0.016 | 0.844 ± 0.010 | 0.914 ± 0.005 | 0.923 ± 0.006 | 0.918 ± 0.005 |
| 32 | 0.901 ± 0.004 | 0.863 ± 0.005 | 0.850 ± 0.008 | 0.856 ± 0.005 | 0.921 ± 0.003 | 0.928 ± 0.004 | 0.924 ± 0.003 | |
| 64 | 0.906 ± 0.003 | 0.881 ± 0.001 | 0.844 ± 0.002 | 0.860 ± 0.001 | 0.919 ± 0.003 | 0.939 ± 0.002 | 0.929 ± 0.002 | |
| 128 | 0.899 ± 0.005 | 0.858 ± 0.006 | 0.850 ± 0.010 | 0.854 ± 0.007 | 0.921 ± 0.004 | 0.925 ± 0.005 | 0.923 ± 0.004 | |
| Parameters | Value | Accuracy | Fake News | Real News | ||||
|---|---|---|---|---|---|---|---|---|
| Precision | Recall | F1 | Precision | Recall | F1 | |||
| num_heads | 2 | 0.898 ± 0.005 | 0.889 ± 0.006 | 0.868 ± 0.011 | 0.878 ± 0.007 | 0.905 ± 0.004 | 0.920 ± 0.005 | 0.912 ± 0.004 |
| 4 | 0.904 ± 0.003 | 0.894 ± 0.005 | 0.878 ± 0.006 | 0.886 ± 0.004 | 0.911 ± 0.003 | 0.923 ± 0.004 | 0.917 ± 0.003 | |
| 8 | 0.909 ± 0.001 | 0.905 ± 0.003 | 0.875 ± 0.003 | 0.890 ± 0.001 | 0.912 ± 0.002 | 0.934 ± 0.002 | 0.923 ± 0.002 | |
| 16 | 0.906 ± 0.004 | 0.895 ± 0.005 | 0.882 ± 0.008 | 0.888 ± 0.005 | 0.914 ± 0.004 | 0.924 ± 0.005 | 0.919 ± 0.004 | |
| dropout | 0 | 0.901 ± 0.005 | 0.904 ± 0.006 | 0.858 ± 0.013 | 0.880 ± 0.008 | 0.899 ± 0.005 | 0.933 ± 0.006 | 0.916 ± 0.005 |
| 0.1 | 0.909 ± 0.001 | 0.905 ± 0.003 | 0.875 ± 0.003 | 0.890 ± 0.001 | 0.912 ± 0.002 | 0.934 ± 0.002 | 0.923 ± 0.002 | |
| 0.2 | 0.905 ± 0.004 | 0.909 ± 0.005 | 0.862 ± 0.008 | 0.885 ± 0.005 | 0.902 ± 0.004 | 0.937 ± 0.004 | 0.919 ± 0.003 | |
| 0.4 | 0.899 ± 0.006 | 0.902 ± 0.008 | 0.855 ± 0.015 | 0.878 ± 0.009 | 0.897 ± 0.005 | 0.931 ± 0.006 | 0.914 ± 0.005 | |
| K | 16 | 0.897 ± 0.006 | 0.892 ± 0.008 | 0.860 ± 0.016 | 0.876 ± 0.010 | 0.901 ± 0.005 | 0.924 ± 0.006 | 0.912 ± 0.005 |
| 32 | 0.905 ± 0.004 | 0.900 ± 0.005 | 0.870 ± 0.007 | 0.885 ± 0.005 | 0.908 ± 0.003 | 0.930 ± 0.004 | 0.919 ± 0.003 | |
| 64 | 0.909 ± 0.001 | 0.905 ± 0.003 | 0.875 ± 0.003 | 0.890 ± 0.001 | 0.912 ± 0.002 | 0.934 ± 0.002 | 0.923 ± 0.002 | |
| 128 | 0.903 ± 0.005 | 0.898 ± 0.006 | 0.868 ± 0.010 | 0.883 ± 0.007 | 0.906 ± 0.004 | 0.928 ± 0.005 | 0.917 ± 0.004 | |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Hao, S.; Li, S.; Lin, R.; Zhang, J.; Wang, X. Attention-Level Causal Intervention Framework for Multimodal Fake News Detection. Big Data Cogn. Comput. 2026, 10, 301. https://doi.org/10.3390/bdcc10090301
Hao S, Li S, Lin R, Zhang J, Wang X. Attention-Level Causal Intervention Framework for Multimodal Fake News Detection. Big Data and Cognitive Computing. 2026; 10(9):301. https://doi.org/10.3390/bdcc10090301
Chicago/Turabian StyleHao, Siqi, Shuohao Li, Rongxin Lin, Jun Zhang, and Xianghan Wang. 2026. "Attention-Level Causal Intervention Framework for Multimodal Fake News Detection" Big Data and Cognitive Computing 10, no. 9: 301. https://doi.org/10.3390/bdcc10090301
APA StyleHao, S., Li, S., Lin, R., Zhang, J., & Wang, X. (2026). Attention-Level Causal Intervention Framework for Multimodal Fake News Detection. Big Data and Cognitive Computing, 10(9), 301. https://doi.org/10.3390/bdcc10090301

