Next Article in Journal
Artificial Intelligence and Control Systems for Industry 4.0 and 5.0: Recent Advances, Knowledge Gaps, and Future Research Directions
Next Article in Special Issue
Heterogeneous Conditional Counter-Inspection: Configurable Error Control and Weak-Filter Recovery for 5G Network Intrusion Detection
Previous Article in Journal
Voltage/VAR Control in Active Distribution Networks via DRL Under False Data Injection Attacks on Distributed PV Systems
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection

by
Shahad Mohammad Bn Dokiey
1,
Tariq M. Khan
1,2,* and
Qazi Emad Ul Haq
1,2
1
Department of Cybersecurity and Digital Forensics, College of Forensic & Investigative Sciences, Naif Arab University for Security Sciences, Riyadh 14812, Saudi Arabia
2
Center of Artificial Intelligence for Security, Naif Arab University for Security Sciences, Riyadh 14812, Saudi Arabia
*
Author to whom correspondence should be addressed.
Future Internet 2026, 18(7), 379; https://doi.org/10.3390/fi18070379
Submission received: 4 June 2026 / Revised: 15 July 2026 / Accepted: 17 July 2026 / Published: 20 July 2026

Abstract

Audio-visual deepfake detection remains challenging under cross-dataset distribution shift, especially when the source-domain training data are severely imbalanced. Existing middle-fusion detectors often rely on softmax-based cross-attention, which can learn sharp source-domain token interactions and may transfer poorly to unseen datasets. This study proposes FocalNet, an audio-visual detector that replaces the cross-attention block of the 2D3MF framework with cross-modal focal modulation. The proposed module aggregates multi-scale temporal context before audio-visual interaction, enabling softmax-free contextual modulation between visual MARLIN features and audio EAT features. We evaluate the method under a strict FakeAVCeleb-to-DFDC protocol, where all training and validation is performed on FakeAVCeleb and the DFDC is used only as an unseen target-domain test set. Compared with the reproduced 2D3MF baseline, which collapses to a single-class prediction pattern on the DFDC, FocalNet achieves substantially stronger zero-shot score separation, with a DFDC ROC-AUC of 0.9324. Thresholded analysis further shows an improved balanced accuracy, macro F1 score, and MCC when the frozen source-domain operating point is applied. The model also preserves practical efficiency, requiring comparable FLOPs and lower per-sample inference time than the reproduced baseline. These findings suggest that cross-modal focal modulation is a promising alternative to attention-based middle fusion for audio-visual deepfake detection under dataset shifts, while broader validation across additional unseen datasets, multi-seed training, and deployment-oriented calibration remain important for future work.
Keywords: audio-visual deepfake detection; cross-dataset generalization; focal modulation; multimodal fusion; class imbalance; confusion matrix; explainable artificial intelligence; FakeAVCeleb; DFDC audio-visual deepfake detection; cross-dataset generalization; focal modulation; multimodal fusion; class imbalance; confusion matrix; explainable artificial intelligence; FakeAVCeleb; DFDC
Graphical Abstract

Share and Cite

MDPI and ACS Style

Dokiey, S.M.B.; Khan, T.M.; Ul Haq, Q.E. Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection. Future Internet 2026, 18, 379. https://doi.org/10.3390/fi18070379

AMA Style

Dokiey SMB, Khan TM, Ul Haq QE. Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection. Future Internet. 2026; 18(7):379. https://doi.org/10.3390/fi18070379

Chicago/Turabian Style

Dokiey, Shahad Mohammad Bn, Tariq M. Khan, and Qazi Emad Ul Haq. 2026. "Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection" Future Internet 18, no. 7: 379. https://doi.org/10.3390/fi18070379

APA Style

Dokiey, S. M. B., Khan, T. M., & Ul Haq, Q. E. (2026). Imbalance-Aware Cross-Modal Focal Modulation for Cross-Dataset Audio-Visual Deepfake Detection. Future Internet, 18(7), 379. https://doi.org/10.3390/fi18070379

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop