Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement
Abstract
1. Introduction
2. Related Work
2.1. Spatial-Domain Deepfake Detection
2.2. Frequency-Domain Deepfake Detection
2.3. Generalizable Deepfake Detection
2.4. Feature Disentanglement
3. Method
3.1. Overview of the Proposed Framework
3.2. Phase-Amplitude Frequency Enhancement Module
3.3. Dual-Branch Feature Disentanglement
3.3.1. Dual-Branch Feature Extraction
3.3.2. Spatial Self-Attention Enhancement
3.4. Multi-Task Loss Optimization
3.4.1. Image-Level Reconstruction Loss
3.4.2. Feature-Level Contrastive Loss
3.4.3. Classification Loss
3.4.4. Overall Training Objective
4. Experiments
4.1. Experimental Settings
4.1.1. Datasets
4.1.2. Implementation Details and Evaluation Metrics
4.2. Cross-Dataset Evaluation
4.3. Cross-Manipulation Evaluation
4.4. Ablation Study
4.5. Visualization Analysis
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Use of Artificial Intelligence
Acknowledgments
Conflicts of Interest
References
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. NeurIPS 2014, 27, 2672–2680. [Google Scholar]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. NeurIPS 2020, 33, 6840–6851. [Google Scholar]
- Verdoliva, L. Media forensics and deepfakes: An overview. IEEE J. Sel. Top. Signal Process. 2020, 14, 910–932. [Google Scholar] [CrossRef] [Scilit]
- Mirsky, Y.; Lee, W. The creation and detection of deepfakes: A survey. ACM Comput. Surv. 2021, 54, 7. [Google Scholar] [CrossRef] [Scilit]
- Bhat, N.A.; Giri, K.J. DeepFake detection in images and videos: A survey on models, datasets, and evaluation metrics. Multimed. Tools Appl. 2026, 85, 233. [Google Scholar] [CrossRef] [Scilit]
- Geirhos, R.; Jacobsen, J.H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; Wichmann, F.A. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2020, 2, 665–673. [Google Scholar] [CrossRef] [Scilit]
- Shi, L.; Zhang, J.; Ji, Z.; Bai, J.; Shan, S. Real face foundation representation learning for generalized deepfake detection. Pattern Recognit. 2025, 161, 111299. [Google Scholar] [CrossRef] [Scilit]
- Liang, J.; Shi, H.; Deng, W. Exploring disentangled content information for face forgery detection. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 128–145. [Google Scholar]
- Yan, Z.; Zhang, Y.; Fan, Y.; Wu, B. UCF: Uncovering common features for generalizable deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 22412–22423. [Google Scholar]
- Durall, R.; Keuper, M.; Keuper, J. Watch your up-convolution: CNN based generative deep neural networks are failing to reproduce spectral distributions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 13–19 June 2020; pp. 7890–7899. [Google Scholar]
- Frank, J.; Eisenhofer, T.; Schönherr, L.; Fischer, A.; Kolossa, D.; Holz, T. Leveraging frequency analysis for deep fake image recognition. In Proceedings of the 37th International Conference on Machine Learning, Virtual, 13–18 July 2020; pp. 3247–3258. [Google Scholar]
- Li, L.; Bao, J.; Zhang, T.; Yang, H.; Chen, D.; Wen, F.; Guo, B. Face X-ray for more general face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 13–19 June 2020; pp. 5001–5010. [Google Scholar]
- Zhao, H.; Zhou, W.; Chen, D.; Wei, T.; Zhang, W.; Yu, N. Multi-attentional deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 19–25 June 2021; pp. 2185–2194. [Google Scholar]
- Zhuang, W.; Chu, Q.; Tan, Z.; Liu, Q.; Yuan, H.; Miao, C.; Luo, Z.; Yu, N. UIA-ViT: Unsupervised inconsistency-aware method based on Vision Transformer for face forgery detection. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 391–407. [Google Scholar]
- Cao, J.; Ma, C.; Yao, T.; Chen, S.; Ding, S.; Yang, X. End-to-end reconstruction-classification learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 4113–4122. [Google Scholar]
- Qian, Y.; Yin, G.; Sheng, L.; Chen, Z.; Shao, J. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In Proceedings of the European Conference on Computer Vision, Virtual, 23–28 August 2020; pp. 86–103. [Google Scholar]
- Masi, I.; Killekar, A.; Mascarenhas, R.M.; Gurudatt, S.P.; AbdAlmageed, W. Two-branch recurrent network for isolating deepfakes in videos. In Proceedings of the European Conference on Computer Vision Workshops, Glasgow, UK, 23–28 August 2020; pp. 667–684. [Google Scholar]
- Liu, H.; Li, X.; Zhou, W.; Chen, Y.; He, Y.; Xue, H.; Zhang, W.; Yu, N. Spatial-phase shallow learning: Rethinking face forgery detection in frequency domain. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 772–781. [Google Scholar]
- Luo, Y.; Zhang, Y.; Yan, J.; Liu, W. Generalizing face forgery detection with high-frequency features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 16317–16326. [Google Scholar]
- Gong, R.; Chen, J.; Zhang, D.; Sangaiah, A.K.; Alenazi, M.J.F. Face Forgery Detection via Multi-Scale and Multi-Domain Features Fusion. IET Image Process. 2025, 19, e70131. [Google Scholar] [CrossRef] [Scilit]
- Qi, Y.; Wen, S.; Zhang, H.; Liang, A.; Chen, H.; Cao, P. Face forgery detection by progressively enhancing spatial and frequency-aware features. Multimed. Syst. 2024, 30, 156. [Google Scholar] [CrossRef] [Scilit]
- Tan, C.; Zhao, Y.; Wei, S.; Gu, G.; Wei, Y. Learning on gradients: Generalized artifacts representation for GAN-generated images detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 12105–12114. [Google Scholar]
- Shiohara, K.; Yamasaki, T. Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18720–18729. [Google Scholar]
- Chen, L.; Zhang, Y.; Song, Y.; Liu, L.; Wang, J. Self-supervised learning of adversarial example towards good generalizations for deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18710–18719. [Google Scholar]
- Zhang, F.; Yang, T.; Cao, L.; Du, K.; Guo, Y.; Song, P.; Shao, C. Deep supervised anomaly detection for generalized face forgery detection. Pattern Recognit. 2026, 169, 111976. [Google Scholar] [CrossRef] [Scilit]
- Liu, B.; Zhang, X.; Ling, H.; Li, Z.; Wang, R.; Zhang, H.; Li, P. AIM-Bone: Texture Discrepancy Generation and Localization for Generalized Deepfake Detection. IEEE Trans. Biom. Behav. Identity Sci. 2025, 7, 422–431. [Google Scholar] [CrossRef] [Scilit]
- Corvi, R.; Cozzolino, D.; Poggi, G.; Nagano, K.; Verdoliva, L. Intriguing properties of synthetic images: From GANs to diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, BC, Canada, 17–24 June 2023; pp. 973–982. [Google Scholar]
- Ojha, U.; Li, Y.; Lee, Y.J. Towards universal fake image detectors that generalize across generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 24480–24489. [Google Scholar]
- Yan, B.; Liu, P.; Yang, Y.; Guo, Y. Self-Supervised Feature Disentanglement for Deepfake Detection. Mathematics 2025, 13, 2024. [Google Scholar] [CrossRef] [Scilit]
- Zou, Z.; Peng, D.; Zhao, Y.; Tian, Z.; Cai, J. FTA-DFI: A framework for generalizable deepfake detection based on distinctive features from various manipulations compared to genuine images. Appl. Intell. 2025, 55, 769. [Google Scholar] [CrossRef] [Scilit]
- Rossler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; Nießner, M. FaceForensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27–28 October 2019; pp. 1–11. [Google Scholar]
- Li, Y.; Yang, X.; Sun, P.; Qi, H.; Lyu, S. Celeb-DF: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 3207–3216. [Google Scholar]
- Google Research. Contributing Data to Deepfake Detection Research. 2019. Available online: https://ai.googleblog.com/2019/09/contributing-data-to-deepfake-detection.html (accessed on 4 January 2026).
- Dolhansky, B.; Bitton, J.; Pflaum, B.; Lu, J.; Howes, R.; Wang, M.; Ferrer, C.C. The DeepFake Detection Challenge Dataset. arXiv 2020, arXiv:2006.07397. [Google Scholar]
- Ni, Y.; Meng, D.; Yu, C.; Quan, C.; Ren, D.; Zhao, Y. Core: Consistent representation learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, New Orleans, LA, USA, 19–20 June 2022; pp. 12–21. [Google Scholar]
- Dang, H.; Liu, F.; Stehouwer, J.; Liu, X.; Jain, A.K. On the detection of digital face manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 5781–5790. [Google Scholar]
- Li, Y.; Lyu, S. Exposing DeepFake Videos By Detecting Face Warping Artifacts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long Beach, CA, USA, 16–17 June 2019; pp. 46–52. [Google Scholar]
- Chollet, F. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1251–1258. [Google Scholar]





| Method | Celeb-DF | Celeb-DF-v2 | DFDCP | DFDC | DFD |
|---|---|---|---|---|---|
| Face X-ray [12] | 0.709 | 0.679 | 0.694 | 0.633 | 0.766 |
| RECCE [15] | 0.768 | 0.732 | 0.742 | 0.713 | 0.812 |
| F3Net [16] | 0.777 | 0.735 | 0.735 | 0.702 | 0.798 |
| UCF [9] | 0.779 | 0.753 | 0.759 | 0.719 | 0.807 |
| CORE [35] | 0.780 | 0.743 | 0.734 | 0.705 | 0.802 |
| FFD [36] | 0.784 | 0.744 | 0.743 | 0.703 | 0.802 |
| FWA [37] | 0.790 | 0.668 | 0.638 | 0.613 | 0.740 |
| SRM [19] | 0.793 | 0.755 | 0.741 | 0.700 | 0.812 |
| SPSL [18] | 0.815 | 0.765 | 0.741 | 0.704 | 0.812 |
| Ours | 0.807 | 0.773 | 0.798 | 0.752 | 0.813 |
| Method | No-DF | No-F2F | No-FS | No-NT | ||||
|---|---|---|---|---|---|---|---|---|
| Celeb-DF | DFDC | Celeb-DF | DFDC | Celeb-DF | DFDC | Celeb-DF | DFDC | |
| Xception [38] | 0.660 | 0.651 | 0.716 | 0.646 | 0.737 | 0.665 | 0.709 | 0.647 |
| CORE [35] | 0.706 | 0.630 | 0.724 | 0.671 | 0.718 | 0.661 | 0.667 | 0.659 |
| RECCE [15] | 0.644 | 0.636 | 0.708 | 0.641 | 0.771 | 0.653 | 0.759 | 0.633 |
| FWA [37] | 0.711 | 0.625 | 0.716 | 0.701 | 0.752 | 0.659 | 0.731 | 0.670 |
| Face X-ray [12] | 0.678 | 0.636 | 0.717 | 0.734 | 0.753 | 0.685 | 0.744 | 0.631 |
| SLADD [24] | 0.645 | 0.738 | 0.726 | 0.675 | 0.715 | 0.729 | 0.734 | 0.759 |
| F3Net [16] | 0.745 | 0.749 | 0.739 | 0.745 | 0.699 | 0.664 | 0.778 | 0.797 |
| SRM [19] | 0.736 | 0.725 | 0.757 | 0.737 | 0.694 | 0.728 | 0.791 | 0.813 |
| SPSL [18] | 0.787 | 0.743 | 0.737 | 0.759 | 0.699 | 0.669 | 0.829 | 0.807 |
| UCF [9] | 0.749 | 0.767 | 0.782 | 0.765 | 0.800 | 0.711 | 0.808 | 0.800 |
| Ours | 0.783 | 0.782 | 0.792 | 0.781 | 0.812 | 0.736 | 0.824 | 0.836 |
| Experiment | Core Modules | FF++ | Celeb-DF |
|---|---|---|---|
| 1 | Baseline | 0.901 | 0.672 |
| 2 | Baseline + D | 0.954 | 0.654 |
| 3 | Baseline + D + A | 0.958 | 0.696 |
| 4 | Baseline + D + F | 0.967 | 0.764 |
| 5 | Baseline + D + F + A | 0.972 | 0.807 |
| Exp. | FF++ | Celeb-DF | |||
|---|---|---|---|---|---|
| 1 | ✓ | 0.956 | 0.685 | ||
| 2 | ✓ | ✓ | 0.963 | 0.772 | |
| 3 | ✓ | ✓ | 0.970 | 0.759 | |
| 4 | ✓ | ✓ | ✓ | 0.972 | 0.807 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Wang, Q.; Feng, J.; Zhou, Y.; Li, M.; Zhang, Z.; Wang, L.; Song, W. Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement. AI 2026, 7, 344. https://doi.org/10.3390/ai7090344
Wang Q, Feng J, Zhou Y, Li M, Zhang Z, Wang L, Song W. Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement. AI. 2026; 7(9):344. https://doi.org/10.3390/ai7090344
Chicago/Turabian StyleWang, Qian, Jiaqi Feng, Yu Zhou, Miao Li, Zhi Zhang, Luyao Wang, and Wenping Song. 2026. "Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement" AI 7, no. 9: 344. https://doi.org/10.3390/ai7090344
APA StyleWang, Q., Feng, J., Zhou, Y., Li, M., Zhang, Z., Wang, L., & Song, W. (2026). Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement. AI, 7(9), 344. https://doi.org/10.3390/ai7090344

