Multi-Domain Perception Transformer for Generalized Forgery Image Detection
Abstract
1. Introduction
- We propose a multi-domain feature fusion Transformer network that jointly leverages spatial, frequency, and wavelet-domain representations—capturing multi-scale characteristics through wavelet transforms—to uncover subtle and otherwise imperceptible forgery traces in deepfake images.
- Features are extracted using EfficientNet independently from the spatial domain, the frequency domain, and the wavelet domain. These features are then fused via a novel Cross-Domain Attention Fusion (CDAF) module, followed by a Swin Transformer that effectively identifies fine-grained forgery artifacts.
- We conduct comprehensive evaluations across multiple benchmark datasets for AI-generated image detection, covering a wide spectrum of state-of-the-art generative models. Extensive experimental results demonstrate that our approach significantly outperforms existing state-of-the-art detectors in both in-domain and cross-domain settings.
2. Related Works
3. Materials and Methods
3.1. Overview
3.2. Multi-Domain Feature Extraction
3.3. Cross-Domain Attention Feature Fusion Module
3.4. Transformer Module
4. Experiments and Results
4.1. Dataset
4.2. Evaluation Metrics and Implementation Details
4.3. Comparisons with State-of-the-Art Methods
5. Discussion and Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial networks. Commun. ACM 2020, 63, 139–144. [Google Scholar] [CrossRef] [Scilit]
- Kingma, D.P.; Welling, M. Auto-encoding variational bayes. arXiv 2013, arXiv:1312.6114. [Google Scholar]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–20 June 2022; pp. 10684–10695. [Google Scholar]
- Luo, Z.; Chen, D.; Zhang, Y.; Huang, Y.; Wang, L.; Shen, Y.; Zhao, D.; Zhou, J.; Tan, T. Videofusion: Decomposed diffusion models for high-quality video generation. arXiv 2023, arXiv:2303.08320. [Google Scholar]
- Tyagi, S.; Yadav, D. MiniNet: A concise CNN for image forgery detection. Evol. Syst. 2023, 14, 545–556. [Google Scholar] [CrossRef] [Scilit]
- Aggarwal, G.; Srivastava, A.K.; Jhajharia, K.; Sharma, N.V.; Singh, G. Detection of deep fake images using convolutional neural networks. In Proceedings of the 2023 3rd International Conference on Technological Advancements in Computational Sciences (ICTACS), Tashkent, Uzbekistan, 1–3 November 2023; pp. 1083–1087. [Google Scholar]
- Muthukumar, A.; Raj, M.T.; Ramalakshmi, R.; Meena, A.; Kaleeswari, P. Fake and propaganda images detection using automated adaptive gaining sharing knowledge algorithm with DenseNet121. J. Ambient. Intell. Humaniz. Comput. 2024, 15, 3519–3531. [Google Scholar] [CrossRef] [Scilit]
- Yang, J.; Xiao, S.; Li, A.; Lan, G.; Wang, H. Detecting fake images by identifying potential texture difference. Future Gener. Comput. Syst. 2021, 125, 127–135. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Yang, Z.; You, W.; Zhou, L.; Chu, B. Fake face images detection and identification of celebrities based on semantic segmentation. IEEE Signal Process. Lett. 2022, 29, 2018–2022. [Google Scholar] [CrossRef] [Scilit]
- Frank, J.; Eisenhofer, T.; Schönherr, L.; Fischer, A.; Kolossa, D.; Holz, T. Leveraging frequency analysis for deep fake image recognition. In Proceedings of the International Conference on Machine Learning, Vienna, Austria, 12–18 July 2020; pp. 3247–3258. [Google Scholar]
- Gu, Q.; Chen, S.; Yao, T.; Chen, Y.; Ding, S.; Yi, R. Exploiting fine-grained face forgery clues via progressive enhancement learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 22 February–1 March 2022; pp. 735–743. [Google Scholar]
- Jeong, Y.; Kim, D.; Ro, Y.; Choi, J. Frepgan: Robust deepfake detection using frequency-level perturbations. In Proceedings of the AAAI Conference on Artificial Intelligence, Virtual, 22 February–1 March 2022; pp. 1060–1068. [Google Scholar]
- Shiohara, K.; Yamasaki, T. Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18720–18729. [Google Scholar]
- Shi, L.; Zhang, J.; Ji, Z.; Bai, J.; Shan, S. Real face foundation representation learning for generalized deepfake detection. Pattern Recognit. 2025, 161, 111299. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Chang, M.-C.; Lyu, S. In ictu oculi: Exposing ai generated fake face videos by detecting eye blinking. arXiv 2018, arXiv:1806.02877. [Google Scholar] [CrossRef] [Scilit]
- Yang, X.; Li, Y.; Lyu, S. Exposing deep fakes using inconsistent head poses. In Proceedings of the ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK, 12–17 May 2019; pp. 8261–8265. [Google Scholar]
- Rossler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; Nießner, M. Faceforensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1–11. [Google Scholar]
- Chollet, F. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1251–1258. [Google Scholar]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Identity mappings in deep residual networks. In Proceedings of the European Conference on Computer Vision, Amsterdam, The Netherlands, 11–14 October 2016; pp. 630–645. [Google Scholar]
- Nguyen, H.H.; Yamagishi, J.; Echizen, I. Capsule-forensics networks for deepfake detection. In Handbook of Digital Face Manipulation and Detection: From DeepFakes to Morphing Attacks; Springer International Publishing: Cham, Switzerland, 2022; pp. 275–301. [Google Scholar]
- Chen, L.; Zhang, Y.; Song, Y.; Liu, L.; Wang, J. Self-supervised learning of adversarial example: Towards good generalizations for deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–20 June 2022; pp. 18710–18719. [Google Scholar]
- Ren, H.; Yan, A.; Ren, X.; Ye, P.-G.; Gao, C.-Z.; Zhou, Z.; Li, J. Ganfinger: Gan-based fingerprint generation for deep neural network ownership verification. arXiv 2023, arXiv:2312.15617. [Google Scholar]
- Zhao, H.; Zhou, W.; Chen, D.; Wei, T.; Zhang, W.; Yu, N. Multi-attentional deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 2185–2194. [Google Scholar]
- Kim, M.; Tariq, S.; Woo, S.S. Fretal: Generalizing deepfake detection using knowledge distillation and representation learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 1001–1012. [Google Scholar]
- Aneja, S.; Nießner, M. Generalized zero and few-shot transfer for facial forgery detection. arXiv 2020, arXiv:2006.11863. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Deng, W. Representative forgery mining for fake face detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 14923–14932. [Google Scholar]
- Dosovitskiy, A. An image is worth 16 × 16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Wodajo, D.; Atnafu, S. Deepfake video detection using convolutional vision transformer. arXiv 2021, arXiv:2102.11126. [Google Scholar] [CrossRef] [Scilit]
- Heo, Y.-J.; Choi, Y.-J.; Lee, Y.-W.; Kim, B.-G. Deepfake detection scheme based on vision transformer and distillation. arXiv 2021, arXiv:2104.01353. [Google Scholar] [CrossRef] [Scilit]
- Park, J. Using the Swin-Transformer for Real & Fake Data Recognition in PC-Model. In Proceedings of the 2024 IEEE Integrated STEM Education Conference (ISEC), Princeton, NJ, USA, 9 March 2024; pp. 1–5. [Google Scholar]
- Heo, Y.-J.; Yeo, W.-H.; Kim, B.-G. Deepfake detection algorithm based on improved vision transformer. Appl. Intell. 2023, 53, 7512–7527. [Google Scholar] [CrossRef] [Scilit]
- Dolhansky, B.; Howes, R.; Pflaum, B.; Baram, N.; Ferrer, C.C. The deepfake detection challenge (dfdc) preview dataset. arXiv 2019, arXiv:1910.08854. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Yang, X.; Sun, P.; Qi, H.; Lyu, S. Celeb-df: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 3207–3216. [Google Scholar]
- Wang, S.-Y.; Wang, O.; Zhang, R.; Owens, A.; Efros, A.A. CNN-generated images are surprisingly easy to spot… for now. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 8695–8704. [Google Scholar]
- Liu, L.; Ren, Y.; Lin, Z.; Zhao, Z. Pseudo numerical methods for diffusion models on manifolds. arXiv 2022, arXiv:2202.09778. [Google Scholar] [CrossRef] [Scilit]
- Dhariwal, P.; Nichol, A. Diffusion models beat gans on image synthesis. Adv. Neural Inf. Process. Syst. 2021, 34, 8780–8794. [Google Scholar]
- Ramesh, A.; Pavlov, M.; Goh, G.; Gray, S.; Voss, C.; Radford, A.; Chen, M.; Sutskever, I. Zero-shot text-to-image generation. In Proceedings of the International Conference on Machine Learning, Virtual, 18–24 July 2021; pp. 8821–8831. [Google Scholar]
- Gu, S.; Chen, D.; Bao, J.; Wen, F.; Zhang, B.; Chen, D.; Yuan, L.; Guo, B. Vector quantized diffusion model for text-to-image synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–20 June 2022; pp. 10696–10706. [Google Scholar]
- Zhao, Y.; Jin, X.; Gao, S.; Wu, L.; Yao, S.; Jiang, Q. TAN-GFD: Generalizing face forgery detection based on texture information and adaptive noise mining. Appl. Intell. 2023, 53, 19007–19027. [Google Scholar] [CrossRef] [Scilit]
- Peng, S.; Zhang, T.; Gao, L.; Zhu, X.; Zhang, H.; Pang, K.; Lei, Z. Wmamba: Wavelet-based mamba for face forgery detection. In Proceedings of the 33rd ACM International Conference on Multimedia, Dublin, Ireland, 27–31 October 2025; pp. 4768–4777. [Google Scholar]






| Models | Celeb-DF | DFDC | ||
|---|---|---|---|---|
| ACC | AP | ACC | AP | |
| Xception | 67.4 | 65.8 | 66.3 | 68.3 |
| F3-Net | 77.2 | 76.7 | 75.7 | 76.0 |
| TAN-GFD [39] | 87.7 | 90.3 | 84.3 | 85.8 |
| WMamba [40] | 92.3 | 91.4 | 90.5 | 90.0 |
| Ours(S+F) | 92.7 | 91.8 | 91.7 | 91.1 |
| Ours(S+F+W) | 95.1 | 94.3 | 93.9 | 93.0 |
| Ours(S+F+W+CDAF) | 96.3 | 95.9 | 95.4 | 94.2 |
| Methods | ProGAN | StyleGAN | StyleGAN2 | BigGAN | CycleGAN | StarGAN | GauGAN | Deepfake | Mean |
|---|---|---|---|---|---|---|---|---|---|
| Wang | 64.6/92.7 | 52.8/80.8 | 75.7/96.3 | 50.7/70.2 | 58.1/79.3 | 51.2/81.7 | 53.6/84.7 | 50.3/51.5 | 57.1/79.7 |
| Fank | 85.7/81.3 | 73.1/68.5 | 75.0/70.9 | 76.9/70.8 | 86.5/80.8 | 85.0/77.0 | 67.3/65.3 | 50.1/55.3 | 75.0/71.2 |
| F3-Net | 87.8/82.4 | 80.3/84.7 | 82.2/87.9 | 65.5/73.4 | 81.2/89.7 | 87.8/90.4 | 57.0/59.5 | 59.9/83.0 | 75.2/81.4 |
| BiHPF | 87.4/89.3 | 71.5/74.1 | 77.0/81.1 | 82.6/80.6 | 86.0/86.6 | 93.8/95.5 | 75.3/84.7 | 53.5/55.8 | 78.4/81.0 |
| FrePGAN | 95.3/97.1 | 82.0/90.9 | 72.2/93.8 | 66.7/69.4 | 69.7/71.1 | 97.3/99.0 | 53.7/55.0 | 62.7/80.1 | 75.0/82.1 |
| UniFD | 98.3/99.8 | 78.5/92.8 | 75.4/96.0 | 89.1/94.7 | 91.9/98.0 | 96.1/99.3 | 92.6/98.3 | 80.8/90.2 | 88.1/96.1 |
| FreqNet | 99.2/99.9 | 90.4/98.0 | 85.8/98.3 | 89.7/96.4 | 96.7/99.1 | 97.5/99.4 | 88.3/98.9 | 81.9/92.7 | 91.2/98.0 |
| Ours | 99.6/99.9 | 95.7/98.5 | 91.7/99.1 | 95.3/98.9 | 98.5/99.7 | 98.8/99.6 | 93.8/99.4 | 92.3/97.9 | 96.0/99.1 |
| Methods | PNDM | Guided | DALL-E | VQ-Diffusion | Mean |
|---|---|---|---|---|---|
| Wang | 50.8/90.3 | 54.9/66.6 | 51.8/61.3 | 50.0/71.0 | 51.8/72.3 |
| Fank | 44.0/38.2 | 53.4/52.7 | 57.1/62.8 | 52.0/66.3 | 51.6/55.0 |
| F3-Net | 72.8/80.5 | 69.7/72.1 | 72.3/80.0 | 91.8/94.7 | 76.7/81.8 |
| UniFD | 75.3/92.5 | 75.7/85.1 | 89.5/96.8 | 83.5/97.7 | 81.0/93.0 |
| FreqNet | 89.3/97.0 | 81.2/92.0 | 94.8/98.3 | 92.0/97.3 | 89.3/96.2 |
| Ours | 94.1/96.3 | 82.0/91.7 | 97.3/99.2 | 95.4/99.1 | 92.2/96.6 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Man, Q.; Gee, S.-J.; Cho, Y.-I. Multi-Domain Perception Transformer for Generalized Forgery Image Detection. Appl. Sci. 2026, 16, 533. https://doi.org/10.3390/app16010533
Man Q, Gee S-J, Cho Y-I. Multi-Domain Perception Transformer for Generalized Forgery Image Detection. Applied Sciences. 2026; 16(1):533. https://doi.org/10.3390/app16010533
Chicago/Turabian StyleMan, Qiaoyue, Seok-Jeong Gee, and Young-Im Cho. 2026. "Multi-Domain Perception Transformer for Generalized Forgery Image Detection" Applied Sciences 16, no. 1: 533. https://doi.org/10.3390/app16010533
APA StyleMan, Q., Gee, S.-J., & Cho, Y.-I. (2026). Multi-Domain Perception Transformer for Generalized Forgery Image Detection. Applied Sciences, 16(1), 533. https://doi.org/10.3390/app16010533

