Adversarial Robustness and Explainability in AI-Generated Face Detection †
Abstract
1. Introduction
2. Related Work
2.1. Adversarial Vulnerabilities in Deepfake Detection
2.2. Explainability in Deepfake Detection
2.3. Adversarial Defense Strategies
2.4. Explainability Methods and Evaluation
3. Proposed Methodology
3.1. Architectural Overview
3.2. Training and Implementation
4. Model Training and Evaluation
4.1. Datasets
4.2. Metrics and Baselines
4.3. Implementation Details
5. Results and Discussion
5.1. Quantitative Results
5.2. Qualitative Explainability
5.3. Discussion
- Xception and ResNet-50 reach 97.53% and 96.64% validation accuracy (F1 97.51% and 96.69%, AUC-ROC 0.9977 and 0.9959) under the same clean-training protocol;
- ViT-B/16 attains 69.08% accuracy, 67.69% F1, and 0.7521 AUC-ROC, indicating a substantial gap on this dataset under the current setup;
- Best validation occurs at epoch 9 for Xception and ViT-B/16, and at epoch 10 for ResNet-50.
5.4. Limitations
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| ARI | Adversarial Robustness Index |
| AUC-ROC | Area Under the Receiver Operating Characteristic Curve |
| CNN | Convolutional Neural Network |
| EF | Explainability Fidelity |
| FGSM | Fast Gradient Sign Method |
| GAN | Generative Adversarial Network |
| Grad-CAM | Gradient-weighted Class Activation Mapping |
| IG | Integrated Gradients |
| LRP | Layer-wise Relevance Propagation |
| PGD | Projected Gradient Descent |
| RED | Robust and Explainable Detection |
| ViT | Vision Transformer |
| XAI | Explainable Artificial Intelligence |
| CLI | Command-Line Interface |
| CSV | Comma-Separated Values |
| FFT | Fast Fourier Transform |
| ML | Machine Learning |
| ROC | Receiver Operating Characteristic |
References
- Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial networks. Adv. Neural Inf. Process. Syst. 2014, 15, 2672–2680. [Google Scholar] [CrossRef]
- Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. Adv. Neural Inf. Process. Syst. 2020, 33, 6840–6851. [Google Scholar] [CrossRef]
- He, K.; Zhang, X.; Ren, S.; Sun, J. Deep residual learning for image recognition. In IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016; pp. 770–778. [Google Scholar] [CrossRef]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16 × 16 words: Transformers for image recognition at scale. arXiv 2021, arXiv:2010.11929. [Google Scholar] [CrossRef]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, Ł.; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017; Volume 30. [Google Scholar] [CrossRef]
- Rössler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; Nießner, M. FaceForensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE International Conference on Computer Vision, Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1–11. [Google Scholar] [CrossRef]
- Lin, L.; Santosh, S.; Wu, M.; Wang, X.; Hu, S. AI-Face: A million-scale demographically annotated AI-generated face dataset and fairness benchmark. In IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2025. [Google Scholar] [CrossRef]
- Goodfellow, I.J.; Shlens, J.; Szegedy, C. Explaining and harnessing adversarial examples. arXiv 2015, arXiv:1412.6572. [Google Scholar] [CrossRef]
- Madry, A.; Makelov, A.; Schmidt, L.; Tsipras, D.; Vladu, A. Towards deep learning models resistant to adversarial attacks. arXiv 2018, arXiv:1706.06083. [Google Scholar] [CrossRef]
- Szegedy, C.; Zaremba, W.; Sutskever, I.; Bruna, J.; Erhan, D.; Goodfellow, I.; Fergus, R. Intriguing properties of neural networks. arXiv 2014, arXiv:1312.6199. [Google Scholar] [CrossRef]
- Moosavi-Dezfooli, S.-M.; Fawzi, A.; Frossard, P. DeepFool: A simple and accurate method to fool deep neural networks. In IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2016; pp. 2574–2582. [Google Scholar] [CrossRef]
- Papernot, N.; McDaniel, P.; Goodfellow, I.; Jha, S.; Celik, Z.B.; Swami, A. Practical black-box attacks against machine learning. In ACM Asia Conference on Computer and Communications Security; Association for Computing Machinery: New York, NY, USA, 2017; pp. 506–519. [Google Scholar] [CrossRef]
- Selvaraju, R.R.; Cogswell, M.; Das, A.; Vedantam, R.; Parikh, D.; Batra, D. Grad-CAM: Visual explanations from deep networks via gradient-based localization. In IEEE International Conference on Computer Vision; IEEE: New York, NY, USA, 2017; pp. 618–626. [Google Scholar] [CrossRef]
- Bach, S.; Binder, A.; Montavon, G.; Klauschen, F.; Müller, K.-R.; Samek, W. On pixel-wise explanations for non-linear classifier decisions by layer-wise relevance propagation. PLoS ONE 2015, 10, e0130140. [Google Scholar] [CrossRef] [PubMed]
- Sundararajan, M.; Taly, A.; Yan, Q. Axiomatic attribution for deep networks. In International Conference on Machine Learning; PMLR: Cambridge, MA, USA, 2017; pp. 3319–3328. [Google Scholar] [CrossRef]
- Chefer, H.; Singh, S.; Guestrin, C. Transformer interpretability beyond attention visualization. In IEEE/CVF Conference on Computer Vision and Pattern Recognition; IEEE: New York, NY, USA, 2021; pp. 782–791. [Google Scholar] [CrossRef]
- Ribeiro, M.T.; Singh, S.; Guestrin, C. “Why should I trust you?” Explaining the predictions of any classifier. In Proceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining, San Francisco, CA, USA, 13–17 August 2016; pp. 1135–1144. [Google Scholar] [CrossRef]
- Lundberg, S.M.; Lee, S.-I. A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30. [Google Scholar] [CrossRef]
- Petsiuk, V.; Das, A.; Saenko, K. RISE: Randomized input sampling for explanation of black-box models. arXiv 2018, arXiv:1806.07421. [Google Scholar] [CrossRef]
- Cohen, J.; Rosenfeld, E.; Kolter, Z. Certified adversarial robustness via randomized smoothing. In International Conference on Machine Learning; IEEE: New York, NY, USA, 2019; pp. 1310–1320. [Google Scholar] [CrossRef]
- Jaccard, P. The distribution of the flora in the alpine zone. New Phytol. 1912, 11, 37–50. [Google Scholar] [CrossRef]
- Chollet, F. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1251–1258. [Google Scholar] [CrossRef]
- Wightman, R. PyTorch Image Models. 2019. Available online: https://github.com/rwightman/pytorch-image-models (accessed on 1 March 2026).
- Russakovsky, O.; Deng, J.; Su, H.; Krause, J.; Satheesh, S.; Ma, S.; Huang, Z.; Karpathy, A.; Khosla, A.; Bernstein, M.; et al. ImageNet large scale visual recognition challenge. Int. J. Comput. Vis. 2015, 115, 211–252. [Google Scholar] [CrossRef]
- Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; Fei-Fei, L. ImageNet: A large-scale hierarchical image database. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Miami, FL, USA, 20–25 June 2009; pp. 248–255. [Google Scholar] [CrossRef]
- Paszke, A.; Gross, S.; Massa, F.; Lerer, A.; Bradbury, J.; Chanan, G.; Killeen, T.; Lin, Z.; Gimelshein, N.; Antiga, L.; et al. PyTorch: An imperative style, high-performance deep learning library. In Advances in Neural Information Processing Systems; Curran Associates, Inc.: Red Hook, NY, USA, 2019; Volume 32. [Google Scholar] [CrossRef]
- Pan, S.J.; Yang, Q. A survey on transfer learning. IEEE Trans. Knowl. Data Eng. 2010, 22, 1345–1359. [Google Scholar] [CrossRef]
- Bishop, C.M. Pattern Recognition and Machine Learning; Springer: Berlin/Heidelberg, Germany, 2006; ISBN 978-0-387-31073-2. [Google Scholar]
- Kingma, D.P.; Ba, J. Adam: A method for stochastic optimization. arXiv 2015, arXiv:1412.6980. [Google Scholar] [CrossRef]
- Pedregosa, F.; Varoquaux, G.; Gramfort, A.; Michel, V.; Thirion, B.; Grisel, O.; Blondel, M.; Prettenhofer, P.; Weiss, R.; Dubourg, V.; et al. Scikit-learn: Machine learning in Python. J. Mach. Learn. Res. 2011, 12, 2825–2830. [Google Scholar] [CrossRef]
- Fawcett, T. An introduction to ROC analysis. Pattern Recognit. Lett. 2006, 27, 861–874. [Google Scholar] [CrossRef]



| Parameter | Value |
|---|---|
| Epochs | 10 |
| Batch size | 32 |
| Learning rate | 0.0003 |
| Max samples per epoch | 2000 |
| Adversarial training | none |
| ε (0–255) | 8 |
| α (PGD) | 2 |
| PGD steps | 5 |
| Seed | 42 |
| Metric | Xception | ResNet50 | ViT-B16 |
|---|---|---|---|
| Best epoch | 9 | 10 | 9 |
| Accuracy (%) | 97.53 | 96.64 | 69.08 |
| F1 (%) | 97.51 | 96.69 | 67.69 |
| AUC-ROC | 0.9977 | 0.9959 | 0.7521 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Kotov, G.; Nakov, P.; Nakov, O. Adversarial Robustness and Explainability in AI-Generated Face Detection. Eng. Proc. 2026, 150, 52. https://doi.org/10.3390/engproc2026150052
Kotov G, Nakov P, Nakov O. Adversarial Robustness and Explainability in AI-Generated Face Detection. Engineering Proceedings. 2026; 150(1):52. https://doi.org/10.3390/engproc2026150052
Chicago/Turabian StyleKotov, Georgi, Plamen Nakov, and Ognyan Nakov. 2026. "Adversarial Robustness and Explainability in AI-Generated Face Detection" Engineering Proceedings 150, no. 1: 52. https://doi.org/10.3390/engproc2026150052
APA StyleKotov, G., Nakov, P., & Nakov, O. (2026). Adversarial Robustness and Explainability in AI-Generated Face Detection. Engineering Proceedings, 150(1), 52. https://doi.org/10.3390/engproc2026150052

