Protecting Digital Identities: Deepfake Face Detection Using Dual-Decoder U-Net Semantic Segmentation
Abstract
1. Introduction
- The dual-decoder architecture enhances the localization of manipulated regions by separating specific features, increasing the robustness against common image distortions, including compression, scaling, cropping, noise, and blurring.
- The proposed architecture generates complementary masks, improving interpretability and providing pixel-level localization of both authentic and forged facial regions.
- The dual-decoder architecture employs attention mechanisms to improve boundary precision by detecting subtle manipulations that conventional single-decoder segmentation models and global classifiers often fail to capture.
- The unified segmentation framework enables multitask learning while maintaining high accuracy for both authentic and manipulated regions.
- The dual-decoder outputs provide structured masks that support subsequent forensic applications, including automated verification, forgery localization, and confidence estimation.
2. Materials and Methods
2.1. Materials
2.2. Proposed Dual-Decoder for Authentic and Manipulated Deepfake Face Segmentation
3. Results
3.1. Quantitative Performance Evaluation Under Image Distortions
3.2. Qualitative Evaluation of Deepfake Segmentation
3.3. Cross-Dataset Validation
3.4. Comparative Analysis of Neural Network Architectures
3.5. Comparison with State-of-the-Art Methods
3.6. Ablation Study
4. Discussion
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| CNN | Convolutional Neural Network |
| SRM | Spatial Rich Model |
| LSTM | Long short-term memory |
| IoU | Intersection over union |
| AUC | Area under the curve |
References
- Zhang, L.; Lu, T.; Du, Y. Overview of facial deepfake video detection methods. Front. Comput. Sci. Technol. 2023, 17, 1. [Google Scholar] [CrossRef]
- Alrashoud, M. Deepfake video detection methods, approaches, and challenges. Alex. Eng. J. 2025, 125, 265–277. [Google Scholar] [CrossRef] [Scilit]
- Gong, L.Y.; Li, X.J. A contemporary survey on deepfake detection: Datasets, algorithms, and challenges. Electronics 2024, 13, 585. [Google Scholar] [CrossRef] [Scilit]
- Zha, R.; Lian, Z.; Li, Q. Centroid-based contrastive consistency learning for transferable deepfake detection. Neurocomputing 2025, 637, 130009. [Google Scholar] [CrossRef] [Scilit]
- Sun, Z.; Han, Y.; Hua, Z.; Ruan, N.; Jia, W. Improving the efficiency and robustness of deepfakes detection through precise geometric features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 20–25 June 2021; pp. 3609–3618. [Google Scholar] [CrossRef] [Scilit]
- Ahmed, N.U.R.; Badshah, A.; Adeel, H.; Tajammul, A.; Daud, A.; Alsahfi, T. Visual deepfake detection: Review of techniques, tools, limitations, and future prospects. IEEE Access 2024, 13, 1923–1961. [Google Scholar] [CrossRef] [Scilit]
- Rana, M.S.; Nobi, M.N.; Murali, B.; Sung, A.H. Deepfake detection: A systematic literature review. IEEE Access 2022, 10, 25494–25513. [Google Scholar] [CrossRef] [Scilit]
- Ni, Y.; Meng, D.; Yu, C.; Quan, C.; Ren, D.; Zhao, Y. CORE: Consistent representation learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), New Orleans, LA, USA, 18–24 June 2022; pp. 12–21. [Google Scholar] [CrossRef] [Scilit]
- Chang, X.; Wu, J.; Yang, T.; Feng, G. DeepFake face image detection based on improved VGG convolutional neural network. In Proceedings of the 39th Chinese Control Conference (CCC), Shenyang, China, 27–29 July 2020. [Google Scholar] [CrossRef] [Scilit]
- Yu, C.-M.; Chen, K.-C.; Chang, C.-T.; Ti, Y.-W. SegNet: A network for detecting deepfake facial videos. Multimed. Syst. 2022, 28, 793–814. [Google Scholar] [CrossRef] [Scilit]
- Wang, R.; Yang, Z.; You, W.; Zhou, L.; Chu, B. Fake face images detection and identification of celebrities based on semantic segmentation. IEEE Signal Process. Lett. 2022, 29, 2018–2022. [Google Scholar] [CrossRef] [Scilit]
- Nirkin, Y.; Wolf, L.; Keller, Y.; Hassner, T. DeepFake detection based on discrepancies between faces and their context. IEEE Trans. Pattern Anal. Mach. Intell. 2022, 44, 6111–6121. [Google Scholar] [CrossRef] [Scilit]
- Abdullah, M.T.; Ali, N.H.M. Deploying facial segmentation landmarks for deepfake detection. J. Al-Qadisiyah Comput. Sci. Math. 2023, 15, 137–149. [Google Scholar] [CrossRef] [Scilit]
- Wang, S.; Zhang, H.; Yang, G.; Guo, Z.; Chen, J. A two-stage fake face image detection algorithm with expanded attention. Multimed. Tools Appl. 2024, 83, 55709–55730. [Google Scholar] [CrossRef] [Scilit]
- Raja Sekar, R.; Dhiliphan Rajkumar, T.; Rao Anne, K. Deep fake detection using an optimal deep learning model with multi-head attention-based feature extraction scheme. Vis. Comput. 2025, 41, 2783–2800. [Google Scholar] [CrossRef] [Scilit]
- Jagam, A.; Patel, N.; Chidirala, B.; Acharya, B. Deepfake detection in facial images using convolutional neural networks. In Proceedings of the Fourth International Conference on Power, Control and Computing Technologies (ICPC2T), Raipur, India, 20–22 January 2025. [Google Scholar] [CrossRef] [Scilit]
- Kurniawan, W.; Kurniasih, A.; Ghani, M.A. Real or deepfake face detection in images and video data using YOLO11 algorithm. J. Artif. Intell. Eng. Appl. 2025, 4, 1514–1521. [Google Scholar] [CrossRef] [Scilit]
- Alsolai, H.; Mahmood, K.; Alshuhail, A.; Ben Miled, A.; Alqahtani, M.; Alshareef, A.; Alallah, F.S.; Alghamdi, B.M. Guardian-AI: A novel deep learning based deepfake detection model in images. Alex. Eng. J. 2025, 126, 507–514. [Google Scholar] [CrossRef] [Scilit]
- Gupta, S.; Hariprasad, Y.; Iyengar, S.S.; Gurappa, S.; Mohanty, P. Enhancing digital security: A novel dual-paradigm approach for robust deepfake detection using pre and post quantum-trained neural networks. Digit. Threat. Res. Pract. 2026, 7, 1–15. [Google Scholar] [CrossRef] [Scilit]
- Le, T.-N.; Nguyen, H.H.; Yamagishi, J.; Echizen, I. OpenForensics: Large-scale challenging dataset for multi-face forgery detection and segmentation in-the-wild. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Montreal, QC, Canada, 10–17 October 2021; pp. 9996–10006. [Google Scholar] [CrossRef] [Scilit]
- Arevalo-Ancona, R.E.; Haro-Mendoza, D.; Cedillo-Hernandez, M.; Gonzalez-Villela, V.J. Advanced dual-branch U-Net decoder for precise and robust surgical instrument and organ segmentation. Biomed. Signal Process. Control 2025, 110, 108348. [Google Scholar] [CrossRef] [Scilit]
- Pandiri, D.N.K.; Murugan, R.; Goel, T. ARM-UNet: Attention residual path modified UNet model to segment the fungal pathogen diseases in potato leaves. Signal Image Video Process. 2024, 19, 80. [Google Scholar] [CrossRef] [Scilit]
- Arora, S.; Banerjee, A.; Katal, N. Enhanced urban driving scene segmentation using modified UNet with residual convolutions and attention guided skip connections. Discov. Artif. Intell. 2025, 5, 198. [Google Scholar] [CrossRef] [Scilit]
- Fan, Y.; Song, J.; Yuan, L.; Jia, Y. HCT-Unet: Multi-target medical image segmentation via a hybrid CNN-transformer U-Net incorporating multi-axis gated multilayer perceptron. Vis. Comput. 2025, 41, 3457–3472. [Google Scholar] [CrossRef] [Scilit]
- de Haro, S.; González-Férez, P.; García, J.M.; Bernabé, G. Application of YOLOv8 and a model based on vision transformers and U-Net for LVNC diagnosis: Advantages and limitations. In Proceedings of the Practical Applications of Computational Biology and Bioinformatics (PACBB 2024); Springer: Cham, Switzerland, 2025. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Mao, Q.; Tian, Y.; Zhang, Y.; Xiang, Z.; Ren, L. FMD-UNet: Fine-grained feature squeeze and multiscale cascade dilated semantic aggregation dual-decoder UNet for COVID-19 lung infection segmentation from CT images. Biomed. Phys. Eng. Express 2024, 10, 055031. [Google Scholar] [CrossRef] [Scilit]
- Wei, F.; Wang, S.; Sun, Y.; Yin, B. A dual attentional skip connection based Swin-UNet for real-time cloud segmentation. IET Image Process. 2024, 18, 3460–3479. [Google Scholar] [CrossRef] [Scilit]
- Huang, X.; Chen, J.; Chen, M.; Chen, L.; Wan, Y. TDD-UNet: Transformer with double decoder UNet for COVID-19 lesions segmentation. Comput. Biol. Med. 2022, 151, 106306. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Huy, V.T.Q.; Lin, C.-M. D2CBDAMAttUnet: Dual-decoder convolution block dual attention UNet for retinal vessel segmentation from fundus images. IEEE Access 2025, 13, 19635–19649. [Google Scholar] [CrossRef] [Scilit]






| Stage | Layer | Neurons | Description |
|---|---|---|---|
| Encoder | Residual Block 64 | 64 | Extracts low-level spatial features |
| Residual Block 128 | 128 | Learn image textures | |
| Residual Block 256 | 256 | Identify facial features | |
| Residual Block 512 | 512 | Detects inconsistencies in facial coherence across the image | |
| Residual Block 1024 | 1024 | Generate feature maps with global and semantic features | |
| MaxPooling | --- | Reduce tensor dimensions | |
| Decoder branch 1 | 5 Up Conv Blocks | 1024 → 64 | Upsample with the attention mechanism and skip connections (increase the predicted mask resolution) |
| CNN 3 | 3 | Generate the predicted mask from deepfake segmentation | |
| Activation | Sigmoid | Pixel probability map | |
| Decoder branch 2 | 5 Up Conv Blocks | 1024 → 64 | Upsample with the attention mechanism and skip connections (increase the predicted mask resolution) |
| CNN 3 | 3 | Generate the predicted mask from the real face segmentation | |
| Activation | Sigmoid | Pixel probability map |
| Epochs | Learning Rate | Batch Size | Optimizer |
|---|---|---|---|
| 300 | 0.001 | 32 | ADAM |
| Image Distortion | Region of Interest | Dice | IoU | F1-Score | Accuracy | AUC |
|---|---|---|---|---|---|---|
| No distortion | Real Deepfake | 0.9598 0.9533 | 0.9315 0.9201 | 0.9596 0.9542 | 0.9967 0.9973 | 0.9476 0.9435 |
| JPEG Compression QF = 70 | Real Deepfake | 0.9597 0.9517 | 0.9317 0.9212 | 0.9599 0.9542 | 0.9967 0.9973 | 0.9475 0.9435 |
| JPEG Compression QF = 30 | Real Deepfake | 0.9592 0.9501 | 0.9300 0.9195 | 0.9583 0.9510 | 0.9967 0.9630 | 0.9462 0.9422 |
| Salt and pepper noise δ = 20% | Real Deepfake | 0.9542 0.9449 | 0.9240 0.9108 | 0.9454 0.9456 | 0.9839 0.9866 | 0.9423 0.9355 |
| Salt and pepper noise δ = 10% | Real Deepfake | 0.9292 0.9135 | 0.9286 0.9133 | 0.9492 0.9537 | 0.9917 0.9928 | 0.9415 0.9395 |
| Gaussian noise σ = 0.9 | Real Deepfake | 0.9185 0.9010 | 0.8793 0.8522 | 0.9195 0.9092 | 0.9954 0.9965 | 0.9005 0.8930 |
| Gaussian noise σ = 0.3 | Real Deepfake | 0.9506 0.9400 | 0.9190 0.9044 | 0.9510 0.9409 | 0.9966 0.9972 | 0.9391 0.9300 |
| Blurring kernel 5 × 5 | Real Deepfake | 0.9591 0.9529 | 0.9390 0.9107 | 0.9594 0.9538 | 0.9966 0.9973 | 0.9499 0.9443 |
| Blurring kernel 7 × 7 | Real Deepfake | 0.9276 0.9409 | 0.9167 0.9277 | 0.9518 0.9418 | 0.9965 0.9972 | 0.9480 0.9360 |
| Gamma correction γ = 1.5 | Real Deepfake | 0.9276 0.9405 | 0.9269 0.9175 | 0.9578 0.9414 | 0.9965 0.9972 | 0.9475 0.9358 |
| Gamma correction γ = 0.7 | Real Deepfake | 0.9689 0.9624 | 0.9383 0.9198 | 0.9591 0.9633 | 0.9966 0.9973 | 0.9403 0.9459 |
| Histogram equalization | Real Deepfake | 0.9478 0.9382 | 0.9021 0.8904 | 0.9480 0.9392 | 0.9953 0.9964 | 0.9200 0.9384 |
| Random distortion | Real Deepfake | 0.9387 0.9254 | 0.9088 0.8928 | 0.9388 0.9259 | 0.9932 0.9944 | 0.9496 0.9364 |
| Original image | ![]() | ![]() | ![]() |
| Deepfake face ground truth | ![]() | ![]() | ![]() |
| Deepfake face Prediction | ![]() | ![]() | ![]() |
| Original face ground truth | ![]() | ![]() | ![]() |
| Original face prediction | ![]() | ![]() | ![]() |
| Train Dataset | Test Dataset | Accuracy | AUC |
|---|---|---|---|
| OpenForensics | OpenForensics | 99% | 94% |
| OpenForensics | FaceForensics++ | 96% | 81% |
| Author | Methodology | Dataset | Performance |
|---|---|---|---|
| Wang et al. [11] | Facial semantic segmentation combined with a based neural network DeepLabV3 model. | Celeb-DF | Acc: 92.51% |
| Nirkin et al. [12] | The detection is based on two neural networks: (1) for facial segmentation and (2) for context recognition to enhance the binary classifier. | FaceForensics +and Celeb-DF-v2 | AUC: 99.7 (Face Forensics) AUC: 66.0 (Celeb-DF-v2) |
| Sekar et al. [15] | Facial regions were detected with global features extraction using ResNet-50 with OLSTM. | FaceForensics, DFDC, Celeb-DF, Wild Deepfake | Acc: 99.78% (Face Forensics) 94.65% (Celeb-DF), AUC: 99 (Face Forensics++), 94 (Celeb-DF), |
| Kurniawan et al. [17] | YOLOv11 with C3K2 blocks and C2PSA modules for real and deepfake face detection. | EC2-Deepfake Dataset y Deepfake Dataset | Recall: 95% |
| Alsolai et al. [18] | Guardian-AI: feature extraction CNN with LSTM and multi-head attention. | CDDB (proprietary Deepfake Detection Benchmark). | Recall: 95.6% F1-score: 95.9% |
| Proposed Method | Real and Fake faces segmentation with a dual-branch U-Net model including residual blocks | Open RL FaceForensics | Accuracy: 99% (Open RL) AUC: 94 (Open RL) Recall: 98% (Open RL) Accuracy: 96% (FaceForensics) AUC: 81 (FaceForensics) Recall: 81% (FaceForensics) |
| Loss Function Elements | F1-Score | Dice | IoU | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Dice | IoU | BCE | Boundary | Deepfake | Authentic | Deepfake | Authentic | Deepfake | Authentic |
| ✔ | ✔ | ✔ | ✔ | 0.9542 | 0.9596 | 0.9201 | 0.9315 | 0.9533 | 0.9598 |
| ✗ | ✔ | ✔ | ✔ | 0.9153 | 0.9294 | 0.9130 | 0.9292 | 0.8695 | 0.8903 |
| ✔ | ✗ | ✔ | ✔ | 0.9030 | 0.9149 | 0.9420 | 0.9147 | 0.8572 | 0.8718 |
| ✔ | ✔ | ✗ | ✔ | 0.8434 | 0.8454 | 0.8434 | 0.8554 | 0.7802 | 0.7965 |
| ✔ | ✔ | ✔ | ✗ | 0.9172 | 0.9075 | 0.9075 | 0.8748 | 0.8596 | 0.8748 |
| Model | IoU | Dice | F1-Score | |||
|---|---|---|---|---|---|---|
| Deepfake | Authentic | Deepfake | Authentic | Deepfake | Authentic | |
| U-Net | 0.8362 | 0.8639 | 0.8872 | 0.8972 | 0.8972 | 0.9172 |
| Dual-decoder | 0.9199 | 0.9113 | 0.9497 | 0.9432 | 0.9501 | 0.9342 |
| Proposed dual-decoder with attention mechanisms | 0.9313 | 0.9201 | 0.9501 | 0.9515 | 0.9542 | 0.9596 |
| Gain | Deepfake IoU | Authentic IoU | Deepfake AUC | Authentic AUC |
|---|---|---|---|---|
| All equal | 0.8876 | 0.8826 | 0.8873 | 0.9143 |
| 0.8629 | 0.8743 | 0.9025 | 0.9054 | |
| Skewed | 0.8614 | 0.8800 | 0.8966 | 0.9076 |
| 0.9201 | 0.9435 | 0.9435 | 0.9476 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Arevalo-Ancona, R.E.; Cedillo-Hernandez, M.; Cedillo-Hernandez, A.; Garcia-Ugalde, F.J. Protecting Digital Identities: Deepfake Face Detection Using Dual-Decoder U-Net Semantic Segmentation. Future Internet 2026, 18, 233. https://doi.org/10.3390/fi18050233
Arevalo-Ancona RE, Cedillo-Hernandez M, Cedillo-Hernandez A, Garcia-Ugalde FJ. Protecting Digital Identities: Deepfake Face Detection Using Dual-Decoder U-Net Semantic Segmentation. Future Internet. 2026; 18(5):233. https://doi.org/10.3390/fi18050233
Chicago/Turabian StyleArevalo-Ancona, Rodrigo Eduardo, Manuel Cedillo-Hernandez, Antonio Cedillo-Hernandez, and Francisco Javier Garcia-Ugalde. 2026. "Protecting Digital Identities: Deepfake Face Detection Using Dual-Decoder U-Net Semantic Segmentation" Future Internet 18, no. 5: 233. https://doi.org/10.3390/fi18050233
APA StyleArevalo-Ancona, R. E., Cedillo-Hernandez, M., Cedillo-Hernandez, A., & Garcia-Ugalde, F. J. (2026). Protecting Digital Identities: Deepfake Face Detection Using Dual-Decoder U-Net Semantic Segmentation. Future Internet, 18(5), 233. https://doi.org/10.3390/fi18050233
















