Deepfakes and Synthetic Media: Generation, Detection, and Governance
Abstract
1. Definition, Classification and Why Deepfake Is Important
- 1.
- Modality deepfakes
- Visual: face swaps, reenactment, lip-sync, attribute editing (age, expression), full-body synthesis, scene relighting.
- Audio: voice cloning, speaker identity conversion, speech-to-speech conversion, text-to-speech impersonation.
- Multimodal: synchronized audio–video generation, avatar systems, “talking head” models with cloned voice.
- 2.
- Manipulation intent deepfakes
- Identity substitution (impersonation, fraud, non-consensual synthetic content).
- Event fabrication (false evidence, fake statements, fake presence).
- Contextual distortion (true footage reframed via synthetic overlays, selective edits, or deceptive narration).
- 3.
- Generation regime deepfakes
- Closed-world generation (trained on a specific target identity).
- Open-world generation (foundation models enabling broad, low-friction synthesis).
1.1. Why Are Deepfakes Important?
- Evidentiary erosion: Over time, authentic recordings become easier to dismiss as fake (e.g., someone might state: “this is not genuine, it could be AI”).
1.2. Why Is Detection Necessary but Not Always Sufficient?
- Policy and compliance controls, including transparency duties for certain AI outputs and platform obligations for risk mitigation and accountability. In the European Union context, deepfakes intersect directly with:
1.3. Human Factors and Cognitive Vulnerabilities
2. Creating Deepfakes: Types of Fakes, Standard Pipelines, and Generative Models
2.1. Common Deepfake Types
- Face swapping: The target’s face is replaced with the source’s face to preserve the pose, lighting, and background. Historically, this was based on GAN families and evolved into high-fidelity models (e.g., the StyleGAN line) [26,27,28]. Such techniques are especially problematic for sensitive communications, including breaking news journalism, legal evidence submissions, and crisis management broadcasts.
- Audio deepfakes (voice cloning/conversion): Speech is generated or transformed to mimic a specific speaker (voice conversion/text-to-speech), often as part of a full audiovisual deepfake [33,34,35,36]. For example, in military contexts, this may support voice-based impersonation or false command messages.
- Text-to-video/full scene synthesis: Video sequences (and not just “doctored” faces) are generated using diffusion-based video generation and text-to-video approaches that expand the threat from “evidence tampering” to “event fabrication” [37,38]. This is crucial as it creates a risk of fabricated military events, scenes, predictions, or operational incidents [39].
2.2. Typical Production Pipeline
- 1.
- Data collection/selection: sufficient variety of poses, expressions, and lighting (especially for identity-specific models).
- This stage is technically important as the model learns to identify specific variations from repeated examples under different poses, illumination conditions, facial expressions, and camera qualities. Limited or biased training material often leads to failures under unseen angles, lighting, or movements, which later become detectable artifacts.
- 2.
- Localization/normalization: face detection, landmarks, alignment, cropping, and photometric normalization.
- The unconstrained visual input of this stage is preprocessed to yield a normalized facial representation. In this stage, the system focuses on localizing the face, estimating landmarks, aligning the facial region, and reducing background variation. Nonetheless, these methods might also result in geometric warping, disruptions at boundaries, or improper blending that forensic detectors can use.
- 3.
- Model training or adaptation: either general-purpose (foundation-style) or tailored to a specific individual/target.
- At this point, the model will learn to map representations of the input data, e.g., facial landmarks, identity embeddings, motion parameters, and acoustic features, to realistic synthetic output. In identity-specific systems, adaptation or fine-tuning enables the model to reproduce the appearance, voice, or movement patterns of a particular target more convincingly at a cost of overfitting to the material available in training.
- 4.
- Inference and temporal consistency: especially in video, temporal consistency is crucial for perceptual plausibility.
- At this stage, when a trained model is evaluated, it creates altered frames or audio segments from what it has learned, typically within a framework of the target pose, expression, or speech signal. Temporal consistency in video deepfakes is challenging in practice because each frame must be consistent with the previous and next one in terms of facial geometry, illumination, lip motion, eye movement, and background continuity.
- 5.
- Post-processing: blending, color matching, denoising/oversampling, and final re-encoding, which often “hides” obvious traces and shifts detection to more subtle statistical cues.
- The last stage of the post-processing aims to ensure the produced region appears visually compatible with the rest of the image or video. This includes boundary smoothing, color difference correction, noise reduction, and adjustment to the platform’s encoding format. While these operations can mask obvious marks, they may also introduce other forensic signs, such as abnormal compression patterns, spectral irregularities, or more subtle inconsistencies between the manipulated region and the original background.
2.3. Major Families of Generative Models
2.4. Transition to Multimodal Models
3. Deepfake Characteristics and Detectable Properties
3.1. Deepfake Categories for Audio and Visual Objects
- Entire face synthesis: the creation of an entirely new face (or even an entire image/scene) that does not correspond to a real, recorded person.
- Identity swap/face swap: replacing one subject’s face with another’s, typically while preserving the “host’s” pose, lighting, and motion.
- Attribute manipulation: changes to specific characteristics (age, gender/expression, morphological features), without necessarily changing the identity.
- Expression swap/reenactment: we “transfer” the expressions/movements of a target from a source so that the mouth, eyebrows, and micro-expressions match another video or a live source. A classic (predating modern deep models but fundamental) family of reenactment techniques is exemplified in projects such as [29].
3.2. Generation-Side Artifacts and Observable Traces
- Spatial/morphological artifacts: Examples include imperfect blending at the face–skin/hair boundaries, unrealistic geometry in teeth/lips, inconsistencies in shading/lighting, or a “mismatched” background in high-frequency details. Such indicators are systematized in approaches that analyze visual artifacts as production “signatures” [59,60].
- Frequency-domain traces: Many generative models leave traces in the frequency spectrum due to resampling and upsampling (e.g., “regularities” not systematically found in natural images). Frequency-domain analysis has been proposed as a complementary “channel” of evidence, especially when spatial cues are attenuated by compression [63,64].
- Physiological cues: A more recent line of thinking capitalizes on the fact that real-time facial video incorporates subtle physiological changes (e.g., remote photoplethysmography based on micro-changes in skin color). Deviations in such signals can serve as an indication of synthetic origin [65,66].
4. Detection and Evaluation of Deepfakes: Methodologies, Data, and Benchmarks
- Authenticity classification (real vs. fake);
- Spatial/temporal localization of alterations;
- Classification of forgery types (e.g., swap, reenactment, synthesis).
4.1. Cues and “Signatures” of Forgery
- Spatial features: Detectors convert local visual inconsistencies into measurable image features, such as texture irregularities, blending boundaries, facial warping, lighting mismatch, or abnormal local details. These features are commonly extracted from face crops, patches, or facial regions and are often used by convolutional neural network (CNN)-based and mesoscopic detectors [59,80,81,82,83].
- Temporal features: Video detectors model the evolution of facial regions across consecutive frames. Instead of only observing that a video contains flickering or unnatural blinking, these methods measure frame-to-frame instability, motion inconsistency, lip-motion irregularity, or temporal feature drift using temporal aggregation, 3D CNNs, recurrent models, or transformers [61,62,84].
- Frequency-domain features: Spectral detectors transform images or video frames into frequency representations and search for abnormal patterns introduced by upsampling, resampling, compression, or generative reconstruction. These features are especially useful when pixel-level artifacts are visually subtle or partly hidden by post-processing [63,64].
- Physiological and acquisition-based features: Some detectors estimate biological or sensor-origin signals, such as rPPG, blink behavior, camera fingerprints, sensor noise, or compression history. These features are used to test whether the content remains consistent with natural human physiology or real camera acquisition [65,67,68].
- Multimodal consistency features: In audiovisual deepfakes, detection may also examine whether the speech signal, speaker identity, lip motion, and facial dynamics agree with each other. This is important because a video may look plausible visually while still showing audiovisual synchronization or speaker-consistency errors [69,79].
4.2. Detector Families: From Frame-Level to Multimodal
- Video-level detectors (spatio-temporal models): They incorporate temporal information (e.g., 3D CNN, temporal aggregation, transformers) and tend to improve detection in cases where the forgery “escapes” spatially but remains temporally inconsistent [84].
- Compact/mesoscopic architectures: for example, Ref. [83] presented a compact architecture for face-forgery detection, targeting meso-level features that remain useful under standard compression.
- Multimodal detection (audio–video): As deepfake production shifts toward full audiovisual synthesis, the integration of audio and video becomes particularly important. Datasets such as the ones used in [79] were designed specifically for evaluating multimodal scenarios (face + voice), while anti-spoofing benchmarks for speech support systematic evaluation of synthetic/transformed speech were also developed at times [69].
4.3. Datasets and Benchmarks: What We Measure and Why It Matters
- Celeb-DF: designed as a more challenging dataset to reduce the “ease” of detection via simplistic artifacts and push for generalization [73].
- DeeperForensics-1.0 focuses on conditions that closely resemble real-world pipelines and highlights the implications of cross-dataset evaluation [74].
- WildDeepfake: compiles an “in-the-wild” collection, highlighting the drop in performance when detectors are applied to real-world internet conditions [75].
- FaceForensics++ serves as a classic evaluation dataset for manipulated facial images and is widely used in experimental comparisons [80].
4.4. Evaluation Protocols and Metrics: From AUROC to Operational Reliability
- 1.
- Intra-dataset vs. cross-dataset evaluation: Cross-dataset performance better captures domain shift and is closer to real-world conditions, where the generator or compression channel is not known in advance [73,74,75,85].
- Intra-dataset evaluation determines whether a detector can detect similar patterns to those seen during training, while cross-dataset evaluation tests whether the learned indications are still valid when the manipulation method, source camera, level of compression, or distribution platform changes. Because of this distinction, many detectors learn dataset-specific artifacts instead of general forensic evidence of manipulation.
- 2.
- Classification metrics and class imbalance: In addition to accuracy, AUROC and Area Under the Precision-Recall Curve (AUPRC) are used (especially in cases of class imbalance), while in anti-spoofing scenarios, metrics such as Equal Error Rate (EER) are standard [69].
- While these metrics technically measure different things, they reflect the detector’s behavior. Area Under the Receiver Operating Characteristic Curve (AUROC) measures the ability to distinguish between real and fake samples over thresholds. AUPRC is informative when the fake content is rare. EER is the operating point at which false acceptance and false rejection are equal. A detector that performs well in tests may be of limited operational use if the threshold has been poorly calibrated.
- 3.
- Robustness tests: Tests under re-encoding, resolution changes, cropping, filtering, and platform transformations are considered essential, as these steps often “neutralize” surface artifacts [71,72,73,74,75,79].
- From a more technical standpoint, these transformations change the pixel, frequency, and compression characteristics of the media. As a result, they can weaken or even remove the artifacts that the detector uses for classification. In other words, robustness testing evaluates whether the detector is solely dependent on fragile surface signatures or if it can maintain effectiveness in realistic degradation and redistribution.
- 4.
- Calibration and decision thresholds: In operational scenarios, the output is not merely a “label,” but a risk score that leads to action (e.g., human review, takedown, posting ban). Therefore, calibration and threshold selection are part of the evaluation.
- Calibration refers to the notion of whether the confidence score of a detector corresponds to the actual probability of a manipulation instead of a high or low number. The next step determines the evidence needed to trigger an action. It is particularly important in high-stakes situations where false positives can harm trust in authentic content and false negatives can allow harmful synthetic media to circulate.
4.5. Recent Detector Models and Reported Benchmark Performance
5. Mitigation, Governance, and Regulatory Compliance for Synthetic Content
- A detector’s performance on controlled data does not guarantee operational reliability under varying dissemination channels, and
5.1. Threat Identification, Risk Modeling, and Transparency Disclosure Controls
5.2. Risk Governance, Management Systems, and Technical Content Transparency
- Testing, evaluation, verification, and validation (TEVV) (testing–evaluation–verification–validation) for detection/labeling/provenance tools;
- Assurance case (documented argumentation that the system is “sufficiently secure/reliable” for a specific use);
- Change management (what changes when the codec, model provider, or platform changes);
- Monitoring and drift management (monitoring of performance degradation/increase in errors).
- Detection signals (probabilistic evidence from Machine Learning detectors);
- Verifiable evidence that can be verified cryptographically or through structured manifests.
5.3. Watermarking Resilience, Operational Integration, and Forensics-Grade Documentation
6. Implications and Scenarios of Abuse: A Case-Based Analysis in Key Application Areas
6.1. Case A—News, Journalism, and Fact-Checking (Newsroom Verification)
6.2. Case B—Financial Fraud, Corporate Security, and Remote Identity Verification (Know Your Customer/Remote Onboarding)
6.3. Case C—Public Sector, Law Enforcement, and Forensic Use (Forensics, Chain of Custody)
6.4. Case D—Education, Academic Integrity, and the Protection of Students
6.5. Synthetic Assessment: Effectiveness as a System Property
7. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
DURC Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| AI | Artificial Intelligence |
| AIMS | AI Management System |
| ATSC | Advanced Television Systems Committee |
| AUPRC | Area Under the Precision-Recall Curve |
| AUROC | Area Under the Receiver Operating Characteristic Curve |
| C2PA | Coalition for Content Provenance and Authenticity |
| CAI | Content Authenticity Initiative |
| CNN | Convolutional Neural Network |
| DDIM | Denoising Diffusion Implicit Model |
| DDPM | Denoising Diffusion Probabilistic Model |
| DFDC | DeepFake Detection Challenge |
| DSA | Digital Services Act |
| DURC | Dual Use Research of Concern |
| EER | Equal Error Rate |
| EU | European Union |
| GAN | Generative Adversarial Network |
| GDPR | General Data Protection Regulation |
| GPT | Generative Pre-trained Transformer |
| ISO/IEC | International Organization for Standardization/International Electrotechnical Commission |
| NIST | National Institute of Standards and Technology |
| OECD | Organisation for Economic Co-operation and Development |
| RGB | Red–Green–Blue |
| rPPG | Remote Photoplethysmography |
| TEVV | Testing, Evaluation, Verification, and Validation |
| UNESCO | United Nations Educational, Scientific and Cultural Organization |
References
- Verdoliva, L. Media Forensics and DeepFakes: An Overview. IEEE J. Sel. Top. Signal Process. 2020, 14, 910–932. [Google Scholar] [CrossRef]
- Akhtar, Z. Deepfakes generation and detection: A short survey. J. Imaging 2023, 9, 18. [Google Scholar] [CrossRef]
- Tolosana, R.; Vera-Rodriguez, R.; Fierrez, J.; Morales, A.; Ortega-Garcia, J. Deepfakes and Beyond: A Survey of Face Manipulation and Fake Detection. Inf. Fusion 2020, 64, 131–148. [Google Scholar] [CrossRef]
- Maxmudjanov, S.; Naimov, A. Deepfake Content Types And Their Generation Methods. Пoтoмки Aль-Φapгaни 2026, 1, 26–31. [Google Scholar] [CrossRef]
- Rainie, L.; Anderson, J.; Vogels, E.A. Experts doubt ethical AI design will be broadly adopted as the norm within the next decade. Pew Res. Cent. 2021, 16, 121–154. [Google Scholar]
- Mahr, D. Sexualized deepfakes as a socio-technical continuation of gendered power. AI Soc. 2026, 1–13. [Google Scholar] [CrossRef]
- Twomey, J. Socially (de) Constructing Deepfakes: Understanding Socio-Technical Interests and Concerns in Deepfake Media. Ph.D. Thesis, University College Cork, Cork, Ireland, 2025. Available online: https://hdl.handle.net/10468/18856 (accessed on 1 June 2026).
- Perriello, L.E. Blurred realities: Legal strategies for the deepfake era. Maastricht J. Eur. Comp. Law 2026, 1023263X261433380. [Google Scholar] [CrossRef]
- Zhang, B. Governing deepfake realities: The truth immunity framework for telecommunications policy. Telecommun. Policy 2026, 50, 103168. [Google Scholar] [CrossRef]
- NIST. Artificial Intelligence Risk Management Framework (AI RMF 1.0); NIST AI 100-1; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2023. [Google Scholar] [CrossRef]
- Soto-Sanfiel, M.T.; Wu, Q. How audiences make sense of deepfake resurrections: A multilevel analysis of realism, ethics, and cultural meaning. Comput. Hum. Behav. 2025, 174, 108822. [Google Scholar] [CrossRef]
- Putra, A.B. The Legal Standing of Deepfake Digital Evidence in Criminal Proceedings: Challenges of Evidentiary Integrity in the AI Era. Fox Justi J. Ilmu Huk. 2026, 16, 283–290. [Google Scholar]
- Cheng, E.K. Deepfakes, Photographs, and Trust in Evidence. Va. Law Rev. Online 2025. [Google Scholar] [CrossRef]
- Peebles, W.; Xie, S. Scalable Diffusion Models with Transformers. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; pp. 4172–4182. [Google Scholar] [CrossRef]
- Barari, S.; Lucas, C.; Munger, K. Political deepfakes are as credible as other fake media and (sometimes) real media. J. Polit. 2025, 87, 510–526. [Google Scholar] [CrossRef]
- Soundarya, B.C.; Gururaj, H.L. Deepfake detection: Critical review of state-of-the-art approaches and future perspectives. Discov. Appl. Sci. 2026, 8, 201. [Google Scholar] [CrossRef]
- Sharma, S.; Selwal, A. Potential of artificial intelligence in deepfake media: From generation to detection mechanisms, state-of-the-art, and challenges. Comput. Sci. Rev. 2026, 60, 100866. [Google Scholar] [CrossRef]
- Erokhin, D.; Komendantova, N. A Review of Tools and Technologies to Combat Deepfakes. Information 2026, 17, 347. [Google Scholar] [CrossRef]
- Dhanapal, R.; Ps, A. Defending Against Adaptive Deepfake Generation and Detection Evasion Attacks. In Proceedings of the 2026 9th International Conference on Inventive Computation Technologies (ICICT), Kirtipur, Nepal, 15–17 April 2026; pp. 1619–1624. [Google Scholar] [CrossRef]
- Modi, K. The Deepfake Conundrum: Assessing Generative AI’s Threat to Digital Reality and Proposing a Multi-Layered Defense Framework. Int. J. Comput. Trends Technol. 2025, 73, 97–103. [Google Scholar] [CrossRef]
- Coalition for Content Provenance and Authenticity (C2PA). Content Credentials: C2PA Technical Specification, Version 2.2; C2PA, 2025. Available online: https://spec.c2pa.org/specifications/specifications/2.2/specs/_attachments/C2PA_Specification.pdf (accessed on 1 June 2026).
- European Parliament; Council of the European Union. Regulation (EU) 2024/1689 (Artificial Intelligence Act). The Official Journal of the European Union: Luxembourg, 2024. Available online: https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng (accessed on 1 June 2026).
- European Parliament; Council of the European Union. Regulation (EU) 2022/2065 (Digital Services Act). The Official Journal of the European Union: Luxembourg, 2022; Volume L 277, pp. 1–102. Available online: https://eur-lex.europa.eu/eli/reg/2022/2065/oj/eng (accessed on 1 June 2026).
- Wulf, A.J.; Seizov, O. “Please understand we cannot provide further information”: Evaluating content and transparency of GDPR-mandated AI disclosures. AI Soc. 2024, 39, 235–256. [Google Scholar] [CrossRef]
- Romero Moreno, F. Generative AI and deepfakes: A human rights approach to tackling harmful content. Int. Rev. Law Comput. Technol. 2024, 38, 297–326. [Google Scholar] [CrossRef]
- Goodfellow, I.J.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative Adversarial Nets. arXiv 2014, arXiv:1406.2661. [Google Scholar] [CrossRef]
- Karras, T.; Laine, S.; Aittala, M.; Hellsten, J.; Lehtinen, J.; Aila, T. Analyzing and Improving the Image Quality of StyleGAN. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020), Seattle, WA, USA, 13–19 June 2020; pp. 8110–8119. [Google Scholar] [CrossRef]
- Waseem, S.; Bakar, S.A.R.S.A.; Ahmed, B.A.; Omar, Z.; Eisa, T.A.E.; Dalam, M.E.E. DeepFake on face and expression swap: A review. IEEE Access. 2023, 11, 117865–117906. [Google Scholar] [CrossRef]
- Thies, J.; Zollhöfer, M.; Stamminger, M.; Theobalt, C.; Nießner, M. Face2Face: Real-Time Face Capture and Reenactment of RGB Videos. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016), Las Vegas, NV, USA, 27–30 June 2016; pp. 2387–2395. [Google Scholar] [CrossRef]
- Dhanyalakshmi, R.; Popirlan, C.I.; Hemanth, D.J. A survey on deep learning based reenactment methods for deepfake applications. IET Image Process. 2024, 18, 4433–4460. [Google Scholar] [CrossRef]
- Prajwal, K.R.; Mukhopadhyay, R.; Namboodiri, V.; Jawahar, C.V. A Lip Sync Expert Is All You Need for Speech to Lip Generation in the Wild. In Proceedings of the 28th ACM International Conference on Multimedia (ACM MM 2020), Virtual Event, 12–16 October 2020; pp. 484–492. [Google Scholar] [CrossRef]
- Xiong, X.; Patel, P.; Fan, Q.; Wadhwa, A.; Selvam, S.; Guo, X.; Qi, L.; Liu, X.; Sengupta, R. Talkingheadbench: A multi-modal benchmark & analysis of talking-head deepfake detection. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Tucson, AZ, USA, 6–10 March 2026; pp. 4139–4149. Available online: https://doi.ieeecomputersociety.org/10.1109/WACV61042.2026.00403 (accessed on 1 June 2026).
- Qian, K.; Zhang, Y.; Chang, S.; Yang, X.; Hasegawa-Johnson, M. AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss. arXiv 2019, arXiv:1905.05879. [Google Scholar] [CrossRef]
- Kong, J.; Kim, J.; Bae, J. HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis. arXiv 2020, arXiv:2010.05646. [Google Scholar] [CrossRef]
- van den Oord, A.; Dieleman, S.; Zen, H.; Simonyan, K.; Vinyals, O.; Graves, A.; Kalchbrenner, N.; Senior, A.; Kavukcuoglu, K. WaveNet: A Generative Model for Raw Audio. arXiv 2016, arXiv:1609.03499. [Google Scholar] [CrossRef]
- Alnaqbi, M.; Ikuesan, R.A. A systematic review of audio deepfake detection techniques for digital investigation. Discov. Comput. 2026, 29, 202. [Google Scholar] [CrossRef]
- Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; Fleet, D.J. Video Diffusion Models. arXiv 2022, arXiv:2204.03458. [Google Scholar] [CrossRef]
- Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; et al. Make-A-Video: Text-to-Video Generation without Text-Video Data. arXiv 2022, arXiv:2209.14792. [Google Scholar] [CrossRef]
- Vu, P.T. A Fusion Model for Precipitation Nowcasting from Radar and Satellite. In Multi-Disciplinary Trends in Artificial Intelligence: 18th International Conference, MIWAI 2025, Ho Chi Minh City, Vietnam, 3–5 December 2025, Proceedings, Part I, 16353, 430; Springer Nature: Berlin/Heidelberg, Germany, 2025. [Google Scholar] [CrossRef]
- Li, Q.; Wang, W.; Du, S.; Peng, B.; Dong, J.; Wang, K.; Sun, Z.; Yang, M.H. Towards High Fidelity Face Swapping: A Comprehensive Survey and New Benchmark. arXiv 2026, arXiv:2605.00883. [Google Scholar] [CrossRef]
- Zhu, J.-Y.; Park, T.; Isola, P.; Efros, A.A. Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks. In Proceedings of the IEEE International Conference on Computer Vision (ICCV 2017), Venice, Italy, 22–29 October 2017; pp. 2242–2251. [Google Scholar] [CrossRef]
- Choi, Y.; Choi, M.; Kim, M.; Ha, J.-W.; Kim, S.; Choo, J. StarGAN: Unified Generative Adversarial Networks for Multi-Domain Image-to-Image Translation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2018), Salt Lake City, UT, USA, 18–22 June 2018; pp. 8789–8797. [Google Scholar] [CrossRef]
- Liu, X.; Xu, R.; Chen, Y. Securing Digital Media Integrity: A Survey of Watermarking and Manipulation Detection for Image Authentication. Authorea 2025. preprint. [Google Scholar] [CrossRef] [PubMed]
- Thies, J.; Zollhöfer, M.; Nießner, M. Deferred Neural Rendering: Image Synthesis using Neural Textures. ACM Trans. Graph. 2019, 38, 66. [Google Scholar] [CrossRef]
- He, Z.; Henderson, P.; Pugeault, N. Beyond Reconstruction: A Physics Based Neural Deferred Shader for Photo-Realistic Rendering. In Proceedings of the International Conference on Artificial Neural Networks, September 2025; Springer Nature Switzerland: Cham, Switzerland, 2025; pp. 378–389. [Google Scholar] [CrossRef]
- Svitov, D.; Dahaghin, M. NBAvatar: Neural Billboards Avatars with Realistic Hand-Face Interaction. arXiv 2026, arXiv:2603.12063. [Google Scholar] [CrossRef]
- Ho, J.; Jain, A.; Abbeel, P. Denoising Diffusion Probabilistic Models. arXiv 2020, arXiv:2006.11239. [Google Scholar] [CrossRef]
- Zhang, X.; Chen, C. Parameter-efficient quantum denoising diffusion probabilistic models with temporal encoding. Future Gener. Comput. Syst. 2026, 174, 107981. [Google Scholar] [CrossRef]
- Song, J.; Meng, C.; Ermon, S. Denoising Diffusion Implicit Models. arXiv 2020, arXiv:2010.02502. [Google Scholar] [CrossRef]
- Zhang, Q.; Tao, M.; Chen, Y. gddim: Generalized denoising diffusion implicit models. arXiv 2022, arXiv:2206.05564. [Google Scholar] [CrossRef]
- Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; Ommer, B. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022), New Orleans, LA, USA, 19–24 June 2022; pp. 10684–10695. [Google Scholar] [CrossRef]
- Lin, Z.; Wang, X.; Zhang, S.; Wang, T.; Guo, Z.; Liu, T. TSCM: Efficient Image Synthesis Using Latent Diffusion. Pattern Recognit. Lett. 2026, 202, 57–62. [Google Scholar] [CrossRef]
- Mildenhall, B.; Srinivasan, P.P.; Tancik, M.; Barron, J.T.; Ramamoorthi, R.; Ng, R. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. arXiv 2020, arXiv:2003.08934. [Google Scholar] [CrossRef]
- Wang, K.; Wei, K.; Li, S.Y. Dynamic view synthesis with topologically-varying neural radiance fields from sparse input views. Neurocomputing 2026, 674, 132942. [Google Scholar] [CrossRef]
- Karypidis, E.; Mouslech, S.G.; Skoulariki, K.; Gazis, A. Comparison Analysis of Traditional Machine Learning and Deep Learning Techniques for Data and Image Classification. WSEAS Trans. Math. 2022, 21, 122–130. [Google Scholar] [CrossRef]
- Altuncu, E.; Franqueira, V.N.L.; Li, S. Deepfake: Definitions, performance metrics and standards, datasets and future directions. Front. Big Data 2024, 7, 1400024. [Google Scholar] [CrossRef] [PubMed]
- Pei, G.; Zhang, J.; Hu, M.; Zhang, Z.; Wang, C.; Wu, Y.; Zhai, G.; Yang, J.; Tao, D. Deepfake generation and detection: A benchmark and survey. ACM Comput. Surv. 2026, 58, 1–41. [Google Scholar] [CrossRef] [PubMed]
- Hossain, S.; Sudarsan, D.; Zaffar, H.; Ahmad, N. Social and psychological impact of deepfakes: A comprehensive bibliometric review. Glob. Knowl. Mem. Commun. 2026, 1–20. [Google Scholar] [CrossRef]
- Matern, F.; Riess, C.; Stamminger, M. Exploiting Visual Artifacts to Expose Deepfakes and Face Manipulations. In Proceedings of the IEEE Winter Applications of Computer Vision Workshops (WACVW), Waikoloa, HI, USA, 7–11 January 2019; pp. 83–92. [Google Scholar] [CrossRef]
- Syed Abu Bakar, S.A.R.; Waseem, S.; Omar, Z.; Bilalashfaqahmed. Exploring the Advancements and Challenges of Deepfake Face-swap: A Survey. Multimed. Tools Appl. 2026, 85, 14. [Google Scholar] [CrossRef]
- Li, Y.; Chang, M.-C.; Lyu, S. In Ictu Oculi: Exposing AI Generated Fake Face Videos by Detecting Eye Blinking. In Proceedings of the IEEE International Workshop on Information Forensics and Security (WIFS), Hong Kong, China, 11–13 December 2018. [Google Scholar] [CrossRef]
- Sar, A.; Roy, S.; Choudhury, T.; Abraham, A. Zero-shot visual deepfake detection: Can ai predict and prevent fake content before it is created? Found. Trends Signal Process. 2025, 19, 212–361. [Google Scholar] [CrossRef]
- Frank, J.; Eisenhofer, T.; Schönherr, L.; Fischer, A.; Kolossa, D.; Holz, T. Leveraging Frequency Analysis for Deep Fake Image Recognition. arXiv 2020, arXiv:2003.08685. [Google Scholar] [CrossRef]
- Hamadene, A.; Allili, M.S. Cross-Model Deepfake Detection Through Contourlet-Based Inter-Channel Spectral Analysis. IEEE Trans. Biom. Behav. Identity Sci. 2026. [Google Scholar] [CrossRef]
- Ciftci, U.A.; Demir, I.; Yin, L. FakeCatcher: Detection of Synthetic Portrait Videos Using Biological Signals. IEEE Trans. Pattern Anal. Mach. Intell. 2020. [Google Scholar] [CrossRef] [PubMed]
- Jędrasiak, K.; Bijoch, J. Physiological and morphometric biomarkers for synthetic media detection. Front. Bioeng. Biotechnol. 2026, 14, 1781235. [Google Scholar] [CrossRef] [PubMed]
- Cozzolino, D.; Verdoliva, L. Noiseprint: A CNN-Based Camera Model Fingerprint. IEEE Trans. Inf. Forensics Secur. 2019, 15, 144–159. [Google Scholar] [CrossRef]
- Siddiqui, N.; Islam, S. An Efficient Feature-Based Framework for Camera Model Identification. IEEE Access 2026, 14, 69972–69997. [Google Scholar] [CrossRef]
- Todisco, M.; Wang, X.; Vestman, V.; Sahidullah, M.; Delgado, H.; Nautsch, A.; Yamagishi, J.; Evans, N. ASVspoof 2019: A Large-Scale Public Database of Synthesized, Converted and Replayed Speech. Comput. Speech Lang. 2021, 64, 101114. [Google Scholar] [CrossRef]
- Langmia, K. Black Communication in the Age of Disinformation: DeepFakes and Synthetic Media; Springer Nature: Cham, Switzerland, 2023. [Google Scholar] [CrossRef]
- Dolhansky, B.; Bitton, J.; Pflaum, B.; Lu, J.; Howes, R.; Wang, M.; Canton Ferrer, C. The DeepFake Detection Challenge (DFDC) Dataset. arXiv 2020, arXiv:2006.07397. [Google Scholar] [CrossRef]
- Meta, A.I. Deepfake Detection Challenge Dataset (DFDC). Available online: https://ai.meta.com/datasets/dfdc/ (accessed on 1 June 2026).
- Li, Y.; Yang, X.; Sun, P.; Qi, H.; Lyu, S. Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020), Seattle, WA, USA, 13–19 June 2020; pp. 3207–3216. [Google Scholar] [CrossRef]
- Jiang, L.; Li, R.; Wu, W.; Qian, C.; Loy, C.C. DeeperForensics-1.0: A Large-Scale Dataset for Real-World Face Forgery Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2020), Seattle, WA, USA, 13–19 June 2020; pp. 2889–2898. [Google Scholar] [CrossRef]
- Zi, B.; Chang, M.; Chen, J.; Ma, X.; Jiang, Y.-G. WildDeepfake: A Challenging Real-World Dataset for Deepfake Detection. In Proceedings of the 28th ACM International Conference on Multimedia (MM ‘20), Virtual Event, 12–16 October 2020; pp. 2382–2390. [Google Scholar] [CrossRef]
- He, Y.; Gan, B.; Chen, S.; Zhou, Y.; Yin, G.; Song, L.; Sheng, L.; Shao, J.; Liu, Z. ForgeryNet: A Versatile Benchmark for Comprehensive Forgery Analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2021), Nashville, TN, USA, 20–25 June 2021; pp. 4360–4369. [Google Scholar] [CrossRef]
- He, Y.; Sheng, L.; Shao, J.; Liu, Z.; Zou, Z.; Guo, Z.; Jiang, S.; Sun, C.; Zhang, G.; Wang, K.; et al. ForgeryNet—Face Forgery Analysis Challenge 2021: Methods and Results. arXiv 2021, arXiv:2112.08325. [Google Scholar] [CrossRef]
- CodaLab. ForgeryNet—Face Forgery Analysis Challenge 2021 (Competition Page). Available online: https://competitions.codalab.org/competitions/33386 (accessed on 1 June 2026).
- Khalid, H.; Tariq, S.; Kim, M.; Woo, S.S. FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset. arXiv 2021, arXiv:2108.05080. [Google Scholar] [CrossRef]
- Rössler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; Nießner, M. FaceForensics++: Learning to Detect Manipulated Facial Images. arXiv 2019. [Google Scholar] [CrossRef]
- Li, Y.; Lyu, S. Exposing DeepFake Videos by Detecting Face Warping Artifacts. arXiv 2018, arXiv:1811.00656. [Google Scholar] [CrossRef]
- Bai, W.; Liu, Y.; Zhang, A.; Wang, Y.; Li, B.; Hu, W.; Zhang, Z. Deepfake Detection via Exploring Degradation Inconsistency. IEEE Trans. Inf. Forensics Secur. 2026, 21, 5627–5642. [Google Scholar] [CrossRef]
- Afchar, D.; Nozick, V.; Yamagishi, J.; Echizen, I. MesoNet: A Compact Facial Video Forgery Detection Network. In Proceedings of the IEEE International Workshop on Information Forensics and Security (WIFS), Hong Kong, China, 11–13 December 2018. [Google Scholar] [CrossRef]
- Alanazi, S.; Asif, S. VIDS-Guard: A Novel Forensics-Aware Multi-Stream Transformer Framework for Robust Deepfake Video Detection. Int. Syst. Appl. 2026, 30, 200664. [Google Scholar] [CrossRef]
- Yan, Z.; Luo, Y.; Lyu, S.; Liu, Q.; Wu, B. Transcending Forgery Specificity with Latent Space Augmentation for Generalizable Deepfake Detection. In Proceedings of the 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 16–22 June 2024; pp. 8984–8994. [Google Scholar] [CrossRef]
- Tariq, R.; Heo, M.; Tariq, S.; Woo, S.; Tariq, S. Through the Lens: Benchmarking Deepfake Detectors Against Moiré-Induced Distortions. arXiv 2025. [Google Scholar] [CrossRef]
- Mesa-Simón, M.; Escobar-Molero, A.; Sáez-Mingorance, B.; Morales, D.P.; Álvarez-Bermejo, J.A.; Romero, F.J. Enabling live video provenance and authenticity: A C2PA-based system with TPM-based security for livestreaming platforms. IEEE Trans. Multimed. 2026, 1–12. [Google Scholar] [CrossRef]
- Content Authenticity Initiative (CAI). Video Provenance and the Ethics of Deepfakes (Blog Post). 2021. Available online: https://contentauthenticity.org/ (accessed on 1 June 2026).
- Nastoska, A.; Jancheska, B.; Rizinski, M.; Trajanov, D. Evaluating trustworthiness in AI: Risks, metrics, and applications across industries. Electronics 2025, 14, 2717. [Google Scholar] [CrossRef]
- Coalition for Content Provenance and Authenticity (C2PA). C2PA Technical Specification, Version 1.3; C2PA: Washington, DC, USA, 2023. Available online: https://spec.c2pa.org/specifications/specifications/1.3/specs/_attachments/C2PA_Specification.pdf (accessed on 1 June 2026).
- Fernandez, P.; Couairon, G.; Jégou, H.; Douze, M.; Furon, T. The Stable Signature: Rooting Watermarks in Latent Diffusion Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023. [Google Scholar] [CrossRef]
- Wen, Y.; Kirchenbauer, J.; Geiping, J.; Goldstein, T. Tree-Ring Watermarks: Fingerprints for Diffusion Images that are Invisible and Robust. arXiv 2023, arXiv:2305.20030. [Google Scholar] [CrossRef]
- Hu, Y.; Jiang, Z.; Guo, M.; Gong, N. Stable Signature is Unstable: Removing Image Watermark from Diffusion Models. arXiv 2024, arXiv:2405.07145. [Google Scholar] [CrossRef]
- Ci, H.; Yang, P.; Song, Y.; Shou, M.Z. RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification. arXiv 2024, arXiv:2404.14055. [Google Scholar] [CrossRef]
- Chandra, B.; Dunietz, J.; Roberts, K. Reducing Risks Posed by Synthetic Content: An Overview of Technical Approaches to Digital Content Transparency (NIST AI 100-4); National Institute of Standards and Technology: Gaithersburg, MD, USA, 2024. [Google Scholar] [CrossRef]
- Guan, H.; Horan, J.; Zhang, A. Guardians of Forensic Evidence: Evaluating Analytic Systems Against AI-Generated Deepfakes; Forensics@NIST 2025; National Institute of Standards and Technology: Gaithersburg, MD, USA, 2025. Available online: https://www.nist.gov/publications/guardians-forensic-evidence-evaluating-analytic-systems-against-ai-generated-deepfakes (accessed on 1 June 2026).
- Chesney, R.; Citron, D. Deep Fakes: A Looming Challenge for Privacy, Democracy, and National Security. Calif. Law Rev. 2019, 107, 1753–1819. [Google Scholar] [CrossRef]
- European Commission. AI Act Service Desk: Article 50—Transparency Obligations for Providers and Deployers of Certain AI Systems; European Commission: Brussels, Belgium, 2025; Available online: https://artificialintelligenceact.eu/article/50/ (accessed on 1 June 2026).
- ISO/IEC 42001:2023; Information Technology—Artificial Intelligence—Management System. International Organization for Standardization (ISO): Geneva, Switzerland, 2023. Available online: https://www.iso.org/obp/ui/en/#iso:std:iso-iec:42001:ed-1:v1:en (accessed on 1 June 2026).
- ISO/IEC 23894:2023; Information Technology—Artificial Intelligence—Guidance on Risk Management. International Organization for Standardization (ISO): Geneva, Switzerland, 2023. Available online: https://www.iso.org/standard/77304.html (accessed on 1 June 2026).
- Advanced Television Systems Committee (ATSC). A/334—Audio Watermark Emission; ATSC: Washington, DC, USA, 2025; Available online: https://www.atsc.org/atsc-documents/a3342016-audio-watermark-emission/ (accessed on 1 June 2026).
- Advanced Television Systems Committee (ATSC). A/335—Video Watermark Emission; ATSC: Washington, DC, USA, 2025; Available online: https://www.atsc.org/atsc-documents/a3352016-video-watermark-emission/ (accessed on 1 June 2026).
- Advanced Television Systems Committee (ATSC). A/336—Content Recovery in Redistribution Scenarios; ATSC: Washington, DC, USA, 2024; Available online: https://www.atsc.org/atsc-documents/a3362017-content-recovery-redistribution-scenarios/ (accessed on 1 June 2026).
- Ganbaatar, U. Do Ethics in AI Still Matter? A Review of the 2021 UNESCO Recommendation on the Ethics of AI. Rev. Faith Int. Aff. 2025, 23, 26–33. [Google Scholar] [CrossRef]
- OECD. Recommendation of the Council on Artificial Intelligence (OECD/LEGAL/0449); OECD: Paris, France, 2019; Available online: https://legalinstruments.oecd.org/en/instruments/oecd-legal-0449 (accessed on 1 June 2026).
- Gaur, L. (Ed.) DeepFakes: Creation, Detection, and Impact; CRC Press: Boca Raton, FL, USA, 2022. [Google Scholar] [CrossRef]
- Schick, N. Deepfakes: The Coming Infocalypse; Twelve (Hachette Book Group): New York, NY, USA, 2020; ISBN 978-1538754313. Available online: https://www.hachettebookgroup.com/titles/nina-schick/deepfakes/9781538754313/?lens=twelve (accessed on 1 June 2026).
- Mooney, C. AI and Deception: Plagiarism, Deep Fakes, and More; ReferencePoint Press: San Diego, CA, USA, 2025; Available online: https://catalog.pcpls.org/Record/23319106 (accessed on 1 June 2026)ISBN 978-1-6782-1066-3.
- Kietzmann, J.; Lee, L.W.; McCarthy, I.P.; Kietzmann, T.C. Deepfakes: Trick or treat? Bus. Horiz. 2020, 63, 135–146. [Google Scholar] [CrossRef]
- Vaccari, C.; Chadwick, A. Deepfakes and Disinformation: Exploring the Impact of Synthetic Political Video on Deception, Uncertainty, and Trust in News. Soc. Media Soc. 2020, 6, 2056305120903408. [Google Scholar] [CrossRef]
- Gregory, J. The Trouble with Deepfakes; Cherry Lake Publishing: Ann Arbor, MI, USA, 2024; Available online: https://cherrylakepublishing.com/shop/show/53868 (accessed on 1 June 2026)ISBN 978-1668946992.
- European Union. Regulation (EU) 2024/1689 (Artificial Intelligence Act)—Transparency Obligations (Art. 50). Available online: https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-50 (accessed on 1 June 2026).
- de Araujo Meirelles Magalhães, F.; Calle, I.M. Consumer Protection in the Digital Age: A Reflection on Regulation 2022/2065 (Digital Services Act—DSA) and Agenda 2030. In International Conference a Digital Europe for Citizens, Data Governance, Digital Markets, Digital Services; Springer Nature: Cham, Switzerland, 2026; pp. 171–188. [Google Scholar] [CrossRef]
- ENISA. Remote ID Proofing Good Practices; European Union Agency for Cybersecurity: Athens/Heraklion, Greece, 2024. Available online: https://www.enisa.europa.eu/sites/default/files/2024-11/Remote%20ID%20Proofing%20Good%20Practices_en_0.pdf (accessed on 1 June 2026).
- SP 800-63A; Digital Identity Guidelines—Identity Proofing and Enrollment. National Institute of Standards and Technology: Gaithersburg, MD, USA, 2017. Available online: https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-63a.pdf (accessed on 1 June 2026).
- European Union. Regulation (EU) 2016/679 (General Data Protection Regulation)—Art. 9 (Special Categories of Personal Data). Available online: https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng (accessed on 1 June 2026).
- Europol Innovation Lab. Facing Reality? Law Enforcement and the Challenge of Deepfakes; Europol: The Hague, The Netherlands, 2022; Available online: https://www.europol.europa.eu/cms/sites/default/files/documents/Europol_Innovation_Lab_Facing_Reality_Law_Enforcement_And_The_Challenge_Of_Deepfakes.pdf (accessed on 1 June 2026).
- UNESCO. Recommendation on the Ethics of Artificial Intelligence; UNESCO: Paris, France, 2021; Available online: https://unesdoc.unesco.org/ark:/48223/pf0000380455 (accessed on 1 June 2026).
- Yan, Z.; Zhang, Y.; Yuan, X.; Lyu, S.; Wu, B. DeepfakeBench: A Comprehensive Benchmark of Deepfake Detection. arXiv 2023, arXiv:2307.01426. [Google Scholar] [CrossRef]
- Mirsky, Y.; Lee, W. The Creation and Detection of Deepfakes: A Survey. ACM Comput. Surv. 2021, 54, 1–41. [Google Scholar] [CrossRef]


| Detection Family | Core Technical Principle | Main Evidence | Strengths | Main Limitations | Indicative References |
|---|---|---|---|---|---|
| Frame-level/image-based detectors | Analyze individual frames as images and learn discriminative spatial features between real and manipulated content. | Texture inconsistencies, blending artifacts, facial boundary errors, lighting mismatch, abnormal local details. | Computationally efficient; useful for large-scale screening; can work even when only single images or extracted frames are available. | May ignore temporal inconsistencies; vulnerable when the fake is visually refined or when only high-quality frames are selected. | [59,80,85] |
| Video-level/spatio-temporal detectors | Model the evolution of visual features across consecutive frames using temporal aggregation, 3D CNNs, recurrent models, or transformers. | Frame-to-frame instability, inconsistent facial motion, unnatural eye movement, lip-movement irregularities, temporal flickering. | Better suited for video deepfakes; can detect manipulations that are not obvious in isolated frames. | Requires more computation and sufficient video length; performance may degrade after compression or frame-rate changes. | [61,62,84] |
| Frequency-domain detectors | Transform images or frames into the frequency domain and search for spectral anomalies introduced by generative upsampling, resampling, or compression. | High-frequency artifacts, spectral regularities, abnormal noise patterns, resampling traces. | Useful when spatial artifacts are visually subtle; complementary to pixel-level detection. | Sensitive to re-encoding, resizing, filtering, and platform transformations. | [63,64] |
| Physiological-signal detectors | Estimate biological signals from facial video and test whether they are consistent with natural human physiology. | Remote photoplethysmography, pulse-related skin color changes, blink behavior, micro-movement patterns. | Can capture cues that generators often fail to reproduce naturally; useful for portrait videos. | Requires sufficient face visibility and video quality; may fail under poor lighting, low resolution, or heavy compression. | [65,66] |
| Device/sensor forensic detectors | Examine whether the media contains traces consistent with real camera or sensor acquisition. | Camera fingerprints, sensor noise, compression history, metadata consistency. | Useful for forensic verification and chain-of-custody analysis. | May be weakened when metadata is removed or when the content is heavily processed. | [67,68] |
| Multimodal audio–video detectors | Compare information across modalities and detect inconsistencies between speech, facial motion, and speaker identity. | Audio–lip synchronization, voice–identity consistency, acoustic features, facial motion alignment. | Important for full audiovisual deepfakes; can detect cases where each modality alone appears plausible. | Requires synchronized audio and video; may be affected by dubbing, noisy audio, or legitimate editing. | [69,79] |
| Provenance-based verification | Verifies origin, editing history, and transformation chain through metadata, cryptographic manifests, signatures, or content credentials. | Content credentials, signed metadata, provenance chains, transformation logs. | Provides verifiable evidence beyond probabilistic detection; useful for institutional workflows. | Depends on adoption, metadata persistence, key management, and platform support. | [21,87,88,89,90] |
| Watermark-based verification | Embeds a detectable signal into generated or edited content so that synthetic media can later be identified or traced. | Embedded watermarks, watermark recovery signals, robustness under transformations. | Supports traceability and accountability at the generation stage. | Can be weakened by compression, cropping, re-encoding, fine-tuning, or removal attacks. | [91,92,93,94] |
| Dataset/Benchmark | Reported Scale | Modality | Manipulation Type | Availability/Access Route | Intended Use |
|---|---|---|---|---|---|
| DFDC | More than 100,000 videos; full version reported as 124 k videos | Video/ audiovisual | Face swapping using multiple manipulation methods | Public research dataset through official/project access routes | Large-scale training and benchmarking of deepfake detectors |
| Celeb-DF | 590 original celebrity videos and 5639 DeepFake videos | Video | High-quality celebrity face swapping | Public research dataset/project and mirror access routes | Challenging evaluation of face-swap detection and generalization |
| DeeperForensics-1.0 | 60,000 videos and 17.6 million frames | Video | End-to-end face swapping with real-world perturbations | Research dataset/project access | Robustness testing under realistic visual degradation |
| WildDeepfake | 7314 face sequences from 707 internet deepfake videos | Video/face sequences | In-the-wild deepfake videos collected from online sources | Public research dataset/GitHub access | Testing detector performance under real-world internet conditions |
| ForgeryNet/ForgeryNet Challenge | 2.9 million images and 221,247 videos | Image and video | Multiple image-level and video-level forgery methods | Research benchmark/ challenge access | Classification, spatial localization, video forgery detection, and temporal localization |
| FaceForensics++ | More than 1.8 million manipulated images from over 1000 videos | Video/frame-level images | DeepFakes, Face2Face, FaceSwap, NeuralTextures | Public benchmark dataset | Standardized evaluation of face manipulation detection under compression settings |
| Model/Approach | Model Family | Benchmark Dataset(s) | Approximate Reported Performance | Evaluation Context |
|---|---|---|---|---|
| DFDT | Vision Transformer-based detector | FaceForensics++, Celeb-DF, WildDeepfake | Around 99% AUC on controlled FaceForensics++/Celeb-DF settings; around 80–81% on WildDeepfake | Frame-level benchmark evaluation |
| SBI | Self-blended image training | FF++, Celeb-DF, DFDC, DFDCP | Generally improves cross-dataset AUC by several percentage points compared with common baselines | Generalization to unseen manipulations |
| FSBI | Frequency-enhanced self-blended images | FF++/Celeb-DF | Around mid-90% AUC on Celeb-DF under cross-dataset evaluation | Frequency-domain and cross-dataset robustness |
| CNN/Transformer comparative models | CNN and transformer architectures | FF++, Google DFD, Celeb-DF, DeeperForensics, DFDC | Often above 90% AUC in controlled intra-dataset settings, but lower under cross-dataset testing | Comparison of CNN-based and transformer-based detectors |
| Deepfake-Eval-2024 evaluated detectors | Open-source and commercial detectors | In-the-wild multimodal 2024 benchmark | Substantial performance drops compared with older academic benchmarks; reported AUC decreases of roughly 45–50% in some modalities | Real-world, multimodal, in-the-wild robustness evaluation |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gazis, A.; Karypidis, E.; Santamouri, K.; Vavouras, T.; Mastorakis, N.E.; Pappas, S. Deepfakes and Synthetic Media: Generation, Detection, and Governance. Encyclopedia 2026, 6, 165. https://doi.org/10.3390/encyclopedia6080165
Gazis A, Karypidis E, Santamouri K, Vavouras T, Mastorakis NE, Pappas S. Deepfakes and Synthetic Media: Generation, Detection, and Governance. Encyclopedia. 2026; 6(8):165. https://doi.org/10.3390/encyclopedia6080165
Chicago/Turabian StyleGazis, Alexandros, Efstathios Karypidis, Kleanthi Santamouri, Theodoros Vavouras, Nikos E. Mastorakis, and Stylianos Pappas. 2026. "Deepfakes and Synthetic Media: Generation, Detection, and Governance" Encyclopedia 6, no. 8: 165. https://doi.org/10.3390/encyclopedia6080165
APA StyleGazis, A., Karypidis, E., Santamouri, K., Vavouras, T., Mastorakis, N. E., & Pappas, S. (2026). Deepfakes and Synthetic Media: Generation, Detection, and Governance. Encyclopedia, 6(8), 165. https://doi.org/10.3390/encyclopedia6080165

