Next Article in Journal
Forecasting Gas-Dynamic Processes and Phenomena in Coal Mines Using Ensemble Model of Artificial Intelligence
Previous Article in Journal
From Prediction to Decision Support: A Critical Review and Six-Layer Framework for Responsible Artificial Intelligence in Sport Science
Previous Article in Special Issue
Investigation into the Spectral Completion Algorithm Leveraging Dense Connection Autoencoders
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement

1
School of Information Engineering, Zhongnan University of Economics and Law, Wuhan 430073, China
2
School of Geospatial Engineering and Science, Sun Yat-sen University, Zhuhai 519082, China
3
State Key Lab for Information Engineering in Surveying, Mapping, and Remote Sensing, Wuhan University, Wuhan 441000, China
4
Information Support Force Engineering University, Wuhan 430035, China
5
College of Electronic Science, National University of Defense Technology, Changsha 410073, China
*
Author to whom correspondence should be addressed.
AI 2026, 7(9), 344; https://doi.org/10.3390/ai7090344
Submission received: 5 June 2026 / Revised: 25 August 2026 / Accepted: 27 August 2026 / Published: 2 September 2026

Abstract

Existing deepfake detectors often perform well on in-domain data but generalize poorly to unseen datasets or manipulation methods. This limitation is largely attributed to their reliance on dataset-specific semantic cues rather than transferable forgery patterns. To address this limitation, we propose a generalizable deepfake detection framework that combines frequency-domain enhancement with feature disentanglement. A Phase-Amplitude Frequency Enhancement (PAFE) module enhances subtle spectral artifacts introduced during deepfake generation. We then feed the enhanced representations into an asymmetric dual-branch architecture that separates content-related information from forgery-related features. The content branch models facial semantics, while the forgery branch extracts discriminative forgery features with reduced content interference. A spatial self-attention module further refines the forgery features. We optimize the framework using image-level reconstruction loss, feature-level contrastive loss, and classification loss. Together, these objectives encourage effective feature disentanglement and improve the discriminability of the learned forgery features. Extensive experiments on several widely used deepfake benchmarks show that the proposed framework achieves competitive detection performance and improved cross-domain generalization compared with existing methods.

1. Introduction

Recent advances in generative models, particularly Generative Adversarial Networks (GANs) [1] and diffusion-based generation frameworks [2], have significantly improved the realism and diversity of synthesized visual content. Modern deepfake technologies can generate highly realistic forged faces and videos that are difficult for human observers to distinguish from authentic media. Although these technologies have facilitated the development of digital entertainment and virtual content creation, they have also introduced serious risks to identity authentication, misinformation dissemination, public opinion manipulation, and multimedia credibility. Therefore, developing robust and reliable deepfake detection methods has become an important research topic in multimedia forensics and cybersecurity [3,4,5].
Despite their good in-domain performance, existing deepfake detectors often generalize poorly to unseen datasets and manipulation methods. One reason is that they may rely on dataset-specific shortcut cues associated with content-related information, such as identity, pose and illumination conditions, rather than transferable forgery evidence [6,7]. Reducing interference from content-related information is therefore important for learning generalizable forgery features. Feature disentanglement can reduce content interference by separating content-related information from forgery-related representations [8,9]. Nevertheless, its effectiveness may remain limited when the forgery cues are too subtle to be discriminative for detection. In addition, frequency-domain information can provide important forgery cues because image synthesis operations may leave spectral artifacts that are difficult to observe in spatial-domain images [10,11]. However, subtle spectral artifacts may be obscured by dominant facial content, while different forgery mechanisms may exhibit different spectral characteristics.
Building on these observations, we propose a generalizable deepfake detection framework that combines frequency-domain enhancement with feature disentanglement, as illustrated in Figure 1b. Figure 1 compares the learning paradigms of conventional spatial-domain detectors and the proposed framework. As shown in Figure 1a, conventional spatial-domain representations may entangle subtle forgery cues with content-related information, which can limit their generalization to unseen domains. In contrast, the proposed framework first employs a learnable PAFE module to adaptively enhance subtle spectral forgery artifacts. An asymmetric dual-branch architecture then separates content representations from forgery-related representations. By reducing interference from content-related information, the learned forgery features contribute to improved generalization under cross-dataset and cross-manipulation settings. To further improve the learned forgery representations, which is critical for feature disentanglement in this work, we incorporate several optimization strategies, including attention-based refinement and multi-task loss optimization, which are detailed in the following sections.

2. Related Work

2.1. Spatial-Domain Deepfake Detection

Spatial-domain analysis represents one of the earliest directions in deepfake detection. Face manipulation techniques, such as face swapping and facial reenactment, often involve local editing, cross-source blending, resampling, and boundary refinement. These operations may disrupt the spatial consistency of authentic images. They can leave pixel-level artifacts, including abnormal boundaries, texture discontinuities, and local structural distortions. Early studies therefore focused on extracting such spatial cues directly from RGB images. Li et al. [12] proposed Face X-ray, which captures foreground–background blending boundaries commonly introduced by face manipulation and detects tampering through statistical discrepancies across these boundaries. Zhao et al. [13] developed a multi-attentional network that uses multiple spatial attention heads to localize forged regions and a texture enhancement block to amplify subtle artifacts in shallow features. To capture spatial inconsistencies without pixel-level annotations, Zhuang et al. [14] proposed UIA-ViT, which uses a Vision Transformer to learn inconsistency-aware representations. Cao et al. [15] developed RECCE, an encoder–decoder framework that models common representations of genuine faces and uses reconstruction differences to guide the detector toward forged regions. As synthesis techniques evolve, these spatial artifacts become less conspicuous, calling for more effective forgery cues.

2.2. Frequency-Domain Deepfake Detection

Frequency-domain deepfake detection aims to identify synthesis traces that are difficult to observe directly in spatial-domain images. Durall et al. [10] found that up-convolutional operations can introduce discrepancies between the spectral distributions of generated and natural images. Frank et al. [11] further showed that these frequency artifacts can help distinguish synthetic images from authentic ones. Representative methods have explored frequency information from different perspectives. Qian et al. [16] proposed F3Net, which combines DCT-based frequency decomposition with local frequency statistics to capture manipulation-related patterns. Masi et al. [17] designed a two-branch network in which a Laplacian-of-Gaussian (LoG) bottleneck suppresses high-level facial content and emphasizes multi-band frequency artifacts. Liu et al. [18] proposed SPSL to combine spatial features with phase-spectrum information, while Luo et al. [19] used SRM high-pass filters and cross-modality attention to capture multi-scale high-frequency noise. Recent studies have also combined spatial and frequency information. Gong et al. [20] fused multi-scale spatial and frequency features, whereas Qi et al. [21] progressively enhanced spatial and frequency-aware cues. Despite this progress, many frequency-based detectors still construct spectral representations using predefined transforms or handcrafted filter priors, such as DCT, phase transforms, LoG, and SRM filters [16,17,18,19]. Some methods introduce learnable parameters over predefined frequency bands. However, their spectral partitioning or initial feature extraction remains constrained by the underlying operators. Therefore, the adaptive enhancement in weak forgery cues distributed across different frequency bands before feature extraction remains underexplored.

2.3. Generalizable Deepfake Detection

Generalizable deepfake detection aims to maintain reliable performance on unseen datasets and manipulation methods. Unlike in-domain detection, it requires models to learn transferable forgery patterns rather than dataset-specific correlations. Existing methods mainly follow two directions: increasing the diversity of training forgeries and learning representations that transfer across domains. LGrad [22] uses a pretrained convolutional neural network to transform input images into gradient maps, which serve as generalized representations of manipulation artifacts. Other methods improve generalization by synthesizing diverse training samples. SBI [23] generates self-blended images from pristine faces to reproduce common blending artifacts, while SLADD [24] dynamically generates challenging samples using different facial regions and blending configurations. Recent face-specific methods have also explored anomaly modeling and texture discrepancies. Zhang et al. [25] formulated generalized face forgery detection as a supervised anomaly detection problem and used artifact-map scores to identify forged videos. Liu et al. [26] proposed AIM-Bone to generate and localize texture discrepancies for detecting unseen forgeries. The development of diffusion models has further expanded the range of unseen image generators. Corvi et al. [27] analyzed the forensic properties of images generated by generative adversarial networks and diffusion models. Ojha et al. [28] used features extracted from a pretrained vision-language model to detect images produced by unseen generative models. Despite this progress, existing methods may still depend on training-specific artifacts, synthesis strategies, or general-purpose visual representations. Learning task-specific representations that capture transferable forgery cues while reducing content-related interference therefore remains important for generalizable deepfake detection.

2.4. Feature Disentanglement

Feature disentanglement aims to separate forgery-related evidence from content information, such as facial identity, appearance, and background, which may introduce dataset bias. Liang et al. [8] proposed a content information removal framework. They used a content consistency constraint and a global representation contrastive constraint to reduce content interference and guide the detector toward manipulation traces. UCF [9] further decomposed image representations into forgery-irrelevant features, method-specific forgery features, and common forgery features. Multi-task supervision and contrastive regularization were used to separate these components, while only the common forgery features were retained for detection. Recent studies have extended feature disentanglement to different learning settings. Yan et al. [29] constructed self-supervised training samples through image self-mixing. Their framework separated forgery-irrelevant content features from forgery-related features using conditional reconstruction and orthogonal constraints. FTA-DFI [30] combined feature separation with frequency-domain trace amplification and adversarial learning. It aimed to suppress identity information and method-specific patterns while identifying differences between authentic and forged faces. These studies indicate that feature disentanglement can reduce content-related interference. However, preserving sufficiently discriminative forgery evidence during feature encoding remains challenging.

3. Method

3.1. Overview of the Proposed Framework

We propose a generalizable deepfake detection framework that combines frequency-domain enhancement with feature disentanglement, as shown in Figure 2.
The framework consists of three main components: the Phase-Amplitude Frequency Enhancement (PAFE) Module, the Dual-branch Feature Disentanglement Module, and the Multi-task Loss Optimization Module. The PAFE Module applies learnable amplitude modulation while preserving phase information. It produces frequency-enhanced inputs for subsequent feature learning. The Dual-branch Feature Disentanglement Module employs a content encoder and a forgery encoder to extract content features and forgery features, respectively. Spatial self-attention is further applied to the forgery features to highlight manipulation-related regions. The Multi-task Loss Optimization Module jointly optimizes image reconstruction, feature-level contrastive learning, and forgery classification. For image reconstruction, content and forgery features are combined through AdaIN mixing and fed into a shared decoder. Self-reconstruction and cross-reconstruction promote the separation of content and forgery information. Meanwhile, the contrastive and classification objectives improve the discriminability of the forgery features. During inference, the input image is processed by the PAFE Module, the forgery encoder, and the spatial self-attention module. The resulting forgery features are then fed into the classifier for real/fake prediction. The following sections describe each component in detail.

3.2. Phase-Amplitude Frequency Enhancement Module

A Phase-Amplitude Frequency Enhancement (PAFE) module is introduced to enhance subtle forgery artifacts in the frequency domain. For each input image x i R C × H × W , where i { 0 , 1 } , a two-dimensional Fourier transform is applied independently to each channel to obtain its amplitude and phase spectra as follows:
F i = F ( x i ) , A i = | F i | , P i = arg ( F i ) ,
where F ( · ) denotes the channel-wise two-dimensional Fourier transform, F i is the resulting complex-valued frequency spectrum, and A i and P i denote its amplitude and phase spectra, respectively. Since the phase spectrum generally preserves more structural information of the image, PAFE retains the original phase spectrum P i and adaptively modulates only the amplitude spectrum A i . A i en denotes the enhanced amplitude spectrum of the i-th input image, obtained by modulating A i with the learnable matrix M, as follows:
A i en = A i M ,
where ⊙ denotes element-wise multiplication. The modulation matrix M R H × W has the same spatial dimensions as each channel of the amplitude spectrum. It is initialized with ones, shared across channels, and jointly optimized with the remaining network parameters. During training, M is continuously updated through backpropagation with respect to the overall loss, whose components are described in a subsequent section. This optimization enables the network to assign larger weights to frequency components containing more pronounced manipulation traces while suppressing uninformative low-frequency background variations. The enhanced spatial image x i en is reconstructed from the modulated amplitude and the original phase.
x i en = F 1 A i en exp ( j P i ) ,
where F 1 ( · ) denotes the channel-wise inverse Fourier transform and j is the imaginary unit. Finally, the enhanced spatial image is concatenated with the original spatial-domain image along the channel dimension:
x i PAFE = Concat x i , x i en ,
where Concat ( · ) denotes channel-wise concatenation. x i PAFE is the final output of the PAFE module. It preserves the information in the original image while emphasizing informative frequency-domain cues through learnable spectral enhancement. Therefore, it will provide a more informative input for subsequent feature disentanglement.

3.3. Dual-Branch Feature Disentanglement

3.3.1. Dual-Branch Feature Extraction

As illustrated in Figure 2, an asymmetric dual-branch framework is introduced to encourage the disentanglement of content features from forgery features. The framework consists of a content encoder E c and a forgery encoder E f . Given a real/fake image pair ( x 0 , x 1 ) , where x 0 and x 1 denote the real/fake samples, respectively. The content encoder E c directly extracts the content feature of x i denoted as F c i from the original spatial-domain image, as follows:
F c i = E c ( x i ) , i { 0 , 1 } ,
The content encoder mainly focuses on semantic information, such as facial structure, identity, and expression. In contrast, the forgery encoder processes the frequency-enhanced x i PAFE generated by the PAFE module:
F f i = E f x i PAFE , i { 0 , 1 } ,
where F f i denotes the forgery features of x i . The forgery encoder is designed to capture manipulation-related patterns from the frequency-enhanced input. In our implementation, both E c and E f use Xception backbones initialized with ImageNet-pretrained weights. The content and forgery features have the same dimensions F c i , F f i R D × H × W , where D = 728 denotes the number of feature channels, and H and W denote the height and width of the feature maps, respectively. By assigning different inputs and learning objectives to the two encoders, the asymmetric architecture guides E c toward semantic content and E f toward forgery-related evidence. During training, the two encoding branches are jointly optimized using the image-level reconstruction loss, feature-level contrastive loss, and classification loss to promote the separation of content and forgery features.

3.3.2. Spatial Self-Attention Enhancement

Deepfake artifacts often appear in localized regions, such as facial boundaries and areas with inconsistent textures. We therefore insert a spatial self-attention module into an intermediate block of the forgery encoder E f to emphasize manipulation-sensitive regions and reduce interference from irrelevant semantic information. Given an intermediate feature map F R D × H × W , three independent 1 × 1 convolutional projections generate the query Q, key K, and value V, where Q and K have channel dimension D / 8 to reduce computational cost, and V maintains dimension D. The spatial dimensions are flattened into N = H × W positions. The attention weight between positions i and j is computed as follows:
S i j = exp Q i K j D / 8 k = 1 N exp Q i K k D / 8 ,
where Q i , K j , and V j denote the query, key, and value vectors at the corresponding spatial positions, and S i j is the normalized attention weight from position i to position j. The attention-enhanced feature at position i is obtained by aggregating the value vectors over all spatial positions.
F a ( i ) = j = 1 N S i j V j .
The resulting vectors are rearranged into F a R D × H × W . A residual connection is then applied to preserve the original feature information.
F s = F + α F a ,
where α is a learnable scaling parameter that controls the contribution of the attention-enhanced feature. The output F s is passed to the subsequent layers of E f to obtain the final forgery feature F f . By modeling dependencies across spatial positions, the module helps E f capture spatially distributed manipulation cues.

3.4. Multi-Task Loss Optimization

The proposed framework is optimized in an end-to-end manner. To facilitate effective feature disentanglement and learn discriminative forgery features, the overall optimization objective is formulated with three complementary loss terms, including the image-level reconstruction loss, the feature-level contrastive loss, and the classification loss.

3.4.1. Image-Level Reconstruction Loss

To encourage the content and forgery features to retain complementary information, we introduce a conditional decoder D for image-level reconstruction. An AdaIN-inspired forgery-guided feature modulation operation is adopted to combine the content feature F c and the forgery feature F f , enabling the decoder to reconstruct images with controlled authenticity attributes. For each sample i { 0 , 1 } , the content feature is normalized as F ^ c i to eliminate statistical biases and facilitate subsequent forgery-driven modulation:
F ^ c i = F c i μ ( F c i ) σ ( F c i ) ,
where μ ( · ) and σ ( · ) denote the mean and standard deviation computed over the spatial positions of each channel. An MLP maps the forgery feature F f j from sample j { 0 , 1 } to the modulation parameters γ ( F f j ) and β ( F f j ) . The content feature from sample i and the forgery feature from sample j are combined as follows.
F mix i , j = γ ( F f j ) F ^ c i + β ( F f j ) ,
where i , j { 0 , 1 } indicate the source samples of the content and forgery features, respectively, and ⊙ denotes element-wise multiplication. Here, 0 and 1 denote the real/fake samples, respectively.
Unlike simple feature concatenation or addition, the proposed modulation operation preserves the distinct roles of content and forgery information. This design supports controlled feature exchange during reconstruction. We therefore feed F mix i , j into the decoder D and define two complementary reconstruction paths: self-reconstruction and cross-reconstruction. As illustrated in Figure 2, self-reconstruction combines the content and forgery features from the same sample, as follows:
x ^ i self = D F mix i , i .
Cross-reconstruction retains the content feature of x i while replacing its forgery feature with that of x 1 i .
x ^ i cross = D F mix i , 1 i .
Thus, x ^ i cross is expected to preserve the facial content of x i while acquiring the authenticity attribute of x 1 i . Its target authenticity label, 1 i , is supervised by the classification loss introduced in a subsequent subsection. Since both reconstruction paths use F c i as their content source, x i serves as the pixel-level reference for content preservation. The image-level reconstruction loss is defined as follows:
L rec = i = 0 1 x ^ i self x i 1 + x ^ i cross x i 1 .
The self-reconstruction term minimizes the reconstruction error between the reconstructed and original images. This encourages the disentangled features to preserve the information needed to reconstruct the input image. In cross-reconstruction, the forgery feature is extracted from the other input image. It may contain identity, appearance, or background information from that image. This information may change the content of the reconstructed image and increase the pixel-level reconstruction error. Minimizing this error helps reduce content leakage into the forgery feature. It therefore promotes the disentanglement of content and forgery features.

3.4.2. Feature-Level Contrastive Loss

Although the image reconstruction constraint promotes the disentanglement of content and forgery information to some extent, relying solely on reconstruction may still allow content information to leak into the forgery feature. To further separate the two representations, we introduce a feature-level contrastive loss. The content and forgery features extracted from the same input image are treated as a negative feature pair, and their distance in the feature space is enlarged to reduce semantic coupling. Specifically, the content encoder produces the content feature map F c , while the forgery encoder produces the final forgery feature map F f after spatial self-attention enhancement and subsequent encoding layers. Global average pooling is first applied to compress the two feature maps into the content feature vector v c and the forgery feature vector v f , respectively.
v c = GAP ( F c ) , v f = GAP ( F f ) .
The Euclidean distance between the two feature vectors d c f is then computed as follows:
d c f = v c v f 2 .
Accordingly, the feature-level contrastive loss L con is formulated as follows:
L con = max 0 , m d c f ,
where m denotes a predefined margin. When the distance between the content and forgery features is smaller than m, the loss penalizes the network and pushes the two representations apart. When their distance is greater than or equal to m, the loss becomes zero. This constraint does not require the construction of additional cross-sample pairs and directly reduces the coupling between content information and forgery-related cues within the same image, thereby promoting more effective feature disentanglement.

3.4.3. Classification Loss

The classification loss directly supervises the authenticity prediction of both the original and cross-reconstructed samples. Let y ^ i and y ^ i cross denote the predicted probabilities of the original image x i and the cross-reconstructed image x ^ i cross being fake, respectively. The binary cross-entropy loss is defined as follows.
l ( y , y ^ ) = y log ( y ^ ) ( 1 y ) log ( 1 y ^ ) .
Since x 0 and x 1 denote the real/fake samples, respectively, their labels satisfy y i = i . The cross-reconstructed image x ^ i cross , which combines F c i with F f 1 i , is assigned the authenticity label 1 i . Accordingly, the classification loss L cls is as follows:
L cls = i = 0 1 l i , y ^ i + l 1 i , y ^ i cross .

3.4.4. Overall Training Objective

The overall training objective L total combines the classification loss, reconstruction loss, and feature-level contrastive loss:
L total = L cls + λ 1 L rec + λ 2 L con ,
where λ 1 and λ 2 balance the contributions of the reconstruction loss and feature-level contrastive loss, respectively. The classification loss provides direct supervision for real/fake classification. The reconstruction loss encourages the two branches to preserve complementary information for image reconstruction. The feature-level contrastive loss reduces the coupling between content and forgery features in the latent space. Joint optimization of these losses promotes effective feature disentanglement and preserves discriminative forgery cues for robust deepfake detection.

4. Experiments

4.1. Experimental Settings

4.1.1. Datasets

Extensive experiments were conducted on several widely used public deepfake benchmark datasets to evaluate the effectiveness and cross-domain generalization capability of the proposed framework. The evaluated datasets include FaceForensics++ (FF++) [31], Celeb-DF and Celeb-DF-v2 [32], the DeepFake Detection Dataset (DFD) [33], the DeepFake Detection Challenge Preview dataset (DFDCP) and the DeepFake Detection Challenge dataset (DFDC) [34]. FaceForensics++ is one of the most commonly used benchmark datasets for deepfake detection. It contains multiple forgery manipulation methods, including DeepFakes (DF), Face2Face (F2F), FaceSwap (FS), and NeuralTextures (NT). Celeb-DF is a challenging deepfake dataset that contains high-quality forged videos with significantly reduced visual artifacts. Celeb-DF-v2 is an improved version of Celeb-DF, containing a larger number of high-fidelity forged videos generated using advanced synthesis techniques. DFDCP and DFDC contain large-scale real-world forged videos generated by diverse manipulation pipelines. These datasets provide more realistic evaluation scenarios for generalizable deepfake detection. The DeepFake Detection Dataset (DFD), released by Google Jigsaw, is an independent benchmark containing real and manipulated videos generated from consented actor footage. It covers diverse acquisition conditions and is widely used to evaluate the robustness of deepfake detection models against unseen forgery patterns.

4.1.2. Implementation Details and Evaluation Metrics

For data preprocessing, Dlib was used for face detection, extraction, and alignment. All extracted face images were resized to a uniform resolution of 256 × 256 pixels. The medium-compression version (c23) of FF++ was used for training to simulate image quality in real-world online environments. For each video, 32 frames were sampled at equal temporal intervals. Several commonly used data augmentation operations were applied during training. Horizontal flipping was performed with a probability of 0.5. Random rotation was applied with a probability of 0.5, with rotation angles ranging from 10 to 10 . Image blurring was applied with a probability of 0.5, using kernel sizes ranging from 3 to 7. Brightness and contrast were randomly adjusted with a probability of 0.5, with both adjustment ranges set to [ 0.1 , 0.1 ] . JPEG compression artifacts were introduced by randomly sampling the quality factor from [ 40 , 100 ] . The proposed framework was implemented using the PyTorch (1.11) deep learning library. During training, the Adam optimizer was adopted for end-to-end optimization, with the exponential decay rates for the first- and second-moment estimates set to β 1 = 0.9 and β 2 = 0.999 , respectively. The initial learning rate was set to 1 × 10 4 and gradually decayed using a cosine annealing strategy. The batch size was set to 32, and the network was trained for 50 epochs. The balancing coefficients in Equation (20) were empirically set to λ 1 = 1.0 and λ 2 = 0.1 to balance the contributions of the reconstruction and feature-level contrastive losses during joint optimization. The margin m in Equation (17) was set to 1.0 . All experiments were conducted on NVIDIA RTX 4090 GPUs. The inference model contains approximately 45.2M parameters and requires 9.8 GFLOPs to process a single image.
Detection performance was evaluated using the Area Under the Receiver Operating Characteristic Curve (AUC). The AUC is a widely adopted evaluation metric for deepfake detection [9], which measures the discriminative capability of a detector across different decision thresholds. A higher AUC value indicates better forgery detection performance. Under cross-dataset evaluation settings, it also reflects stronger cross-domain generalization capability.

4.2. Cross-Dataset Evaluation

To evaluate model generalization across different data distributions, we conduct cross-dataset experiments. The model is trained on FF++ and evaluated on datasets from multiple sources, including Celeb-DF, Celeb-DF-v2, DFDCP, DFDC, and DFD. These datasets differ in manipulation methods, video compression levels, and acquisition conditions. This setting introduces diverse distribution shifts and provides a more realistic assessment of model generalization in practical scenarios. The results are reported in Table 1.
As shown in Table 1, most methods exhibit considerable performance variation across the unseen datasets. For example, Face X-ray achieves AUC values of 0.709 on Celeb-DF and 0.633 on DFDC. This variation implies that its performance is sensitive to shifts in the target distribution. In comparison, methods based on frequency-domain features or high-level semantic representations show stronger cross-domain performance. SPSL and SRM maintain relatively consistent performance across several test datasets. In particular, SPSL achieves the highest AUC of 0.815 on Celeb-DF.
Compared with these methods, our method achieves competitive results across the evaluated datasets. It obtains the highest AUC on Celeb-DF-v2, DFDCP, DFDC, and DFD, with scores of 0.773, 0.798, 0.752, and 0.813, respectively. The advantages are more pronounced on DFDCP and DFDC. On Celeb-DF, our method ranks second with an AUC of 0.807. These results imply that our method captures relatively stable forgery cues across different data distributions. It also adapts to variations in manipulation methods and complex distribution shifts, demonstrating relatively good cross-dataset generalization and detection stability.

4.3. Cross-Manipulation Evaluation

To further evaluate the generalization capability of the proposed framework to unseen manipulation types, leave-one-manipulation-out experiments are conducted on FaceForensics++ (FF++). In each setting, one of the four manipulation types is excluded from training. The remaining three types are used to train the model. The trained model is then evaluated on the external benchmark datasets Celeb-DF and DFDC. This setting approximates the detection of manipulation types that are unavailable during training. The results are reported in Table 2.
As shown in Table 2, most baseline methods exhibit considerable performance variation across the leave-one-manipulation-out settings. Xception and CORE obtain AUC values ranging from 0.63 to 0.74. This variation suggests that conventional single-encoder classifiers may overfit to the manipulation types observed during training. Face X-ray, which detects blending boundaries, also shows considerable variation across the evaluation settings. This result shows that reliance on a fixed forensic cue may limit generalization to unseen manipulation types.
In comparison, the proposed framework achieves the best results in most leave-one-manipulation-out settings and the second-best results in a small number of settings. It therefore shows competitive and relatively consistent overall performance. These results suggest that the proposed framework is less dependent on manipulation-specific patterns. They are also consistent with learning forgery-related representations shared across different manipulation types. The proposed framework therefore shows relatively good generalization to the unseen manipulation types considered in this experiment.

4.4. Ablation Study

To investigate the contribution of each component, detailed ablation studies have been conducted. The experiments are performed on FaceForensics++ (FF++) and Celeb-DF. FF++ is adopted to evaluate the in-domain detection performance, while Celeb-DF is used as an unseen benchmark to assess cross-domain generalization capability. This evaluation protocol provides a comprehensive analysis of both detection effectiveness and generalization ability. The experimental results are reported in Table 3.
The baseline model Xception consists only of the backbone feature extraction network, without incorporating the proposed feature enhancement and representation learning strategies. As shown in Table 3, the baseline model achieves competitive performance on the in-domain FF++ dataset. However, its performance on Celeb-DF is relatively limited. This verifies that directly learning representations from RGB images may not be sufficient to handle the distribution variations in unseen datasets. After introducing the feature disentanglement module, the performance on FF++ is improved. This improvement shows that separating content features from forgery features helps the network focus on more discriminative forgery-related information. However, the cross-domain performance shows a slight decline at this stage. This observation suggests that feature disentanglement alone may not fully address the distribution differences between datasets. With the further integration of the spatial self-attention module, the model achieves improvements in both FF++ and Celeb-DF. The results show that spatial self-attention enhancement provides additional guidance for forgery feature learning and helps capture more informative forgery-related regions. After incorporating the proposed PAFE module, the model obtains further performance improvements, particularly on the cross-domain Celeb-DF benchmark. This result shows that frequency-domain enhancement can provide complementary spectral forgery cues and improve the robustness of the learned features. Finally, the complete model combining all proposed components achieves the best performance under both in-domain and cross-domain evaluation settings. These results indicate that frequency-domain enhancement, feature disentanglement, and spatial self-attention enhancement provide complementary advantages. Their joint optimization enables the model to learn more robust and transferable forgery features.
To further evaluate the contribution of each constraint in the multi-task loss optimization, we conduct a loss-function ablation study, as shown in Table 4.
The results show that the three loss constraints contribute differently to feature disentanglement. In Experiment 1, the classification-only setting using L c l s appears to be more susceptible to domain overfitting. Since L c l s mainly optimizes the decision boundary on the training data, the absence of explicit disentanglement supervision may cause the forgery branch to retain dataset-specific content cues, such as facial backgrounds and illumination conditions. In Experiment 2, adding the image-level reconstruction loss L r e c leads to a modest improvement in FF++ and a clearer gain on Celeb-DF. This implies that L r e c helps reduce interference from identity-related content in unseen domains. In Experiment 3, the feature-level contrastive loss L c o n provides a larger improvement in FF++, although its performance on Celeb-DF remains slightly below that of the reconstruction-based setting. This result implies that L c o n helps reduce the entanglement between forgery and content features and improves feature discriminability. In Experiment 4, jointly optimizing L c l s , L r e c , and L c o n yields the best results, with AUC scores of 0.972 on FF++ and 0.807 on Celeb-DF. Overall, these objectives appear to provide complementary supervision for in-domain detection and cross-domain generalization.

4.5. Visualization Analysis

To provide a more intuitive assessment of feature disentanglement, the disentangled content features F c and forgery features F f are visualized using t-SNE, as shown in Figure 3. In Figure 3a, the content features of real/fake samples overlap considerably. This indicates that the two classes have similar distributions of semantic information, such as facial content and backgrounds. These content features may therefore interfere with real/fake classification. In contrast, Figure 3b shows a clear separation between the forgery features of real/fake samples. This result indicates that the disentangled forgery features are highly discriminative between the two classes. Moreover, the forgery features of fake samples form multiple clusters. The number of clusters matches the number of manipulation types in the dataset, suggesting a potential direction for fine-grained manipulation type detection.
To further analyze the effect of the proposed frequency-domain enhancement strategy, we visualize the Fourier spectra of authentic and manipulated faces, as shown in Figure 4.
For authentic faces, the spectral energy is mainly concentrated in the low-frequency region and gradually decreases as the frequency increases. This distribution is generally consistent with the statistics of natural images. Complex backgrounds may introduce additional spectral variations, as illustrated by the third authentic image in the first row and its corresponding spectrum in the second row. In contrast, manipulated faces often exhibit frequency patterns that deviate from natural image statistics. Several examples in the fourth row show anomalous responses in the mid-to-high-frequency regions. These responses include abrupt changes in spectral energy and spectral discontinuities (e.g., the first, fourth, and fifth spectra), as well as cross-shaped or regularly distributed bright patterns (e.g., the first, second, and third spectra). These anomalous spectral patterns provide potentially useful forgery-related cues for forgery feature learning.
To further analyze the spatial cues used for forgery detection, the model responses are visualized using Grad-CAM, as shown in Figure 5. Strong responses are mainly observed around the eyes, mouth, facial contours, and hairline, whereas most background regions produce relatively weak responses. The periocular region contains rich edges and fine-grained textures. Face alignment, warping, and blending may introduce local texture anomalies or geometric distortions in this region. This may partly explain the strong responses around the eyes. The responses around the mouth may be associated with local structural changes introduced by facial expression manipulation. Those along the facial contours and hairline may reflect discontinuities caused by boundary blending. Overall, this response distribution suggests that the model relies more on local manipulation-related anomalies than on irrelevant scene content such as backgrounds. These observations provide some qualitative support for the effectiveness of feature disentanglement and the interpretability of the learned forgery features.

5. Conclusions

In this work, we propose a generalizable deepfake detection framework that combines PAFE with dual-branch feature disentanglement architecture. We introduce spatial self-attention into the forgery encoder to refine forgery-related representations. Classification loss, image-level reconstruction loss, and feature-level contrastive loss jointly guide the training of the disentanglement encoders. Experiments on multiple deepfake benchmarks validate the effectiveness of the proposed components and demonstrate competitive performance under cross-dataset and cross-manipulation settings.
The current framework primarily targets manipulation-based deepfakes, while its performance on fully-synthetic content (e.g., diffusion-generated faces) remains to be further evaluated. Future work will extend learnable frequency-domain modeling to capture artifacts produced by diverse forgery mechanisms. We will also explore self-supervised representation learning and domain generalization to improve robustness to unseen manipulation types and generative models.

Author Contributions

Conceptualization, Q.W.; methodology, Q.W., J.F. and Y.Z.; validation, J.F.; formal analysis, M.L., Z.Z. and W.S.; writing—original draft preparation, Q.W., J.F. and Y.Z.; writing—review and editing, W.S., Q.W., J.F., Y.Z., M.L., L.W. and Z.Z.; Funding acquisition, W.S. All authors have read and agreed to the published version of the manuscript.

Funding

This work was supported in part by the National Natural Science Foundation of China under Grant 62406334, and the Postgraduate Education and Teaching Reform Project of Zhongnan University of Economics and Law under Grant RCPY202518.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data presented in this study are available on request from the corresponding author and the first author due to privacy or ethics restrictions.

Use of Artificial Intelligence

We confirm that generative AI tools were used in the preparation of this manuscript. Specifically, these were used solely for the purpose of language polishing. No AI tools were used to analyze or interpret scientific content, results, or conclusions. All scientific content, data analysis, and intellectual contributions are entirely the authors’ work.

Acknowledgments

The authors would like to express their deepest appreciation to the editors and anonymous reviewers for their constructive comments and meticulous examination of this manuscript. We are also grateful to the providers of the benchmark datasets for providing the necessary datasets. Finally, we extend our thanks to all those who contributed to this work in various capacities, including technical support and intellectual discussions.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; Bengio, Y. Generative adversarial nets. NeurIPS 2014, 27, 2672–2680. [Google Scholar]
  2. Ho, J.; Jain, A.; Abbeel, P. Denoising diffusion probabilistic models. NeurIPS 2020, 33, 6840–6851. [Google Scholar]
  3. Verdoliva, L. Media forensics and deepfakes: An overview. IEEE J. Sel. Top. Signal Process. 2020, 14, 910–932. [Google Scholar] [CrossRef] [Scilit]
  4. Mirsky, Y.; Lee, W. The creation and detection of deepfakes: A survey. ACM Comput. Surv. 2021, 54, 7. [Google Scholar] [CrossRef] [Scilit]
  5. Bhat, N.A.; Giri, K.J. DeepFake detection in images and videos: A survey on models, datasets, and evaluation metrics. Multimed. Tools Appl. 2026, 85, 233. [Google Scholar] [CrossRef] [Scilit]
  6. Geirhos, R.; Jacobsen, J.H.; Michaelis, C.; Zemel, R.; Brendel, W.; Bethge, M.; Wichmann, F.A. Shortcut learning in deep neural networks. Nat. Mach. Intell. 2020, 2, 665–673. [Google Scholar] [CrossRef] [Scilit]
  7. Shi, L.; Zhang, J.; Ji, Z.; Bai, J.; Shan, S. Real face foundation representation learning for generalized deepfake detection. Pattern Recognit. 2025, 161, 111299. [Google Scholar] [CrossRef] [Scilit]
  8. Liang, J.; Shi, H.; Deng, W. Exploring disentangled content information for face forgery detection. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 128–145. [Google Scholar]
  9. Yan, Z.; Zhang, Y.; Fan, Y.; Wu, B. UCF: Uncovering common features for generalizable deepfake detection. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, 1–6 October 2023; pp. 22412–22423. [Google Scholar]
  10. Durall, R.; Keuper, M.; Keuper, J. Watch your up-convolution: CNN based generative deep neural networks are failing to reproduce spectral distributions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 13–19 June 2020; pp. 7890–7899. [Google Scholar]
  11. Frank, J.; Eisenhofer, T.; Schönherr, L.; Fischer, A.; Kolossa, D.; Holz, T. Leveraging frequency analysis for deep fake image recognition. In Proceedings of the 37th International Conference on Machine Learning, Virtual, 13–18 July 2020; pp. 3247–3258. [Google Scholar]
  12. Li, L.; Bao, J.; Zhang, T.; Yang, H.; Chen, D.; Wen, F.; Guo, B. Face X-ray for more general face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 13–19 June 2020; pp. 5001–5010. [Google Scholar]
  13. Zhao, H.; Zhou, W.; Chen, D.; Wei, T.; Zhang, W.; Yu, N. Multi-attentional deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Virtual, 19–25 June 2021; pp. 2185–2194. [Google Scholar]
  14. Zhuang, W.; Chu, Q.; Tan, Z.; Liu, Q.; Yuan, H.; Miao, C.; Luo, Z.; Yu, N. UIA-ViT: Unsupervised inconsistency-aware method based on Vision Transformer for face forgery detection. In Proceedings of the European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; pp. 391–407. [Google Scholar]
  15. Cao, J.; Ma, C.; Yao, T.; Chen, S.; Ding, S.; Yang, X. End-to-end reconstruction-classification learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 4113–4122. [Google Scholar]
  16. Qian, Y.; Yin, G.; Sheng, L.; Chen, Z.; Shao, J. Thinking in frequency: Face forgery detection by mining frequency-aware clues. In Proceedings of the European Conference on Computer Vision, Virtual, 23–28 August 2020; pp. 86–103. [Google Scholar]
  17. Masi, I.; Killekar, A.; Mascarenhas, R.M.; Gurudatt, S.P.; AbdAlmageed, W. Two-branch recurrent network for isolating deepfakes in videos. In Proceedings of the European Conference on Computer Vision Workshops, Glasgow, UK, 23–28 August 2020; pp. 667–684. [Google Scholar]
  18. Liu, H.; Li, X.; Zhou, W.; Chen, Y.; He, Y.; Xue, H.; Zhang, W.; Yu, N. Spatial-phase shallow learning: Rethinking face forgery detection in frequency domain. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 19–25 June 2021; pp. 772–781. [Google Scholar]
  19. Luo, Y.; Zhang, Y.; Yan, J.; Liu, W. Generalizing face forgery detection with high-frequency features. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nashville, TN, USA, 20–25 June 2021; pp. 16317–16326. [Google Scholar]
  20. Gong, R.; Chen, J.; Zhang, D.; Sangaiah, A.K.; Alenazi, M.J.F. Face Forgery Detection via Multi-Scale and Multi-Domain Features Fusion. IET Image Process. 2025, 19, e70131. [Google Scholar] [CrossRef] [Scilit]
  21. Qi, Y.; Wen, S.; Zhang, H.; Liang, A.; Chen, H.; Cao, P. Face forgery detection by progressively enhancing spatial and frequency-aware features. Multimed. Syst. 2024, 30, 156. [Google Scholar] [CrossRef] [Scilit]
  22. Tan, C.; Zhao, Y.; Wei, S.; Gu, G.; Wei, Y. Learning on gradients: Generalized artifacts representation for GAN-generated images detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 12105–12114. [Google Scholar]
  23. Shiohara, K.; Yamasaki, T. Detecting deepfakes with self-blended images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18720–18729. [Google Scholar]
  24. Chen, L.; Zhang, Y.; Song, Y.; Liu, L.; Wang, J. Self-supervised learning of adversarial example towards good generalizations for deepfake detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 18–24 June 2022; pp. 18710–18719. [Google Scholar]
  25. Zhang, F.; Yang, T.; Cao, L.; Du, K.; Guo, Y.; Song, P.; Shao, C. Deep supervised anomaly detection for generalized face forgery detection. Pattern Recognit. 2026, 169, 111976. [Google Scholar] [CrossRef] [Scilit]
  26. Liu, B.; Zhang, X.; Ling, H.; Li, Z.; Wang, R.; Zhang, H.; Li, P. AIM-Bone: Texture Discrepancy Generation and Localization for Generalized Deepfake Detection. IEEE Trans. Biom. Behav. Identity Sci. 2025, 7, 422–431. [Google Scholar] [CrossRef] [Scilit]
  27. Corvi, R.; Cozzolino, D.; Poggi, G.; Nagano, K.; Verdoliva, L. Intriguing properties of synthetic images: From GANs to diffusion models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Vancouver, BC, Canada, 17–24 June 2023; pp. 973–982. [Google Scholar]
  28. Ojha, U.; Li, Y.; Lee, Y.J. Towards universal fake image detectors that generalize across generative models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Vancouver, BC, Canada, 17–24 June 2023; pp. 24480–24489. [Google Scholar]
  29. Yan, B.; Liu, P.; Yang, Y.; Guo, Y. Self-Supervised Feature Disentanglement for Deepfake Detection. Mathematics 2025, 13, 2024. [Google Scholar] [CrossRef] [Scilit]
  30. Zou, Z.; Peng, D.; Zhao, Y.; Tian, Z.; Cai, J. FTA-DFI: A framework for generalizable deepfake detection based on distinctive features from various manipulations compared to genuine images. Appl. Intell. 2025, 55, 769. [Google Scholar] [CrossRef] [Scilit]
  31. Rossler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; Nießner, M. FaceForensics++: Learning to detect manipulated facial images. In Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Republic of Korea, 27–28 October 2019; pp. 1–11. [Google Scholar]
  32. Li, Y.; Yang, X.; Sun, P.; Qi, H.; Lyu, S. Celeb-DF: A large-scale challenging dataset for deepfake forensics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 3207–3216. [Google Scholar]
  33. Google Research. Contributing Data to Deepfake Detection Research. 2019. Available online: https://ai.googleblog.com/2019/09/contributing-data-to-deepfake-detection.html (accessed on 4 January 2026).
  34. Dolhansky, B.; Bitton, J.; Pflaum, B.; Lu, J.; Howes, R.; Wang, M.; Ferrer, C.C. The DeepFake Detection Challenge Dataset. arXiv 2020, arXiv:2006.07397. [Google Scholar]
  35. Ni, Y.; Meng, D.; Yu, C.; Quan, C.; Ren, D.; Zhao, Y. Core: Consistent representation learning for face forgery detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, New Orleans, LA, USA, 19–20 June 2022; pp. 12–21. [Google Scholar]
  36. Dang, H.; Liu, F.; Stehouwer, J.; Liu, X.; Jain, A.K. On the detection of digital face manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Seattle, WA, USA, 13–19 June 2020; pp. 5781–5790. [Google Scholar]
  37. Li, Y.; Lyu, S. Exposing DeepFake Videos By Detecting Face Warping Artifacts. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Long Beach, CA, USA, 16–17 June 2019; pp. 46–52. [Google Scholar]
  38. Chollet, F. Xception: Deep learning with depthwise separable convolutions. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, HI, USA, 21–26 July 2017; pp. 1251–1258. [Google Scholar]
Figure 1. Motivation of the proposed framework. (a) Existing spatial-domain detection methods. (b) The proposed framework based on frequency-domain enhancement and feature disentanglement.
Figure 1. Motivation of the proposed framework. (a) Existing spatial-domain detection methods. (b) The proposed framework based on frequency-domain enhancement and feature disentanglement.
Ai 07 00344 g001
Figure 2. Overview of the proposed generalizable deepfake detection framework. During training, the Phase-Amplitude Frequency Enhancement (PAFE) module enhances frequency-domain forgery cues, while the Dual-branch Feature Disentanglement Module separates content and forgery features. Reconstruction, contrastive, and classification losses jointly optimize the framework. During inference, the PAFE module, forgery encoder, spatial self-attention module, and classifier are used for real/fake prediction.
Figure 2. Overview of the proposed generalizable deepfake detection framework. During training, the Phase-Amplitude Frequency Enhancement (PAFE) module enhances frequency-domain forgery cues, while the Dual-branch Feature Disentanglement Module separates content and forgery features. Reconstruction, contrastive, and classification losses jointly optimize the framework. During inference, the PAFE module, forgery encoder, spatial self-attention module, and classifier are used for real/fake prediction.
Ai 07 00344 g002
Figure 3. t-SNE visualization of the learned features. Subfigures (a) and (b) illustrate content features and forgery features, respectively. (Blue dots: real face samples; Red dots: fake face samples).
Figure 3. t-SNE visualization of the learned features. Subfigures (a) and (b) illustrate content features and forgery features, respectively. (Blue dots: real face samples; Red dots: fake face samples).
Ai 07 00344 g003
Figure 4. Enhanced Fourier spectra. From top to bottom: authentic face images with their corresponding enhanced Fourier spectra, followed by fake face images and their respective enhanced Fourier spectra.
Figure 4. Enhanced Fourier spectra. From top to bottom: authentic face images with their corresponding enhanced Fourier spectra, followed by fake face images and their respective enhanced Fourier spectra.
Ai 07 00344 g004
Figure 5. Forgery-aware attention heatmaps generated by Grad-CAM (Warm-colored regions (red/yellow) correspond to manipulation-sensitive facial areas that the model focuses on for real/fake classification, while cool-colored blue regions represent irrelevant background and normal facial areas.)
Figure 5. Forgery-aware attention heatmaps generated by Grad-CAM (Warm-colored regions (red/yellow) correspond to manipulation-sensitive facial areas that the model focuses on for real/fake classification, while cool-colored blue regions represent irrelevant background and normal facial areas.)
Ai 07 00344 g005
Table 1. Cross-dataset evaluation results in terms of AUC. All models are trained on FF++ and evaluated on unseen datasets without additional fine-tuning. The best and second-best results are highlighted in bold and underlined, respectively.
Table 1. Cross-dataset evaluation results in terms of AUC. All models are trained on FF++ and evaluated on unseen datasets without additional fine-tuning. The best and second-best results are highlighted in bold and underlined, respectively.
MethodCeleb-DFCeleb-DF-v2DFDCPDFDCDFD
Face X-ray [12]0.7090.6790.6940.6330.766
RECCE [15]0.7680.7320.7420.7130.812
F3Net [16]0.7770.7350.7350.7020.798
UCF [9]0.7790.7530.7590.7190.807
CORE [35]0.7800.7430.7340.7050.802
FFD [36]0.7840.7440.7430.7030.802
FWA [37]0.7900.6680.6380.6130.740
SRM [19]0.7930.7550.7410.7000.812
SPSL [18]0.8150.7650.7410.7040.812
Ours0.8070.7730.7980.7520.813
Table 2. Cross-manipulation evaluation results in terms of AUC under the leave-one-manipulation-out setting. The best and second-best results are highlighted in bold and underlined, respectively.
Table 2. Cross-manipulation evaluation results in terms of AUC under the leave-one-manipulation-out setting. The best and second-best results are highlighted in bold and underlined, respectively.
MethodNo-DFNo-F2FNo-FSNo-NT
Celeb-DFDFDCCeleb-DFDFDCCeleb-DFDFDCCeleb-DFDFDC
Xception [38]0.6600.6510.7160.6460.7370.6650.7090.647
CORE [35]0.7060.6300.7240.6710.7180.6610.6670.659
RECCE [15]0.6440.6360.7080.6410.7710.6530.7590.633
FWA [37]0.7110.6250.7160.7010.7520.6590.7310.670
Face X-ray [12]0.6780.6360.7170.7340.7530.6850.7440.631
SLADD [24]0.6450.7380.7260.6750.7150.7290.7340.759
F3Net [16]0.7450.7490.7390.7450.6990.6640.7780.797
SRM [19]0.7360.7250.7570.7370.6940.7280.7910.813
SPSL [18]0.7870.7430.7370.7590.6990.6690.8290.807
UCF [9]0.7490.7670.7820.7650.8000.7110.8080.800
Ours0.7830.7820.7920.7810.8120.7360.8240.836
Table 3. Component-wise ablation results in terms of AUC. The baseline uses Xception as the backbone. D, F, and A denote the Dual-branch Feature Disentanglement Module, PAFE module, and the spatial self-attention module, respectively. The best results are highlighted in bold.
Table 3. Component-wise ablation results in terms of AUC. The baseline uses Xception as the backbone. D, F, and A denote the Dual-branch Feature Disentanglement Module, PAFE module, and the spatial self-attention module, respectively. The best results are highlighted in bold.
ExperimentCore ModulesFF++Celeb-DF
1Baseline0.9010.672
2Baseline + D0.9540.654
3Baseline + D + A0.9580.696
4Baseline + D + F0.9670.764
5Baseline + D + F + A0.9720.807
Table 4. Loss-function ablation results in terms of AUC. L c l s , L r e c , and L c o n denote the classification loss, image-level reconstruction loss, and feature-level contrastive loss, respectively. The best results are highlighted in bold. The check mark indicates that this loss term is adopted in the current optimization strategy.
Table 4. Loss-function ablation results in terms of AUC. L c l s , L r e c , and L c o n denote the classification loss, image-level reconstruction loss, and feature-level contrastive loss, respectively. The best results are highlighted in bold. The check mark indicates that this loss term is adopted in the current optimization strategy.
Exp. L cls L rec L con FF++Celeb-DF
1 0.9560.685
2 0.9630.772
3 0.9700.759
40.9720.807
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, Q.; Feng, J.; Zhou, Y.; Li, M.; Zhang, Z.; Wang, L.; Song, W. Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement. AI 2026, 7, 344. https://doi.org/10.3390/ai7090344

AMA Style

Wang Q, Feng J, Zhou Y, Li M, Zhang Z, Wang L, Song W. Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement. AI. 2026; 7(9):344. https://doi.org/10.3390/ai7090344

Chicago/Turabian Style

Wang, Qian, Jiaqi Feng, Yu Zhou, Miao Li, Zhi Zhang, Luyao Wang, and Wenping Song. 2026. "Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement" AI 7, no. 9: 344. https://doi.org/10.3390/ai7090344

APA Style

Wang, Q., Feng, J., Zhou, Y., Li, M., Zhang, Z., Wang, L., & Song, W. (2026). Generalizable Deepfake Detection via Frequency-Domain Enhancement and Feature Disentanglement. AI, 7(9), 344. https://doi.org/10.3390/ai7090344

Article Metrics

Back to TopTop