Next Article in Journal
A Markerless RGB-Based Dataset of Continuous Hand Joint Kinematics in Functional Grasping Tasks
Previous Article in Journal
Contexere—Systematic Tracking and Referencing of Digital Artefacts for Postgraduate Students and Early Career Researchers
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

DeepFakeX: A Comprehensive Multimodal Deepfake Dataset for Research and Analysis

1
Department of Computer Science, National University of Computer and Emerging Sciences, Karachi 75030, Pakistan
2
Department of Computer Science, Salim Habib University, Karachi 74900, Pakistan
3
Center for Research in Computer Vision, University of Central Florida, Orlando, FL 32826, USA
*
Author to whom correspondence should be addressed.
Data 2026, 11(6), 141; https://doi.org/10.3390/data11060141
Submission received: 20 April 2026 / Revised: 22 May 2026 / Accepted: 8 June 2026 / Published: 11 June 2026

Abstract

The expanding capabilities of deep learning-based media synthesis have intensified concerns regarding the authenticity of digital content and the reliability of forensic analysis tools. In response to these challenges, this work introduces DeepFakeX, a collection of 800 synthetically generated videos available under controlled access for research purposes. The dataset encompasses four distinct categories of AI-driven synthesis: facial identity replacement, audio track substitution, neural voice cloning, and combined audiovisual alteration. Unlike existing deepfake datasets that predominantly focus on facial synthesis, DeepFakeX covers a broader range of manipulation modalities, reflecting the diversity of synthetic media encountered in real-world settings. All deepfakes were generated using state-of-the-art, publicly available tools. Standardized post-processing procedures were applied to each video to ensure uniformity in terms of quality, duration and encoding format. DeepFakeX also emphasizes diversity in gender, age, ethnicity, and language. Video contexts span speeches, informational videos, movie clips, news broadcasts, and interviews that reflect content scenarios commonly encountered in real-world online environments. The dataset includes videos in both English and Urdu. The dataset’s quality and structural variability were assessed through visual and audio analyses using the Structural Similarity Index Measure (SSIM), Mel-Frequency Cepstral Coefficients (MFCCs), and Principal Component Analysis (PCA). The evaluation results revealed substantial variability within each manipulation category, along with clearly distinguishable patterns specific to each modality. DeepFakeX has been developed to facilitate rigorous and transparent research in deepfake detection, cross-modal forensic analysis, and AI-driven media forensics. It is hosted on Zenodo under controlled access for research use.

1. Summary

The development of AI-generated media content, and deepfakes in particular, poses a multifaceted challenge concerning media authenticity, disinformation mitigation, and digital forensics [1]. Most datasets curated for deepfake detection have significant limitations in terms of modality coverage, manipulation type diversity, and demographic representation. DeepFakeX was initially developed during the evaluation of a deepfake detection model. During early testing, an initial set of 200 deepfake videos was generated to assess model performance across four manipulation modalities, and this subset was subsequently documented in a peer-reviewed article [2].
Building on this foundation, the dataset was expanded to 800 videos, offering a resource that the research community can utilize for the development and benchmarking of deepfake detection models. It consists of four deepfake types: face-swapped, audio-swapped, voice-cloned, and combined audiovisual. The deepfake content was created using a range of publicly accessible software frameworks and mobile applications, enabling realistic reproduction of diverse real-world media scenarios. DeepFakeX covers a wide range of video scenarios, genders, languages, ethnicities, and resolutions, making it broadly applicable to deepfake research. This data article outlines the dataset structure and describes how the data were obtained, providing the foundation for rigorous academic use and ensuring reusability across deepfake detection, multimedia forensics, and related domains.
The first step involved exploratory evaluation using structural and spectral measures, including SSIM [3], while acoustic characteristics [4] were assessed using extracted audio features. In addition, PCA was applied to MFCCs [5] to analyze clustering behavior within the feature space. These analyses confirmed the clear separability of the investigated categories and revealed the existence of modality-specific traits in the synthesized videos. DeepFakeX serves as a resource for scholarly research in the areas of deepfake detection, synthetic content assessment, and audiovisual forensic investigation. The dataset is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license and hosted on the Zenodo platform. The dataset is intended to foster open science practices, research transparency, and responsible progress in artificial intelligence.

2. Data Description

The design and structure of DeepFakeX are described in the subsequent subsections. This section presents a detailed account of the dataset’s organization, storage format, file naming conventions, and the diversity of deepfake modalities included. While the dataset is broadly applicable across multiple research domains, it is particularly suited to audiovisual forensic evaluation.
A machine-readable metadata file in CSV format is also provided with the dataset. This file contains key attributes for each video sample, including the video identifier, filename, deepfake category, generation method, language, contextual scenario, resolution, and duration, enabling easier dataset exploration and reproducibility of experimental studies.

2.1. Dataset Overview

The DeepFakeX dataset contains a total of 800 synthetically manipulated video samples. These samples are systematically organized into four primary categories based on the type of artificial alteration applied:
  • Identity replacement through facial manipulation.
  • Audio track substitution while preserving the original visuals.
  • Artificially generated speech imitating the original speaker.
  • Simultaneous manipulation of both facial identity and audio content.
Each category represents a distinct form of synthetic modification, enabling comprehensive evaluation across different deepfake modalities. All samples were created using commonly accessible, state-of-the-art desktop software, mobile applications, and AI-based tools. All media files were encoded in MP4 format using the H.264 codec, with video resolutions of either 720p or 1080p, durations ranging from 4 to 12 s, and frame rates of 25 or 30 frames per second. The audio signals were normalized across the dataset to ensure consistent loudness levels and perceptual audio quality.
The dataset follows a structured directory hierarchy, where compressed archives are grouped by deepfake category. Each folder contains video files that follow a standardized filename structure, where the assigned label clearly identifies the corresponding category of deepfake.

2.2. Directory Organization and File Naming Convention

DeepFakeX is organized into four primary directories (provided as compressed archives), with each directory dedicated to a specific deepfake type:
  • deepfake_face_swapped/—facial identity altered while preserving the original audio track.
  • deepfake_audio_swapped/—audio content modified while retaining the original visual stream.
  • deepfake_voice_cloned/—synthetic voice generation applied to the original visual content.
  • deepfake_combined_face_audio_swapped/—facial identity and audio content are both synthetically altered.
Table 1 lists the prefixes used to identify each deepfake category in the filename structure.
The name of every video file has an appended prefix that represents the category of deepfake, followed by an identifier. For example, ff_047.mp4 denotes the forty-seventh face-swapped video. Because deepfake types are already encoded in the filenames, separate annotation files are not required.

2.3. Modal Diversity

DeepFakeX reflects the diversity of user-generated media encountered in real-world online environments. The videos feature subjects diversified across the following dimensions:
  • Ethnicities: South Asian, African, East Asian, and Caucasian.
  • Gender Identity: Male, female, and transgender individuals are represented.
  • Age groups: Young, middle-aged, and older adults.
  • Languages: Videos are primarily in English, with a subset in Urdu.
  • Contexts: Video contexts include speeches, informational videos, interviews, news coverage, vlogs, and film excerpts.
Each sample was carefully selected to exclude sensitive content, including material involving minors, political leaders, religious figures, election-related events, and sexually explicit material. The deepfake videos are predominantly characterized by a single speaker in a frontal, camera-facing orientation with clear, intelligible speech.
Table 2 summarizes the ethnicity, gender, age, and content-context distributions within the dataset.

2.4. Data Specifications and Availability

  • Video format: All files are provided in .mp4.
  • Audio configuration: Mono or stereo tracks with normalized volume levels.
  • Spatial resolution: Videos are available in 720p or 1080p formats.
  • Video duration: Each video ranges between 4 and 12 s.
  • Usage License: Distributed under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license.

3. Methods

3.1. Deepfake Video Generation and Processing Pipeline

The pipeline used to construct DeepFakeX consists of several sequential stages, each designed to ensure consistency, diversity, and quality of the generated samples. DeepFakeX was developed through the following steps:
(1)
Source content selection;
(2)
Use of deepfake creation applications;
(3)
Post-processing operations for uniformity;
(4)
Dataset organization and file structuring.
The synthetic videos in DeepFakeX were generated using automated processes from publicly available source content featuring public figures (celebrities, content creators, and public-domain media), in line with established practice in the deepfake research literature (e.g., FakeAVCeleb, Celeb-DF, FaceForensics++). The original source videos are not redistributed; only the synthetic derivatives are included in the public release. The research does not involve human subjects in an experimental capacity, does not collect personal data from private individuals, and follows established ethical frameworks (IEEE Code of Ethics, ACM Code of Ethics, and the core principles of the EU General Data Protection Regulation), particularly the legitimate interests basis (GDPR Article 6(1)(f)) for research use of publicly available material.

3.2. Deepfake Generation Tools and Software Platforms

Table 3 summarizes the desktop-based software, mobile applications, and tools that were utilized to generate each deepfake category.
All generation frameworks and software tools were executed on a single local machine without relying on cloud services or distributed computing infrastructure. This setup ensured full control over the synthesis process, supporting consistency and replicability throughout dataset construction.

3.3. Post-Processing and Dataset Curation

To improve the quality and consistency of the created deepfake videos, the following post-processing were applied:
  • Format Standardization: The videos were encoded in MP4 format by utilizing the H.264 codec.
  • Resolution Adjustment: The video files were scaled to either 720p or 1080p resolution.
  • Frame Rate Normalization: The frame rate was standardized to 25 or 30 frames per second to ensure temporal consistency.
  • Audio Level Normalization: Audio levels were calibrated to achieve balanced and uniform audio output across the dataset.
  • File Naming Protocol: Each file was assigned a standardized prefix (e.g., ff_, fa_, vc_, ffa_) identifying its deepfake category.
In addition, every generated sample was manually reviewed to identify and remove those exhibiting visible artifacts, generation failures, or insufficient synthesis quality. Only samples demonstrating both visual and acoustic coherence were retained in the final dataset.

3.4. Dataset Validation and Quality Assessment

To analyze the structural diversity and audiovisual characteristics of the DeepFakeX dataset, a set of exploratory analytical procedures was employed. For comparative analysis, a set of original video sources served as reference material for computing visual and audio feature statistics. In particular, the Structural Similarity Index Measure (SSIM) [3] was used to quantify visual similarity between original and manipulated frames. SSIM measures perceptual similarity based on luminance, contrast, and structural coherence, producing values between 0 (no similarity) and 1 (perfect match). SSIM has been widely adopted in image and video quality assessment tasks, including deepfake detection and image manipulation analysis, and is used in this work to characterize the visual variability of synthetic content across manipulation categories.
These reference clips were used only to provide a baseline comparison between authentic and manipulated content. The original videos are not included in the released DeepFakeX dataset, as they were used solely during the validation stage and are excluded due to licensing and ethical considerations.

3.4.1. Visual Integrity Evaluation by Structural Similarity Index Measure (SSIM)

To characterize the visual integrity and diversity of the DeepFakeX dataset, we conducted two complementary SSIM analyses: (i) an inter-category analysis measuring structural deviation between real source videos and their corresponding manipulated counterparts, and (ii) an intra-category analysis measuring within-category visual diversity. SSIM produces values between 0 (no similarity) and 1 (perfect match) based on perceptual differences in luminance, contrast, and structural coherence; lower SSIM values indicate greater visual dissimilarity.
Inter-Category SSIM Analysis: For each manipulation category, ten real source videos were selected and compared against ten corresponding manipulated videos, yielding 100 pairwise comparisons per category. SSIM was computed on the first 100 frames of each video, with frames aligned by temporal index. The average inter-category SSIM scores were approximately 0.41 for real vs. face-swapped (range 0.19–0.90), 0.35 for real vs. audio-swapped, 0.39 for real vs. voice-cloned (range 0.18–0.66), and 0.33 for real vs. face-and-audio-swapped (range 0.15–0.61). These results indicate that all four manipulation types introduce substantial structural deviation from authentic source content, with combined face-and-audio-swapped videos showing the greatest deviation which is consistent with the compounded nature of dual-modality manipulation.
Intra-Category SSIM Analysis: To assess natural visual diversity within each category, all pairwise comparisons among ten randomly selected videos per category were computed (45 comparisons per category). The intra-category mean SSIM scores were:
-
Real: 0.37 (range 0.15–0.61).
-
Face-swapped: 0.32 (range 0.16–0.66).
-
Audio-swapped: 0.35 (range 0.04–0.63).
-
Voice-cloned: 0.36 (range 0.19–0.98).
-
Face-and-audio-swapped: 0.44 (range 0.33–0.77).
These intra-category values capture the inherent visual diversity of each category, reflecting variation in scenes, subjects, lighting, and manipulation techniques, rather than near-identity between a video and itself.
Interpretation: Figure 1 reports the intra-category mean SSIM values for each of the five categories. The “real” category value (~0.37) represents intra-category SSIM (real videos compared to other real videos in the dataset), not real-vs-itself (which would trivially yield ~1.0). The moderate value reflects the natural visual diversity across different real-world scenes and subjects and serves as a diversity baseline against which the intra-category values for manipulated categories can be interpreted. Notably, the face-and-audio-swapped category shows the highest intra-category SSIM (0.44), suggesting more uniform manipulation patterns within this category compared with single-modality manipulations.
The wide range of SSIM values across both inter- and intra-category comparisons confirms that DeepFakeX captures a realistic spectrum of synthesis complexity and visual diversity, supporting its use as a benchmark for evaluating detection performance across varying manipulation severities. In addition to SSIM-based characterization, all synthetic videos were subjected to manual visual inspection by the research team during dataset construction to verify acceptable visual quality and to exclude generation failures (e.g., severely distorted faces, audio synchronization errors) from the final release.

3.4.2. Audio Feature Analysis

To analyze the audio content of voice-cloned and audio-swapped deepfakes systematically, three major audio features were extracted:
  • Spectral centroid and Spectral Rolloff [17] capture spectral brightness distribution and overall energy concentration.
  • Zero-crossing rate (ZCR) [5] measures waveform complexity and the rate of signal sign changes, reflecting transient characteristics.
  • Mel-Frequency Cepstral Coefficients (MFCCs) [5] encode phonetic patterns and vocal-tract properties, and are widely used in speech and speaker recognition tasks.
Figure 2 presents boxplots comparing these features between original and synthetic samples. The empirical findings demonstrated that:
  • Audio generated through audio-swapping and voice-cloning techniques exhibited reduced spectral centroid and zero-crossing rate values, indicating more uniform and less complex waveform characteristics.
  • The MFCCs extracted from synthetic clips were more uniform compared to real audio; however, inter-class differences remained sufficiently distinct to support effective classification.
  • Real audio recordings exhibited a wider range of feature values, reflecting natural variation in pitch, timbre, and articulation.

3.4.3. Audio Feature Correlation Heatmap

Figure 3 presents a correlation heatmap [18] computed to examine the relationships among spectral audio features. The heatmap revealed the following:
  • Strong correlations among the spectral attributes of synthetic audio reflect the structured, algorithm-driven nature of the generation pipeline.
  • Real audio exhibited weaker and more diffuse inter-feature correlations, consistent with the natural variability inherent in human speech.

3.4.4. Spectrogram Comparison

Mel spectrograms [17] were computed for representative samples from each category and are presented in Figure 4. These representations reveal differences in spectral complexity and temporal dynamics across the deepfake categories.
  • Real speech exhibited complex harmonic transitions and greater diversity in energy distribution across the time-frequency domain.
  • Synthetic audio displayed flatter spectral profiles, characterized by smoother frequency bands and reduced temporal variation, consistent with the constrained output of generative synthesis pipelines.

3.4.5. Principal Component Analysis of MFCC Features

Principal Component Analysis (PCA) was applied to MFCC feature vectors to examine the distribution and clustering of samples across deepfake categories in the feature space.
  • Synthetic samples exhibited tight clustering, suggesting relative homogeneity introduced by the synthesis pipeline.
  • Real samples were more dispersed, reflecting the natural variability present in human speech.

3.4.6. Data Quality Assurance and Noise Mitigation

All audiovisual recordings were subjected to a manual inspection procedure designed to omit samples that exhibit the following:
  • Visible artifacts such as facial distortions and lip-synchronization mismatches.
  • Misalignment between vocal and facial expressions.
  • High levels of background noise or extended silent segments.
The dataset was curated to exclude duplicates and to ensure that every retained sample contributes distinct structural or acoustic information.
These analyses are intended to provide exploratory insights into the structural and acoustic characteristics of the dataset rather than a definitive validation of its quality.
Figure 5 presents the PCA projection of MFCC features across the real, audio-swapped, and voice-cloned categories, illustrating the clustering behavior described above.

3.4.7. Benchmark Evaluation Using State-of-the-Art Detection Models

To further assess the practical utility and research relevance of the DeepFakeX dataset, benchmarking experiments were conducted using a diverse set of state-of-the-art (SOTA) deepfake detection models. These models span visual-only, audio-only, lip-synchronization, and multimodal detection approaches, enabling assessment of model behavior across different manipulation categories. All detectors were evaluated in an inference-only setting, without additional training or fine-tuning on DeepFakeX.
Evaluation Subset: Each baseline was applied only to the DeepFakeX subset corresponding to the manipulation type the baseline was designed to detect, paired with real reference videos. Specifically:
-
Visual-only detectors targeting face-swap manipulations (MesoNet, GenConViT, AltFreezing, FreqNet, CNN-based Ensemble) were evaluated on the 500 face-swapped videos plus the 200 real reference videos.
-
Audio-only detectors (Synthetic Speech Detection, FakeSound, RawNet2Vocoder) were evaluated on the 100 audio-swapped or 100 voice-cloned videos (as appropriate to each model’s published scope) plus the 200 real reference videos.
-
Audiovisual and lip-synchronization detectors (SyncNet, SyncNet with Frequency, LIPINC) were evaluated on the audio-swapped, voice-cloned, and combined face-and-audio-swapped subsets plus the 200 real reference videos.
-
Multimodal detectors (FakeAVCeleb Ensemble, REFEREE, AVT2-DWF, Not Made for Each Other) were evaluated on all 800 synthetic videos plus the 200 real reference videos.
Train/Test Protocol: As DeepFakeX is positioned as an evaluation benchmark rather than a training resource in this work, all baseline models were applied in an inference-only configuration using their publicly released pretrained weights. No fine-tuning or re-training on DeepFakeX was performed for any baseline.
Preprocessing: Each baseline was evaluated using its published preprocessing pipeline to preserve its native performance characteristics. Common preprocessing steps included frame extraction at the model’s native frame rate, face detection and cropping (typically 224 × 224 for visual-only and multimodal models), and audio normalization with model-specific spectrogram parameters (e.g., Mel spectrograms with model-specified channel counts and analysis windows). Where a baseline required a specific input format, that baseline’s published preprocessing was applied; otherwise standard frame and audio extraction was used.
Decision Threshold: All baselines used the standard 0.5 sigmoid decision threshold to classify a video as authentic or manipulated. No per-model threshold tuning was performed, to avoid introducing evaluator bias.
Metric Computation: The F1-score reported in Table 4 was computed separately for each baseline on its applicable DeepFakeX subset (as listed under “Applied Deepfake Types”), rather than as a pooled or averaged value across all manipulation categories. This per-category reporting provides a more informative comparison of each baseline’s detection capability within the manipulation type it was designed for, and it avoids artificially deflating unimodal baselines on manipulation types outside their intended scope.
Table 4 summarizes the benchmarking results across different detection modalities and applicable deepfake types.
The benchmarking results demonstrate that multimodal approaches generally achieve stronger performance across combined manipulation types, whereas unimodal detectors show varying effectiveness depending on the manipulation modality. We note that “FakeAVCeleb Ensemble” in Table 4 refers to the ensemble detection baseline from the FakeAVCeleb paper [19] (combining Xception, MesoInception4, and Meso4), distinct from the FakeAVCeleb dataset discussed in the dataset comparison (Table 5).
Visual-only models perform competitively on face-swapped videos, while audio-based models show improved performance on voice-cloned and audio-swapped samples. These results underscore the importance of multimodal detection approaches and confirm the suitability of DeepFakeX as a challenging benchmark for evaluating cross-modal robustness and generalization.

3.5. Ethical Considerations

The DeepFakeX dataset was developed using generative AI techniques applied to publicly accessible source material featuring public figures (celebrities, content creators, and public-domain media), consistent with the established practice of comparable deepfake datasets including FakeAVCeleb [19], Celeb-DF, and FaceForensics++. We acknowledge that public availability does not, in itself, eliminate the need for explicit ethical and legal grounding. Accordingly, the ethical positioning of DeepFakeX rests on the following:
Legal basis: The use of publicly available source content is grounded in the legitimate-interests basis under Article 6(1)(f) of the EU General Data Protection Regulation [32], supplemented by the data-minimization principle of GDPR Article 5(1)(c), operationalized by the non-redistribution of source videos—a stronger safeguard than is typical of comparable datasets.
Licensing: The synthetic content is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license and hosted on Zenodo (DOI: 10.5281/zenodo.14202454) under a controlled-access mechanism by which researchers may request access subject to creator review.
Ethical review: The research does not involve human subjects in an experimental capacity, does not collect personal data from private individuals, and does not involve interaction with the people whose public-figure likenesses appear in the source content. Under these conditions, formal institutional ethics committee review was not required at our institution. The project nonetheless adhered throughout to the principles of established ethical frameworks: the IEEE Code of Ethics [33] (Principle I.5: respect for privacy; Principle I.7: avoid injury), the ACM Code of Ethics [34] (Principle 1.2: avoid harm), and the GDPR principles cited above.
Operational safeguards: Specific exclusion criteria were applied during dataset construction: content involving minors, political leaders, religious figures, election-related material, sexually explicit content, and private (non-public) individuals was strictly excluded.
All videos in the dataset are synthetically created and are intended solely for academic and research purposes, consistent with MDPI’s ethical principles for research publication and with established principles of responsible AI development.

4. Dataset Novelty and Research Relevance

DeepFakeX seeks to overcome limitations in current deepfake datasets by offering carefully structured and systematically curated resources that will enable future research in the field of deepfake detection. Although popular databases like FaceForensics++ [35], DFDC [36], and Celeb-DF [37] have made significant progress in advancing the field, they are subject to certain limitations, including the following:
Modality Imbalance: Most existing datasets focus predominantly on face-swap artifacts, with limited coverage of audio-based manipulations such as voice cloning or audio swapping. DeepFakeX addresses this by incorporating four deepfake types, i.e., face-swapped, audio-swapped, voice-cloned, and combined face-and-audio-swapped—enabling comprehensive evaluation across both unimodal and multimodal detection frameworks. The higher proportion of face-swapped samples reflects the greater prevalence of facial identity manipulation in real-world deepfake media. The dataset is used exclusively as an evaluation benchmark in this work; detection models were not trained on DeepFakeX but applied in an inference-only configuration using their publicly released pretrained weights (Section 3.4.7) so the asymmetric sample distribution does not introduce class-imbalance bias during model training. For benchmarking, per-category metrics are reported separately for each manipulation type in Table 4, ensuring fair evaluation regardless of sample size. Users intending to fine-tune or train detection models on DeepFakeX should consider stratified sampling or class-weighted loss to ensure balanced representation across categories.
Demographic Homogeneity: Existing datasets frequently exhibit imbalances in demographic representation across gender, age group, ethnicity, and language, which limits the generalizability of trained detection models. DeepFakeX was curated to achieve balanced representation across genders (male, female, and transgender) and includes speakers from different ethnicities, with content in both English and Urdu. This demographic and linguistic diversity is expected to enhance the generalizability of models trained on the dataset.
Contextual Limitations: Most existing datasets feature paid actors or scripted, studio-created content. In comparison, DeepFakeX consists of real-life scenarios, such as news-style broadcasts, public addresses, and educational content, which reflect the type of media regularly seen on digital platforms.
To further characterize the dataset, analytical evaluations were conducted using SSIM, MFCC, and PCA, demonstrating structural variability and modality-specific patterns across deepfake categories.
By addressing these limitations, DeepFakeX is designed to support:
  • Evaluation of deepfake detection performance across many modalities (visual, audio, and multimodal).
  • Generalization testing across diverse languages, demographic groups, and real-world scenarios.
  • Development of bias-aware and explainable deepfake detection models.
  • Enhancement of audiovisual coherence analysis for digital forensic investigation.
By enabling the examination of manipulated synthetic media under diverse, real-world conditions, DeepFakeX contributes to ethical artificial intelligence, integrated audiovisual forensic research, and the design of equitable detection frameworks.
To further position DeepFakeX within the existing landscape of deepfake datasets, a comparative analysis is presented in Table 5. The comparison highlights key characteristics, including manipulation types, modality coverage, and diversity-related attributes. Unlike many existing datasets that primarily focus on visual manipulations, DeepFakeX incorporates multiple forms of synthetic media generation and supports both audio and visual modalities, enabling comprehensive multimodal analysis.
Table 5. Comparative analysis of DeepFakeX with existing deepfake datasets based on manipulation types, modality coverage, and diversity attributes.
Table 5. Comparative analysis of DeepFakeX with existing deepfake datasets based on manipulation types, modality coverage, and diversity attributes.
DatasetFace SwapAudio SwapVoice CloningMultimodalDemographic DiversityLanguage DiversityRealistic Scenarios
FaceForensics++ [35]XXXMediumLowPartial
Celeb-DF v2 [37] XXXMediumLowPartial
DFDC [36]ΔΔHighLowPartial
DeepFake-TIMIT [38] XXXLowLowNo
FakeAVCeleb [19]XHighMediumPartial
DeepFakeX (Proposed)HighMediumFull
Note: Diversity ratings are defined by objective criteria applied uniformly across all entries: (1) Demographic Diversity—High: ≥3 ethnicities with balanced gender and varied age coverage; Medium: 2 ethnicities or partial balance; Low: predominantly homogeneous. (2) Language Diversity—High: ≥3 languages with balanced representation; Medium: 2 languages; Low: single-language/English-dominated. (3) Realistic Scenarios—(Full): multiple real-world contexts (e.g., news, interviews, tutorials, speeches); Partial: limited or indirect real-world support; No: controlled or laboratory settings only. Symbols: ✓ supports manipulation type; X absent; Δ partial support.
As shown in Table 5, most existing datasets predominantly focus on face-swapped visual manipulations and provide limited or no support for audio-based or multimodal deepfake scenarios. In contrast, DeepFakeX integrates multiple manipulation types, including face swapping, audio swapping, and voice cloning, within a unified framework. This enables more comprehensive evaluation of both unimodal and multimodal deepfake detection approaches.
Furthermore, DeepFakeX emphasizes diversity in terms of demographic attributes and linguistic variation, which are often underrepresented in existing datasets. The inclusion of multiple languages and varied real-world contexts enhances its applicability for realistic and robust deepfake detection research. These characteristics collectively distinguish DeepFakeX as a more versatile and application-oriented resource compared to existing datasets.

5. Limitations

While DeepFakeX provides a valuable multimodal resource for deepfake detection research, several limitations should be acknowledged.
Dataset Scale and Language Coverage: DeepFakeX comprises 800 synthetic videos across four manipulation categories, supplemented with 200 real reference videos. While this scale is comparable to several established deepfake datasets, larger collections remain valuable for training data-intensive detection models. Similarly, the current release covers two languages, English and Urdu, which represents an advance over predominantly English language existing datasets, but broader multilingual expansion is a priority for future releases.
Generative Model Coverage: The synthetic content was produced using a selected set of state-of-the-art generation tools current at the time of dataset construction (2024). Deepfakes generated by emerging or future generative architectures may exhibit different artifact signatures, and detection models trained exclusively on DeepFakeX may require periodic re-training to maintain performance against evolving manipulation techniques.
Demographic Representation: While DeepFakeX offers broader demographic representation than several predecessor datasets covering multiple ethnicities, gender distributions, and content contexts—demographic coverage remains finite. Subtle demographic biases may persist, and ongoing efforts to expand representation are planned.
Operational Deployment: DeepFakeX is intended as a research dataset for the development and benchmarking of deepfake detection methods. Real-world operational deployment of detectors trained on DeepFakeX, including testing under live-streaming conditions, adversarial perturbation scenarios, and integration with content-moderation pipelines is beyond the scope of the current work and represents a direction for applied research.
These limitations identify natural directions for future work and underscore the importance of ongoing dataset development in this fast-moving research area.

6. User Notes

The DeepFakeX dataset is hosted on Zenodo and can be accessed via the DOI: https://doi.org/10.5281/zenodo.14202454. It is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. Due to ethical and licensing considerations associated with the distribution of synthetic media containing identifiable human faces and voices, the dataset is provided under controlled access. Researchers interested in using the dataset for academic and research purposes may request access through the Zenodo platform. All access requests are reviewed by the dataset creators to support responsible use of the data.
The source media used to generate synthetic samples were obtained from publicly accessible online content and were used solely for research and dataset construction purposes. These original materials are not redistributed as part of DeepFakeX.
To improve transparency and reproducibility, a machine-readable metadata file in CSV format is also provided with the dataset. This file contains descriptive information for each video sample, including identifiers, deepfake category, generation method, language, context, resolution, and duration. The metadata file allows researchers to systematically analyze, filter, and organize the dataset for experimental studies.
No additional annotation files are included, as the dataset organization, folder hierarchy, filename prefixes, and accompanying metadata file clearly indicate the corresponding deepfake category and associated attributes for each sample. The dataset can be utilized by researchers to:
  • Train and evaluate deepfake detection systems across visual, audio, and multimodal frameworks.
  • Examine model robustness and generalization across diverse synthesis techniques.
  • Conduct research in digital media forensics, synthetic media detection, and content authenticity verification.
The majority of videos feature frontal facial views and clearly audible speech, with demographic diversity across gender, age, and ethnicity, enhancing suitability for algorithm development and experimental studies involving audiovisual data processing. In addition, the inclusion of face-swapped, audio-swapped, voice-cloned, and combined audiovisual manipulations makes the dataset particularly valuable for multimodal deepfake detection research.
Researchers using DeepFakeX are requested to cite this manuscript in their publications and may also use the provided metadata attributes and extracted feature statistics (e.g., SSIM trends and MFCC characteristics) as a foundation for further validation, benchmarking, or modeling efforts.

Author Contributions

S.S.: Conceptualization, Methodology, Software, Validation, Formal analysis, Investigation, Resources, Data Curation, Writing—Original Draft Preparation, Writing—Review and Editing, Visualization, Project Administration. J.A.S.: Conceptualization, Resources, Writing—Original Draft Preparation, Writing—Review and Editing, Supervision, Project Administration. R.Q.: Supervision, Writing—Review and Editing, Methodology Advisory. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The DeepFakeX dataset is available on Zenodo at https://doi.org/10.5281/zenodo.14202454 under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. Access is provided under controlled access and requires approval by the authors. Interested researchers may request access directly through the Zenodo request feature.

Acknowledgments

The authors acknowledge the use of publicly available deepfake generation tools and applications for the creation of synthetic video data. All data processing and video generation were independently performed by the authors for academic and non-commercial purposes.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
MFCCMel-Frequency Cepstral Coefficients
SSIMStructural Similarity Index Measure
ZCRZero Crossing Rate
PCAPrincipal Component Analysis
H.264Advanced Video Coding Standard
CC BY-NCCreative Commons Attribution NonCommercial License

References

  1. Salman, S.; Shamsi, J.A. Comparison of Deepfakes Detection Techniques. In Proceedings of the 2023 3rd International Conference on Artificial Intelligence (ICAI), Islamabad, Pakistan, 22 February 2023; pp. 227–232. [Google Scholar]
  2. Salman, S.; Shamsi, J.A.; Qureshi, R. Deep Fake Generation and Detection: Issues, Challenges, and Solutions. IT Prof. 2023, 25, 52–59. [Google Scholar] [CrossRef] [Scilit]
  3. Wang, Z.; Bovik, A.C.; Sheikh, H.R.; Simoncelli, E.P. Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Trans. Image Process. 2004, 13, 600–612. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Zafar, F.; Khan, T.A.; Akbar, S.; Ubaid, M.T.; Javaid, S.; Kadir, K.A. A Hybrid Deep Learning Framework for Deepfake Detection Using Temporal and Spatial Features. IEEE Access 2025, 13, 79560–79570. [Google Scholar] [CrossRef] [Scilit]
  5. Hamza, A.; Javed, A.R.R.; Iqbal, F.; Kryvinska, N.; Almadhor, A.S.; Jalil, Z.; Borghol, R. Deepfake Audio Detection via MFCC Features Using Machine Learning. IEEE Access 2022, 10, 134018–134028. [Google Scholar] [CrossRef] [Scilit]
  6. Perov, I.; Gao, D.; Chervoniy, N.; Liu, K.; Marangonda, S.; Umé, C.; Dpfks; Facenheim, C.S.; RP, L.; Jiang, J.; et al. DeepFaceLab: Integrated, Flexible and Extensible Face-Swapping Framework. arXiv 2021, arXiv:2005.05535. [Google Scholar]
  7. Chen, R.; Chen, X.; Ni, B.; Ge, Y. SimSwap: An Efficient Framework for High Fidelity Face Swapping. In Proceedings of the Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020; pp. 2003–2011. [Google Scholar]
  8. Siarohin, A.; Lathuilière, S.; Tulyakov, S.; Ricci, E.; Sebe, N. First Order Motion Model for Image Animation. In Proceedings of the 33rd International Conference on Neural Information Processing Systems; Curran Associates Inc.: Red Hook, NY, USA, 2020. [Google Scholar]
  9. Swapface. Available online: https://www.swapface.org/#/home (accessed on 20 August 2024).
  10. Reface—AI Face Swap App & Video Face Swaps. Available online: https://reface.ai/ (accessed on 20 August 2024).
  11. Prajwal, K.R.; Mukhopadhyay, R.; Namboodiri, V.P.; Jawahar, C.V. A Lip Sync Expert Is All You Need for Speech to Lip Generation In the Wild. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020; pp. 484–492. [Google Scholar]
  12. Pika. Available online: https://pika.art/ (accessed on 30 May 2024).
  13. Lalamu Studio. Available online: https://www.canva.com/your-apps/AAF5prtvncA/lalamu-studio?q=Lalamu+Studio (accessed on 20 August 2024).
  14. Free Text to Speech & AI Voice Generator|ElevenLabs. Available online: https://elevenlabs.io/ (accessed on 15 April 2025).
  15. PlayHT. Available online: https://playhtai.com/ (accessed on 15 April 2025).
  16. Gooey.AI—The Best of Private & Open Source AI. Available online: https://gooey.ai/ (accessed on 15 April 2025).
  17. Kumar, V.; Kapoor, A.; Chaudhary, R.R.; Gupta, L.; Khokhar, D. Preserving Integrity: A Binary Classification Approach to Unmasking Artificially Generated Voices in the Age of Deepkakes. In Proceedings of the 2024 11th International Conference on Computing for Sustainable Global Development (INDIACom), New Delhi, India, 28 February–1 March 2024; pp. 1449–1454. [Google Scholar]
  18. Amerini, I.; Barni, M.; Battiato, S.; Bestagini, P.; Boato, G.; Bruni, V.; Caldelli, R.; De Natale, F.; De Nicola, R.; Guarnera, L.; et al. Deepfake Media Forensics: Status and Future Challenges. J. Imaging 2025, 11, 73. [Google Scholar] [CrossRef] [Scilit]
  19. Khalid, H.; Tariq, S.; Kim, M.; Woo, S.S. FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset. arXiv 2022, arXiv:2108.05080. [Google Scholar]
  20. Xie, Z.; Li, B.; Xu, X.; Liang, Z.; Yu, K.; Wu, M. FakeSound: Deepfake General Audio Detection. arXiv 2024, arXiv:2406.08052. [Google Scholar] [CrossRef] [Scilit]
  21. Boo, H.; Lee, E.; Lee, J. Referee: Reference-Aware Audiovisual Deepfake Detection. arXiv 2025, arXiv:2510.27475. [Google Scholar]
  22. Tak, H.; Patino, J.; Todisco, M.; Nautsch, A.; Evans, N.; Larcher, A. End-to-End Anti-Spoofing with RawNet2. In Proceedings of the ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Toronto, ON, Canada, 6–11 June 2021; pp. 6369–6373. [Google Scholar]
  23. Datta, S.K.; Jia, S.; Lyu, S. Exposing Lip-Syncing Deepfakes from Mouth Inconsistencies. In Proceedings of the 2024 IEEE International Conference on Multimedia and Expo (ICME), Niagara Falls, ON, Canada, 15–19 July 2024. [Google Scholar]
  24. Wang, Z.; Bao, J.; Zhou, W.; Wang, W.; Li, H. AltFreezing for More General Video Face Forgery Detection. In Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Vancouver, BC, Canada, 17–24 June 2023; pp. 4129–4138. [Google Scholar]
  25. Wang, R.; Ye, D.; Tang, L.; Zhang, Y.; Deng, J. AVT2-DWF: Improving Deepfake Detection with Audio-Visual Fusion and Dynamic Weighting Strategies. IEEE Signal Process. Lett. 2024, 31, 1960–1964. [Google Scholar] [CrossRef] [Scilit]
  26. Tan, C.; Zhao, Y.; Wei, S.; Gu, G.; Liu, P.; Wei, Y. Frequency-Aware Deepfake Detection: Improving Generalizability through Frequency Space Domain Learning. Proc. AAAI Conf. Artif. Intell. 2024, 38, 5052–5060. [Google Scholar] [CrossRef] [Scilit]
  27. Afchar, D.; Nozick, V.; Yamagishi, J.; Echizen, I. MesoNet: A Compact Facial Video Forgery Detection Network. In Proceedings of the 2018 IEEE International Workshop on Information Forensics and Security (WIFS), Hong Kong, China, 11–13 December 2018; pp. 1–7. [Google Scholar]
  28. Bonettini, N.; Cannas, E.D.; Mandelli, S.; Bondi, L.; Bestagini, P.; Tubaro, S. Video Face Manipulation Detection Through Ensemble of CNNs. In Proceedings of the 2020 25th International Conference on Pattern Recognition (ICPR), Milan, Italy, 10–15 January 2021. [Google Scholar]
  29. Chugh, K.; Gupta, P.; Dhall, A.; Subramanian, R. Not Made for Each Other- Audio-Visual Dissonance-Based Deepfake Detection and Localization. In Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA, 12–16 October 2020; pp. 439–447. [Google Scholar]
  30. Towards End-to-End Synthetic Speech Detection|IEEE Journals & Magazine|IEEE Xplore. Available online: https://ieeexplore.ieee.org/document/9456037 (accessed on 31 May 2024).
  31. Deressa, D.W.; Mareen, H.; Lambert, P.; Atnafu, S.; Akhtar, Z.; Van Wallendael, G. Deepfake Video Detection Using Generative Convolutional Vision Transformer. arXiv 2023, arXiv:2307.07036. [Google Scholar] [CrossRef] [Scilit]
  32. European Parliament; Council of the European Union. Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the Protection of Natural Persons with Regard to the Processing of Personal Data and on the Free Movement of Such Data (General Data Protection Regulation). Off. J. Eur. Union 2016, L119, 1–88. [Google Scholar]
  33. IEEE. IEEE Code of Ethics; IEEE: Piscataway, NJ, USA, 2024. [Google Scholar]
  34. Association for Computing Machinery. ACM Code of Ethics and Professional Conduct; ACM: New York, NY, USA, 2018. [Google Scholar]
  35. Rossler, A.; Cozzolino, D.; Verdoliva, L.; Riess, C.; Thies, J.; Niessner, M. FaceForensics++: Learning to Detect Manipulated Facial Images. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 27 October–2 November 2019; pp. 1–11. [Google Scholar]
  36. Dolhansky, B.; Bitton, J.; Pflaum, B.; Lu, J.; Howes, R.; Wang, M.; Ferrer, C.C. The DeepFake Detection Challenge (DFDC) Dataset. arXiv 2020, arXiv:2006.07397. [Google Scholar] [CrossRef] [Scilit]
  37. Li, Y.; Yang, X.; Sun, P.; Qi, H.; Lyu, S. Celeb-DF: A Large-Scale Challenging Dataset for DeepFake Forensics. In Proceedings of the 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, WA, USA, 13–19 June 2020; pp. 3204–3213. [Google Scholar]
  38. Korshunov, P.; Marcel, S. DeepFakes: A New Threat to Face Recognition? Assessment and Detection. arXiv 2018, arXiv:1812.08685. [Google Scholar] [CrossRef] [Scilit]
Figure 1. Mean intra-category SSIM across video categories, computed via pairwise comparisons among 10 videos per category (45 pairs each), capturing within-category visual diversity rather than self-similarity.
Figure 1. Mean intra-category SSIM across video categories, computed via pairwise comparisons among 10 videos per category (45 pairs each), capturing within-category visual diversity rather than self-similarity.
Data 11 00141 g001
Figure 2. Boxplots of spectral centroid, ZCR, and MFCCs for real, audio-swapped, and voice-cloned samples, showing greater spectral and phonetic variability in real speech.
Figure 2. Boxplots of spectral centroid, ZCR, and MFCCs for real, audio-swapped, and voice-cloned samples, showing greater spectral and phonetic variability in real speech.
Data 11 00141 g002
Figure 3. Correlation heatmap of MFCC and spectral audio features. Strong positive correlations are observed among spectral features, while MFCCs show more varied relationships across the analyzed samples.
Figure 3. Correlation heatmap of MFCC and spectral audio features. Strong positive correlations are observed among spectral features, while MFCCs show more varied relationships across the analyzed samples.
Data 11 00141 g003
Figure 4. Mel spectrogram comparison of real, audio-swapped, and voice-cloned samples. Real audio exhibits richer spectral variability and smoother harmonic transitions, whereas synthetic samples show more uniform and repetitive patterns indicative of generation artifacts.
Figure 4. Mel spectrogram comparison of real, audio-swapped, and voice-cloned samples. Real audio exhibits richer spectral variability and smoother harmonic transitions, whereas synthetic samples show more uniform and repetitive patterns indicative of generation artifacts.
Data 11 00141 g004
Figure 5. PCA projection of MFCC features for real, audio-swapped, and voice-cloned samples. Real audio exhibits greater dispersion, while synthetic samples form relatively tighter clusters, indicating reduced variability introduced by generative processes.
Figure 5. PCA projection of MFCC features for real, audio-swapped, and voice-cloned samples. Real audio exhibits greater dispersion, while synthetic samples form relatively tighter clusters, indicating reduced variability introduced by generative processes.
Data 11 00141 g005
Table 1. Video prefixes according to the corresponding deepfake category.
Table 1. Video prefixes according to the corresponding deepfake category.
PrefixType of Deepfake
ff_Face-swapped
fa_Audio-swapped
vc_Voice-cloned
ffa_Face- + audio-swapped
Table 2. Demographic and contextual attribute distribution in DeepFakeX.
Table 2. Demographic and contextual attribute distribution in DeepFakeX.
AttributeDemographic CategoryProportion (%)
EthnicitySouth Asian37.8
Caucasian32.7
East Asian15.3
African14.3
Gender IdentityMale51.0
Female47.0
Transgender2.0
Age GroupYoung-aged19.2
Middle-aged60.6
Older Adults20.2
Video ContextNews segments12.7
Speeches22.4
Informational/Video clips34.7
Interviews23.5
Film Excerpts7.1
Table 3. Deepfake generation tools and corresponding video distribution across the four deepfake types.
Table 3. Deepfake generation tools and corresponding video distribution across the four deepfake types.
Deepfake CategoryGeneration Tools/PlatformsNumber of Videos
Face-swappedDeepFaceLab [6], SimSwap [7], First Order
Motion Model [8], SwapFace [9], Reface [10]
500
Audio-swappedWav2Lip [11], Pika Labs [12], Lalamu Studio [13]100
Voice-clonedElevenLabs [14], Play.ht [15], Gooey.ai [16], Lalamu Studio [13]100
Combined (Face + Audio) swappedCombination of face-swap and audio-swap tools listed above100
Table 4. Benchmark performance of state-of-the-art deepfake detection models evaluated on DeepFakeX.
Table 4. Benchmark performance of state-of-the-art deepfake detection models evaluated on DeepFakeX.
ModelDetection ModalityApplied Deepfake TypesF1-Score
FakeAVCeleb Ensemble [19]Audio–VisualAll deepfake types0.81
FakeSound [20]Audio-onlyAudio-swapped, voice-cloned0.69
REFEREE [21]Audio–VisualAll deepfake types0.64
RawNet2Vocoder [22]Audio-onlyVoice-cloned0.76
LIPINC [23]Visual-onlyAudio-swapped0.75
AltFreezing [24]Visual-onlyFace-swapped0.75
AVT2-DWF [25]Audio–VisualAll deepfake types0.80
FreqNet [26]Visual-onlyFace-swapped0.70
SyncNet [11]Audio–VisualAudio-swapped0.20
SyncNet with Frequency [2]Audio–VisualAll deepfake types0.69
MesoNet [27]Visual-onlyFace-swapped0.54
CNN-based Ensemble (Face Manipulation) [28]Visual-onlyFace-swapped0.35
Not Made for Each Other [29]Audio–VisualAll deepfake types0.62
Synthetic Speech Detection [30]Audio-onlyAudio-swapped0.54
GenConViT (Hybrid CNN-Transformer) [31]Visual-onlyFace-swapped0.68
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Salman, S.; Shamsi, J.A.; Qureshi, R. DeepFakeX: A Comprehensive Multimodal Deepfake Dataset for Research and Analysis. Data 2026, 11, 141. https://doi.org/10.3390/data11060141

AMA Style

Salman S, Shamsi JA, Qureshi R. DeepFakeX: A Comprehensive Multimodal Deepfake Dataset for Research and Analysis. Data. 2026; 11(6):141. https://doi.org/10.3390/data11060141

Chicago/Turabian Style

Salman, Sonia, Jawwad Ahmed Shamsi, and Rizwan Qureshi. 2026. "DeepFakeX: A Comprehensive Multimodal Deepfake Dataset for Research and Analysis" Data 11, no. 6: 141. https://doi.org/10.3390/data11060141

APA Style

Salman, S., Shamsi, J. A., & Qureshi, R. (2026). DeepFakeX: A Comprehensive Multimodal Deepfake Dataset for Research and Analysis. Data, 11(6), 141. https://doi.org/10.3390/data11060141

Article Metrics

Back to TopTop