Next Article in Journal
TC-GAFnet: A Time-Aware Contrastive Grouped Attention Fusion Network with Empirical Priors for Adolescent Vision Prediction
Previous Article in Journal
Data-Mining-Based Detector Calibrations for High-Energy Physics
Previous Article in Special Issue
LLM-Augmented Ensemble Reasoning for Adversarial-Aware Power Quality Monitoring in Smart Grids
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Evaluating Explainable Artificial Intelligence in EEG-Based Deception Detection

1
Information Technology Department, College of Computer and Information Sciences, King Saud University, Riyadh 11543, Saudi Arabia
2
Next Generation Connectivity and Wireless Sensors Institute, King Abdulaziz City for Science and Technology (KACST), Riyadh 11442, Saudi Arabia
3
Computer Science Department, College of Computer and Information Sciences, Imam Mohammad Ibn Saud Islamic University (IMSIU), Riyadh 11432, Saudi Arabia
*
Author to whom correspondence should be addressed.
Electronics 2026, 15(17), 4041; https://doi.org/10.3390/electronics15174041
Submission received: 16 August 2026 / Revised: 30 August 2026 / Accepted: 2 September 2026 / Published: 7 September 2026

Abstract

Electroencephalography (EEG) offers a promising basis for objective deception detection, as it captures covert neural responses that may not be accessible through behavior or self-report. In this context, explainable artificial intelligence (XAI) has the potential to both improve model performance and increase trust by clarifying which EEG features and temporal–spatial patterns drive classification decisions. This work presents a systematic review of EEG-based deception detection with a focus on how machine learning and deep learning methods are currently applied and to what extent XAI methods are integrated into these pipelines. The review organizes recent studies by dataset type, feature extraction methodology, and classification strategy, lie detection paradigms and their experimental and analytical designs. The goal is to identify methodological limitations and research gaps, synthesize current challenges and emerging trends, and propose directions for future work on interpretable EEG-based deception detection. The findings indicate a growing reliance on deep learning architectures but a limited and unsystematic use of XAI, highlighting the need for more principled integration of interpretability into deception detection research.

1. Introduction

Deception detection has long been used in criminal investigations, forensic science and security screening. Traditional approaches—including behavioral assessment and polygraph testing—infer deceit intent from observable physiological and behavioral responses. However, these reactions are not uniquely indicative of deception; they are equally susceptible to stress, anxiety, and other extraneous psychological factors. Given these inherent ambiguities, researchers have increasingly turned to neural correlates as a complementary source of evidence [1].
Electroencephalography (EEG) has emerged as one of the most prevalent neuroimaging techniques in this domain, owing to its capacity for non-invasive measurement of neural activity with millisecond-level temporal precision—a property that renders it particularly well-suited to investigating the rapid cognitive processes underlying deceptive behavior [2].
Machine learning (ML) and deep learning (DL) techniques have been increasingly employed in lie and deception detection. Early investigations relied upon handcrafted EEG features processed through conventional classifiers. In contrast, recent studies have shifted toward deep learning architectures that learn discriminative feature representations directly from raw EEG signals [3].
Despite this progress, the field suffers from a critical lack of standardization. Studies vary extensively in their experimental designs—from recording settings to preprocessing, feature extraction and evaluation—and this variability severely obstructs the generalizability of findings and renders direct performance comparisons across the literature largely untenable.
This study presents a methodological survey of recent EEG-based deception detection research. The surveyed literature is categorized according to their objective, dataset characteristics, and computational techniques. A notable gap across this body of work is the complete absence of explainable artificial intelligence (XAI). To address this gap, we undertake a parallel review of XAI implementations in broader EEG contexts, distilling interpretability strategies that may prove transferable to deception detection settings.
The contributions of this research are twofold: (i) a structured, unified classification scheme and (ii) a baseline framework integrating wavelet-based feature extraction with support vector machine (SVM) and random forest (RF) classifiers, using a public benchmark dataset.
The remainder of this paper is organized as follows. Section 2 presents the background. Section 3 describes the methodology. Section 4 introduces the proposed classification schema. Section 5 discusses the reviewed studies. Section 6 presents the experimental evaluation, and Section 7 concludes the paper.

2. Background

Deception is a complex cognitive activity that engages multiple brain regions, including the prefrontal cortex, anterior cingulate cortex, and both parietal and temporal lobes. Brain–computer interface (BCI) technology has diverse applications, including uses in forensics, neurorehabilitation, assistive technology, and gaming. BCIs consist of several key components, such as signal acquisition, preprocessing, feature extraction, and classification. This section aims to discuss the fundamentals of deception detection systems alongside BCI technology.

2.1. Deception from a Neuroscientific Perspective

Deception is a complex cognitive function that engages multiple brain regions and neural mechanisms, including executive control, response inhibition, memory retrieval, and decision-making. Neuroimaging studies utilizing EEG have identified key neural correlates linked to deceptive behavior. This process is orchestrated by a network of regions such as the prefrontal cortex (PFC), responsible for executive control; the anterior cingulate cortex (ACC), involved in conflict detection; the parietal cortex, associated with memory manipulation; and the temporal lobes, linked to language and memory processing [4].
The prefrontal cortex (PFC) plays an essential role in executive functions and inhibition during deceitful behaviors. Within the PFC, the dorsolateral PFC is believed to play an essential role in planning, working memory and cognitive flexibility; furthermore, it displays increased beta (12–30 Hz) and gamma activity during deceptive responses production. The ventrolateral PFC plays a significant role in inhibiting truthful responses through response suppression, with changes observed at frequencies between 2 Hz and 4 Hz during lying. Meanwhile, the orbitofrontal cortex (OFC) plays an important role in moral reasoning and impulse control—its reduced activity being observed among individuals who exhibit pathological lying tendencies, suggesting impaired emotional regulation [3].
The anterior cingulate cortex (ACCR) plays an important part in conflict monitoring between truthful and deceptive responses, with increased theta (4–8 Hz) activity signaling increased emotional tension and emotional tension associated with deception. The parietal cortex supports episodic memory retrieval processes as well as attentional processes, which increase activation when false memories are created; precuneus and posterior cingulate cortex (PCC) play roles in self-referential thinking as well as mental imagery, which facilitate detailed deceptive narratives [4]. The temporal lobes play an essential role in language processing and the manipulation of episodic memory, particularly the medial temporal lobe (MTL), which assists recall of personal experiences used to construct plausible false narratives, and the superior temporal sulcus (STS), which processes social cues and detects inconsistencies in verbal communication—key elements in monitoring deceptive speech [4].
EEG studies have identified consistent neural markers associated with deception, such as increased theta, beta, and gamma power, as well as specific event-related potential (ERP) components like the P300 and N200, which contribute to our knowledge about the neural basis of deception as well as potential applications within cognitive neuroscience research, forensic investigations, or neurolaw [4]. Event-related potentials (ERPs) offer additional insight into deception-related brain activity [4]. The P300 component, typically seen from 300–600 milliseconds after a stimulus, shows increased amplitude during deception due to greater emotional and decision-making demands; conversely, the N200 shows reduced amplitude during lying, which may reflect conflict detection mechanisms within the prefrontal cortex (PFC) and ACCR areas of the brain [3].
Spectrum analysis of EEG signals reveals that, during deception, theta activity (4–8 Hz; prevalent among frontal and ACCR regions) increases, reflecting greater cognitive effort and suppressing truthful responses. Beta activity (12–30 Hz; common among prefrontal and parietal regions), on the other hand, correlates to sustained engagement and mental effort with peak activity at the dorsolateral prefrontal cortex (DLPFC) and inferior parietal lobule (IPL). Gamma activity (>30 Hz; prevalent across both frontal and temporal lobes) correlates with complex memory retrieval or narrative construction during deceptive responses [3].

2.2. Brain–Computer Interface (BCI)

Brain–computer interface (BCI) technology facilitates direct communication between the brain and an external device bypassing traditional neuromuscular pathways. BCIs operate by monitoring neural activity. Active BCIs require users to intentionally produce brain activity to control an external system (e.g., moving prosthetic limbs); passive BCIs monitor subconscious brain signals for cognitive assessment purposes (e.g., detecting deception or fatigue). BCIs have many uses, from neurorehabilitation and assistive technology for disabled individuals to cognitive enhancement, gaming and lie detection systems [5].
EEG-based BCIs can help detect deception by analyzing ERP, frequency band changes (theta, beta and gamma) and connectivity patterns to discriminate between truth and falsity. Recently developed ML techniques have significantly enhanced BCI accuracy, making it a reliable solution for use in forensic science, security analysis, medical diagnostics as well as diagnostic medicine [5].
BCIs can serve multiple practical functions beyond deception detection, from neurorehabilitation to security [6]. Neurorehabilitation uses BCIs to assist individuals with motor disabilities such as paralysis in controlling robotic limbs or wheelchairs using brain signals [7]. BCIs are also being explored as potential polygraph replacement tests since EEG-based deception detection has shown increased accuracy and resistance against countermeasures when compared with traditional lie detectors. Finally, gaming industry users use BCIs to control virtual environments [8].
BCI systems consist of various basic components that enable communication between the mind and external devices, including: signal acquisition, signal processing, classification, and application interface.
Signal acquisition is the cornerstone of BCI systems, where neural activity is captured and converted to electrical signals for further processing. EEG, electrocorticography (ECoG), and implanting electrodes into the brain are methods to use sensors or electrodes to detect brain signals non-invasively or semi-invasively; a more invasive method through implanted electrodes is also possible. Signal acquisition quality depends on resolution, sensitivity and placement of electrodes. Since brain signals can often be weak and susceptible to interference from external noise and physiological artifacts (e.g., muscle movements and eye blinks), sophisticated filtering and amplification techniques must be utilized in order to increase signal clarity [9].
Signal processing is an integral component of BCIs, as it converts raw neural signals into meaningful data for interpretation and control. This process entails two major steps: preprocessing and feature extraction. Preprocessing involves filtering out noise such as muscle movements and electrical interference to improve signal quality, while feature extraction identifies relevant signal characteristics such as frequency bands, amplitude variations or specific brainwave patterns associated with user intentions. Efficient signal processing ensures accuracy, speed, and reliability for an effective BCI system that effectively utilizes brain signals for controlling external devices and applications [10].
Preprocessing of EEG data is intended to improve its quality by removing noise and artefacts that may obscure neural activity. This stage includes denoising and filtering. EEG signals can often become corrupted by artifacts such as blinks, cardiac activity, and muscle contractions that interfere with signal integrity. Advanced denoising techniques are employed to suppress or remove these interferences so that EEG analyses more accurately reflect true brain activity. Filtering can separate relevant brain activity from other unwanted signals, and several filtration methods such as high-pass, low-pass, and band-pass filters and notch filters are used. These filters remove unwanted frequency components, such as high-frequency muscle noise or low-frequency drift, and retain the desired frequency range of neural signals [10].
Once processing has taken place, feature extraction is used to isolate and quantify specific attributes of EEG signals which correspond with user intentions or cognitive states. Such features could include spectral power, frequency band ratios, statistical descriptors or time-domain metrics, such as the Fourier transform or wavelet transform, which decompose signals into their component frequency components, making it possible to identify oscillations (e.g., alpha, beta, theta and gamma bands) [10].
Classification is used to interpret the neural signals by analyzing the extracted features after preprocessing and converting them into actionable outputs. ML encompasses a diverse set of symbolic and statistical approaches that enable systems to learn from experience and adapt over time without explicit programming changes. BCI applications utilizing ML offer improved performance across a range of fields, including neurological diagnosis, seizure prediction and motor rehabilitation, by recognizing complex, non-linear patterns present in EEG/neuroimaging data. Support vector machines (SVMs) and artificial neural networks (ANNs) have become widely employed for use in supervised settings, while unsupervised and ensemble approaches provide greater robustness [9]. DL, an advanced extension of ML, uses multilayered neural networks, such as CNNs, RNNs, and Transformer architectures, to autonomously discover hierarchical representations from large datasets [9].
Application interfaces play a vital role in enhancing user interaction by providing real-time feedback based on the brain’s signals. This feedback can be visual (e.g., screen display), auditory (e.g., sound cues), or haptic (e.g., vibrations or force feedback), allowing users to perceive the outcome of their mental actions. Effective application interfaces rely on their ability to quickly respond to user commands, creating a seamless and intuitive experience for its users [11]. A well-designed application interface enhances usability and accessibility of BCI systems by providing feedback mechanisms for user training, reduces errors, and enhances the overall performance of the system. In adaptive BCIs, feedback is also used to refine ML models, ensuring better alignment between user intentions and system responses. Ultimately, an effective feedback system is essential for creating an intuitive and responsive BCI experience [12].

3. Methodology

A systematic search was conducted across the Web of Science Core Collection database and the Saudi Digital Library in accordance with the PRISMA guidelines. The search was restricted to publications from 2019 to 2025 to focus on recent studies related to EEG-based classification, machine learning, and deep learning. The search strategy used combinations of the keywords “EEG,” “deception detection,” and “lie detection.” The search retrieved 76 records; among these, 17 records were classified under the categories of Computer Science and Artificial Intelligence.
During the initial screening stage, 59 records were excluded because they did not focus on EEG-based lie or deception detection, did not employ a machine learning classifier, did not involve model training or evaluation using a labeled EEG dataset, used non-EEG modalities such as fMRI or GSR, or were review or survey papers without original experimental data. After applying the eligibility criteria, 12 studies were identified as eligible.
The eligibility criteria required that each study: (1) address lie or deception detection as the primary objective; (2) employ a machine learning or deep learning classifier; and (3) train and evaluate the model using an actual labeled EEG dataset. In addition, two relevant papers were identified through supplementary searching in the Saudi Digital Library. Therefore, a total of 14 EEG-based lie and deception detection studies were included in the main review. The flow diagram of the systematic search and selection process is presented in Figure 1.
During the review process, it was observed that none of the included EEG-based lie or deception detection studies applied explainable artificial intelligence (XAI) methods to interpret model decisions. Therefore, a supplementary search was conducted using the same databases and publication period to identify EEG studies that applied XAI techniques in related neurological tasks. This search identified 14 additional EEG-XAI studies, which are discussed in a dedicated section to contextualize the explainability gap in EEG-based lie and deception detection research. Accordingly, the total number of studies analyzed in this review was 28.
This literature review provides a focused examination of recent EEG-based lie and deception detection studies, with particular attention to experimental design, feature extraction methods, classification approaches, evaluation protocols, and the use of explainability techniques. Although the terms lie detection and deception detection are sometimes used interchangeably in the literature, this review categorizes the included studies according to their stated objective and methodological focus rather than by title alone. In this review, lie detection refers to studies that primarily classify truthful and deceptive responses, whereas deception detection refers to studies that examine neural patterns associated with deceptive behavior more broadly.
A total of 14 EEG-based lie and deception detection studies were included in the main review. Of these, 10 studies focused on lie detection, while four studies addressed deception detection, reflecting differences in research objectives, experimental paradigms, and analytical approaches across the reviewed literature.

4. Classification Schema

Scientific literature often uses the terms “lie detection” and “deception detection” interchangeably, yet they refer to two cognitively and experimentally distinct constructs.
  • Lie detection involves categorizing responses as either truthful or non-truthful based on direct answers to structured questions (e.g., yes/no, multiple choice). EEG-based studies in this category typically aim for binary classification, often using event-related potentials such as the P300 component.
  • Deception detection examines the cognitive and neural processes underlying deceptive behavior. This line of research may involve open-ended responses, emotional or intentional manipulation, functional brain network analysis, and more complex signal processing techniques (e.g., effective connectivity).
Importantly, some studies titled “deception detection” actually focus on lie detection (e.g., using ERP-based classification to distinguish truthful from deceptive responses). Therefore, this review categorizes each study based on computational approaches and methodological objective, rather than relying solely on the title.
A systematic classification scheme was developed to organize the 14 reviewed studies according to their methodological focus. First, studies were categorized by their stated objective into two primary types: lie detection (explicit identification of truthful responses) and deception detection (broader analysis of deceptive behaviors).
Within lie detection studies, we distinguished between classical machine learning approaches (e.g., SVM, random forests, logistic regression) and deep learning approaches (e.g., CNNs, RNNs, LSTMs). Among deception detection studies, we further separated classical machine learning approaches from connectivity or graph-theory analyses (e.g., functional connectivity, EEG coherence, network metrics).
In addition to the objective type and dataset characteristics, the reviewed studies were classified according to their computational methods. Two aspects were considered: feature extraction techniques and classification techniques. This classification facilitates the comparison of the computational methods used in EEG-based lie and deception detection studies.

5. Results and Findings

5.1. Objective Type

The reviewed studies were first classified according to their research objective as either lie detection or deception detection studies. Table 1 summarizes the reviewed studies in terms of objective, dataset, feature extraction method, classification approach, best reported accuracy, and evaluation protocol.
The evaluation protocol was included to indicate the validation setting used to obtain each reported accuracy. Since the reviewed studies vary in dataset size, feature representation, classifier design, and validation strategy, the reported accuracies should not be considered directly comparable. This issue is particularly relevant in EEG-based lie and deception detection, where trial-level or window-level data splits may include samples from the same participant in both training and testing sets, leading to optimistic performance estimates. Participant-wise validation protocols provide a more appropriate assessment of generalization to unseen participants.

5.1.1. Lie Detection Studies

Ten of the reviewed experiments focused on detecting whether a subject was lying, predominantly using binary decision tasks. The experiments were classified into two broad methodological categories: classical machine learning and deep learning.
Classical Machine Learning Approaches
Turnip et al. [16] developed an EEG-based deception detection method using a support vector machine (SVM) applied to the P300 component. EEG data were collected from 11 male participants (mean age 24 ± 3 years) during a mock-theft interview. Preprocessing included band-pass filtering and independent component analysis (ICA) to isolate the P300 from artifacts. Statistical features (minimum, maximum, mode, median, mean) were extracted from P300 amplitudes, and an SVM classifier achieved 70.83% accuracy.
Saini et al. [18] used EEG data recorded from 33 participants at nine electrode locations to classify stimuli into probe, irrelevant and target groups using EEG classification techniques such as wavelet decomposition or empirical mode decomposition for intrinsic mode functions to extract features of each stimulus type. A support vector machine with RBF kernel was optimized using grid search, with 10-fold cross validation achieving 99.44% accuracy—surpassing traditional classifiers like ANN and KNN because of its ability to account for non-stationary EEG signals.
Lakshan et al. [17] developed the PREDICTOR system for real-time deception detection to support criminal investigations. The system combined EEG-based lie detection, emotion recognition, and attentiveness monitoring. EEG data were recorded using a MUSE 2 headband from 30 participants. Features were extracted using Fisher’s linear discriminant analysis and approximate entropy. Among the classifiers tested (k-NN and random forest), random forest achieved 87% accuracy on mixed-gender data.
Deep Learning Approaches
  • Amber et al. [3] proposed a CNN-based method for P300 deception detection to categorize brain signals as guilty or innocent. EEG data from 34 participants were preprocessed, normalized, and transformed into 100 × 100 2D images. A CNN model trained on 70% of the data (20% testing, 10% validation) achieved 99.6% accuracy, outperforming SVM.
  • Baghel et al. [13] used a CNN for lie detection on two datasets: the Dryad dataset (30 participants, with honest and lying stimulus categories) and a custom dataset (50 samples from 10 participants). After time-domain filtering, the CNN with an Adadelta optimizer achieved 84.44% accuracy on Dryad and 82.00% on the custom dataset.
  • Mai and Nguyen [19] built a BCI system for lie detection using EEG. Eleven features in time, frequency, and time–frequency domains were extracted such as spectral entropy and variance; continuous wavelet transform (CWT) produced 8-dimensional images. Four models (CNN, gated recurrent unit (GRU), long short-term memory (LSTM), and multilayer perceptron (MLP)) were evaluated. CNN outperformed others with 96.51% average accuracy on CWT-generated images, reaching 98.02% on subject-dependent data.
Rahmani et al. [20] proposed EEG-based lie detection model using Type-2 fuzzy sets (TF-2) and deep graph convolutional networks (GCNs). Features were extracted as dynamic information from six graph convolutional layers. The model achieved 98.2% accuracy, with precision and sensitivity exceeding 98%.
Aslan et al. [14] introduced a hybrid model combining LSTM with neural circuit policy (NCP) to classify lies from truths. EEG data (13 channels, Emotiv EPOC+) were processed using discrete wavelet transform (DWT). The hybrid model achieved 97.88% accuracy, 96.18% precision, and 99.34% recall.
Aslan et al. [15] also developed the LieWaves dataset (5-channel Emotiv Insight). Preprocessing included band-pass filtering (0.55–45 Hz), automatic and tunable artifact removal (ATAR), and overlapping sliding window augmentation. Features were extracted using DWT, FFT, and statistical methods. Among CNN, LSTM, and CNN + LSTM, the DWT + LSTM combination achieved 99.88% accuracy.
Hamza et al. [21] developed an EEG-based lie detection system employing DL models for distinguishing truthful statements from deceptive ones. They used an OpenBCI Ultracortex “Mark IV” headset (14 channels) with 10 participants in a mock-robbery scenario. Features were extracted via FFT. CNN, LSTM, and multilayer perceptron’s (MLP) classifiers were compared; CNN achieved 99.96% accuracy on their dataset and 99.36% on the Dryad dataset.

5.1.2. Deception Detection Studies

While the above studies focus on binary lie classification, a second stream of research examines the broader cognitive and neural mechanisms of deception, often using connectivity analysis or multimodal signals. Four studies examined the cognitive and neural mechanisms of deception, often using binary classification but also exploring brain connectivity or multimodal physiological responses. Based on analysis method, these studies are divided as follows.
Classical Machine Learning Approaches
Daneshi Kohan et al. [22] used EEG connectivity analysis for real-time lie detection during open-ended interviews. Data were collected from 40 participants (32-channel Electrocap, 512 Hz) responding to neutral, truthful, and deceptive questions. Preprocessing included ICA and noise reduction. Features (coherence, dDTF, GPDC) were extracted from five frequency bands across nine brain regions. Principal component analysis (PCA) reduced dimensionality, and linear discriminant analysis (LDA) achieved 86.25% accuracy.
Another study, Daneshi Kohan et al. [24], integrated EEG and PPG signals with efficient connectivity techniques to identify interview fraud. The feature extraction process involved applying a wavelet technique to EEG and PPG signals, followed by connectivity analysis using methods like generalized partial directed coherence (gPDC) and direct directed transfer function (dDTF). Using leave-one-subject-out (LOO) evaluation, they obtained 84.14% average accuracy.
Connectivity/Graph-Theory Analysis
Gao et al. [25] investigated effective connectivity (EC) and graph theory for deception detection in terms of attention, working memory and conflict monitoring. Data from 30 participants (15 innocent, 15 guilty) were recorded with a 64-channel EEG system using a Guilty Knowledge Test protocol. Cortical activity was estimated via sLORETA across 24 ROIs. EC was measured using partial directed coherence (PDC). Graph metrics (in/out-degree, clustering coefficient, efficiency) were extracted, and an SVM achieved 99.06% accuracy.
Wei et al. [23] proposed weighted-directed functional brain networks (WDFBNs) using normalized phase transfer entropy (dPTE). The purpose of the study was to examine the directionality of brain network connections during misleading situations. A 32-channel EEG system (500 Hz) recorded data. dPTE features were extracted across delta, theta, alpha, and beta bands. The CatBoost classifier achieved 92.83% (delta), 94.17% (theta), 85.93% (alpha), and 92.25% (beta) accuracy.

5.2. Dataset Characteristics

Based on the datasets used in the reviewed studies, the datasets were classified into two categories: public datasets and self-collected datasets. Public datasets are openly available and can be used by different researchers for benchmarking and comparative evaluation. In contrast, self-collected datasets are gathered specifically for a particular study and are not publicly accessible. Among the reviewed studies, four studies utilized public datasets, namely Dryad, Bag-of-Lies, and LieWaves, whereas the remaining eleven studies relied on self-collected datasets. A summary of the datasets used in the reviewed studies is presented in Table 2.
Among the fourteen studies reviewed, three used publicly available datasets, namely Dryad, Bag-of-Lies, and LieWaves, whereas the remaining eleven studies used self-collected datasets. The Dryad dataset was recorded using a 12-channel EEG system. The Bag-of-Lies dataset was collected using a 13-channel Emotiv EPOC+ headset, while the LieWaves dataset was acquired with a 5-channel Emotiv Insight device. In contrast, the self-collected datasets differed in several aspects, including the number of participants, EEG configurations, sampling rates, and experimental procedures. The sample size varied from five participants in Mai and Nguyen [19] to 80 participants in Wei et al. [23].

5.3. Computational Methods

Beyond the taxonomic distinction between lie detection and deception detection, we also analyzed the computational methods employed across all studies: feature extraction techniques, classification algorithms, and explainable AI tools.

5.3.1. Feature Extraction Techniques

Feature extraction and selection play a critical role in EEG-based deception detection, directly influencing the performance of ML and DL models. By converting raw EEG signals into informative representations, they improve classification accuracy as well as allow the models to identify subtle patterns of deceptive behavior. Table 3 summarizes the feature extraction techniques employed across the reviewed studies.
Baghel et al. [13] used end-to-end automatic feature extraction via a CNN. Aslan et al. [14] applied DWT, supplemented with FFT and statistical features, before classification with an LSTM–neural circuit policy (NCP) hybrid.
In a study using traditional ML pipelines [16,17], Turnip et al. [16] applied independent component analysis (ICA) to isolate the P300 component, then extracted statistical features (mean, mode, min, max, median) from the resulting signals. Lakshan et al. [17] used Fisher’s linear discriminant analysis and approximate entropy to capture EEG signal complexity. Meanwhile, Saini et al. [18] utilized empirical mode decomposition (EMD), whereas Daneshi Kohan et al. [24] adopted a wavelet-based effective connectivity framework, both aiming to capture the time–frequency characteristics of non-stationary EEG signals.
Connectivity-based extraction was prominent in recent research. Daneshi Kohan et al. [22,24] used PDC, dDTF, and GPDC to measure effective connectivity. Gao et al. [25] additionally computed graph-theory metrics (in/out-degree, clustering coefficient, efficiency) to characterize topological features of brain networks. sLORETA was used for cortical activity estimation.
In deep learning contexts, Wei et al. [23] used dPTE to assess directional information flow. Amber et al. [3] preprocessed P300 ERP components into 2D images for CNN training. Mai and Nguyen [19] combined CWT with spectral entropy and variance features. Rahmani et al. [20] used graph convolutional networks with Type-2 fuzzy sets for dynamic feature learning. Hamza et al. [21] applied FFT for frequency-domain preprocessing before CNN/LSTM/MLP classification.
Overall, while earlier studies favored time-domain and frequency-domain methods with conventional ML classifiers, recent trends show a shift toward hybrid deep neural architectures and brain connectivity analyses, reflecting increasing sophistication in feature extraction for EEG-based deception detection.

5.3.2. Classification Techniques

Classification techniques play a pivotal role in EEG-based lie detection systems, as they directly determine their ability to discriminate between truthful and deceptive responses based on extracted EEG features. Comparative analysis of reviewed studies indicates an array of classification paradigms being employed such as ML, DL, and fuzzy logic (FL)-based approaches; all offering distinct strengths when applied as deception detection models. Table 4 lists the classification techniques employed across the reviewed studies.
SVM remains one of the most favored ML classifiers due to its strong performance on high-dimensional data and its ability to handle non-linear class boundaries. For example, Turnip et al. [16] used SVM effectively on non-stationary EEG signals.
DL models have gained increasing recognition due to their advanced abilities in learning hierarchically. DL models enable these sophisticated algorithms to identify patterns in EEG data that would otherwise be difficult or impossible for an engineer or human feature engineer to detect manually. CNNs were widely adopted for deception detection and lie detection tasks [3,13,15,19,21].
Hybrid LSTM-based model networks [14,15] leverage both spatial and temporal features of EEG signals. Aslan et al. [14,15] underscored this trend.
Rahmani et al. [20] developed more complex DL structures combining deep graph convolutional networks (GCNs) with Type-2 fuzzy sets (TF-2). This illustrates the increasing sophistication beyond traditional network designs to incorporate modeling of uncertainty and relational data structures inherent to EEG signal networks. Daneshi Kohan et al. [24] used wavelet-based connectivity features. Their EEG-PPG fusion model achieved 84.14% accuracy, highlighting the utility of multimodal physiological fusion.
In summary, while SVM remains a robust baseline, deep learning models—especially hybrid CNN-LSTM and graph-based networks—represent the cutting edge. These models, combined with effective feature extraction, significantly improve accuracy and robustness.

5.3.3. XAI and EEG-Based Analysis

None of the reviewed lie and deception detection studies employed XAI. To address this gap, EEG-based XAI studies were reviewed separately. The review focuses on the explainability methods used in EEG analysis and discusses their potential use in future EEG-based lie and deception detection systems. Table 5 presents a summary of the reviewed studies, including the application domain, learning model, explainability method, and the main findings.
SHAP was the most widely used XAI method in the reviewed studies. Arpaia et al. [26,27] applied it to identify frequency-band features driving cognitive workload and tES effect classification. Vieira et al. [28] and Pérez-Velasco et al. [29] used SHAP for both explanation and channel reduction in seizure detection and motor imagery decoding. Sylvester et al. [30] applied SHAP to localize N170 ERP components, and Khan et al. [31] used it to identify frequency-band features most influential for dementia subtype classification.
LIME and ELI5: Islam et al. [32] applied both ELI5 and LIME to explain EEG-based stroke prediction, while Hussain et al. [33] applied LIME to activity recognition. Both studies found that low-frequency spectral features (delta and theta bands) were the most influential inputs.
Grad-CAM and Integrated Gradients: Gagliardi et al. [34], Khan et al. [35], and Chaudary et al. [36] applied Grad-CAM to CNN-based emotion recognition models, producing spatially localized explanations consistent with known emotion-related neural patterns. Khan et al. [35] additionally used integrated gradients to provide complementary temporal explanations.
Class Activation Mapping (CAM): Loaiza-Arias et al. [37] used CAM to visualize spatio-frequency patterns in motor imagery EEG classification, supported by QMIP-CCA to correlate physiological data with MI performance.
Built-in Interpretability: Ye et al. [38] embedded interpretability directly into the EEG-GMACN architecture using graph weights and mutual attention, identifying key electrodes and conveying prediction uncertainty without post hoc tools.
TCAV: Brenner et al. [39] applied TCAV to EEG abnormality detection, generating explanations based on clinically meaningful concepts rather than low-level features.
Table 5. XAI methods applied to EEG studies.
Table 5. XAI methods applied to EEG studies.
ReferenceEEG Task TypeML ModelXAI MethodKey Interpretability
Findings
Loaiza-Arias et al. [37]Motor imageryShallow/multimodal DLClass activation mapping (CAM); QMIP-CCASpatio-frequency MI patterns visualized; physiological data correlated with MI performance.
Ye et al. [38]EEG/BCI classificationGraph mutual-attention CNNBuilt-in graph weights and attentionKey electrodes highlighted via learned graph weights; attention scores reflect prediction uncertainty.
Arpaia et al. [26]Cognitive workload (fine motor)ML classifier with feature selectionSHAPDelta power (C3, Fz) and theta power (Fz) identified as top features; both decreased under high load.
Arpaia et al. [27]tES effect identification (MS)SFS + SVMSHAPDelta and theta band features identified as key markers of TES neuromodulation effects in MS patients.
Islam et al. [32]Stroke predictionAdaptive gradient boostingELI5 and LIMEDelta and theta spectral features were key contributors to stroke vs. control predictions.
Vieira et al. [28]Seizure detectionFeature-reduced classifiersSHAPSix temporal features and five channels maintained >95% accuracy.
Pérez-Velasco et al. [29]Motor imagery (inter-subject MI decoding)EEGSym DLSHAP Frontal electrodes (F7, F8) and first ~1500 ms drove decoding; reduced 8-electrode setup validated.
Hussain et al. [33]Human activity recognitionRandom forest, gradient boostingLIMEModel decisions aligned with spectral band knowledge; key features identified per activity class.
Sylvester et al. [30]ERP (face perception; N170)Convolutional neural network (CNN) classifierSHAP Occipital N170 cluster identified; importance scores quantify ERP relevance across time and electrodes.
Brenner et al. [39]EEG abnormality detectionXceptionTime DLTCAVConcept-based explanations using clinically meaningful EEG concepts rather than raw features.
Khan et al. [31]Dementia classification (AD/FTD)TCN + LSTMSHAPFrequency-band power features identified as most influential per dementia subtype.
Gagliardi et al. [34]Fine-grained emotion recognitionExplainable CNNGrad-CAMBrain–heart features validated; activations consistent with known emotion-related neural patterns.
Khan et al. [35]Emotion recognition (SEED)EEG-ConvNet (CNN)Grad-CAM and integrated gradients (IGs)Subject-dependent approaches improve emotion recognition accuracy.
Chaudary et al. [36]Emotion recognition (SEED)CNN ensemble (model souping)Grad-CAMSpatially localized explanations consistent with known neurological correlates of emotion.

6. Findings

A review of 14 research studies on EEG-based deception detection and 13 studies on XAI applications to EEG analysis provided a comprehensive overview of current while also revealing several critical challenges: lack of dataset standardization, limited generalization across heterogeneous populations, and insufficient validation in real-world scenarios.

6.1. Methodological Concerns in Deception Detection Studies

EEG-based deception detection has shown promising advances, yet significant methodological issues persist.
Many studies report high classification accuracies without employing rigorous cross-validation practices, raising concerns about overfitting and the reliability of findings. This problem is particularly acute in DL approaches, where trial-and-error training often replaces systematic hyperparameter optimization. The reported high accuracies should be interpreted with caution given the small sample sizes prevalent in the field. With 20–30 participants per study, models have high capacity to memorize idiosyncratic neural patterns specific to the training sample, leading to optimistic performance estimates that do not generalize to new individuals or experimental conditions [35]. Furthermore, accuracy as a sole metric is insufficient—particularly in deception detection, where class distributions are often imbalanced and the costs of false positives and false negatives are asymmetrical. We recommend that future studies report precision, recall, F1-score, and AUC alongside accuracy and provide confidence intervals for all estimates.
Another notable limitation is the prevalence of subject-dependent evaluations, where training and testing occur on the same participants. While this approach artificially inflates performance metrics [35], it severely limits applicability in real-world, subject-independent scenarios—a context that remains largely understudied. For deception detection to transition from laboratory settings to forensic or legal applications, models must demonstrate generalization across unseen subjects, demographics, and experimental conditions.
Beyond evaluations practices, a fundamental limitation across the reviewed deception detection literature is the reliance on small, custom datasets—most studies employ fewer than 50 participants, often 20–30 subjects. This sample size is insufficient to capture inter-subject variability (age, gender, cognitive style, cultural background), and models trained on such datasets are unlikely to generalize to real-world forensic settings. By contrast, adjacent EEG domains such as emotion recognition have benefited from large-scale benchmarks like FACED (123 subjects) and EEGEmotions-27, enabling more robust evaluation. While emerging resources such as LieWaves (27 subjects) represent progress, their sample sizes remain modest.

6.2. The Interpretability Gap

Beyond methodological considerations, a major research gap lies in the lack of interpretability in deception detection models. Most studies prioritize classification accuracy as the sole success metric, with little attention to why a model reaches a particular decision. Current systems largely function as black-box solutions, which diminishes their scientific value and, more critically, limits their admissibility and trustworthiness in sensitive contexts such as criminal investigations or court proceedings.
This situation contrasts sharply with EEG emotion recognition research, where XAI techniques have been increasingly adopted [34,35,36]. The overemphasis on accuracy reflects a broader tendency to prioritize predictive performance over scientific understanding. While high accuracy is desirable, a model that achieves 95% accuracy on a small dataset but lacks interpretability provides limited scientific insight into the neurophysiological processes underlying deception—and limited practical utility in forensic settings where explainability is legally and ethically required. Bridging this interpretability gap in deception detection is essential for developing models that are not only accurate but also transparent, trustworthy, and suitable for real-world legal and forensic applications [36].

6.3. Insights from EEG and XAI Research

Beyond technical performance, the integration of XAI into deception detection is an ethical imperative. Forensic applications demand transparency, accountability, and the ability to explain decisions to judges, juries, and legal counsel. XAI methods thus serve not only scientific but also legal and ethical functions. Based on the reviewed EEG-based XAI studies, we propose a structured roadmap organized by the type of insight provided (Table 6).
  • Spectral Explanations (Frequency-Band Attribution): SHAP, LIME, and ELI5 consistently highlight band-power features (e.g., delta, theta, gamma, high beta) as key predictors across diverse EEG tasks [26,27,31,32,33]. In deception detection, these methods can quantify which frequency bands most influence classification decisions, helping validate neurophysiological theories of deception (e.g., frontal theta increases, parietal alpha suppression).
  • Spatial Explanations (Electrode-Level Attributions): CAM, Grad-CAM, and SHAP-based topographies localize informative electrodes, enabling channel reduction and compact system design [28,29,34,35,36,37,38]. For deception detection, spatial explanations could identify electrodes (e.g., Fz, Cz, Pz) most discriminative for deceptive responses, guiding optimal sensor placement.
  • Temporal Explanations (Time-Window Localization): Saliency maps and SHAP temporal attributions localize relevant prediction windows (e.g., the first 1500 ms in motor imagery) [29]. In deception detection, temporal explanations could identify critical time windows (e.g., P300 latency, response-locked epochs) associated with deceptive versus truthful responses.
  • Concept-Level Explanations: Testing with concept activation vectors (TCAV) enables concept-based explanations using clinically or cognitively meaningful features [39]. In deception detection, TCAV could test hypotheses about specific ERP components (e.g., P300, N400) or cognitive processes (e.g., response inhibition, cognitive load) without relying on low-level feature attributions.
  • Deployment-Oriented Reduction: XAI-guided feature and channel selection preserves performance while enabling portable systems (e.g., eight-electrode MI montage; six features + five channels for seizure detection) [28,29]. This is directly applicable to deception detection, where wearable, low-channel systems are desirable for forensic field deployment.
Interpretability findings from EEG and XAI studies offer a roadmap for addressing the above shortcomings. Several concrete strategies have emerged:
  • Confidence Quantification: Attention weights, calibration measures, and SHAP-based importance scores quantify model confidence and enhance trust in EEG classifiers.
  • Channel Reduction: AI-driven electrode and feature selection has enabled reduced configurations (e.g., eight electrodes, six features) that maintain high performance while supporting lightweight, portable systems [28,29].
  • Spatio-Temporal Attribution: Attributions at electrode, frequency, and latency levels generate biologically meaningful insights, supporting hypotheses about cortical networks and neural dynamics [26,27,28,29,30,31,34,35,36,37,38]. Applying these methods to deception detection could clarify which EEG features (e.g., P300 latency, frontal theta activity) truly underpin deceptive versus truthful responses [3,16,22,23,25].
The SHAP analyses identified frequency bands and electrodes that align with known EEG correlates of deception. Frontal theta power has been identfied as a key predictor in adjacent cognitive tasks such as workload and memory [26,27,31] consistent with its proposed role in cognitive control during deception. Delta/theta features were linked to P300 attenuation, a well-established ERP marker of cognitive load during deception [3,4,16]. Spatial explanations localized frontal (Fz, F7, F8) and centroparietal (Cz, Pz) electrodes, corresponding to prefrontal executive control and attention networks [22,23,25]. These convergences suggest that SHAP-identified features reflect genuine deception-related neural signals, not spurious correlations.
Building on the proposed roadmap, we propose specific recommendations for deception detection:
  • SHAP for global feature attribution to identify frequency bands and electrodes most discriminative for deception.
  • LIME or SHAP local explanations to justify individual trial classifications for forensic admissibility.
  • Grad-CAM for CNN-based detectors to produce spatio-temporal heatmaps localizing ERP components and time windows.
  • TCAV to test concept-level hypotheses (e.g., “cognitive load” or “response inhibition” as drivers of classification).
  • Multimethod convergence (SHAP + Grad-CAM + LIME) to cross-validate explanations and strengthen confidence in identified neural markers.
The effective implementation of the XAI roadmap depends critically on the availability of large, diverse, and publicly accessible datasets. We recommend that future research prioritize the collection of deception detection datasets with: (1) >50 participants to capture inter-subject variability, (2) diverse demographic representation, (3) ecologically valid paradigms, (4) multiple sessions per subject, and (5) standardized protocols. Such resources would enable rigorous cross-validation, robust model training, and meaningful comparison of XAI explanations cross studies.

6.4. Ethical and Legal Considerations

The forensic deployment of EEG-based deception detection systems raises profound ethical and legal challenges that extend well beyond technical performance. These concerns must be addressed proactively for the technology to achieve legitimate forensic application.
First, privacy and mental autonomy are implicated because EEG signals provide direct access to neural activity. Unlike traditional polygraphy, EEG-based deception detection constitutes a form of “brain-reading” that implicates fundamental rights to mental privacy and cognitive liberty. The potential for function creep—where data collected for one purpose are repurposed for surveillance or prediction—further compounds these risks.
Second, informed consent and coercion are particularly acute in criminal justice settings. Individuals may be pressured to undergo testing, raising tensions with the privilege against self-incrimination. Even when consent is nominally obtained, power asymmetry calls into question its voluntariness.
Third, algorithmic bias and fairness must be addressed. Models trained on small, non-representative datasets are vulnerable to systematic biases that could disproportionately affect demographic groups. The reviewed studies provide little evidence of subgroup validation, and XAI methods, while helpful for transparency, do not guarantee fairness. Bias audits are essential.
Fourth, stigma and social harm arise from false positives, which can lead to wrongful accusation and lasting reputational damage. The availability of such technology could erode public trust and create a chilling effect on civil liberties.
Fifth, misuse and mission creep remain concerns, as technologies developed for forensic use could be repurposed for workplace surveillance, border security, or political control. Clear boundaries on permissible use are needed.
Finally, regulatory and legal frameworks lag behind technical development. Jurisdictions are beginning to address neurorights, but legal safeguards risk being outpaced. We recommend engaging legal scholars and ethicists to co-develop governance frameworks that protect rights while enabling legitimate applications.
In summary, while EEG-based deception detection has achieved high reported accuracies, these figures often stem from optimistic, subject-dependent, and non-reproducible designs, compounded by reliance on small datasets that limit generalizability. Moreover, the ethical and legal dimensions of forensic brain-reading remain critically underexplored. Adopting rigorous cross-validation, subject-independent evaluation, the XAI methods outlined above, and the development of large-scale public benchmark datasets would address technical shortcomings.
However, these advances must be accompanied by robust ethical frameworks, informed consent protocols, bias audits, and regulatory oversight to ensure that EEG-based deception detection serves justice rather than undermining fundamental rights. Only through such an integrated approach can this technology transition from laboratory research to a scientifically grounded, interpretable, transparent, and forensically credible tool.

7. Experimental Evaluation

As part of a systematic literature review, an initial experimental evaluation has been performed in order to gauge the efficacy of machine learning models often employed as EEG-based deception detectors.

7.1. Dataset Description

We used the LieWaves dataset [15], a publicly available EEG corpus for deception detection. The data were collected from 27 health participants (students and faculty, mean age 23.1 years) using Insight, Emotiv Inc, San Francisco, United States. a 5-channel headset (AF3, T7, Pz, T8, AF4; 128 Hz sampling rate) during truth-telling and lying tasks. Stimuli were presented as projected videos in a quiet room.
The LieWaves dataset offers high-quality EEG data corresponding to both deceptive and truthful responses for machine learning models designed for lie detection. Due to these unique qualities, this dataset makes a good basis for building and testing machine learning models for lie detection [15], providing researchers in cognitive neuroscience and forensic computing a reliable resource for their future studies.
The dataset is balanced across conditions and has received ethics approval (Firat University Non-Invasive Research Ethics Committee). For this research, we employed the full dataset to test whether a simple, single-feature classifier could detect deception above chance—a proof-of-concept before more complex modeling—and the characteristics of this dataset are summarized in Table 7.

7.2. Implementation

The experimental pipeline was implemented in Python 3.9 using scikit-learn 1.3.0 and consisted of four sequential stages: preprocessing, feature extraction, classification and evaluation.

7.2.1. Preprocessing

EEG signals were then segmented using an overlapping sliding window (OSW) to capture temporal dynamics. Specifically, windows of 2 s duration with a 1 s step size (50% overlap) were used. This approach improves temporal representation and increases the number of training samples. Within each window, z-score normalization was applied to each feature distribution to scale the features, promote stable distributions and enhance model robustness.

7.2.2. Feature Extraction

Each normalized segment was processed using the discrete wavelet transform (DWT) with the Daubechies 4 (db4) wavelet at decomposition level 5. From the resulting wavelet coefficients, the following statistical measures were extracted: mean, standard deviation, variance, energy, root mean square (RMS), skewness, kurtosis, maximum, minimum, median, interquartile range (IQR), zero-crossing rate (ZCR), and entropy.
Wavelet packet (WP) decomposition was also evaluated by extracting node energy and relative energy features from the wavelet packet coefficients. In addition, spectral features were obtained using Welch’s power spectral density (PSD) with a Hanning window of 256 samples and 50% overlap, computing band power for the delta (0.5–4 Hz), theta (4–8 Hz), alpha (8–13 Hz), beta (13–30 Hz), and gamma (30–45 Hz). Band power values were computed for the delta, theta, alpha, beta, and gamma frequency bands. The DWT, WP, and PSD features were combined into a single feature vector and normalized using the StandardScaler (zero mean, unit variance) before classifier training. This complete pipeline is detailed above to ensure reproducibility.

7.2.3. Classification Algorithms

Support vector machine (SVM) was used with RBF kernel. Two versions of SVMs have been used: baseline and optimized. The regularization parameter (C) and kernel parameter (γ) were optimized using GridSearchCV with five-fold cross-validation (Table 8) to avoid data leakage during hyperparameter tuning. The optimized configuration (C = 10, γ = 0.01) was then used for final evaluation.
Random forest was constructed using an ensemble of decision trees trained on bootstrap samples of the training data, with a random subset of features considered at each split. The final class label was determined by majority voting across all trees. In this study, the random forest classifier was configured with 500 decision trees and square-root feature selection at each split.

7.2.4. Model Evaluation

To evaluate model performance, five-fold stratified cross-validation was employed. The dataset was divided into five folds while preserving the proportional distribution of truthful and deceptive samples across the folds. In each iteration, four folds were used for model training and the remaining fold was used for testing, with each fold serving once as the test set. This procedure was repeated across all five folds to provide a more robust estimate of model performance than a single train–test split. The final performance was reported as the mean across the five folds. Model performance was assessed using widely adopted classification metrics, including accuracy, precision, recall, and F1-score.
Accuracy refers to the ratio of correctly classified instances (both true and false) relative to total instances as seen in Equation (1), where TP is True Positive, TN is True Negative, FP is False Positive, and FN is False Negative.
Accuracy = T P + T N TP + TN + FP + FN
Precision measures the proportion of correctly identified deceptive instances among all instances classified as deceptive as shown in Equation (2).
Precision = T P T P + F P
Recall (also referred to as sensitivity) measures the proportion of correctly identified deceptive instances relative to all actual instances as shown in Equation (3).
Recall = T P T P + F N
F1-score is the harmonic mean of precision and recall, providing a balanced measure that accounts for both false positives and false negatives as shown in Equation (4).
F1-score = 2 × P r e c i s i o n × R e c a l l P r e c i s i o n + R e c a l l

7.2.5. Results and Discussion

The performance of the proposed model was evaluated using five-fold cross-validation. The results, averaged across all folds, are presented in Table 9. The initial SVM model achieved 62.62% accuracy. After incorporating enhanced feature extraction techniques and hyperparameter optimization, its accuracy increased dramatically to 84.75%. The random forest model demonstrated strong performance, achieving 83.75% accuracy without extensive parameter tuning, indicating its effectiveness for EEG-based deception detection tasks.
The marked improvement in the SVM model underscores the critical role of feature engineering and hyperparameter optimization in EEG classification. Model performance is highly dependent on the quality and diversity of extracted features. These results show that traditional ML approaches can be competitive when combined with appropriate preprocessing and feature extraction. When properly tuned, the SVM outperformed the random forest, suggesting kernel-based methods are particularly adept at capturing complex EEG patterns. The random forest’s strong performance aligns with its known ability to handle high-dimensional, non-linear data, consistent with existing literature.
SHapley Additive Explanations (SHAP) was used to further analyze the behavior of an optimized SVM model by assessing its contribution of EEG features extracted during extraction to classification decisions. As shown in Figure 2, global SHAP analysis indicates that DWT maximum, DWT minimum, DWT relative energy, DWT zero-crossing rate (ZCR), and wavelet packet energy provided useful information in distinguishing deceptive from truthful responses in EEG responses.
Figure 3 displays the SHAP summary plot, which illustrates both magnitude and direction of each feature’s contribution across evaluated samples. Its variation shows how individual features may have differential influence depending on which EEG segment was being examined—this ties in well with the non-linear nature of SVM models like the optimized SVM model. Overall, SHAP results provide greater transparency of the model decision-making process as they confirm predictions are driven primarily by meaningful characteristics extracted from EEG signals.
Despite these encouraging results, several limitations must be noted. The evaluation was limited to traditional models using handcrafted features, which may not fully capture the spatio-temporal nature of EEG signals. Additionally, performance is influenced by preprocessing and feature extraction choices, limiting generalizability across conditions. Future work will incorporate deep learning models for automatic feature learning, along with more robust preprocessing and richer feature representations, toward a reliable EEG-based deception detection system.

8. Conclusions

This study reviewed recent EEG-based lie and deception detection research, focusing on research objectives, datasets, EEG acquisition settings, feature extraction methods and classification techniques. Based on its findings, a baseline framework consisting of wavelet-based feature extraction combined with support vector machine (SVM) and random forest (RF) classifiers was created and evaluated for EEG-based lie detection.
The review showed that none of the identified lie or deception detection studies incorporated XAI. Therefore, EEG-based XAI studies were reviewed separately to examine the explainability methods currently used in EEG analysis. The reviewed studies were classified according to spectral, spatial, temporal, and concept-based explanation approaches, providing an overview of methods that may be adapted for deception detection.
The reviewed studies also revealed several limitations. Most studies were based on self-collected datasets with relatively small numbers of participants, and substantial differences exist in EEG acquisition protocols, preprocessing pipelines, and evaluation procedures. As a result, comparing the reported results across studies remains difficult, and the limited number of publicly available benchmark datasets continues to affect reproducibility.
Future work will extend the proposed framework by investigating deep learning models for automatic feature learning from EEG signals. In addition, the XAI methods reviewed in this study will be integrated into the proposed framework to examine their use in EEG-based lie and deception detection.

Author Contributions

Conceptualization, M.A. (Mashael Aldayel), M.A. (Maryam Alkanhal) and A.A.-N.; methodology, M.A. (Mashael Aldayel), M.A. (Maryam Alkanhal) and A.A.-N.; formal analysis, M.A. (Mashael Aldayel), M.A. (Maryam Alkanhal) and A.A.-N.; investigation; writing—original draft preparation, M.A. (Maryam Alkanhal); writing—review and editing, M.A. (Mashael Aldayel), M.A. (Maryam Alkanhal) and A.A.-N.; visualization, M.A. (Maryam Alkanhal); supervision, M.A. (Mashael Aldayel) and A.A.-N.; funding acquisition, M.A. (Mashael Aldayel). All authors have read and agreed to the published version of the manuscript.

Funding

This scientific paper is derived from a research grant funded by the Research, Development, and Innovation Authority (RDIA)—Kingdom of Saudi Arabia—with grant number (13461-imamu-2023-IMIU-R-3-1-HW-).

Data Availability Statement

The LieWaves dataset is available in the Mendeley Data repository at https://data.mendeley.com/datasets/5gzxb2bzs2/2, reference number DOI: 10.17632/5gzxb2bzs2.2. This dataset was originally presented in Aslan et al. (2024) [15].

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Denno, D.W. The Place for Neuroscience in Criminal Law. In Philosophical Foundations of Law and Neuroscience; Oxford University: New York, NY, USA, 2016; pp. 69–83. [Google Scholar]
  2. Simbolon, A.I.; Turnip, A.; Hutahaean, J.; Siagian, Y.; Irawati, N. An Experiment of Lie Detection Based EEG-P300 Classified by SVM Algorithm. In 2015 International Conference on Automation, Cognitive Science, Optics, Micro Electro-Mechanical System, and Information Technology; IEEE: Piscataway, NJ, USA, 2016. [Google Scholar]
  3. Amber, F.; Yousaf, A.; Imran, M.; Khurshid, K. P300 Based Deception Detection Using Convolutional Neural Network. In Proceedings of the 2019 2nd International Conference on Communication, Computing and Digital Systems, C-CODE 2019; Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2019; pp. 201–204. [Google Scholar]
  4. Happel, M. Neuroscience and the Detection of Deception. Rev. Policy Res. 2005, 22, 667–685. [Google Scholar] [CrossRef] [Scilit]
  5. Wolpaw, J.; Wolpaw, E. Brain-Computer Interfaces: Principles and Practice; Oxford University: New York, NY, USA, 2012. [Google Scholar]
  6. Meijer, E.; Ben-Shakhar, G.; Verschuere, B.; Donchin, E. A Comment on Farwell (2012): Brain Fingerprinting: A Comprehensive Tutorial Review of Detection of Concealed Information with Event-Related Brain Potentials. Cogn. Neurodyn. 2013, 7, 155–158. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  7. Luis, L.; Gomez-Gil, J. Brain Computer Interfaces, a Review. Sensors 2012, 12, 1211–1279. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  8. Ou, C.Z.; Lin, B.S.; Chang, C.J.; Lin, C.T. Brain Computer Interface-Based Smart Environmental Control System. In Proceedings of the 2012 8th International Conference on Intelligent Information Hiding and Multimedia Signal Processing, IIH-MSP 2012; IEEE: Piscataway, NJ, USA, 2012; pp. 281–284. [Google Scholar]
  9. Aggarwal, S.; Chugh, N. Review of Machine Learning Techniques for EEG Based Brain Computer Interface. Arch. Comput. Methods Eng. 2022, 29, 3001–3020. [Google Scholar] [CrossRef] [Scilit]
  10. Chavarriaga, R.; Carey, C.; Contreras-Vidal, J.L.; McKinney, Z.; Bianchi, L. Standardization of Neurotechnology for Brain-Machine Interfacing: State of the Art and Recommendations. IEEE Open J. Eng. Med. Biol. 2021, 2, 71–73. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  11. Arpaia, P.; Esposito, A.; Mancino, F.; Moccaldi, N.; Natalizio, A. Active and Passive Brain-Computer Interfaces Integrated with Extended Reality for Applications in Health 4.0. In International Conference on Augmented Reality, Virtual Reality and Computer Graphics; Springer International Publishing: Cham, Switzerland, 2021; pp. 392–405. [Google Scholar]
  12. Orlandi, S.; House, S.C.; Karlsson, P.; Saab, R.; Chau, T. Brain-Computer Interfaces for Children with Complex Communication Needs and Limited Mobility: A Systematic Review. Front. Hum. Neurosci. 2021, 15, 643294. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  13. Baghel, N.; Singh, D.; Dutta, M.K.; Burget, R.; Myska, V. Truth Identification from EEG Signal by Using Neural Network: Lie Detection. In 2020 International Conference on Telecommunications and Signal Processing (TSP); IEEE: Piscataway, NJ, USA, 2020; pp. 550–555. [Google Scholar]
  14. Aslan, M.; Baykara, M.; Alakuş, T.B. LSTMNCP: Lie Detection from EEG Signals with Novel Hybrid Deep Learning Method. Multimed. Tools Appl. 2024, 83, 31655–31671. [Google Scholar] [CrossRef] [Scilit]
  15. Aslan, M.; Baykara, M.; Alakus, T.B. LieWaves: Dataset for Lie Detection Based on EEG Signals and Wavelets. Med. Biol. Eng. Comput. 2024, 62, 1571–1588. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  16. Turnip, A.; Amri, M.F.; Fakrurroja, H.; Simbolon, A.I.; Suhendra, M.A.; Kusumandari, D.E. Deception Detection of EEG-P300 Component Classified by SVM Method. In Proceedings of the 6th International Conference on Software and Computer Applications; ACM: New York, NY, USA, 2017; pp. 299–303. [Google Scholar]
  17. Lakshan, I.; Wickramasinghe, L.; Disala, S.; Chandrasegar, S.; Haddela, P.S. Real Time Deception Detection for Criminal Investigation. In Proceedings of the 2019 National Information Technology Conference, NITC 2019; Institute of Electrical and Electronics Engineers Inc.: New York, NY, USA, 2019. [Google Scholar]
  18. Saini, N.; Bhardwaj, S.; Agarwal, R. Classification of EEG Signals Using Hybrid Combination of Features for Lie Detection. Neural Comput. Appl. 2020, 32, 3777–3787. [Google Scholar] [CrossRef] [Scilit]
  19. Mai, N.-D.; Nguyen, H. Deception Detection Using a Multichannel Custom-Design EEG System and Multiple Variants of Neural Network. In International Conference on Intelligent Human Computer Interaction; Springer International Publishing: Cham, Switzerland, 2021; pp. 104–109. [Google Scholar]
  20. Rahmani, M.; Mohajelin, F.; Khaleghi, N.; Sheykhivand, S.; Danishvar, S. An Automatic Lie Detection Model Using EEG Signals Based on the Combination of Type 2 Fuzzy Sets and Deep Graph Convolutional Networks. Sensors 2024, 24, 3598. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Hamza, H.W.; Al-Hamadani, A.A.; Al-Qazzaz, N.K. EEG Signals Classification Using Novel Acquisition Protocol for Lie Detection System. Ing. Syst. D’information 2025, 30, 157–167. [Google Scholar] [CrossRef] [Scilit]
  22. Daneshi Kohan, M.; Motie NasrAbadi, A.; Sharifi, A.; Bagher Shamsollahi, M. Interview Based Connectivity Analysis of EEG in Order to Detect Deception. Med. Hypotheses 2020, 136, 109517. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  23. Wei, S.; Gao, J.; Yang, Y.; Xiong, N.; Zhang, J.; Song, J.; Kang, Q.; Li, Y.; Lv, H. Analysis of Weight-Directed Functional Brain Networks in the Deception State Based on EEG Signal. IEEE J. Biomed. Health Inform. 2023, 27, 4736–4747. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  24. Daneshi Kohan, M.; Motie Nasrabadi, A.; Shamsollahi, M.B.; Sharifi, A. EEG/PPG Effective Connectivity Fusion for Analyzing Deception in Interview. Signal Image Video Process. 2020, 14, 907–914. [Google Scholar] [CrossRef] [Scilit]
  25. Gao, J.; Min, X.; Kang, Q.; Si, H.; Zhan, H.; Manyande, A.; Tian, X.; Dong, Y.; Zheng, H.; Song, J. Effective Connectivity in Cortical Networks During Deception: A Lie Detection Study Based on EEG. IEEE J. Biomed. Health Inform. 2022, 26, 3755–3766. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Arpaia, P.; De Luca, M.; Della Calce, A.; Carone, G.; Castelli, N.; Duran, D.; Gargiulo, L.; Moccaldi, N.; Nalin, M.; Perin, A.; et al. EXplainable Artificial Intelligence Improves EEG-Based Cognitive Workload Assessment Induced by Fine Motor Activity in Neurosurgeons. In Proceedings of the 2024 IEEE International Conference on Metrology for eXtended Reality, Artificial Intelligence and Neural Engineering (MetroXRAINE); IEEE: Piscataway, NJ, USA, 2024; pp. 588–593. [Google Scholar]
  27. Arpaia, P.; Ammendola, L.; Cropano, M.; De Luca, M.; Della Calce, A.; Gargiulo, L.; Lus, G.; Maffei, L.; Malangone, D.; Moccaldi, N.; et al. Identification of EEG Features of Transcranial Electrical Stimulation (TES) Based on EXplainable Artificial Intelligence (XAI). In 2024 IEEE International Conference on Metrology for eXtended Reality, Artificial Intelligence and Neural Engineering (MetroXRAINE); IEEE: Piscataway, NJ, USA, 2024; pp. 1153–1158. [Google Scholar]
  28. Vieira, J.C.; Guedes, L.A.; Santos, M.R.; Sanchez-Gendriz, I. Using Explainable Artificial Intelligence to Obtain Efficient Seizure-Detection Models Based on Electroencephalography Signals. Sensors 2023, 23, 9871. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Pérez-Velasco, S.; Marcos-Martínez, D.; Santamaría-Vázquez, E.; Martínez-Cagigal, V.; Moreno-Calderón, S.; Hornero, R. Unraveling Motor Imagery Brain Patterns Using Explainable Artificial Intelligence Based on Shapley Values. Comput. Methods Programs Biomed. 2024, 246, 108048. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Sylvester, S.; Sagehorn, M.; Gruber, T.; Atzmueller, M.; Schöne, B. SHAP Value-Based ERP Analysis (SHERPA): Increasing the Sensitivity of EEG Signals with Explainable AI Methods. Behav. Res. Methods 2024, 56, 6067–6081. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  31. Khan, W.; Khan, M.S.; Qasem, S.N.; Ghaban, W.; Saeed, F.; Hanif, M.; Ahmad, J. An Explainable and Efficient Deep Learning Framework for EEG-Based Diagnosis of Alzheimer’s Disease and Frontotemporal Dementia. Front. Med. 2025, 12, 1590201. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  32. Islam, M.S.; Hussain, I.; Rahman, M.M.; Park, S.J.; Hossain, M.A. Explainable Artificial Intelligence Model for Stroke Prediction Using EEG Signal. Sensors 2022, 22, 9859. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  33. Hussain, I.; Jany, R.; Boyer, R.; Azad, A.K.M.; Alyami, S.A.; Park, S.J.; Hasan, M.M.; Hossain, M.A. An Explainable EEG-Based Human Activity Recognition Model Using Machine-Learning Approach and LIME. Sensors 2023, 23, 7452. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  34. Gagliardi, G.; Alfeo, A.L.; Catrambone, V.; Cimino, M.G.C.A.; De Vos, M.; Valenzal, G. Fine-Grained Emotion Recognition Using Brain-Heart Interplay Measurements and EXplainable Convolutional Neural Networks. In Proceedings of the International IEEE/EMBS Conference on Neural Engineering, NER; IEEE: Piscataway, NJ, USA, 2023; Volume 2023-April. [Google Scholar]
  35. Khan, S.A.; Chaudary, E.; Mumtaz, W. EEG-ConvNet: Convolutional Networks for EEG-Based Subject-Dependent Emotion Recognition. Comput. Electr. Eng. 2024, 116, 109178. [Google Scholar] [CrossRef] [Scilit]
  36. Chaudary, E.; Khan, S.A.; Mumtaz, W. EEG-CNN-Souping: Interpretable Emotion Recognition from EEG Signals Using EEG-CNN-Souping Model and Explainable AI. Comput. Electr. Eng. 2025, 123, 110189. [Google Scholar] [CrossRef] [Scilit]
  37. Loaiza-Arias, M.; Álvarez-Meza, A.M.; Cárdenas-Peña, D.; Orozco-Gutierrez, Á.Á.; Castellanos-Dominguez, G. Multimodal Explainability Using Class Activation Maps and Canonical Correlation for MI-EEG Deep Learning Classification. Appl. Sci. 2024, 14, 11208. [Google Scholar] [CrossRef] [Scilit]
  38. Ye, H.; Goerttler, S.; He, F. EEG-GMACN: Interpretable EEG Graph Mutual Attention Convolutional Network. In 2024 46th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC); IEEE: Piscataway, NJ, USA, 2024. [Google Scholar]
  39. Brenner, A.; Knispel, F.; Fischer, F.; Rossmanith, P.; Weber, Y.; Koch, H.; Röhrig, R.; Varghese, J.; Kutafina, E. Concept-Based AI Interpretability in Physiological Time-Series Data: Example of Abnormality Detection in Electroencephalography. Comput. Methods Programs Biomed. 2024, 257, 108448. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Figure 1. PRISMA flow diagram of the study selection process.
Figure 1. PRISMA flow diagram of the study selection process.
Electronics 15 04041 g001
Figure 2. Feature importance via SHAP analysis for Lie class.
Figure 2. Feature importance via SHAP analysis for Lie class.
Electronics 15 04041 g002
Figure 3. SHAP summary graph for Lie class.
Figure 3. SHAP summary graph for Lie class.
Electronics 15 04041 g003
Table 1. Comparative analysis of EEG-based lie and deception detection studies.
Table 1. Comparative analysis of EEG-based lie and deception detection studies.
ObjectiveReferenceDatasetFeature ExtractionClassificationBest
Accuracy
Evaluation
Protocol
Lie DetectionBaghel et al. [13]Custom dataset (50 samples)Time-domain filters, CNN for feature learning from raw EEG dataCNN84.4% Dryad
82%
Custom
Holdout
Aslan et al. [14]Bag-of-Lies dataset Discrete wavelet transform (DWT)LSTM + NCP97.88%Holdout, subject-dependent
Aslan et al. [15]LieWaves datasetOSW, DWT, FFT, statistical methodsCNN, LSTM,
CNN-LSTM
99.88%.Holdout, subject-dependent
Turnip et al. [16] Custom dataset
(11 participants)
P300 component, statistical features (mean, median, mode)SVM70.83%.Holdout
Lakshan et al. [17] Custom dataset
(30 participants)
Various frequency features, feature engineeringk-NN,
random forest
87%Not reported
Saini et al. [18]Custom dataset
(33 participants)
Wavelet decomposition, EMDSVM99.44%Holdout
Amber et al. [3]Custom dataset
(34 participants)
Preprocessed EEG data into 2D imagesCNN99.6%Not reported
Mai and Nguyen [19]Custom dataset
(5 participants)
Continuous wavelet transform (CWT)CNN, GRU, LSTM, MLP96.51%70/20/10 random holdout
Rahmani et al. [20]Custom dataset
(20 participants)
Dynamic EEG features, graph convolutional networksGCN + TF-298.2%5-fold cross-validation (CV)
Hamza et al. [21]Custom dataset
(10 participants)
FFT for frequency-based featuresCNN, LSTM, MLP99.96%Holdout
Deception DetectionDaneshi
Kohan et al.
[22]
Custom dataset
(40 participants)
ICA, coherence, dDTF, GPDCLDA86.25%Leave-one-person-out
Wei et al.
[23]
Custom dataset
(80 participants)
Normalized phase transfer entropy (dPTE), graph analysisCatBoost, SVM and linear regression94.17%Nested 6-fold CV (outer) + 9-fold CV (inner)
Daneshi Kohan et al. [24]Custom dataset
(41 participants)
Wavelet-based technique, connectivity analysis (gPDC, dDTF)Signal fusion (EEG + PPG)84.14%Leave-one-subject-out
Gao et al.
[25]
Custom dataset
(30 participants)
Effective connectivity (EC), PDC, graph theorySVM99.06%15-fold subject-wise CV (outer) + nested 10-fold CV (inner)
Table 2. Dataset characteristics of the reviewed studies.
Table 2. Dataset characteristics of the reviewed studies.
Dataset TypeDataset NameParticipantsReferences
PublicDryad30[13,21]
Bag-of-Lies35[14]
LieWaves27[15]
Self-Collected<10[19]
10–15[13,16,21]
16–20[20]
30–35[3,4,17,18,25]
≥40[22,23]
Table 3. Feature extraction techniques used in related studies.
Table 3. Feature extraction techniques used in related studies.
MethodNo. of StudiesReferences
Wavelet decomposition (DWT/CWT)3[14,15,19]
Independent component analysis (ICA)1[22]
Statistical features3[15,16,17]
FFT-based features2[15,21]
Graph-theoretic features (PDC, dPTE, GPDC)3[22,23,25]
Automatic feature learning via CNN, GCN3[3,13,20]
Table 4. Classification algorithms used in related studies.
Table 4. Classification algorithms used in related studies.
MethodNo. of StudiesReferences
Support Vector Machine (SVM)4[16,18,23,25]
Convolutional Neural Network (CNN)5[3,13,15,19,21]
Long Short-Term Memory (LSTM)3[15,19,21]
Hybrid Models (CNN-LSTM, LSTM-NCP…)2[14,15]
Graph Convolutional Network (GCN) + Type-2 Fuzzy Sets1[20]
Random Forest (RF)1[17]
CatBoost1[23]
Signal Fusion (EEG + PPG)1[24]
Table 6. Recommended XAI Workflow for EEG Deception Detection.
Table 6. Recommended XAI Workflow for EEG Deception Detection.
StageXAI TechniquePurposeExample Output
Feature SelectionSHAP (global)Rank electrodes and frequency bandsFrontal theta and parietal alpha are top predictors
Model TrainingRule extractionGenerate global interpretable rulesIf frontal theta > threshold → deceptive
Per-Trial ExplanationLIME, SHAP (local)Explain individual predictionsHigh frontal theta at 300–500 ms drove this classification
Cross-Subject ValidationAveraged LIME, SHAP aggregationIdentify stable neural patternsConsistent P300 differences across subjects
VerificationMultimethod comparison (SHAP + Grad-CAM + LIME)Cross-validate explanationsSHAP and Grad-CAM converge on same electrodes
Table 7. Dataset Description.
Table 7. Dataset Description.
Participants27 healthy subjects (students and faculty), average age 23.1 years
EEG DeviceEmotiv Insight (5 channels: AF3, T7, Pz, T8, AF4)
Stimuli TypeVisual stimuli consisting of 10 unique prayer beads shown in randomized order
Experimental
Protocol
Each participant completed one truthful and one deceptive session. Each session lasted 75 s and consisted of 25 image trials (2 s per image) separated by black-screen intervals. Participants responded “YES” or “NO” according to their assigned role (truthful or deceptive).
Data Volume9600 EEG samples per session per channel
Table 8. SVM Hyperparameter Search Space.
Table 8. SVM Hyperparameter Search Space.
HyperparameterSearch Values
KernelRBF
C{0.1, 1, 10, 100}
Gamma{scale, 0.1, 0.01, 0.001}
Table 9. Performance Comparison of Machine Learning Models.
Table 9. Performance Comparison of Machine Learning Models.
ModelAccuracyPrecisionRecallF1-Score
SVM (Baseline)62.62%62.68%62.62%62.58%
SVM (Optimized)84.75%84.81%84.75%84.74%
Random Forest83.75%83.80%83.75%83.74%
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Aldayel, M.; Alkanhal, M.; Al-Nafjan, A. Evaluating Explainable Artificial Intelligence in EEG-Based Deception Detection. Electronics 2026, 15, 4041. https://doi.org/10.3390/electronics15174041

AMA Style

Aldayel M, Alkanhal M, Al-Nafjan A. Evaluating Explainable Artificial Intelligence in EEG-Based Deception Detection. Electronics. 2026; 15(17):4041. https://doi.org/10.3390/electronics15174041

Chicago/Turabian Style

Aldayel, Mashael, Maryam Alkanhal, and Abeer Al-Nafjan. 2026. "Evaluating Explainable Artificial Intelligence in EEG-Based Deception Detection" Electronics 15, no. 17: 4041. https://doi.org/10.3390/electronics15174041

APA Style

Aldayel, M., Alkanhal, M., & Al-Nafjan, A. (2026). Evaluating Explainable Artificial Intelligence in EEG-Based Deception Detection. Electronics, 15(17), 4041. https://doi.org/10.3390/electronics15174041

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop