Next Article in Journal
Scalable Optimization of Ultra-Dense Heterogeneous Networks Using Stochastic Geometry and Deep Learning Techniques
Previous Article in Journal
Intelligent Car Park Occupancy Monitoring System Based on Parking Slot and Vehicle Detection Using DJI Mini 3 Aerial Imagery and YOLOv11
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Systematic Review

Lost in Thought: An End-to-End Systematic Review on Imagined Speech Decoding Through Electroencephalographic Readings

by
Luis Felipe Estrella-Ibarra
,
Luis Roberto García-Noguez
,
Jesús Carlos Pedraza-Ortega
,
Juan Manuel Ramos-Arreguín
and
Saul Tovar-Arriaga
*
Faculty of Engineering, Autonomous University of Queretaro, Santiago de Queretaro 76010, Mexico
*
Author to whom correspondence should be addressed.
Submission received: 31 December 2025 / Revised: 29 January 2026 / Accepted: 2 February 2026 / Published: 13 February 2026
(This article belongs to the Section Medical & Healthcare AI)

Abstract

Many fields, including psychology, neuroscience, linguistics, computational modeling, and even philosophy, have been investigating the neuroscience of language for many years. Even so, a lack of comprehensive, interdisciplinary guidelines remains for research projects that aim to decode or model language from brain activity. Electroencephalography (EEG) is unique among neuroimaging methods in that it is a non-invasive technique. This review provides a comprehensive examination of the fundamental elements of imagined speech decoding using EEG, offering a tour of the most recent developments and perspectives in linguistic, neurological, and computational approaches over the past decade. It highlights essential findings such as the consistent involvement of sensory–motor brain regions, the strong influence of language abstraction and selection, and the superior classification performance attained with spectral and temporal features. This study was conducted and reported in accordance with the PRISMA 2020 guidelines for systematic reviews.

1. Introduction

Driven by the first neurological findings, research on language in the brain has developed into an integrated discipline molded by neuroscience, psychology, linguistics, computational modeling, and even philosophical inquiry. Anchored in the foundational work of Paul Broca and Carl Wernicke, early studies were essentially shaped by the identification of two key regions involved in linguistic processing: the inferior frontal gyrus and the posterior superior temporal gyrus—now officially known as Broca’s and Wernicke’s areas—essential regions for speech production and comprehension [1].
Although these results lay the foundation for understanding language-related brain activity, subsequent clinical and neuroimaging research has demonstrated that these two areas alone cannot adequately explain the intricacy of language processing. Modern studies reveal that speech production originates from the coordinated activity of extensive brain systems that extend beyond the boundaries of classical language regions, encompassing both cortical and subcortical areas interconnected by dynamic pathways [1].
Consequently, the conventional question of where language takes place within the brain has become too constrained; language is now understood as a distributed process in which meaning, structure, and articulation develop from the interaction of numerous brain regions that act together in flexible, task-dependent ways. Nevertheless, there is substantial room for further study, as the fundamental nature of language and its associated cognitive processes remains a subject of ongoing debate [1].
Within this broader inquiry, inner speech, also known as silent verbal thought, has attracted a growing interest [2]. Although it exhibits no observable behavior to the external eye, its importance in coordinating linguistic intention, situated at the junction of cognition and language planning under motor inhibition, has received significant focus. This form of verbal thought is particularly important for individuals who retain the cognitive ability to formulate speech but lack the physical capacity to articulate it aloud due to motor impairments [3].
With advancements in neuroimaging, its empirical study has been increasingly enabled; EEG is a common acquisition method due to its safety, non-invasiveness, and high temporal resolution. This technology is fundamental to the broader objective of building Brain–Computer Interfaces (BCIs), systems designed to decode cerebral signals and translate them into communicative output, thereby restoring a communication channel when speech is no longer attainable. Nonetheless, despite its extensive application, EEG is constrained by low signal-to-noise ratio, intrinsically low spatial resolution, and inadequate sensitivity to deep or localized brain activity [4].
Despite these challenges, new advancements present an encouraging sight. While early studies were limited to classifying simple linguistic units, such as “yes” and “no”, with a maximum of 63% accuracy [5], research since 2020 has shown that increasingly larger vocabularies can now be decoded with significantly improved accuracy, particularly by means of deep learning approaches, even generalizing to the point of decoding words not included in the original training vocabulary [6], and in some cases, extending beyond classification to approximations of actual voice generation [7,8,9].
Despite substantial developments, imagined speech decoding still has several major constraints, most of which derive from a lack of consistent multidisciplinary integration.
First, results depend on the vocabulary and participant selection design. Many studies prioritize algorithmic development and signal processing techniques while overlooking the need for consistent reporting and comparison of key factors related to linguistics, cognitive demands, and participant diversity. This oversight limits the comparability and generalizability of findings across studies, despite the fact that the very essence of speech inherently requires an integrative approach.
Although generalizable capability may differ across algorithmic techniques, classification accuracy remains subject-dependent, regardless of the decoding method applied [10]. Not only can noise arise from elements including concentration artifacts, variations in attentional engagement, age related neural variability, and health state altering cognitive and electrophysiological stability [11,12,13], but also from sociolinguistic and socioeconomic background, which have been demonstrated to influence early brain development, resting state neural activity, and cognitive resilience, with lower socioeconomic level associated with reduced frontal gamma power even in infancy, potentially increasing vulnerability to later language and attention difficulties [14,15].
Meanwhile, imagining speech in a second language can also activate translation systems, adding more layers of complexity to the neural signal as bilingual people recruit more executive attention systems and show dynamic neuroplastic adaptations, especially in subcortical areas such the basal ganglia and thalamus, reflecting the cognitive demands of language selection, inhibition, and constant conflict monitoring between linguistic systems [16,17].
Semantically or phonetically identical phrases may also elicit overlapping brain responses; cognitive load also varies depending on whether participants perceive phonemes, individual words, or complete sentences. Further challenging classical modular conceptions of language and supporting interactive neurocognitive models where affective states, attention, and prior expectations influence linguistic representations at multiple levels, from early lexical access to sentence integration.
Such observations underscore a broader shift reflected in recent evidence, which shows that emotional and contextual factors modulate the processing of semantic and syntactic information in dynamic ways [18,19].
Second, the electrode count and arrangement differ greatly among EEG equipment. While research-grade caps have 32–256+ electrodes, which leads to higher spatial resolution alongside finer brain activity, low-density consumer headsets usually cover only frontal or temporal areas [20,21]. In addition, influencing signal quality is the type of electrode; Ag/AgCl gel or saline electrodes have low impedance and clearer signals than dry electrodes, which cause more noise [22,23].
In EEG systems, sampling rates are also highly dependent on device class. While clinical and research high-end EEG amplifiers sample at greater rates, like several hundred Hz up to 1024 Hz, to capture fast brain dynamics, consumer-grade headsets usually sample in the range of merely a few hundred hertz [21].
Usually spanning 0.1 to 70 Hz, most systems remove both slow drifts and high-frequency noise using a band-pass filter, therefore guaranteeing signal purity.
Generally tuned dependent on the regional power supply at 50 Hz or 60 Hz, a notch filter is employed to eliminate interference from the electrical mains [24]. The exact cutoff parameters are often selected to strike a balance between noise reduction and signal preservation, and they may vary depending on the EEG hardware and the experimental protocol. Moreover, as optimal timing varies between phonemes, lexical tones, and internal cognitive activity, the positioning of EEG analysis windows can significantly influence decoding accuracy [25].
At last, decoding techniques range from handcrafted feature pipelines paired with classical classifiers, such as support vector machines and linear discriminant analysis, to modern deep learning models [26], and recent generative frameworks aim to reconstruct intelligible speech directly from imagined speech EEG [7,8,9].
Despite these advances, the field still has numerous unresolved problems and notable performance disparities.
Seven main components comprise this review, each offering insight into a key aspect of imagined speech decoding:
  • Introduction;
  • Search strategy and study selection;
  • Participants and vocabulary design;
  • EEG Acquisition;
  • Feature processing and model approach;
  • General discussion;
  • Conclusion.
Unlike reviews that primarily focus on isolated technical aspects, this work aims to integrate methodological, neurophysiological, and linguistic perspectives.
Structuring the conversation among participants, acquisition techniques, feature types, and modeling approaches helps the review not only outline the technical foundations but also encourages reflection on how each element influences general decoding performance. This holistic strategy aims to highlight the interactions among linguistic, neurological, and computational aspects, thereby providing a broader platform for future investigations and system development.

2. Search Strategy and Study Selection

This review synthesizes and analyzes the linguistic traits, neurological underpinnings, and methodological strategies, including classification algorithms and feature extraction methods, applied in EEG-based imagined speech decoding in a systematic and organized manner. Beyond merely charting the current state of the field, this review aims to identify trends across studies and investigate whether variables, such as the type of linguistic unit, electrode configurations, frequency bands, or modeling strategies, are consistently associated with improved classification accuracy. This approach aims to identify potential elements causing performance fluctuations and direct subsequent improvements in the development of imagined speech decoding systems.
Focusing on papers from 2016 to 2025, the literature search was conducted across five main bibliographic databases: Scopus, Web of Science, IEEE Xplore, PubMed, and Google Scholar. Considered qualified were only conference proceedings and peer-reviewed journal entries.
Formulated as follows, the search query was meant to catch a broad spectrum of studies on imagined speech using EEG:
(Imagined speech AND (EEG OR electroencephalography)).
OR (“Covert speech” AND (EEG OR electroencephalography)).
The search turned back 3451 papers overall. Specifically, 135 items were obtained from IEEE Xplore, 110 from PubMed, 278 from Scopus, 173 from Web of Science, and 2595 from Google Scholar. Following the application of inclusion and exclusion rules and the elimination of duplicates, a total of 45 studies were retained for thorough investigation. Studies assessed in full text but excluded were removed mainly because they were not EEG-based imagined/covert speech decoding studies, did not include an imagined/covert speech condition, or did not report sufficient methodological details for comparison.
The PRISMA methodology was employed to systematically and methodically document the selection process. The Preferred Reporting Items for Systematic Reviews and Meta-Analyses, or PRISMA for short, provides a methodical approach to help identify, screen, assess, and ultimately include records in a review. No formal sensitivity analyses were performed, as the objective of the synthesis was descriptive and comparative rather than inferential. Figure 1 shows a PRISMA flow diagram capturing this procedure. This systematic review was retrospectively registered in the Open Science Framework (OSF) and is available at https://osf.io/kvhdc/overview?view_only=991f5c9739154d44a8f8dafbfd9c9dbe (accessed on 1 February 2026), and the PRISMA checklist used to guide the review process is provided as Supplementary Material [27].
From every kept paper, the following fields were extracted: (1) Participants, (2) Vocabulary size, (3) Repetitions per subject, (4) Language & word list, (5) Best accuracy, (6) EEG device & base-array size, (7) Exact electrode subset (or all electrodes), (8) Sampling rate, (9) Filtering & artefact removal, (10) Feature-extraction technique, (11) Feature type fed to the model, (12) Classifier.
A formal study-level risk of bias assessment was not performed. The included literature is predominantly methodological and engineering-oriented, with substantial heterogeneity in experimental designs, datasets, preprocessing pipelines, feature definitions, model architectures, and validation strategies, which limits the applicability of standardized clinical risk-of-bias tools. In addition, the assessment of reporting bias due to missing or selectively reported results was not feasible, as performance metrics, negative findings, and evaluation protocols were inconsistently reported across studies. For these reasons, no formal assessment of the overall certainty or confidence of the body of evidence was conducted. Consequently, the findings of this review should be interpreted as qualitative trends and design considerations rather than definitive quantitative estimates of effect.

3. Participant and Language Factors

One begins to understand the decoding of imagined speech by considering the individuals behind the signals. Before delving into the development of brain signal features and classification models, it is essential to understand that every brain signal reflects a highly personal cognitive and linguistic process shaped by the participant’s language proficiency, cognitive strategies, and neurological individuality. These inner signals are the consequence of deliberate, context-dependent activity, an attempt to communicate without articulation, not just random electrical fluctuations. Decoding imagined speech is thus not only a technical chore but also a process based on the diversity, complexity, and adaptability of human cognition.
Risk of bias within studies was evaluated qualitatively across the included literature. Given the methodological and engineering-oriented nature of imagined speech EEG research, formal clinical risk-of-bias tools were not applicable. Instead, potential sources of bias were identified based on study design characteristics, including small sample sizes, subject-dependent evaluation strategies, limited vocabulary sizes, and heterogeneous preprocessing and validation protocols. These factors were considered when interpreting performance differences across studies rather than used as exclusion criteria.

3.1. Participant-Related Factors

Commonly involving 4 to 27 participants, EEG-based imagined speech research raises legitimate questions regarding statistical robustness and the generalizability of results when small cohorts are used. Although some imagined speech stimuli may elicit similar brain activations across individuals, increasing the number of participants does not always result in better generalization across cognitive pathways, leading to improved accuracy. Tailored, subject-dependent models help to produce better outcomes. For example, Vorontsova et al. [28] developed a model using data from 256 individuals. They reported that subject-independent testing yielded an accuracy of just 13.39%, whereas subject-dependent training and testing on a single person achieved an accuracy of 84.51%. Similarly, Liyanagoonawardena et al. [10] confirmed that in all tested configurations, subject-dependent models consistently outperformed subject-independent methods. These results suggest a potential shift in the direction of research: should the goal be to develop dynamic, flexible systems that acquire robustness through subject-specific learning, or should efforts continue toward building models that generalize across individuals? Most studies consider participants as a relatively uniform population, usually reporting individual-level classification accuracy without addressing the broader spectrum of participant-related variability. Interestingly, although none of the research explicitly supports their selection of participant age range, the average age across studies is often about 25. This is a pertinent choice since the 25–35 age range signifies a moment of total cognitive development before the beginning of age-related physiologic decline, usually starting in the 30 s, which has been associated with gradual changes in neural efficiency, speech motor control, and cognitive-linguistic processing that may introduce increased variability in speech production and its neural representations, potentially affecting decoding performance [29,30]. This demographic consistency supports the idea that, if generalization is already challenging in a subject-independent model used on a relatively homogeneous group, it becomes even more problematic when larger, more diverse populations are included. Age, socioeconomic background, cognitive diversity, and language experience pose significant challenges that greatly complicate population-level generalization.
The participant’s native language and its correspondence with the language used in the stimuli while obtaining brain signals of imagined speech is one underreported feature in imagined speech research. Although English is the most often utilized language in these studies, participants’ linguistic backgrounds are generally kept vague. Fortunately, some studies examine bilingualism or specifically note whether subjects speak the language used natively. For example, Advait Balaji et al. [31] combined English and Hindi vocabularies in a four-word imagined speech task, achieving 75.38% accuracy with eight participants, all of whom had high proficiency in English. Also, Larocco et al. [32] worked with 44 English phonemes using 16 participants, each repeating the set 15 times. Though none of the subjects spoke English natively, the model achieved an accuracy of 98%. The phoneme-level character of the work, which emphasizes sound articulation over syntactic or semantic understanding, may help explain this outstanding performance. In some situations, the impact of language proficiency may be mitigated, allowing non-native speakers to function adequately even with limited fluency. This does not, however, diminish the value of proficiency as a variable; instead, it suggests that its impact can depend on the linguistic level targeted by the task. Thus, in imagined speech research, both language proficiency and the type of linguistic units involved should be considered as interacting elements.
Direct cross-linguistic comparisons remain rare, even though few studies have investigated non-English languages, such as Mandarin Chinese [25,33], Spanish [10,34,35,36,37,38,39,40,41,42,43,44,45], Dutch [46], Bengali [47], and Russian [28]. Given that brain representations, tonal characteristics, and phonological systems differ greatly between languages, this gap is crucial. For Mandarin lexical tones, for example, tracking pitch variations with semantic significance presents special difficulties for BCI systems [33]. Although this is reasonable, given that the field is still mostly exploratory and focused on proof-of-concept implementations rather than fully practical, language-independent systems, what is more alarming is the general absence of reporting on participant-related features that could significantly affect decoding performance.
Beyond the regularly stated characteristic of age, there is a clear duty to extend the dialogue to include additional elements such as socioeconomic level, educational background, health issues, and larger demographic characteristics. Aspects that define how language is learned, absorbed, and applied during the lifetime, hence producing what can be understood as subject-specific evolutionary language paths. Ignoring these not only reduces the inclusiveness of decoding systems but also runs the danger of aggravating already existing technological inequities.
This highlights the previously mentioned issue once more: is the difficulty of imagined speech decoding a matter of reaching generalization, or does it instead require the creation of flexible, subject-aware systems that reflect the diversity and complexity of human language experience?

3.2. Language-Related Factors

No less crucial is the issue of which speech units are intended for decoding. Currently, the most commonly used units in research are phonemes and words, and the decision of which to use significantly impacts model accuracy. Due to their higher level of abstraction, words provide challenges. Unlike phonemes, which are essentially auditory and articulatory, as briefly discussed, words have semantic meaning; therefore, they engage more general and individualized brain networks [48]. Although it’s intuitively known, even without experimental validation, that this semantic activation is not homogeneous among individuals, it’s crucial to keep in mind that people’s experiences, cultural background, and personal associations define how they internalize meaning and approach language. The brain creates organized networks of associations throughout time that allow relational knowledge and effective communication akin to internal graphs [49,50]. This resembles how natural language processing algorithms operate, linking words and concepts through contextual relationships.
Therefore, even if that were not the intended objective, picturing a semantically loaded statement like “My grandmother knit me a…,” would naturally lead many people to associate the word “sweater.” On the other hand, a more generic sentence like “He went to the mall to buy…,” can set off a broader spectrum of possible completions. This kind of anticipatory semantic activation can cause what would be regarded as semantic noise, or the inadvertent activation of related concepts, therefore interfering with the exact decoding of imagined words. Although these effects should ideally be minimized in final BCI applications, they should nonetheless be taken into account in experimental studies aimed at validating the precise capture of language intention from brain signals.
This phenomenon is related to the well-known psycholinguistic theory known as the priming effect, which posits that, in response to a specific stimulus—in this case, language—the brain tends to activate particular actions or responses. Studies have demonstrated that repeated exposure to specific words can alter later brain activity and even influence behavior, thereby underscoring the strong entrenchment of language in cognitive functioning [51,52,53,54].
From the acoustic standpoint, both word-level and phoneme-level decoding face difficulties related to phonetic similarity. Few studies specifically address this; however, Varshney et al. [55] chose a group of six terms while purposefully evaluating their phonetic distribution to reduce confusion. Such design decisions highlight the need for careful stimulus selection to isolate the impact of phonetic overlap on decoding performance.
Apart from semantics and phonetics, another understudied aspect of imagined speech processing is the concreteness of a word. Whereas abstract words like “freedom” depend more on language and contextual interpretation, concrete words like “water” or “toilet” could generate stronger and more consistent semantic connotations. Cognitive science has extensively studied the concreteness effect, finding that concrete words often elicit faster and more consistent brain responses [56]. Nevertheless, EEG-based BCI research rarely annotates or controls for word concreteness, hence lacking knowledge of how it could influence decoding accuracy.
Table 1 compares all datasets used for imagined speech decoding, including participant count, language, vocabulary size, and trials per subject, thereby illustrating this language choice. The overview of each dataset displays the average accuracy achieved across studies using it, as well as the highest and lowest accuracy values reported among those studies. Given that wide gaps between the highest and lowest accuracy numbers could indicate notable advances or deviations in certain elements of the decoding technique, this synthesis enables a more accurate estimation of overall performance.
Based on the overview of the datasets utilized over the past ten years, it is clear that there is a strong inclination toward vocabulary consisting of directional directives, such as “up”, “down”, “left”, and “right”, in both English and Spanish. Many of these studies appear to have chosen such vocabularies with the intention of enabling basic interaction and communication, not yet through full language generation or reconstruction, but rather through identifiable commands that allow users to trigger or express intentions in a simplified and structured manner. This suggests an early phase of imagined speech BCI development, where communication is approached through functional command sets rather than complete linguistic expression.
Other lines of research, such as the one carried out by Seo-Hyun Lee et al. [75], align more directly with the aims of speech restoration. Their dataset consists of words deliberately chosen from communication boards routinely used in hospitals all around for patients with paralysis or aphasia. This neurological disorder usually results from brain damage or stroke and compromises language production or comprehension. Their dataset also includes general-purpose words pertinent to daily life, as well as frequently used expressions in BCI systems. This captures the fundamental needs of people who retain the cognitive drive to communicate but lack the muscular control necessary for effective communication.
This leads back to a fundamental question that has already been asked theoretically: how does the linguistic unit itself affect decoding performance? Studies were arranged according to vocabulary size ranges to investigate this. After that, the average classification accuracy was examined across these groups, and performance between word-level and phoneme-level tasks was compared. Figure 2 shows that for most language ranges, phoneme-based classifications consistently produced better average accuracies.
Nevertheless, the unequal representation of linguistic unit types across vocabulary size categories may lead one to suspect potential comparability issues in this analysis. To explore this possibility, a more focused examination was conducted by restricting the analysis to studies employing exactly five linguistic units, allowing for a clearer assessment of whether such discrepancies influence the results. In this instance, a fairer comparison was possible, as the number of studies utilizing word-level and phoneme-level units was equal. As shown in Figure 3, the corresponding boxplot indicates that phoneme-based classification once again exhibits superior performance. No outliers were observed in either group, supporting the hypothesis that this performance difference is consistent.
Conversely, when considering vocabulary abstraction in relation to word-level linguistic units, it is clear, as already indicated, that the predominant trend across research favors concrete directional commands, which leave little room for semantic ambiguity. However, some research adopts more interpretive or abstract words, such as “cooperate” or “independent”, which might vary in semantic activation more.
All studies using word-level units were divided into two groups: one comprised concrete directional commands, and the other mixed-concept vocabularies, which included a wider range of terms that were not strictly action-oriented and may involve higher degrees of abstraction, to further investigate this dimension. This clustering enabled a comparison of how conceptual clarity affects categorization performance.
Figure 4 shows that, although the performance difference is not significant, the average accuracy for concrete-directed vocabularies is consistently better than for mixed concepts. Although the margin seems small, it represents a tendency consistent with theoretical expectation: words with more exact, limited meanings may generate more consistent brain patterns, hence improving the decoding accuracy.
This is not, of course, a rigorous assessment of lexical concreteness in the strict psycholinguistic sense. The comparison was not based on controlled contrasts between abstract and concrete sets of words. Consequently, it cannot be asserted that vocabulary abstraction alone determines performance. Still, both the conceptual justification and other studies confirm this theory [56], suggesting that the abstraction levels in word stimuli have a significant influence on decoding difficulties, particularly in semantically complex or subjective material.

4. Signal Acquisition and Brain Dynamics in Imagined Speech

After participant-related variability and language features have been addressed, attention must shift to the brain activity and acquisition mechanisms through which imagined speech is recorded. The fidelity of the recorded signal depends not only on the underlying brain activity evoked by the linguistic stimuli but also on how that activity is registered, more especially, in which brain region, within what time frame, and through what recording configuration. As a non-invasive technology, EEG remains the most readily available instrument for deciphering inner speech; nonetheless, it relies on well-designed methodological choices to ensure the best data quality.

4.1. Neurological Basis of Inner Speech Processing

By first gaining an accurate understanding of what brain activity is and how it is recorded using techniques such as EEG, one can lay the foundations for imagined speech decoding within a neuroscientific framework. The brain activity refers to the electrical synaptic potential produced by communicating neurons. Small electrical impulses are generated as neurons respond to stimuli or internal processes, such as language. When coordinated across neuronal populations, they result in measurable voltage fluctuations. EEG, or electroencephalography, is a non-invasive technique that captures these fluctuations through electrodes placed on the scalp [84].
Although EEG offers excellent temporal resolution for capturing rapid fluctuations in brain activity, it has relatively low spatial resolution. This is because the electrodes placed on the scalp only record aggregate activity over large portions of the brain, without the capacity to localize fine-grained neuronal activations. In addition, the layered structure of the brain and the resistance of the skull and scalp dilute the original millivolt-level signals, therefore producing diffused microvolt signals by the time they reach the electrodes. These constraints make it more challenging to correlate specific language processes to localized brain activity—a necessary first step toward raising the accuracy of imagined speech decoding [85,86,87].
More intrusive methods, such as electrocorticography (ECoG), on the other hand, position electrodes exactly on the cortical surface, thereby enabling better signal clarity and spatial resolution. Still, these techniques require surgical intervention [88]. EEG’s non-invasiveness, safety, and ease of access help explain why, despite its limitations, it is still extensively utilized.
Underneath this stimulus reception now is the fact that brain activity reflects reinforced paths and learnt associations. Experience and repetition help specific brain circuits grow stronger, thereby increasing the likelihood that particular clusters of neurons will activate in response to familiar ideas or stimuli [89,90]. This is the basis for cognitive abilities, including language. When one thinks of or pictures a word, particular brain pathways are triggered in predictable patterns. It is this repeatability that imagined speech decoding seeks to capture, and it is also where previously mentioned intersubject variability becomes relevant.
Language processing has long been associated with specific cortical areas, most notably Wernicke’s area and Broca’s area. While Wernicke’s region, in the posterior superior temporal gyrus, is linked with language understanding, Broca’s area, in the inferior frontal gyrus, is connected with language production [91,92]. These associations date back to 19th-century clinical cases, beginning with the work of Paul Broca, who linked the observed loss of language production to a lesion in a specific brain region of a deceased patient with severe speech impairment. Although based on a single example, this finding laid the foundation for the contemporary neurological view of language, a perspective that would subsequently be developed by Carl Wernicke and others [93].
It is now clearly known, nonetheless, that language processing transcends the domains of Broca’s and Wernicke’s alone. Word retrieval, semantic processing, and phonological planning interact extensively over several cortical lobes [92]. Research shows activation in several brain areas—not only across different studies but also within the same study using identical stimuli—highlighting that variability results from both individual differences in neural responses and vocabulary.
This variability raises the crucial issue of what should be done when observed brain activity reflects associations with specific words, and those exact words may elicit different patterns of activation across individuals, or even within the same individual, outside a controlled environment, rather than reflecting generalizable mechanisms of language production. This is a very pertinent issue, taking into consideration how to progress toward practical solutions using a small number of strategically placed electrodes and lower the demand for highly dense electrode arrays.
Should decoding accuracy depend on such localized, vocabulary-driven activations, the model may end up learning to identify particular words in a specific type of individual rather than decoding language as a more general cognitive process. Although this method might work well for command-based BCIs, it is not sufficient for systems meant to support fully developed language processing. For those developing models meant to generalize across vocabularies and contexts, it is essential to avoid conflating universal representations of language with word-specific brain patterns observed in particular and controlled populations.

4.2. EEG Acquisition Systems and Signal Processing Parameters

Given this variability, the selection of EEG electrodes becomes critical. Different electrode layouts and signal processing methods may significantly affect the recorded brain activity, as well as its interpretation [94]. The two main hardware-related factors directly influencing data quality are the number of electrodes and the sampling rate, both of which are inherent features of the EEG system used.
As previously mentioned, EEG has inadequate spatial resolution compared to more invasive techniques. However, fewer electrodes reduce this spatial accuracy even further, as each electrode records only the electrical activity of the area directly under it. While low-density designs can overlook significant spatial information, higher-density electrode systems can provide more comprehensive coverage of cortical activity over the scalp.
Conversely, the sampling rate refers to the frequency at which the device records voltage readings over time for each electrode. Sampling rate controls the finely tuned tracking of the electrical variations in the brain that EEG detects. The ability to capture faster variations in brain activity is made possible by a higher sampling rate, hence enhancing temporal resolution. On the other hand, reduced sample rates run the risk of missing fast fluctuations, thereby producing an inadequate view of the signal’s dynamics.
Usually including between 32 and more than 256 electrodes, research-grade EEG devices have sampling rates ranging from 500 Hz to over 2000 Hz. By contrast, consumer-grade systems typically operate at much lower rates, often between 128 and 512 Hz, and are equipped with only 14 to 32 electrodes. Given the enormous amount of data generated at higher frequencies, studies reduce later computational expenses by applying down-sampling, typically decreasing the signal to around 250–512 Hz. This, however, compromises temporal precision, thereby influencing the accuracy with which fast-changing brain dynamics are recorded. These limitations in both spatial and temporal resolution must be taken into account when interpreting results or attempting to generalize findings from such systems [20,21].
To clean and isolate meaningful brain signals, it is necessary to apply a variety of filtering and preprocessing techniques aimed not only at reducing noise, but also at mitigating the effects of sensor noise, channel interference, environmental disturbances, and other physiological and non-physiological artifacts, thereby enhancing overall signal quality. One type of filter often used to accomplish this is the notch filter, commonly used to filter out external interference coming from power lines, which, depending on the local power grid, typically appears as a constant signal at 50 or 60 Hz [95]. In some instances, further notch filters are utilized at harmonic frequencies—multiples of the base frequency—such as 120 Hz or 180 Hz, which can further contaminate the recording.
Another commonly used filter is the bandpass filter, which rejects both low-frequency drifts and high-frequency artifacts, thereby isolating specific portions of the frequency spectrum. This is particularly important since distinct cognitive and physiological processes are connected to different frequency intervals, as will be discussed later in the conversation. By focusing on the target frequency range, it is then possible to investigate in more detail the elements of brain activity that are most likely to represent the cognitive process underlying the task of interest, which in this case is language. Butterworth filter designs are commonly used for these filters due to their smooth and consistent frequency response. Unlike filters that introduce abrupt distortions near cutoff frequencies, this approach allows for smooth transitions between retained and attenuated frequencies, thereby enabling signal preservation. This type of filtering can be realized in the form of Infinite Impulse Response (IIR) filters, whose recursive-feedback operation allows for computationally efficient filtering with fewer operations, allowing for real-time processing but at the expense of potential phase distortion, which may affect the temporal structure and alignment of EEG signals relevant for imagined speech decoding, as well as lower numerical stability. Alternatively, Finite Impulse Response (FIR) filters which operate without feedback, providing a linear phase response and improved numerical stability.
In addition to bandpass filters, there are also high-pass and low-pass filters, which allow for more specific targeting of the undesirable components of the signal. Low-pass filters eliminate fast, transient noise by retaining only the low-frequency components, and high-pass filters suppress slow shifts in the signal level by allowing only higher frequencies to pass through [84,94,95].
Beyond filtering, advanced signal separation methods are often applied. Independent Component Analysis (ICA) is a popular mathematical tool for disentangling the various contributing sources of recorded EEG signals by breaking down the time series into components that are statistically independent. This allows for the removal of common sources of contamination, such as eye blinks, muscle activity, or cardiac signals. Another often-used method is the Common Average Reference (CAR), in which the signal at each electrode is re-referenced by subtracting the average signal computed across all electrodes. This technique enhances the spatial resolution of EEG data and helps reduce widespread noise that affects all channels similarly, thereby mitigating some of the limitations inherent to EEG spatial resolution [95,96].
As briefly mentioned, frequency ranges are associated with specific cognitive and behavioral states. Delta waves, which occur at frequencies below 4 Hz, are known to be associated with deep sleep and non-conscious neural processes. Theta waves, ranging from 4 to 7 Hz, are typically detected during drowsiness, a state in which memory encoding and emotional processing are involved. Alpha activity, spanning 8 to 12 Hz, tends to dominate during relaxed wakefulness and introspective thought, and it is typically attenuated during tasks that require focused attention. Beta waves, which range from 13 to 30 Hz, are associated with active cognitive processes, such as problem-solving, decision-making, and motor planning. Lastly, gamma waves, exceeding 30 Hz, are thought to be involved in higher-order cognitive functions such as perception, language comprehension, and conscious awareness [97].
Each study must tailor its filtering and preprocessing pipeline to maximize signal quality. To summarize, Table 2 provides an overview of the signal acquisition and preprocessing settings used in the various datasets described above. This includes the number of electrodes and the sampling rate used in each study, as well as the filtering techniques and frequency ranges associated with the highest reported accuracies. It also considers whether utilizing a subset of electrodes outperforms the whole array in imagined speech decoding.
Understanding the brain mechanisms underlying both language reception and intention is essential. As previously stated, the well-known language-related areas—Broca’s and Wernicke’s regions—are important but not enough on their own to fully explain the intricacy of language processing [92].
Language represents a symbiotic relationship between semantics, phonetics, and motor capabilities. It is thus time to consider the theory that motor planning could be a natural and inseparable aspect of language itself. For instance, imagining a word like “dog” or “please” that calls for particular tongue or mouth gestures could at first seem simple. Still, the motor activation needed to say these words remains in the cognitive processes, even when articulation is inhibited [98]. This becomes evident when contrasting the act of hearing a word with the process of imagining or mentally articulating it. While auditory reception is relatively consistent among people, the motor intention associated with internal language representation varies significantly.
This variability helps to explain why speaking in a non-native language can be challenging: the necessary motor patterns vary and need to be acquired. Studies have amply demonstrated that the motor cortex plays a major role in imagined speech, suggesting that motor engagement continues even in the absence of movement [99,100,101]. However, limiting knowledge of language perception and intention to individuals who can speak and communicate orally runs the risk of ignoring the broader spectrum of how language behaves across various populations, an awareness that may be essential to understanding how communication can be generalized beyond speech. Language intent and its motor synergy remain true, regardless of the population or communication channel. Whether expressed aloud or through sign language, paying attention to intent is essential for understanding how language emerges from and is rooted in lived human experience.
For instance, research has demonstrated that in deaf individuals who communicate using sign language, classical language-related areas, such as Broca’s region, are activated not only during overt signing but also during the observation and planning of signs, confirming its role in motor planning during the linguistic process [102].
It is not illogical, therefore, to speculate that if a deaf person is thinking about a word in sign language, their brain may recruit the manual motor regions in the same way as their oral motor regions are coaxed into action during inner speech in a hearing person. This suggests that motor planning is a fundamental part of language processing, independent of the effector used to produce it. Furthermore, just as imagined speech is known to involve the transient activation of motor processes that are suppressed to prevent articulation, imagined signing may similarly involve inhibitory modulation of covert motor execution.
Driven by the individual’s primary mode of communication, the underlying motor–language synergy likely persists even in the absence of overt movement.
This synthesis identified that multiple studies, regardless of the total number of electrodes available in their EEG systems, achieved better performance by selecting a specific and limited subset of channels. Some of these were chosen intentionally during the experimental design as a methodological decision, while others were obtained as a result of the post hoc determination of the electrodes most correlated with imagined speech.
The cortical distribution of selected electrodes was examined based on the best-performing studies in each dataset that employed electrode selection. This analysis enabled the grouping of studies according to the brain region most predominantly targeted by the selected electrodes. Classification accuracies for linguistic units—phonemes and words—were then averaged across the studies within each group. These results are presented in Table 3, where each row corresponds to one of these study groups, categorized by the predominant cortical region covered. The table reports the proportion of electrodes within each selection that were located in the dominant region, not the percentage of the cortical area itself, but rather the relative number of selected electrodes situated within that region, and the mean classification accuracy reported across the studies. Among these, the group with a predominance of sensorimotor cortex coverage exhibited the highest average performance. This average performance trend can also be visually appreciated in Figure 5, which maps the mean classification accuracy to each EEG electrode based on its inclusion across studies using selective electrode configurations.
Particular attention is drawn to the contrast between the KARA ONE [65] and Coretto [73] datasets. Both adopted distinct approaches regarding the involvement of the motor cortex. The Coretto dataset used six electrodes deliberately positioned as close as possible to language-related cortical areas, while maintaining sufficient distance from speech-related muscles to minimize myoelectric noise. This configuration resulted in a balanced cortical distribution, with 33.3% of the selected electrodes covering the frontal, central, and parietal lobes, respectively. In contrast, KARA ONE recordings were acquired using a full 64-electrode cap, from which ten electrodes, frequently referenced in studies using this dataset, were identified as most correlated with imagined speech. Of these, 70% were located over central and sensorimotor regions.
This spatial disparity was also reflected in performance outcomes: on average, Coretto yielded high accuracy in multiclass classification (95.39%) but performed poorly in phoneme (39.16%) and word (30.1%) classification. Conversely, KARA ONE achieved substantially higher accuracy in phoneme (86.73%) and word (91.12%) classification, yet showed weaker performance in the multiclass setting (40.9%).
These results demonstrate a more complex interplay among coverage of cortical regions, unit type, and classification performance in imagined speech decoding. Although the Coretto dataset achieved the best average accuracy in multi-class classification (95.39%), it struggled in word-level decoding, even when considering only the best-performing study, which reached only 34% for word classification. This possibly reflects the difficulty in decoding semantically richer units with a limited and spatially distributed plurality of electrodes, which may be insufficient to capture the entire scope of cortical activity enlisted for processing this specific vocabulary. For phonemes, however, the best-performing study still achieved a strong 81.69%, indicating a better match to simpler, lower-level structural aspects of the language.
On the other hand, the best studies that applied KARA ONE in this literature regularly presented high classification accuracy results for all linguistic units: 86.73% (phonemes), 92.68% (words), and 77.37% (multiclass tasks). These results were achieved despite the reduced spatial coverage, which focused predominantly on sensorimotor regions. The inclusion of more complex phonemes and words with overlapping phonetic structures in the KARA ONE vocabulary may have favored a decoding strategy grounded in articulatory planning signals, reinforcing the utility of motor-related cortical activation for imagined speech decoding.
This contrast highlights the trade-off between semantic flexibility and motor stability. A broader use of the cortex, as in the case of Coretto, may allow for the capture of semantic variation across linguistic units. Nonetheless, it can also introduce content-specific variability and may require broader cortical coverage to reliably access semantic activations, which are unlikely to be consistently localized or predictable across participants or even words. In contrast, sensorimotor-based studies, such as those using KARA ONE, can exploit articulatory regularities that generalize more effectively across speakers and linguistic units. Yet, this comes with the inherent challenges of increased vulnerability to non-linguistic motor planning noise and the limited spatial resolution of EEG, which can complicate the isolation of relevant signals.
Ultimately, these results thus indicate that no single cortical mechanism is optimal across all decoding tasks. Far from favoring one regional strategy over the other, it may be most fruitful to consider how the different types of information are dynamically combined. In EEG, where spatial resolution is inherently limited and susceptibility to non-speech-related motor noise is elevated, this tradeoff may call for complementary motor planning features with linguistic context modeling, much like how natural language models predict speech based on prior intent. For invasive systems, where finer-grained spatial data is available, motor signals might serve as a more reliable and generalizable anchor. In either case, decoding imagined speech requires understanding not just what is said, but how language itself is lived through planning, meaning, and embodiment. Clearly, the field remains open to further validation and exploration of these possibilities.
On the other hand, returning to the frequency-based framework of EEG and its association with cognitive activity, it is essential to consider how the so-called frequency bands—specific ranges of brainwave frequencies—relate to language processing. While high-frequency bands, particularly the gamma band [100], have been repeatedly implicated in language tasks, this review includes few reports that contain band-specific analysis. More generally, ranges are treated as broad windows of possible frequencies, which may include or exclude some bands without attending to their individual contributions.
To counteract this bias, performance results were pooled and averaged across studies based on whether they incorporated any specific frequency bands, regardless of the inclusion of other frequency content. As illustrated in Figure 6, the lower the frequency bands, the lower the average accuracy. These results should not be taken as conclusive, given the large number of confounding factors, such as vocabulary, model architecture, preprocessing, and, more importantly, the inconsistent reporting of precision scores tied to specific frequency bands, which prevents a clear assessment of their individual contributions. At least for the current data, higher frequency ranges, particularly those above theta, are more commonly observed in studies with higher decoding performance, which still provides a potential cue to target these bands in imagined speech decoding tasks.

5. Feature Extraction Pipelines and Modelling Strategies in Imagined Speech EEG

As already discussed, language and brain implications provide the necessary context for understanding inner speech representations. However, decoding imagined speech is ultimately a pattern recognition problem: raw scalp potentials must be transformed into informative features that a classifier or regressor can utilize. It is therefore imperative to take a closer look at the nature of the input data and the extraction mechanisms that define the starting point for any analysis.

5.1. Extraction Methods

EEG feature extraction draws on four complementary representations that together capture when, how, and where neural activity unfolds [103,104]. In the time domain, the signal is sliced into short windows and treated purely as a voltage-versus-time waveform. Simple statistics or autoregressive coefficients are computed almost instantaneously, which keeps processing lightweight for real-time or low-power pipelines. However, these descriptors can miss oscillatory details when deeper patterns depend on frequency content. Converting the same waveform to the frequency domain exposes the power carried by canonical rhythms—delta through gamma—via transforms such as FFT or band-power estimates. These measures often deliver high decoding accuracy because rhythmic signatures are highly discriminative; however, repeated spectral calculations inflate computation, and fixed windows struggle with the non-stationary shifts common in EEG [104]. A richer view emerges from time-frequency representations, such as the short-time Fourier transform or wavelets, which map how band-specific power evolves moment by moment. This joint description consistently enhances performance by capturing transient bursts and cross-frequency interactions, although it must strike a balance between temporal and spectral resolution and typically demands more resources. Finally, spatial filtering leverages the multichannel layout of the scalp. Methods like common spatial patterns and their variants learn patterns of covariance that isolate task-relevant sources, yielding some of the strongest discrimination across applications, but at the cost of greater algorithmic complexity and data requirements [103,104].
Across the reviewed studies, the four categories of features—Temporal, Spectral, Spatial, and Temporal–Spectral—have been extracted using various methods and demonstrate different levels of classification performance. The choice of feature type influences not only the model’s accuracy but also its interpretability and computational requirements. Therefore, understanding their theoretical basis and extraction mechanisms is essential when comparing studies or designing new experiments.

5.2. Modeling Approaches

After features are extracted from EEG signals, a critical step in imagined speech decoding is to implement machine learning and deep learning models for identifying the neural patterns corresponding to linguistic intent. These strategies differ not only in architecture but also in the level of abstraction and type of patterns they are designed to capture.
Traditional machine learning (ML) models such as Random Forest, Support Vector Machines (SVM), K-Nearest Neighbors (KNN), Linear Discriminant Analysis (LDA), and Extreme Learning Machines, operate by learning explicit decision boundaries in a predefined feature space [5,32,38,55,57,64,80]. These models rely on handcrafted features such as time-domain statistics, spectral band-power measures, or time–frequency descriptors manually designed based on domain knowledge, and are thus very interpretable and computationally efficient, particularly when one deals with a small to moderate number of features. For example, SVMs excel at finding hyperplanes that discriminate between multivariate patterns in high-dimensional input spaces. At the same time, Random Forests effectively handle non-linear relationships and enhance generalization by combining decision trees. Although such approaches may require preprocessing and the appropriate selection of features, their structure provides a transparent view of which variables have the most significant strength towards classification.
Deep learning (DL) models, on the other hand, automatically learn abstract representations directly from the data, reducing the need for manual feature engineering. Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Long Short-Term Memory (LSTM) networks, and Transformer-based architectures have been applied to imagined speech tasks with increasing success [28,39,42,43,45,47,67,77,78,79,81,83,105]. These models can capture both spatial and temporal dependencies in EEG signals, making them suitable for decoding complex brain dynamics over time. CNNs, in particular, are well-suited for learning spatial patterns in EEG time-frequency images [62], while LSTM and GRU layers are commonly used to capture temporal evolution and dependencies. More advanced combinations, such as ResNet18 with GRUs, Multi-Receptive Field CNNs, or 3D-CNN-LSTM hybrids, allow the integration of multiscale patterns and deep temporal encoding [28,81]. Furthermore, approaches like Denoising Diffusion Probabilistic Models (DDPMs) or models involving Autoencoders and Conditional Latent representations have recently emerged as promising alternatives capable of modeling the high variability in neural data [77].
The performance of these models varies depending on all the factors previously discussed, from language characteristics to the type of extracted features. Table 4 summarizes the most effective combinations of feature types and models reported across the reviewed datasets, highlighting the configurations that yielded the highest decoding accuracy in each case.
To explore potential trends in decoding performance, the reviewed papers were divided into two main branches: those related to Machine Learning (ML) techniques and those utilizing Deep Learning (DL) approaches. Of these, 56% were using DL models and 44% were using ML methods. Despite this relatively balanced distribution, DL methods demonstrated a slight but consistent superiority, achieving a mean classification accuracy of 66.16%, compared with 60.78% for ML. One potential advantage of DL as a classifier in this context may be its ability to better model complex, high-dimensional patterns in neural data.
To better understand how performance varies, a second level of stratification was introduced. Within each algorithmic category, studies were further grouped based on the type of features used for classification: temporal, spectral, spatial, or a combination of temporal and spectral features. This stratification is illustrated in Figure 7, which compares the distribution of classification accuracies for each feature type within ML and DL approaches.
As shown in the figure, Temporal–Spectral features, understood as features that jointly encode temporal and spectral information either through time–frequency representations or through combined temporal and spectral descriptors, consistently provided the highest classification accuracies across both modeling paradigms, within a narrow range. In contrast, using temporal or spectral features in isolation generally resulted in lower and more variable performance, especially for ML models. Spatial features showed the lowest performance in ML approaches, while in DL, temporal-only features exhibited both reduced accuracy and increased dispersion.
These results indicate that though model architecture contributes to decoding imagined speech, feature representation design and dimensionality can have a greater influence on classification success. One reasonable theory is that imagined speech depends on frequency-specific cognitive processes captured by spectral aspects as well as time-sensitive brain activations reflected in temporal patterns. The integration of both dimensions likely provides a more complete and discriminative representation of neural activity, better suited to the complexity and variability of internal language generation.

6. Discussion

Language constitutes a fundamental pillar in the construction of reality, capable of shaping perceptions, fostering understanding, and deeply influencing the trajectory of human experience. The mere act of writing this article is made possible by it. Communication should be a birthright; however, due to greater circumstances, it is not. Cultural and social factors can turn it into a privilege, and language differences can serve to rank or exclude. Beyond the social dimension, motor impairments further limit access to communication. There are individuals with intact cognitive capabilities—fully able to form linguistic intentions—who are unable to express them due to motor constraints. This raises an unavoidable question: What kind of life can someone aspire to if they are denied the freedom to communicate, to express, to be?
This is why the study and development of systems and interfaces designed to restore communication capacity are of vital importance. Returning this capability is, in essence, restoring a fundamental human right. The challenge is not simply repairing but bypassing the need to decode what the biological machine can’t. It is about constructing alternative pathways that allow the return of speech—and with it, the ability to exist in society—to those who have lost it.
Throughout this review, numerous breakthroughs and advances have been presented. Advancing this field will require a coordinated, multidisciplinary effort that integrates linguistic, neurological, and computational perspectives. Only by understanding each of these dimensions can studies design effective systems for decoding imagined speech. Without such a comprehensive perspective, one risks getting lost in the task of improving classification scores on a specific dataset, instead of formulating the right questions and reflections toward the ultimate goal of restoring communication. By adopting this broader range, progress could be greatly accelerated.
During analysis, it became clear that system performance is not solely based on the complexity of classification models; one of the most critical factors is the nature of the features extracted from the signal. The use of spectral and temporal information combined yielded better performance than isolated feature types.
Additionally, the choice of electrodes and brain regions covered, particularly those related to the sensorimotor cortex, proved highly relevant. This region may be key in designing models that generalize language itself, rather than simply recognizing specific vocabularies.
While achieving high accuracy by leveraging the semantic or emotional salience of a limited set of words may be appealing, such models would probably fail when confronted with different languages, new vocabularies, or diversified communicative settings. Therefore, regardless of the population involved in communication, attention must be paid to the motor symbiosis inherent in all kinds of communication. This reflects not only in speech but also in sign language and imagined articulation.
Lastly, careful consideration must be given to how linguistic units are represented and approached. While word-level classification may appear intuitive, decoding the entire vocabulary space is extremely complex. These findings suggest that phoneme-level decoding offers greater potential. A progressive system that decodes phonemes over short windows and incrementally links them into likely words, based on both current and prior phoneme sequences, could offer a structured, scalable path toward real-time imagined speech decoding.

7. Conclusions

This work emphasizes the importance of a multidisciplinary approach in future research on imagined speech decoding. The combination of spectral features and temporal features has achieved better performance than methods relying on a single type of feature. In addition, the sensorimotor cortex emerged as an important region, especially for constructing models intended to generalize language representations rather than identifying particular words. Motor–language symbiosis appears to be a consistent phenomenon across individuals and modalities. To achieve broader progress, it is also necessary to reevaluate the linguistic unit of classification. Phoneme-level decoding methods are more promising in terms of applicability and generalization compared to word-level decoding. A progressive decoding pipeline, focused on short phoneme windows and predictive modeling of word probabilities, might help deliver more powerful, expressive, and robust BCI systems that could return the ability to speak—and to be—to those who need it most.

Supplementary Materials

The following supporting information can be downloaded at: https://www.mdpi.com/article/10.3390/ai7020075/s1. The PRISMA checklist supporting this systematic review is available as Supplementary Material. Reference [27] is cited in Supplementary Materials.

Author Contributions

Conceptualization, L.F.E.-I. and S.T.-A.; methodology, L.F.E.-I. and S.T.-A.; validation, S.T.-A. and L.R.G.-N.; formal analysis, L.F.E.-I.; investigation, L.F.E.-I.; resources, S.T.-A. and J.M.R.-A.; writing—original draft preparation, L.F.E.-I.; writing—review and editing, L.F.E.-I., S.T.-A. and L.R.G.-N.; visualization, L.F.E.-I.; supervision, S.T.-A. and J.C.P.-O.; project administration, S.T.-A.; funding acquisition, S.T.-A. and J.C.P.-O. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

No new data were created or analyzed in this study. Data sharing is not applicable to this article.

Acknowledgments

We extend our appreciation to CONAHCYT for providing the necessary scholarship to conduct this research.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Tremblay, P.; Brambati, S.M. A historical perspective on the neurobiology of speech and language: From the 19th century to the present. Front. Psychol. 2024, 15, 1420133. [Google Scholar] [CrossRef] [Scilit]
  2. Cooney, C.; Folli, R.; Coyle, D. Opportunities, pitfalls and trade-offs in designing protocols for measuring the neural correlates of speech. Neurosci. Biobehav. Rev. 2022, 140, 104783. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Bhat, C.G. Automatic Recognition and Assessment of Dysarthric Speech; Radboud University Press: Nijmegan, The Netherlands, 2025. [Google Scholar] [CrossRef] [Scilit]
  4. Herbet, G.; Duffau, H. Revisiting the Functional Anatomy of the Human Brain: Toward a Meta-Networking Theory of Cerebral Functions. Physiol. Rev. 2020, 100, 1181–1228. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Hashim, N.; Ali, A.; Mohd-Isa, W.-N. Word-Based Classification of Imagined Speech Using EEG. In Computational Science and Technology; Alfred, R., Iida, H., Ibrahim, A.A.A., Lim, Y., Eds.; Lecture Notes in Electrical Engineering; Springer: Singapore, 2018; Volume 488, pp. 195–204. [Google Scholar] [CrossRef] [Scilit]
  6. Wang, Z.; Ji, H. Open Vocabulary Electroencephalography-To-Text Decoding and Zero-shot Sentiment Classification. arXiv 2024, arXiv:2112.02690. [Google Scholar] [CrossRef] [Scilit]
  7. Lee, Y.-E.; Kim, S.-H.; Lee, S.-H.; Lee, J.-S.; Kim, S.; Lee, S.-W. Speech Synthesis from Brain Signals Based on Generative Model. In 2023 11th International Winter Conference on Brain-Computer Interface (BCI), Gangwon, Republic of Korea; IEEE: New York, NY, USA, 2023; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  8. Lee, Y.-E.; Lee, S.-H.; Kim, S.; Lee, J.-S.; Kim, D.-S. Enhanced Generative Adversarial Networks for Unseen Word Generation from EEG Signals. In 2024 12th International Winter Conference on Brain-Computer Interface (BCI), Gangwon, Republic of Korea; IEEE: New York, NY, USA, 2024; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  9. Lee, Y.-E.; Lee, S.-H.; Kim, S.-H.; Lee, S.-W. Towards Voice Reconstruction from EEG during Imagined Speech. arXiv 2023, arXiv:2301.07173. [Google Scholar] [CrossRef] [Scilit]
  10. Liyanagoonawardena, S.N.; Palihakkara, A.T.; Mudalige, D.N.; Gamage, C.J.U.; De Silva, A.C.; Chang, T. Significance of Subject-Dependency for Imagined Speech Classification Using EEG. In 2023 IEEE 17th International Conference on Industrial and Information Systems (ICIIS), Peradeniya, Sri Lanka; IEEE: New York, NY, USA, 2023; pp. 541–546. [Google Scholar] [CrossRef] [Scilit]
  11. Rashmi, C.R.; Shantala, C.P. EEG artifacts detection and removal techniques for brain computer interface applications: A systematic review. Int. J. Adv. Technol. Eng. Explor. 2022, 9, 354. [Google Scholar] [CrossRef] [Scilit]
  12. Katmah, R.; Al-Shargie, F.; Tariq, U.; Babiloni, F.; Al-Mughairbi, F.; Al-Nashash, H. A Review on Mental Stress Assessment Methods Using EEG Signals. Sensors 2021, 21, 5043. [Google Scholar] [CrossRef] [Scilit]
  13. Buján, A.; Sampaio, A.; Pinal, D. Resting-state electroencephalographic correlates of cognitive reserve: Moderating the age-related worsening in cognitive function. Front. Aging Neurosci. 2022, 14, 854928. [Google Scholar] [CrossRef] [Scilit]
  14. Tomalski, P.; Moore, D.G.; Ribeiro, H.; Axelsson, E.L.; Murphy, E.; Karmiloff-Smith, A.; Johnson, M.H.; Kushnerenko, E. Socioeconomic status and functional brain development–associations in early infancy. Dev. Sci. 2013, 16, 676–687. [Google Scholar] [CrossRef] [Scilit]
  15. Nolvi, S.; Merz, E.C.; Kataja, E.-L.; Parsons, C.E. Prenatal Stress and the Developing Brain: Postnatal Environments Promoting Resilience. Biol. Psychiatry 2023, 93, 942–952. [Google Scholar] [CrossRef] [Scilit]
  16. Bialystok, E. The bilingual adaptation: How minds accommodate experience. Psychol. Bull. 2017, 143, 233–262. [Google Scholar] [CrossRef] [Scilit]
  17. Korenar, M.; Treffers-Daller, J.; Pliatsikas, C. Dynamic effects of bilingualism on brain structure map onto general principles of experience-based neuroplasticity. Sci. Rep. 2023, 13, 3428. [Google Scholar] [CrossRef] [Scilit]
  18. Chwilla, D.J. Context effects in language comprehension: The role of emotional state and attention on semantic and syntactic processing. Front. Hum. Neurosci. 2022, 16, 1014547. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  19. Hinojosa, J.A.; Moreno, E.M.; Ferré, P. Affective neurolinguistics: Towards a framework for reconciling language and emotion. Lang. Cogn. Neurosci. 2020, 35, 813–839. [Google Scholar] [CrossRef] [Scilit]
  20. Stoyell, S.M.; Wilmskoetter, J.; Dobrota, M.-A.; Chinappen, D.M.; Bonilha, L.; Mintz, M.; Brinkmann, B.H.; Herman, S.T.; Peters, J.M.; Vulliemoz, S.; et al. High-Density EEG in Current Clinical Practice and Opportunities for the Future. J. Clin. Neurophysiol. 2021, 38, 112–123. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  21. Sabio, J.; Williams, N.S.; McArthur, G.M.; Badcock, N.A. A scoping review on the use of consumer-grade EEG devices for research. PLoS ONE 2024, 19, e0291186. [Google Scholar] [CrossRef] [Scilit]
  22. Liu, Q.; Yang, L.; Zhang, Z.; Yang, H.; Zhang, Y.; Wu, J. The Feature, Performance, and Prospect of Advanced Electrodes for Electroencephalogram. Biosensors 2023, 13, 101. [Google Scholar] [CrossRef] [Scilit]
  23. Fiedler, P.; Graichen, U.; Zimmer, E.; Haueisen, J. Simultaneous Dry and Gel-Based High-Density Electroencephalography Recordings. Sensors 2023, 23, 9745. [Google Scholar] [CrossRef] [Scilit]
  24. Tan, J.-L.; Liang, Z.-F.; Zhang, R.; Dong, Y.-Q.; Li, G.-H.; Zhang, M.; Wang, H.; Xu, N. Suppressing of Power Line Artifact from Electroencephalogram Measurements Using Sparsity in Frequency Domain. Front. Neurosci. 2021, 15, 780373. [Google Scholar] [CrossRef] [Scilit]
  25. Li, M.; Liao, S.; Pun, S.H.; Chen, F. Effects of EEG Analysis Window Location on Classifying Spoken Mandarin Monosyllables. In 2023 11th International IEEE/EMBS Conference on Neural Engineering (NER), Baltimore, MD, USA; IEEE: New York, NY, USA, 2023; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  26. Lopez-Bernal, D.; Balderas, D.; Ponce, P.; Molina, A. A State-of-the-Art Review of EEG-Based Imagined Speech Decoding. Front. Hum. Neurosci. 2022, 16, 867281. [Google Scholar] [CrossRef] [Scilit]
  27. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: An updated guideline for reporting systematic reviews. BMJ 2021, 372, n71. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  28. Vorontsova, D.; Menshikov, I.; Zubov, A.; Orlov, K.; Rikunov, P.; Zvereva, E.; Flitman, L.; Lanikin, A.; Sokolova, A.; Markov, S.; et al. Silent EEG-Speech Recognition Using Convolutional and Recurrent Neural Network with 85% Accuracy of 9 Words Classification. Sensors 2021, 21, 6744. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  29. Yang, Y.C.; Walsh, C.E.; Shartle, K.; Stebbins, R.C.; Aiello, A.E.; Belsky, D.W.; Harris, K.M.; Chanti-Ketterl, M.; Plassman, B.L. An Early and Unequal Decline: Life Course Trajectories of Cognitive Aging in the United States. J. Aging Health 2024, 36, 230–245. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  30. Ferguson, H.J.; Brunsdon, V.E.A.; Bradford, E.E.F. The developmental trajectories of executive function from adolescence to old age. Sci. Rep. 2021, 11, 1382. [Google Scholar] [CrossRef] [Scilit]
  31. Balaji, A.; Haldar, A.; Patil, K.; Ruthvik, T.S.; Ca, V.; Jartarkar, M.; Baths, V. EEG-based classification of bilingual unspoken speech using ANN. In 2017 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society (EMBC), Seogwipo; IEEE: New York, NY, USA, 2017; pp. 1022–1025. [Google Scholar] [CrossRef] [Scilit]
  32. LaRocco, J.; Tahmina, Q.; Lecian, S.; Moore, J.; Helbig, C.; Gupta, S. Evaluation of an English language phoneme-based imagined speech brain computer interface with low-cost electroencephalography. Front. Neuroinform. 2023, 17, 1306277. [Google Scholar] [CrossRef] [Scilit]
  33. Guo, Z.; Zhang, H.; Chen, F. Impacts of imagined lexical tone on Mandarin speech imagery BCI performance. In 2023 11th International IEEE/EMBS Conference on Neural Engineering (NER), Baltimore, MD, USA; IEEE: New York, NY, USA, 2023; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
  34. Torres-García, A.A.; Reyes-García, C.A.; Villaseñor-Pineda, L.; García-Aguilar, G. Implementing a fuzzy inference system in a multi-objective EEG channel selection model for imagined speech classification. Expert Syst. Appl. 2016, 59, 1–12. [Google Scholar] [CrossRef] [Scilit]
  35. García-Salinas, J.S.; Villaseñor-Pineda, L.; Reyes-García, C.A.; Torres-García, A.A. Transfer learning in imagined speech EEG-based BCIs. Biomed. Signal Process. Control 2019, 50, 151–157. [Google Scholar] [CrossRef] [Scilit]
  36. Jiménez-Guarneros, M.; Gómez-Gil, P. Standardization-refinement domain adaptation method for cross-subject EEG-based classification in imagined speech recognition. Pattern Recognit. Lett. 2021, 141, 54–60. [Google Scholar] [CrossRef] [Scilit]
  37. García-Salinas, J.S.; Torres-García, A.A.; Reyes-Garćia, C.A.; Villaseñor-Pineda, L. Intra-subject class-incremental deep learning approach for EEG-based imagined speech recognition. Biomed. Signal Process. Control 2023, 81, 104433. [Google Scholar] [CrossRef] [Scilit]
  38. Hernández-Del-Toro, T.; Reyes-García, C.A.; Villaseñor-Pineda, L. Toward asynchronous EEG-based BCI: Detecting imagined words segments in continuous EEG signals. Biomed. Signal Process. Control 2021, 65, 102351. [Google Scholar] [CrossRef] [Scilit]
  39. Tamm, M.-O.; Muhammad, Y.; Muhammad, N. Classification of Vowels from Imagined Speech with Convolutional Neural Networks. Computers 2020, 9, 46. [Google Scholar] [CrossRef] [Scilit]
  40. Carvalho, V.R.; Mendes, E.M.A.M.; Fallah, A.; Sejnowski, T.J.; Comstock, L.; Lainscsek, C. Decoding imagined speech with delay differential analysis. Front. Hum. Neurosci. 2024, 18, 1398065. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  41. Cooney, C.; Korik, A.; Folli, R.; Coyle, D. Evaluation of Hyperparameter Optimization in Machine and Deep Learning Methods for Decoding Imagined Speech EEG. Sensors 2020, 20, 4629. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Lee, D.-Y.; Lee, M.; Lee, S.-W. Classification of Imagined Speech Using Siamese Neural Network. In 2020 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Toronto, ON, Canada; IEEE: New York, NY, USA, 2020; pp. 2979–2984. [Google Scholar] [CrossRef] [Scilit]
  43. Mahapatra, N.C.; Bhuyan, P. Decoding of Imagined Speech Neural EEG Signals Using Deep Reinforcement Learning Technique. In 2022 International Conference on Advancements in Smart, Secure and Intelligent Computing (ASSIC), Bhubaneswar, India; IEEE: New York, NY, USA, 2022; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  44. Das, A.; Soni, P.; Huang, M.-C.; Lin, F.; Xu, W. Multimodal speech recognition using EEG and audio signals: A novel approach for enhancing ASR systems. Smart Health 2024, 32, 100477. [Google Scholar] [CrossRef] [Scilit]
  45. Cooney, C.; Folli, R.; Coyle, D. Optimizing Layers Improves CNN Generalization and Transfer Learning for Imagined Speech Decoding from EEG. In 2019 IEEE International Conference on Systems, Man and Cybernetics (SMC), Bari, Italy; IEEE: New York, NY, USA, 2019; pp. 1311–1316. [Google Scholar] [CrossRef] [Scilit]
  46. Bras, C.; Patel, T.; Scharenborg, O. Using articulated speech EEG signals for imagined speech decoding. In Interspeech 2024; ISCA: Kos Island, Greece, 2024; pp. 407–411. [Google Scholar] [CrossRef] [Scilit]
  47. Ghosh, R.; Sinha, N.; Phadikar, S. Identification of Imagined Bengali Vowels from EEG Signals Using Activity Map and Convolutional Neural Network. In Brain-Computer Interface, 1st ed.; Sumithra, M.G., Dhanaraj, R.K., Milanova, M., Balusamy, B., Venkatesan, C., Eds.; Wiley: Hoboken, NJ, USA, 2023; pp. 231–254. [Google Scholar] [CrossRef] [Scilit]
  48. Gillis, M.; Vanthornhout, J.; Simon, J.Z.; Francart, T.; Brodbeck, C. Neural Markers of Speech Comprehension: Measuring EEG Tracking of Linguistic Speech Representations, Controlling the Speech Acoustics. J. Neurosci. 2021, 41, 10316–10329. [Google Scholar] [CrossRef] [Scilit]
  49. Steinberg, J.; Sompolinsky, H. Associative memory of structured knowledge. Sci. Rep. 2022, 12, 21808. [Google Scholar] [CrossRef] [Scilit]
  50. Jung, J.; Ralph, M.A.L. Distinct but cooperating brain networks supporting semantic cognition. Cereb. Cortex 2023, 33, 2021–2036. [Google Scholar] [CrossRef] [Scilit]
  51. Huang, X.; Wong, B.W.L.; Ng, H.T.-Y.; Sommer, W.; Dimigen, O.; Maurer, U. Neural mechanism underlying preview effects and masked priming effects in visual word processing. Atten. Percept. Psychophys. 2025, 87, 5–24. [Google Scholar] [CrossRef] [Scilit]
  52. Levari, T.; Snedeker, J. Understanding words in context: A naturalistic EEG study of children’s lexical processing. J. Mem. Lang. 2024, 137, 104512. [Google Scholar] [CrossRef] [Scilit]
  53. Grisoni, L.; Boux, I.P.; Pulvermüller, F. Predictive brain activity shows congruent semantic specificity in language comprehension and production. J. Neurosci. 2024, 44, e1723232023. [Google Scholar] [CrossRef] [Scilit]
  54. Angulo-Chavira, A.Q.; Castellón-Flores, A.M.; Ciria, A.; Arias-Trejo, N. Sentence-final completion norms for 2925 Mexican Spanish sentence contexts. Behav. Res. Methods 2023, 56, 2486–2498. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  55. Varshney, Y.V.; Khan, A. Imagined Speech Classification Using Six Phonetically Distributed Words. Front. Signal Process. 2022, 2, 760643. [Google Scholar] [CrossRef] [Scilit]
  56. Muraki, E.J.; Doyle, A.; Protzner, A.B.; Pexman, P.M. Context matters: How do task demands modulate the recruitment of sensorimotor information during language processing? Front. Hum. Neurosci. 2023, 16, 976954. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  57. Min, B.; Kim, J.; Park, H.; Lee, B. Vowel Imagery Decoding toward Silent Speech BCI Using Extreme Learning Machine with Electroencephalogram. BioMed Res. Int. 2016, 2016, 2618265. [Google Scholar] [CrossRef] [Scilit]
  58. Nguyen, C.H.; Karavas, G.K.; Artemiadis, P. Inferring imagined speech using EEG signals: A new approach using Riemannian manifold features. J. Neural Eng. 2018, 15, 016002. [Google Scholar] [CrossRef] [Scilit]
  59. Saha, P.; Fels, S. Hierarchical Deep Feature Learning for Decoding Imagined Speech from EEG. In Proceedings of the AAAI Conference on Artificial Intelligence; AAAI Press: Washington, DC, USA, 2019; Volume 33, pp. 10019–10020. [Google Scholar] [CrossRef] [Scilit]
  60. Mohan, A.; Anand, R. Multi-modal deep learning architecture for enhanced feature extraction and classification of imagined speech words. Eng. Res. Express 2025, 7, 015323. [Google Scholar] [CrossRef] [Scilit]
  61. Mohan, A.; Anand, R. Classification of Imagined Speech Signals Using Functional Connectivity Graphs and Machine Learning Models. Brain Topogr. 2025, 38, 25. [Google Scholar] [CrossRef] [Scilit]
  62. Bisla, M.; Anand, R.S. Transfer Learning Enabled Imagined Speech Interpretation Using Phase-Based Brain Functional Connectivity and Power Analysis. IEEE Access 2024, 12, 108399–108413. [Google Scholar] [CrossRef] [Scilit]
  63. Panachakel, J.T.; Ganesan, R.A. Decoding Imagined Speech from EEG Using Transfer Learning. IEEE Access 2021, 9, 135371–135383. [Google Scholar] [CrossRef] [Scilit]
  64. Qureshi, M.N.I.; Min, B.; Park, H.-J.; Cho, D.; Choi, W.; Lee, B. Multiclass Classification of Word Imagination Speech with Hybrid Connectivity Features. IEEE Trans. Biomed. Eng. 2018, 65, 2168–2177. [Google Scholar] [CrossRef] [Scilit]
  65. Zhao, S.; Rudzicz, F. Classifying phonological categories in imagined and articulated speech. In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), South Brisbane, Queensland, Australia; IEEE: New York, NY, USA, 2015; pp. 992–996. [Google Scholar] [CrossRef] [Scilit]
  66. Cooney, C.; Folli, R.; Coyle, D. Mel Frequency Cepstral Coefficients Enhance Imagined Speech Decoding Accuracy from EEG. In 2018 29th Irish Signals and Systems Conference (ISSC), Belfast; IEEE: New York, NY, USA, 2018; pp. 1–7. [Google Scholar] [CrossRef] [Scilit]
  67. Saha, P.; Fels, S.; Abdul-Mageed, M. Deep Learning the EEG Manifold for Phonological Categorization from Active Thoughts. In ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Brighton, UK; IEEE: New York, NY, USA, 2019; pp. 2762–2766. [Google Scholar] [CrossRef] [Scilit]
  68. Bakhshali, M.A.; Khademi, M.; Ebrahimi-Moghadam, A.; Moghimi, S. EEG signal classification of imagined speech based on Riemannian distance of correntropy spectral density. Biomed. Signal Process. Control 2020, 59, 101899. [Google Scholar] [CrossRef] [Scilit]
  69. Mini, P.P.; Thomas, T.; Gopikakumari, R. EEG based direct speech BCI system using a fusion of SMRT and MFCC/LPCC features with ANN classifier. Biomed. Signal Process. Control 2021, 68, 102625. [Google Scholar] [CrossRef] [Scilit]
  70. Mohan, A.; Anand, R.S. Analysis of Variation in Electrode Location and frequency on Imagined Speech Task Prediction. In 2023 7th International Conference on Computer Applications in Electrical Engineering-Recent Advances (CERA), Roorkee, India; IEEE: New York, NY, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
  71. Rousis, G.; Kalaganis, F.P.; Nikolopoulos, S.; Kompatsiaris, I.; Petrantonakis, P.C. Combining EEGNet with SPDNet Towards an End-To-End Architecture for Imagined Speech Decoding. In 2024 32nd European Signal Processing Conference (EUSIPCO), Lyon, France; IEEE: New York, NY, USA, 2024; pp. 1531–1535. [Google Scholar] [CrossRef] [Scilit]
  72. Sharon, R.; Sur, M.; Murthy, H. Harnessing the Multi-Phasal Nature of Speech-EEG for Enhancing Imagined Speech Recognition. IEEE Open J. Signal Process. 2025, 6, 78–88. [Google Scholar] [CrossRef] [Scilit]
  73. Coretto, G.A.P.; Gareis, I.E.; Rufiner, H.L. Open access database of EEG signals recorded during imagined speech. In 12th International Symposium on Medical Information Processing and Analysis, Tandil, Argentina; Romero, E., Lepore, N., Brieva, J., Larrabide, I., Eds.; SPIE: Bellingham, WA, USA, 2017; p. 1016002. [Google Scholar] [CrossRef] [Scilit]
  74. Lee, S.-H.; Lee, M.; Lee, S.-W. Neural Decoding of Imagined Speech and Visual Imagery as Intuitive Paradigms for BCI Communication. IEEE Trans. Neural Syst. Rehabil. Eng. 2020, 28, 2647–2659. [Google Scholar] [CrossRef] [Scilit]
  75. Lee, S.-H.; Lee, M.; Jeong, J.-H.; Lee, S.-W. Towards an EEG-based Intuitive BCI Communication System Using Imagined Speech and Visual Imagery. In 2019 IEEE International Conference on Systems, Man and Cybernetics (SMC), Bari, Italy; IEEE: New York, NY, USA, 2019; pp. 4409–4414. [Google Scholar] [CrossRef] [Scilit]
  76. Lee, S.-H.; Lee, M.; Lee, S.-W. EEG Representations of Spatial and Temporal Features in Imagined Speech and Overt Speech. In Pattern Recognition; Palaiahnakote, S., Di Baja, G.S., Wang, L., Yan, W.Q., Eds.; Lecture Notes in Computer Science; Springer International Publishing: Cham, Switzerland, 2020; Volume 12047, pp. 387–400. [Google Scholar] [CrossRef] [Scilit]
  77. Kim, S.; Lee, Y.-E.; Lee, S.-H.; Lee, S.-W. Diff-E: Diffusion-based Learning for Decoding Imagined Speech EEG. In Interspeech 2023; ISCA: Kos Island, Greece, 2023; pp. 1159–1163. [Google Scholar] [CrossRef] [Scilit]
  78. Lee, D.-H.; Kim, S.-J.; Han, H.-T.; Lee, S.-W. TOINet: Transfer Learning from Overt Speech- to Imagined Speech-Based EEG Signals with Convolutional Autoencoder. In 2023 IEEE International Conference on Systems, Man, and Cybernetics (SMC), Honolulu, Oahu, HI, USA; IEEE: New York, NY, USA, 2023; pp. 4441–4446. [Google Scholar] [CrossRef] [Scilit]
  79. Park, H.; Lee, B. Multiclass classification of imagined speech EEG using noise-assisted multivariate empirical mode decomposition and multireceptive field convolutional neural network. Front. Hum. Neurosci. 2023, 17, 1186594. [Google Scholar] [CrossRef] [Scilit]
  80. Wu, S.; Bhadra, K.; Giraud, A.-L.; Marchesotti, S. Adaptive LDA Classifier Enhances Real-Time Control of an EEG Brain–Computer Interface for Decoding Imagined Syllables. Brain Sci. 2024, 14, 196. [Google Scholar] [CrossRef] [Scilit]
  81. Alharbi, Y.F.; Alotaibi, Y.A. Decoding Imagined Speech from EEG Data: A Hybrid Deep Learning Approach to Capturing Spatial and Temporal Features. Life 2024, 14, 1501. [Google Scholar] [CrossRef] [Scilit]
  82. Alharbi, Y.F.; Alotaibi, Y.A. Imagined Speech Recognition and the Role of Brain Areas Based on Topographical Maps of EEG Signal. In 2024 47th International Conference on Telecommunications and Signal Processing (TSP), Prague, Czech Republic; IEEE: New York, NY, USA, 2024; pp. 274–279. [Google Scholar] [CrossRef] [Scilit]
  83. Tiwari, S.; Goel, S.; Bhardwaj, A. Classification of imagined speech of vowels from EEG signals using multi-headed CNNs feature fusion network. Digit. Signal Process. 2024, 148, 104447. [Google Scholar] [CrossRef] [Scilit]
  84. Chaddad, A.; Wu, Y.; Kateb, R.; Bouridane, A. Electroencephalography Signal Processing: A Comprehensive Review and Analysis of Methods and Techniques. Sensors 2023, 23, 6434. [Google Scholar] [CrossRef] [Scilit]
  85. Cao, J.; Huppert, T.J.; Grover, P.; Kainerstorfer, J.M. Enhanced spatiotemporal resolution imaging of neuronal activity using joint electroencephalography and diffuse optical tomography. Neurophotonics 2021, 8, 015002. [Google Scholar] [CrossRef] [Scilit]
  86. Schreiner, L.; Jordan, M.; Sieghartsleitner, S.; Kapeller, C.; Pretl, H.; Kamada, K.; Asman, P.; Ince, N.F.; Miller, K.J.; Guger, C. Mapping of the central sulcus using non-invasive ultra-high-density brain recordings. Sci. Rep. 2024, 14, 6527. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  87. Ebrahiminia, F.; Cichy, R.M.; Khaligh-Razavi, S.-M. A multivariate comparison of electroencephalogram and functional magnetic resonance imaging to electrocorticogram using visual object representations in humans. Front. Neurosci. 2022, 16, 983602. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  88. Hnazaee, M.F.; Wittevrongel, B.; Khachatryan, E.; Libert, A.; Carrette, E.; Dauwe, I.; Meurs, A.; Boon, P.; Van Roost, D.; Van Hulle, M.M. Localization of deep brain activity with scalp and subdural EEG. NeuroImage 2020, 223, 117344. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  89. Yuste, R.; Cossart, R.; Yaksi, E. Neuronal ensembles: Building blocks of neural circuits. Neuron 2024, 112, 875–892. [Google Scholar] [CrossRef] [Scilit]
  90. Alotaibi, S.; Alamri, S.; Alsaleh, A.; Meyer, G. Neural adaptations in short-term learning of sign language revealed by fMRI and DTI. Sci. Rep. 2025, 15, 5345. [Google Scholar] [CrossRef] [Scilit]
  91. Hertrich, I.; Dietrich, S.; Ackermann, H. The Margins of the Language Network in the Brain. Front. Commun. 2020, 5, 519955. [Google Scholar] [CrossRef] [Scilit]
  92. Forkel, S.J.; Hagoort, P. Redefining language networks: Connectivity beyond localised regions. Brain Struct. Funct. 2024, 229, 2073–2078. [Google Scholar] [CrossRef] [Scilit]
  93. Berker, E.A.; Berker, A.H.; Smith, A. Translation of Broca’s 1865 Report Localization of Speech in the Third Left Frontal Convolution. Arch. Neurol. 1986, 43, 1065–1072. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  94. Karpiel, I.; Kurasz, Z.; Kurasz, R.; Duch, K. The Influence of Filters on EEG-ERP Testing: Analysis of Motor Cortex in Healthy Subjects. Sensors 2021, 21, 7711. [Google Scholar] [CrossRef] [Scilit]
  95. Das, R.K.; Martin, A.; Zurales, T.; Dowling, D.; Khan, A. A Survey on EEG Data Analysis Software. Sci 2023, 5, 23. [Google Scholar] [CrossRef] [Scilit]
  96. Zhang, L.; Wang, P.; Zhang, R.; Chen, M.; Shi, L.; Gao, J.; Hu, Y. The Influence of Different EEG References on Scalp EEG Functional Network Analysis During Hand Movement Tasks. Front. Hum. Neurosci. 2020, 14, 367. [Google Scholar] [CrossRef] [Scilit]
  97. Attar, E.T. Review of electroencephalography signals approaches for mental stress assessment. Neurosciences 2022, 27, 209–215. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  98. Nalborczyk, L.; Debarnot, U.; Longcamp, M.; Guillot, A.; Alario, F.X. The role of motor inhibition during covert speech production. Front. Hum. Neurosci. 2022, 16, 804832. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  99. Wandelt, S.K.; Bjånes, D.A.; Pejsa, K.; Lee, B.; Liu, C.; Andersen, R.A. Representation of internal speech by single neurons in human supramarginal gyrus. Nat. Hum. Behav. 2024, 8, 1136–1149. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  100. De Borman, A.; Wittevrongel, B.; Dauwe, I.; Carrette, E.; Meurs, A.; Van Roost, D.; Boon, P.; Van Hulle , M.M. Imagined speech event detection from electrocorticography and its transfer between speech modes and subjects. Commun. Biol. 2024, 7, 818. [Google Scholar] [CrossRef] [Scilit]
  101. Zheng, Y.; Zhang, J.; Yang, Y.; Xu, M. Neural representation of sensorimotor features in language-motor areas during auditory and visual perception. Commun. Biol. 2025, 8, 41. [Google Scholar] [CrossRef] [Scilit]
  102. Chibawanye, I.E.; Silva, E.; Noll, K.; Bradshaw, M.; Connelly, K.; Liu, H.; Tummala, R.; Ferson, D.; Lang, F.F.; Kumar, V.A. Language mapping during awake brain surgery in a deaf patient with a brain tumor: Illustrative case. J. Neurosurg. Case Lessons 2025, 9, CASE2597. [Google Scholar] [CrossRef] [Scilit]
  103. Rossi, E.; Soares, S.M.P.; Prystauka, Y.; Nakamura, M.; Rothman, J. Riding the (brain) waves! Using neural oscillations to inform bilingualism research. Biling. Lang. Cogn. 2023, 26, 202–215. [Google Scholar] [CrossRef] [Scilit]
  104. Singh, A.K.; Krishnan, S. Trends in EEG signal feature extraction applications. Front. Artif. Intell. 2023, 5, 1072801. [Google Scholar] [CrossRef] [Scilit]
  105. Lee, Y.-E.; Lee, S.-H. EEG-Transformer: Self-attention from Transformer Architecture for Decoding EEG of Imagined Speech. In 2022 10th International Winter Conference on Brain-Computer Interface (BCI), Gangwon-do, Republic of Korea; IEEE: New York, NY, USA, 2022; pp. 1–4. [Google Scholar] [CrossRef] [Scilit]
Figure 1. PRISMA Flow diagram of study selection process.
Figure 1. PRISMA Flow diagram of study selection process.
Ai 07 00075 g001
Figure 2. Comparison of classification accuracy by vocabulary size range, distinguishing between phoneme-based and word-based linguistic units.
Figure 2. Comparison of classification accuracy by vocabulary size range, distinguishing between phoneme-based and word-based linguistic units.
Ai 07 00075 g002
Figure 3. Comparison of classification accuracy for a fixed vocabulary size (5), showing performance differences between phoneme-level and word-level decoding.
Figure 3. Comparison of classification accuracy for a fixed vocabulary size (5), showing performance differences between phoneme-level and word-level decoding.
Ai 07 00075 g003
Figure 4. Comparison of classification accuracy based on vocabulary abstraction, contrasting concrete directional commands with abstract or mixed semantic categories.
Figure 4. Comparison of classification accuracy based on vocabulary abstraction, contrasting concrete directional commands with abstract or mixed semantic categories.
Ai 07 00075 g004
Figure 5. Mean imagined speech classification accuracy mapped to individual 10–20 EEG electrodes, based solely on studies that employed a selected subset of electrodes rather than full-head configurations.
Figure 5. Mean imagined speech classification accuracy mapped to individual 10–20 EEG electrodes, based solely on studies that employed a selected subset of electrodes rather than full-head configurations.
Ai 07 00075 g005
Figure 6. Comparison of average classification accuracy across EEG frequency bands.
Figure 6. Comparison of average classification accuracy across EEG frequency bands.
Ai 07 00075 g006
Figure 7. Comparison of classification performance according to the type of extracted features, showing differences between spectral, temporal, spatial, and combined (temporal–spectral) representations. (a) Machine Learning (ML) models and (b) Deep Learning (DL) models.
Figure 7. Comparison of classification performance according to the type of extracted features, showing differences between spectral, temporal, spatial, and combined (temporal–spectral) representations. (a) Machine Learning (ML) models and (b) Deep Learning (DL) models.
Ai 07 00075 g007
Table 1. Linguistic properties and accuracy ranges per dataset in imagined speech studies.
Table 1. Linguistic properties and accuracy ranges per dataset in imagined speech studies.
Dataset No. ParticipantsVocab SizeClass TrialsLanguageVocabulary ContentMean ACC by Linguistic UnitLowest ACC by Linguistic UnitHighest ACC by Linguistic Unit
By Torres-García et al. [34] used in [34,35,36,37,38]275 words33Spanish“arriba”,”abajo”,”izquierda”, “derecha”, “seleccionar”63.42% (multiclass)58.84%68.18%
Min et al. [57]55 phonemes50Korean/a/,/e/,/i/,/o/,/u/88.84% (Binary)Not ApplicableNot Applicable
Noramiza Hashim et al. [5]42 words50English“Yes”, “No”58% (Binary)Not ApplicableNot Applicable
Advait Balaji et al. [31]84 words10English and Hindi“yes”, “no”, “haan”, “na”75.38% (multiclass)Not ApplicableNot Applicable
By Nguyen et al. [58] used in [36,37,58,59,60,61,62,63]155 words and 3 phonemes100English/a/,/i/,/u/, “in”, “out”, “up”, “cooperate”, “independent”phonemes: 71.82%
Short words: 72.57%
Long words: 79.34%
phonemes: 44%
Short words: 42%
Long words: 62.99%
phonemes: 94.53%
Short words: 95.02%
Long words: 95.85%
Naveed et al. [64]85 words100English“go”, “back”, “left”, “right”, “stop”40.3% (multiclass)Not ApplicableNot Applicable
KARA ONE [65] used in [40,62,66,67,68,69,70,71,72]147 phonemes and 4 words12English/iy/,/uw/,/piy/,/tuy/,/diy/,/m/,/n/, “pat”, “pot”, “knew”, “gnaw”multiclass: 40.9%
words: 91.12%
phonemes: 86.73%
C/V: 89.19%
nasal:78.22%
bilabial: 77.81%
/iy/: 80.08%
/uw/: 88.13%
multiclass: 20.80%
words: 89.56%
phonemes: 86.73%
C/V: 85.23%
nasal:72.10%
bilabial: 75.55%
/iy/: 73.30%
/uw/: 81.99%
multiclass: 77.37%
words: 92.68%,
phonemes: 86.73%
C/V: 95.06%
nasal:90.43%
bilabial: 89.62%
/iy/: 90.08%
/uw/: 97.61%
Coretto [73] used in [10,39,40,41,42,43,44,45]155 phonemes and 6 words40Spanish/a/,/e/,/i/,/o/,/u/, “arriba”, “abajo”, “derecha”, “izquierda”, “adelante”, “atrás”multiclass:95.39%
phonemes:39.16%
words:30.1%
multiclass:95.39%
phonemes:24.77%
words:24.90%
multiclass:95.39%
phonemes:81.69%
words:34%
By Seo-Hyun Lee et al. used in [40,74,75,76,77]2212 words100English“ambulance”, “light”, “TV”, “water”, “pain”, “hello”, “toilet”, “clock”, “yes”, “stop”, “help me”, “thank you”, rest-state40% (Multiclass)16.20%60.63%
Hernández-Del-Toro et al. [38]D2:27 and D3:205 wordsD2:32 and D3:40Spanish“up”, “down”, “left”, “right”, “select”D2: 70% (multiclass) and D3: 65% (multiclass)D2: 65%D3: 70%
Vorontsova et al. [28]2688 words~6 to 18Russian“forward”, “backward”, “up”, “down”, “help”, “take”, “stop”, “release”Model trained and tested on the same individual: 84.51% (multiclass)Model trained and tested on 256 subjects: 13.39%Model trained and tested on the same individual: 84.51%
Varshney et al. [55]156 words50English“could”, “yard”, “give”, “him”, “there”, “toe”28.61% (multiclass)Not ApplicableNot Applicable
Rajdeep Ghosh et al. [47]225 phonemes5BengaliAi 07 00075 i00168.9% (multiclass)Not ApplicableNot Applicable
Dae-Hyeok Lee et al. [78]84 words50Korean/Ba/,/Ku/,/He/,/Li/48.41% (multiclass)Not ApplicableNot Applicable
Xin Zhang et al. [79]95 phonemes50Korean/a/,/e/,/i/,/o/,/u/73.09% (multiclass)Not ApplicableNot Applicable
Mingtao Li et al. [25]1070 phonemes5Mandarin ChineseCombination of 5 consonants (/b_/,/m_/,/f_/,/l_/,/j_/) with 4 vowels (/a/,/u/,/i/,/y/), each in 4 different tonesphonemes: 70.7%
Tones: 54.9%
Tones: 54.9%phonemes: 70.7%
Larocco et al. [32]1644 phonemes15English/a/,/b/,/c/,/d/,/e/,/f/,/g/,/h/,/i/,/j/,/k/,/l/,/m/,/n/,/o/,/p/,/q/,/r/,/s/,/t/,/u/,/v/,/w/,/x/,/y/,/z/,/θ/,/ð/,/ŋ/,/ʃ/,/t͡ʃ/,/d͡ʒ/,/j/,/w/,/h/,/ʔ/98% (multiclass)Not ApplicableNot Applicable
DAIS [46]205 phonemes and 10 words20Dutch/a/,/e/,/i/,/o/,/u/, “taal”, “laat”, “leeg”, “geel”, “niet”, “tien”, “toon”, “noot”, “soep”, “poes”phonemes: 27.5%Not ApplicableNot Applicable
Wu et al. [80]202 phonemes45Not mentioned/fO/,/gi/55.7% (Binary)Not ApplicableNot Applicable
BCI2020 [81,82]155 words70English“help me”, “hello”, “stop”, “thank you”, “yes”Binary classification: 76.9%
3-class classification: 59.6%
5-class classification:44.7%
Binary classification: 76%
3-class classification: 59.5%
5-class classification:44.7%
Binary classification: 77.8%
3-class classification: 59.7%
5-class classification:44.7%
Tiwari et al. [83]165 phonemes90English/a/,/e/,/i/,/o/,/u/98.98% (multiclass)Not ApplicableNot Applicable
Table 2. EEG Acquisition settings and preprocessing parameters with associated accuracy outcomes.
Table 2. EEG Acquisition settings and preprocessing parameters with associated accuracy outcomes.
Dataset Number of ElectrodesSampling Rate (Hz)Highest ACC Preprocessing PipelineHighest ACC SAMPLING Cut-Off ConfigurationHighest ACC Electrode Configuration
By Torres-García et al. [34] used in [34,35,36,37,38]14128 HzCommon average reference, Notch filter at 50 Hz and 60 Hz0–64 HzT7, T8, P8, FC6, F8, P7, FC5
Min et al. [57]64250 HzInfinite impulse response Butterworth, Notch filter 59/61 Hz30–70 Hzusing all
Noramiza Hashim et al. [5]14128 HzNotch filter at 50 Hz and 60 Hz and Butterworth HPF, Sinc0.16–43 HzAF3, F7, F3, FC5, T7, P7
Advait Balaji et al. [31]32250 HzNotch filter 50/60 Hz0–40 HzF7, T7, P7, P3, C3, Fp1, FpZ, Fp2, F3, Fz, F4
By Nguyen et al. [58] used in [36,37,58,59,60,61,62,63]64256 HzBandpass Butterworth 8–70 Hz (5th order), Notch 60 Hz4–80 Hzusing all
Naveed et al. [64]641000 HzFinite impulse response filter 0.5/100 Hz0.5–100 Hzusing all
KARA ONE [65] used in [40,62,66,67,68,69,70,71,72]641000 HzFinite impulse response filter 5–60 Hz (20th order)5–60 Hz FC6, FT8, C5, CP2, CP3, T7, CP5, C3, CP1, C4
Coretto [73] used in [10,39,40,41,42,43,44,45]61024 HzFinite impulse response filter 2/40 Hz and Independent Component Analysis (ICA)2–40 HzF3, F4, C3, C4, P3, P4
By Seo-Hyun Lee et al. used in [40,74,75,76,77]64250 HzCommon average reference, Bandpass filter 0.5–125 Hz, Notch filter at 60 Hz and 120 Hz0.5–125 Hzusing all
Hernández-Del-Toro et al. [38]14128 HzCommon Average Reference0–64 Hzusing all
Vorontsova et al. [28]40500 HzIndependent Component Analysis (ICA)5–49 Hzusing all
Varshney et al. [55]64512 HzZero-phase band-pass filter (0.01–250 Hz), notch filter (48–52 Hz), and Independent Component Analysis (ICA)2–64 Hzusing all
Rajdeep Ghosh et al. [47]64512 HzBandpass Butterworth 0–60 Hz, Notch 50 Hz0–60 Hzusing all
Dae-Hyeok Lee et al. [78]581000 HzNotch filter 60 HzFull-band EEGusing all
Xin Zhang et al. [79]64250 HzInfinite impulse response Butterworth filter 59–61 Hz (4th order)0–61 Hzusing all
Mingtao Li et al. [25]641000 HzBandpass 0.5–70 Hz, notch filter 49–51 Hz, and Independent Component Analysis (ICA)0.5–70 Hzusing all
Larocco et al. [32]16250 HzButterworth filter 0.1–125 Hz (4th order), Notch 60 Hz1–100 Hzusing all
DAIS [46]621024 HzButterworth filter 1–40 Hz (2nd order)1–40 Hzusing all
Wu et al. [80]64512 HzCommon Average Reference, Notch 50 Hz1–70 Hzusing all
BCI2020 [81,82]64256 HzNot reportedFull-band EEGFp1, AF3, Fp2, AF4, AF7, AF8, F1, Fz, F7, F5, F3, F4, F6, F8
Tiwari et al. [83]14128 HzInfinite impulse response Butterworth filter >45 Hz (5th order), Notch 50 Hz8–12 Hzusing all
Table 3. Classification accuracy by the most represented brain lobe in each study.
Table 3. Classification accuracy by the most represented brain lobe in each study.
Brain LobeMean Main Lobe Electrode CoverageMean ACC (%)
Sensory/Motor Cortex [40,62,66,67,68,69,70,71,72]70%88.92%
Frontal [5,31,81,82]71%59.36%
Balanced [10,39,40,41,42,43,44,45]33.3%44.2%
Table 4. Best performing feature extraction methods and models across datasets.
Table 4. Best performing feature extraction methods and models across datasets.
DatasetHighest ACC Extraction TechniqueHighest ACC Feature TypeHighest ACC Model
By Torres-García et al. [34] used in [34,35,36,37,38]Discrete Wavelet Transform (DWT) Temporal–SpectralRandom Forest
Min et al. [57]mean, variance, standard deviation, and skewnessTemporalExtreme learning machine with a linear kernel
Noramiza Hashim et al. [5]Mel Frequency Cepstral Coefficients (MFCC)SpectralKNN
Advait Balaji et al. [31]Fast Fourier Transform (FFT)SpectralArtificial Neural Networks (ANN)
By Nguyen et al. [58] used in [36,37,58,59,60,61,62,63]Phase Lag Index (PLI), Intersite Phase Clustering (ISPC), Power Spectrum Analysis (PSA)SpectralConvolutional Neural Network with Transfer Learning (DenseNet-121)
Naveed et al. [64]Phase-only features (PoF) processed through Covariance Matrix (COV) and Maximum Linear Cross-Correlation (MaxCOR)SpatialExtreme Learning Machine
KARA ONE [65] used in [40,62,66,67,68,69,70,71,72]Mel-Frequency Cepstral Coefficients, Linear Predictive Cepstral Coefficients, and Sequency-Mapped Real Transform (MFCC, LPCC, SMRT).SpectralArtificial Neural Networks (ANN)
Coretto [73] used in [10,39,40,41,42,43,44,45]Power Spectrum Cross-Covariance Matrix using FFT (multiclass), Raw temporal EEG (vowels), Delay Differential Analysis (words)Spectral (multiclass), Temporal (vowels and words)Multimodal Neural Network (multiclass), Deep Reinforcement Learning (vowels), Dynamical Ergodicity Delay Differential Analysis + Support Vector Machine(words)
By Seo-Hyun Lee et al. used in [40,74,75,76,77]Raw temporal EEGTemporalDenoising Diffusion Probabilistic Model (DDPM) with Conditional Autoencoder (CAE) and Linear Classifier (LC)
Hernández-Del-Toro et al. [38]Discrete Wavelet Transform (DWT), Empirical Mode Decomposition (EMD), and Cleaned signal (post-preprocessing)Temporal–SpectralD2: Random Forest and D3: Support Vector Machine
Vorontsova et al. [28]Fast Fourier Transform (FFT)SpectralConvolutional Neural Network (ResNet18 + 2GRU)
Varshney et al. [55]Discrete wavelet transform (DWT)Temporal–SpectralSupport Vector Machine (SVM)
Rajdeep Ghosh et al. [47]Activity map (AM)Temporal–SpectralConvolutional neural network (CNN)
Dae-Hyeok Lee et al. [78]Raw temporal EEGTemporalConvolutional Neural Network with Autoencoder Architecture
Xin Zhang et al. [79]Noise-Assisted Multivariate Empirical Mode Decomposition (Noise-Assisted MEMD), followed by statistical feature extraction (mean, absolute mean, variance, standard deviation, skewness, kurtosis)SpectralMulti-Receptive Field Convolutional Neural Network (MRF-CNN)
Mingtao Li et al. [25]Riemannian manifold projection of covariance matricesSpatial Linear Discriminant Analysis (LDA)
Larocco et al. [32]Welch’s Power Spectral Density (PSD), Temporal Average, Percent IntensityTemporal–SpectralSupport Vector Machine (SVM)
DAIS [46]Raw temporal EEGTemporalConvolutional neural network (CNN)
Wu et al. [80]power spectral density (PSD)SpectralAdaptive Linear Discriminant Analysis classifier (LDA)
BCI2020 [81,82]Topographic brain mapsspatialThree-Dimensional Convolutional Neural Network with Long Short-Term Memory
(3DCNN-LSTM)
Tiwari et al. [83]Hilbert-Huang Transform (HHT)Temporal–spectralMulti-headed 1D Convolutional Neural Network (CNN)
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Estrella-Ibarra, L.F.; García-Noguez, L.R.; Pedraza-Ortega, J.C.; Ramos-Arreguín, J.M.; Tovar-Arriaga, S. Lost in Thought: An End-to-End Systematic Review on Imagined Speech Decoding Through Electroencephalographic Readings. AI 2026, 7, 75. https://doi.org/10.3390/ai7020075

AMA Style

Estrella-Ibarra LF, García-Noguez LR, Pedraza-Ortega JC, Ramos-Arreguín JM, Tovar-Arriaga S. Lost in Thought: An End-to-End Systematic Review on Imagined Speech Decoding Through Electroencephalographic Readings. AI. 2026; 7(2):75. https://doi.org/10.3390/ai7020075

Chicago/Turabian Style

Estrella-Ibarra, Luis Felipe, Luis Roberto García-Noguez, Jesús Carlos Pedraza-Ortega, Juan Manuel Ramos-Arreguín, and Saul Tovar-Arriaga. 2026. "Lost in Thought: An End-to-End Systematic Review on Imagined Speech Decoding Through Electroencephalographic Readings" AI 7, no. 2: 75. https://doi.org/10.3390/ai7020075

APA Style

Estrella-Ibarra, L. F., García-Noguez, L. R., Pedraza-Ortega, J. C., Ramos-Arreguín, J. M., & Tovar-Arriaga, S. (2026). Lost in Thought: An End-to-End Systematic Review on Imagined Speech Decoding Through Electroencephalographic Readings. AI, 7(2), 75. https://doi.org/10.3390/ai7020075

Article Metrics

Back to TopTop