Abstract
Steady-state visual evoked potential (SSVEP)-based brain–computer interfaces (BCIs) are consttablerained by three core challenges: data scarcity, individual variability, and temporal instability. To mitigate these issues, this study proposes an individual-adapted SSVEP decoding model that integrates hybrid augmentation, collaborative discrimination, and a Double-Area network within an ensemble learning framework, thus covering data expansion, feature discrimination, and multi-band harmonic extraction. The model enlarges the training set through hybrid augmentation using Source Aliasing Matrix Estimation (SAME) and Phase-Locked Time-Shift (PLTS), enhances feature robustness and class separability via collaborative discrimination using ensemble task-related component analysis with Kendall’s coefficient (eTRCA-K) and task-discriminant component analysis (TDCA), and effectively captures multi-band harmonics through a Double-Area network based on multi-reference least-squares transformations. Experimental results on the Benchmark dataset demonstrate that the proposed framework achieves strong adaptation capability and improved generalization, consistently outperforming existing methods across all data lengths and achieving an average accuracy of 88%. These findings highlight the model’s potential to further enhance SSVEP decoding performance and support high-performance target detection in BCI applications.
1. Introduction
As one of the cutting-edge directions, brain–computer interfaces (BCI) provide a direct communication pathway between the human brain and external devices without relying on muscles or the peripheral nervous system. Among various BCI techniques, electroencephalogram (EEG)-based BCI have attracted extensive attention due to their ease of use and low cost. EEG-based BCI encompass several experimental paradigms, including motor imagery (MI)-based BCI, event-related potentials (ERP)-based BCI, and steady-state visual evoked potential (SSVEP)-based BCI. Among them, SSVEP are elicited by periodic visual stimuli (e.g., flickering lights or oscillatory patterns) and exhibit strong stability, ease of detection, high information transfer rate (ITR), and minimal training requirements. Consequently, SSVEP-based BCI have demonstrated irreplaceable practical value in applications such as assistive communication, medical rehabilitation, and intelligent interaction.
SSVEP recognition algorithms in their early stage primarily relied on short-time Fourier transform (STFT) techniques to extract spectral features for identifying frequency components. However, these approaches exhibited strong dependence on data length, and overly short segments failed to guarantee reliable accuracy. As a result, STFT-based methods were gradually replaced by algorithms with higher computational efficiency and better short-term adaptability. In 2007, Lin et al. [1] introduced the canonical correlation analysis (CCA) method, which achieved promising performance in SSVEP recognition. CCA identifies the target stimulus by computing the correlation between multichannel EEG signals and sinusoidal reference signals corresponding to different stimulus frequencies. Given its notable performance, subsequent studies explored various improvements to CCA. For instance, Chen et al. [2] developed the filter bank CCA (FBCCA), which decomposes EEG signals into multiple sub-band components using a set of bandpass filters and then applies CCA to each sub-band, significantly enhancing SSVEP decoding accuracy. Zhang et al. [3] further improved CCA by training subject-specific spatial filters to better distinguish target-related signals from background noise.
Nakanishi et al. [4] proposed task-related component analysis (TRCA), which constructs spatial filters by identifying the optimal covariance structure across multiple trials, enabling extraction of task-relevant components. Owing to its high accuracy and information transfer rate (ITR), TRCA has received widespread attention. Building upon TRCA, many studies have introduced extensions and improvements. For example, Wong et al. [5] proposed a cross-stimulus learning approach capable of learning target-related features while generalizing to other stimuli, thereby enhancing the adaptability and robustness of TRCA. Based on the observation that “under the same mixing coefficients, spatial filters for different stimuli exhibit similarity,” Xu et al. [6] introduced the ensemble task-related component analysis (eTRCA). Compared with TRCA, eTRCA demonstrates better classification performance across multi-frequency filters. To address TRCA’s insufficient utilization of temporal information and redundancy in filter banks, Liu et al. [7] proposed the task discriminant component analysis (TDCA) framework, which further improves the overall performance and applicability of SSVEP-based BCI.
In recent years, with the rapid advancement of deep learning in image, speech, and text processing, neural network-based deep learning approaches have also gained significant traction in SSVEP decoding. Convolutional neural networks (CNN) remain the most widely used deep learning models for SSVEP decoding [8]. For example, Lawhern et al. [9] proposed EEGNet, a compact CNN employing depthwise separable convolutions to construct an EEG-specific architecture that integrates spatial and temporal filtering mechanisms, achieving substantial performance improvements in multiclass EEG classification tasks. Subsequently, Waytowich et al. [10] refined EEGNet and applied it to asynchronous SSVEP-based BCI systems. Guney et al. [11] introduced a deep neural network framework tailored for SSVEP sub-band processing. By performing convolutions across harmonic sub-bands, channels, and temporal dimensions, followed by classification via fully connected layers, the framework significantly improved ITR. Pan et al. [12] proposed SSVEPNet for frequency classification, combining 1D convolution and LSTM modules along with spectral normalization and label smoothing techniques, achieving strong performance. In summary, both traditional and deep learning-based methods have achieved considerable progress in SSVEP decoding. However, these approaches typically rely on a single classification model, which limits generalization capability and poses higher risks of overfitting and reduced stability.
To address the aforementioned issues, Deng et al. [13] proposed an algorithm named TRCA-Net, which integrates TRCA with a CNN to enhance the classification performance, stability, and applicability of SSVEP signals. Li et al. [14] introduced Dis-ComNet, a network capable of extracting complex spatiotemporal features from new data and achieving high-accuracy classification by calibrating network parameters using only a very small number of samples. Deng et al. [15] proposed OS-SSVEP, a network that combines multiple SSVEP decoding approaches and employs data augmentation-based parameter calibration with limited samples, significantly improving decoding accuracy and generalization capability. The integration of multiple SSVEP decoding strategies and the use of few-sample parameter calibration have demonstrated strong potential in improving model generalization and applicability. However, the data augmentation strategies used in the above calibration processes lack extraction of temporal dynamic features, limiting the ability of the constructed models to capture the steady-state characteristics of SSVEP signals as they evolve over time.
To address the above issues, this study proposes a hybrid augmentation framework based on Source Aliasing Matrix Estimation (SAME) [16] and Phase-Locked Time-Shift (PLTS) [17], combined with a multi-path ensemble strategy that integrates traditional approaches with deep learning methods to enhance SSVEP decoding performance, robustness, and individual adaptability. The main contributions of this work are as follows:
- (1)
- This paper presents a five-path ensemble SSVEP decoding framework that leverages the complementarity between deep learning and traditional techniques: the Double-Area network, SAME + eTRCA-k, SAME + TDCA, PLTS + eTRCA-k, and PLTS + TDCA. Among these, the Double-Area network adopts a three-level sub-band partitioning scheme to ensure complete harmonic coverage of SSVEP. Through convolutional filtering operations combined with multi-reference least-squares pathways, the model integrates nonlinear and linear features, thereby improving the decoding performance.
- (2)
- This study employs a hybrid augmentation strategy based on SAME and PLTS to address the challenge of limited calibration data. SAME generates augmented samples through alias matrix mapping and adaptive noise generation, enabling precise preservation of individual static characteristics (e.g., waveform and amplitude), ensuring that subject-specific response signatures are maintained. PLTS exploits the periodic nature of SSVEP to generate temporally dynamic samples via optimal temporal-window shifts, supplementing temporal information. The combination of SAME and PLTS simultaneously maintains individual specificity and temporal stability, producing augmented data that more closely resemble real SSVEP responses and ensuring effective calibration even with few samples.
- (3)
- Within the ensemble learning framework, this study incorporates collaborative discrimination strategies using eTRCA with Kendall’s coefficient (eTRCA-k) and TDCA. eTRCA-k enhances the accuracy of nonlinear harmonic associations across individuals by introducing Kendall’s coefficient, reducing the influence of outliers on feature matching. TDCA constructs a filtering matrix by maximizing inter-class differences and minimizing intra-class variance, strengthening the extraction of individual-specific features such as harmonic distribution patterns. The collaborative discrimination of eTRCA-k and TDCA focuses on fundamental SSVEP properties—multi-harmonic structure, strong periodicity, and substantial inter-individual variability—thereby improving overall decoding accuracy.
2. Materials and Methods
2.1. Benchmark Dataset
This dataset contains SSVEP-based BCI recordings from 35 healthy subjects. The experimental paradigm includes 40 stimulus targets, with frequency intervals of 0.2 Hz between adjacent targets. For each subject, the experiment consists of six blocks, each containing 40 trials (corresponding to the 40 characters presented in random order). Each trial lasts 6 s, including 0.5 s of cue presentation, 5 s of fixed-frequency stimulation corresponding to the target character, and 0.5 s of rest. EEG signals were collected using a 64-channel SynAmps2 system (Neuroscan Inc., El Paso, TX, USA) at an initial sampling rate of 1000 Hz and later downsampled to 250 Hz. More details can be found in reference [18].
2.2. BETA Dataset
This dataset contains SSVEP-based BCI recordings from 70 healthy subjects. The 64-channel EEG data were recorded by SynAmps2 (Neuroscan Inc., El Paso, TX, USA). The sampling rate was set 1000 Hz, and the pass-band of the hardware filter was 0.15–200 Hz. The number of stimulus targets, frequency range, and inter-frequency spacing are consistent with those of the Benchmark dataset. It is important to note that this dataset includes four blocks. For the first 15 subjects, each trial lasts 3 s (0.5 s for cue presentation, 2 s of fixed-frequency stimulation, and 0.5 s of rest). For the remaining 55 subjects, each trial lasts 4 s (0.5 s for cue presentation, 3 s of fixed-frequency stimulation, and 0.5 s of rest). Additional details can be found in reference [19].
2.3. Data Preprocessing
Following previous studies [18], nine channels highly relevant to SSVEP recognition were selected: Pz, PO5, PO3, POz, PO4, PO6, O1, Oz, and O2. Considering the visual latency of SSVEP responses, EEG analysis was conducted using data starting 140 ms after stimulus onset (corresponding to a visual latency of 140 ms).
2.4. Proposed Method
As shown in Figure 1, the proposed procedure consists of the following six steps:
Figure 1.
Technical roadmap of the proposed method.
Step 1: Data Splitting. For dataset partitioning, a leave-one-subject-out cross-validation scheme is employed. Using the Benchmark dataset as an example, which includes 35 subjects, data from 34 subjects are used as the source-domain data in each iteration, while the remaining subject serves as the target-domain data. This procedure is repeated 35 times to ensure that every subject is used exactly once as the target-domain subject.
Step 2: Training the Double-Area network using source-domain data. The Double-Area network is designed to exploit the complementary characteristics of SSVEP representations, including multi-band harmonic information and stimulus-related periodic and phase-locked characteristics. Accordingly, the input EEG is processed through two parallel pathways with different representation mechanisms.
The first pathway is the Filter Path. SSVEP responses contain discriminative information not only at the fundamental stimulus frequency but also at multiple harmonic frequencies, and the relative strengths of these components may vary across participants and data lengths. Therefore, using only a single frequency band may not sufficiently exploit the complementary harmonic information. To address this issue, the Filter Path decomposes the EEG signals into three overlapping frequency bands (5–90 Hz, 14–90 Hz, and 22–90 Hz). The 5–90 Hz band preserves the fundamental and broad harmonic components, whereas the 14–90 Hz and 22–90 Hz bands progressively emphasize higher-frequency harmonic information. This hierarchical filter-bank design enables the network to learn complementary representations from different harmonic ranges rather than treating the entire spectrum uniformly.
After bandpass filtering, spatial features are extracted using 120 convolution kernels with size of Nc × 1, where Nc denotes the number of EEG channels. Subsequently, convolutional layers with kernel sizes 1 × 2 and 1 × 10 are applied to capture temporal-scale features, and the feature representations from the different sub-bands are then fused. The kernel size Nc × 1 indicates that spatial filtering is performed across all EEG channels, while the 1 × 2 and 1 × 10 kernels represent temporal convolution operations with different receptive fields. The number of convolution kernels (120) denotes the number of feature maps generated in each convolutional layer.
The second pathway is the Multi-Reference Least-Squares (MRLS) Path. Unlike the Filter Path, which learns data-driven representations from multiple harmonic bands, the MRLS Path incorporates prior knowledge of the periodic and phase-locked properties of SSVEP signals. Because SSVEP responses are synchronized with the visual stimulation, class-specific sine–cosine templates provide explicit references for the expected periodic response patterns. The MRLS transformation maps the EEG signals into this stimulus-related reference space through least-squares fitting, thereby emphasizing components consistent with the stimulus frequency and phase while reducing the influence of unrelated background activity. This reference-guided linear representation complements the nonlinear features learned by the Filter Path and provides additional stability when only limited calibration data are available.
The least-squares transform is a commonly used statistical method that constructs a linear relationship model between variables. Its primary objective is to estimate the parameters of the regression model by minimizing the sum of squared residuals, thereby quantifying the discrepancy between predicted values and actual observations. When and represent two two-dimensional variables, the goal of the least-squares transform is to determine a transformation matrix by solving Equation (1):
The solution to Equation (1) can be obtained using Equation (2):
In this study, the Multi-Reference Least-Squares Transform is applied to each trial corresponding to each stimulus to generate a new multi-channel signal. Let denote the i-th trial of the t stimulus (corresponding to the i-th block of data for the t-th character), and let represent the sine–cosine template associated with the stimulus frequency of the t-th stimulus. The objective of the MRLS Transform (MRLST) is to map the SSVEP signals from different subjects onto the sine–cosine template using Equation (3).
represents the average signal across all trials of the t-th stimulus. The sine–cosine template corresponding to the t-th stimulus can be obtained using Equation (4).
Here, denotes the number of harmonics (dimensionless), denotes the frequency of the t-th stimulus (Hz), denotes the stimulus phase (rad), denotes the number of sampling points (points), and denotes the sampling rate (Hz). Once is obtained, the transformation is applied to each trial signal under each stimulus frequency to generate new data, as shown in Equation (5).
For example, in the Benchmark dataset, where there are 40 stimulus classes, 40 corresponding sine–cosine templates are constructed. The data processed by the MRLST are then passed through a convolutional layer with kernel size 1 × 1 for initial feature extraction, followed by a convolutional layer with 120 kernels (size = 2Nh × 1) to extract spatial information across channels, and finally through convolutional layers with 120 kernels of size 1 × 2 and 1 × 10 to capture temporal-scale features. To better leverage the information from both the filter path and the MRLS path, this study fuses the outputs of the first, second, and third convolutional layers from the Filter Path with the outputs of the second, third, and fourth convolutional layers from the MRLS path, and applies an additional convolution to the fused features. Ultimately, the outputs from the two paths are fed into fully connected layers for classification, and the final prediction is obtained by computing a weighted sum of their classification scores.
Step 3: Target-Domain Data Partitioning and Labeled-Data Augmentation. The target-domain data of each subject are first divided into Calibration data and Independent Test data. For each stimulus class, only one sample is selected as Calibration data, while all remaining samples are treated as Independent Test data (for example, in the Benchmark target-domain dataset, each class contains 1 Calibration sample and 5 Test samples). This process is iterated six times to perform cross-validation, ensuring that each sample serves as Calibration data exactly once. The Calibration data are then augmented using the SAME augmentation and PLTS augmentation methods. The detailed procedure is as follows:
- (1)
- SAME Augmentation
For the Calibration sample of a given stimulus class and the corresponding sine–cosine template (with t denoting the stimulus class), the transformation matrix is estimated by minimizing the mean squared error, as shown in Equation (6):
where denotes the Frobenius norm. The solution to Equation (6) can be obtained by solving a least-squares problem The artificially reconstructed template signal is defined in Equation (7):
To generate multiple augmented signals, random noise is added to . To ensure that the noise intensity added to each channel is proportional to the signal strength of , let denote the signal of the c channel of and denote the variance of the c-th channel signal. The covariance matrix of the noise can then be computed as shown in Equation (8):
where diag( ) denotes arranging the elements along the diagonal, and Nc represents the number of channels. Finally, the artificially generated signals are produced according to Equation (9):
denotes the number of generated signals. The generated follow a Gaussian distribution [17] with a zero mean vector and a covariance matrix , . The parameter is used to control the noise intensity and is set to 0.05.
- (2)
- PLTS Augmentation
To fully exploit the experimental data, in addition to the original samples (extracted from time d with a duration of seconds), we assume that the data within each trial’s temporal window (where denotes the end time of the trial) can also be used for training the SSVEP model. Since subjects continuously fixate on the flickering stimulus throughout the entire trial, SSVEP responses are persistently elicited during this period. Specifically, the data extracted from any sub-interval (where denotes the starting point and denotes the ending point for generating new samples) within the trial should exhibit similar SSVEP characteristics with and can therefore serve as additional training samples. The optimal sub-interval for generating new data can be determined by locating the sampling point a that yields the maximum similarity, as defined in Equation (10).
The similarity is measured by the correlation between and , as defined in Equation (11).
Here, and denote the mean values of and , respectively, and represents the number of selected data points (points). A small positive constant is added to each variance term in the denominator to avoid division by zero when either input signal has zero or extremely small variance. Using a grid-search strategy, the starting point a is iteratively assigned values from the range . For each candidate a, the similarity between and is computed, and the sampling position that yields the highest similarity is chosen as the optimal a.
Based on SSVEP, the similarity between newly extracted samples and the original data exhibits a periodic pattern as the sampling position a varies, and this period equals the reciprocal of the stimulus frequency of the trial [17]. Therefore, for the n class of trial data (with a stimulus frequency of Hz), in addition to the original samples, time-shift-based augmentation can be employed to generate additional samples. The step size for this augmentation is computed as shown in Equation (12).
The newly generated samples extracted within the interval of seconds can thus be used as additional training data. Accordingly, the sampling positions a are computed as defined in Equation (13).
Here, denotes the number of generated signals.
Step 4: The datasets enhanced by the different methods are then separately used to train the eTRCA-K and TDCA models.
- (1)
- The training procedure of eTRCA-K is generally consistent with that of TRCA and it yields the average signal for each stimulus class. The main difference is that eTRCA-K integrates multiple spatial filters trained by TRCA into , and employs Kendall correlation analysis instead of canonical correlation analysis for classification.
- (2)
- The training process of the TDCA model follows the description in reference [7]. After TDCA training, the optimal template parameter Q and the center for each stimulus frequency are obtained.
Step 5: Classification of the Test Data. For the test data, classification is performed through five pathways: the trained Double-Area fusion network, SAME + eTRCA-K, SAME + TDCA, PLTS + eTRCA-K, and PLTS + TDCA. The classification procedures of eTRCA-K and TDCA are described as follows:
- (1)
- eTRCA-K
The correlation coefficient between the test data and the average training signal of each class is computed as defined in Equation (14).
Here, denotes the two-dimensional correlation analysis between the two signals, where the Kendall rank correlation coefficient is used.
- (2)
- TDCA
For the test data , the Pearson correlation coefficient between the sample and the class-specific center is computed, as given in Equation (15).
Step 6: Fusion and Classification of the Five Pathways. Let the feature vector of the k-th pathway be , where k = 1, 2, …5 corresponding to the outputs of the Double-Area network, SAME + eTRCA-K, SAME + TDCA, PLTS + eTRCA-K, and PLTS + TDCA, respectively. The fusion of the five pathways proceeds as follows: each output is first linearly normalized, as defined in Equation (16).
Here, and denote the minimum and maximum values of , respectively, and t represents the t-th stimulus class. A small positive constant is used to avoid division by zero when either input signal has zero or extremely small variance. The normalized feature vectors from the five pathways are then fused, as defined in Equation (17).
The class t corresponding to the maximum value of is identified as the predicted class of the test sample. Following score normalization, the group-balanced fusion ensures comparable contributions from the different pathways without introducing additional learnable parameters.
2.5. Evaluation Metrics and Baseline Methods
- (1)
- Evaluation Metrics
In this study, the performance of SSVEP recognition is evaluated primarily using two metrics: classification accuracy and ITR. Accuracy (ACC) represents the percentage of correctly identified stimulus classes relative to the total number of stimuli. ITR is the most commonly used performance metric in SSVEP, and is computed as shown in Equation (18).
In the formula, N denotes the number of stimulus frequencies, P is the average classification accuracy, and T is the average time (in seconds) required for recognition. In this study, the average time T includes a gaze shift interval of 0.5 s.
- (2)
- Baseline Methods
To evaluate the performance of the proposed method, it is compared with seven state-of-the-art SSVEP decoding algorithms: msCCA [5], FBCCA [2], TRCA [4], LST [20], Compact CNN [21], Ensemble DNN [22], and OS-SSVEP [15].
2.6. Experimental Settings
All experiments were conducted on a computer equipped with an AMD Ryzen 7 5800H CPU, 16 GB of system memory, and an NVIDIA GeForce RTX 3050 Laptop GPU with 4 GB of video memory. The operating system was 64-bit Windows 10. The software environment consisted of Python 3.10.0, PyTorch 1.12.1, CUDA 11.3, and cuDNN 8.3.2. Deep learning models were evaluated in FP32 on the GPU, whereas the conventional matrix-based methods were executed on the CPU.
- (1)
- Computational Complexity Analysis
As shown in Table 1, the proposed method contains 0.642 M trainable parameters and requires 108.866 M neural-branch FLOPs per trial. Its standardized fitting/training time and inference latency were 0.921 s and 356.377 ms per trial, respectively, which were higher than those of OS-SSVEP (0.527 s and 177.235 ms). This additional cost mainly arises from combining the complementary SAME and PLTS decision pathways. Nevertheless, the inference latency remained below the 0.5-s EEG acquisition window, and the proposed method required substantially fewer stored parameters than Ensemble DNN (0.642 M versus 15.881 M). Together with the numerically highest classification accuracy at all six data lengths, these results indicate that the proposed method favors decoding performance and complementary feature utilization rather than minimum computational cost, yielding a modest but consistent performance improvement at the expense of additional inference time.
Table 1.
Computational complexity comparison of the proposed method and baseline methods on the Benchmark dataset.
- (2)
- Sensitivity analysis
A preliminary local sensitivity analysis was conducted on the Benchmark dataset using data lengths of 0.5, 0.8, and 1.0 s. In the sensitivity analysis, the noise-intensity coefficient was varied among 0.03, 0.05, and 0.07; the number of SAME-generated samples among 1, 3, and 5; and the number of PLTS-generated samples among 4, 6, and 8. The default settings were a SAME noise-intensity coefficient of 0.05, three SAME-generated samples, and six PLTS-generated samples. Only one parameter was varied at a time, while the remaining parameters were fixed at their default values. The same trained model, data partition, and random seed were maintained across all configurations. The observed within-condition performance variations were 1.50, 0.50, and 0.50 percentage points for data lengths of 0.5, 0.8, and 1.0 s, respectively. Thus, the maximum variation did not exceed 1.50 percentage points, indicating that the proposed framework was relatively stable within the evaluated parameter ranges.
3. Results
3.1. Comparison Between the Proposed Model and State-of-the-Art Methods
As shown in Table 2, for the Benchmark dataset, at data lengths of 0.5–0.6 s, the classification accuracy ranks from highest to lowest as follows: the proposed method, OS-SSVEP, TRCA, Ensemble DNN, msCCA, LST, Compact CNN, and FBCCA; at data lengths of 0.7–0.9 s, the ranking is as follows: the proposed method, OS-SSVEP, TRCA, Ensemble DNN, LST, msCCA, Compact CNN, and FBCCA; at a data length of 1 s, the ranking is as follows: the proposed method, OS-SSVEP, TRCA, LST, Ensemble DNN, msCCA, FBCCA, and Compact CNN. The proposed method achieves the highest ACC on the Benchmark dataset, and all other methods are significantly different from it across all data lengths.
Table 2.
Comparison of the ACC (%) on the Benchmark dataset.
As shown in Table 3, for the Benchmark dataset, at data lengths of 0.5–0.6 s, the ITR ranks from highest to lowest as follows: the proposed method, OS-SSVEP, TRCA, Ensemble DNN, msCCA, LST, Compact CNN, and FBCCA; at 0.7 s, the ranking is as follows: the proposed method, OS-SSVEP, TRCA, Ensemble DNN, LST, msCCA, Compact CNN, and FBCCA; at 0.8–0.9 s, the ranking is as follows: the proposed method, OS-SSVEP, TRCA, LST, Ensemble DNN, msCCA, Compact CNN, and FBCCA; at 1 s, the ranking is as follows: the proposed method, OS-SSVEP, TRCA, LST, Ensemble DNN, msCCA, FBCCA, and Compact CNN. These results indicate that the proposed method achieves the highest accuracy and ITR across all data lengths. A similar pattern is observed for ITR; the proposed method obtains the highest ITR on the Benchmark dataset.
Table 3.
Comparison of ITR (bits/min) on the Benchmark dataset.
3.2. Effect of Data Length on Different Models
- (1)
- Comparison of Models at Different Data Lengths
As shown in Figure 2, for all methods, classification accuracy increases with data length, which is consistent with the decoding characteristics of SSVEP. For the Benchmark dataset, considering the concentration of accuracy values over the range of 0.5–1 s, the proposed method exhibits the most concentrated accuracy across all data lengths, followed by OS-SSVEP, while FBCCA performs the worst. Furthermore, the difference in accuracy between the 1 s and 0.5 s data from largest to smallest is as follows: FBCCA (44.7), LST (39.96), Compact CNN (30.98), msCCA (27.76), TRCA (26.74), Ensemble DNN (26.29), OS-SSVEP (17.52), and the proposed method (15.48). This demonstrates that the impact of data length on the accuracy of the proposed method is minimal.
Figure 2.
Comparison of the ACC (%) of different methods on the Benchmark dataset.
As shown in Table 3, for all methods, the ITR generally increases and then decreases with data length (although this trend is not fully observed for some methods within the 0.5–1 s range), which is consistent with the periodic steady-state characteristics of SSVEP responses. For the Benchmark dataset, as illustrated in Figure 3, the data lengths at which the maximum ITR is achieved for the proposed method, OS-SSVEP, TRCA, Ensemble DNN, LST, msCCA, FBCCA, and Compact CNN are 0.6 s, 0.7 s, 0.7 s, 0.9 s, 0.9 s, 0.9 s, 1 s, and 1 s, respectively. These results indicate that the proposed method reaches the highest ITR in the shortest time, followed by OS-SSVEP and TRCA, then Ensemble DNN, LST, and msCCA, and finally FBCCA and Compact CNN.
Figure 3.
Comparison of ITR (bits/min) achieved by different methods on the Benchmark dataset.
- (2)
- Comparison of Single-Branch Pathways at Different Data Lengths
As shown in Table 4, the proposed method clearly outperforms single-pathway classification. On the Benchmark dataset, the relative importance of the five pathways varies across different data lengths. At 0.5 s and 0.6 s, the classification accuracy ranks from highest to lowest as follows: Double-Area network, TDCA with PLTS, TDCA with SAME, eTRCA-K with PLTS, and eTRCA-K with SAME; at 0.7 s: Double-Area network, TDCA with SAME, TDCA with PLTS, eTRCA-K with PLTS, and eTRCA-K with SAME; at 0.8 s: TDCA with SAME, TDCA with PLTS, eTRCA-K with SAME, Double-Area network, and eTRCA-K with PLTS; at 0.9 s: TDCA with SAME, TDCA with PLTS, Double-Area network, eTRCA-K with SAME, and eTRCA-K with PLTS; at 1.0 s: TDCA with SAME, TDCA with PLTS, eTRCA-K with SAME, Double-Area network, and eTRCA-K with PLTS. These results indicate that different pathways are suited to different data lengths. For shorter data segments, the Double-Area network contributes more significantly to decoding, whereas for longer segments, the importance of SAME-based augmentation becomes more pronounced. Overall, TDCA exhibits substantially better classification performance than eTRCA-K. The proposed method significantly outperforms every single-branch pathway in ACC across all data lengths on the Benchmark dataset.
Table 4.
ACC (%) of different single-branch pathways on the Benchmark dataset.
As shown in Table 5, the ITR achieved by the proposed method is clearly superior to that of single-pathway classification. On the Benchmark dataset, the relative importance of the five pathways varies with data length. At 0.5 s and 0.6 s, the ITR ranks from highest to lowest as follows: Double-Area network, TDCA + PLTS, TDCA + SAME, eTRCA-K + PLTS, and eTRCA-K + SAME; at 0.7 s: Double-Area network, TDCA + SAME, TDCA + PLTS, eTRCA-K + PLTS, and eTRCA-K + SAME; at 0.8 s: TDCA + SAME, TDCA + PLTS, eTRCA-K + SAME, Double-Area network, and eTRCA-K + PLTS; at 0.9 s: TDCA + SAME, TDCA + PLTS, Double-Area network, eTRCA-K + SAME, and eTRCA-K + PLTS; at 1.0 s: TDCA + SAME, TDCA + PLTS, eTRCA-K + SAME, Double-Area network, and eTRCA-K + PLTS. These results indicate that different pathways are suitable for different data lengths. For shorter data segments, the Double-Area network achieves better decoding performance, whereas for longer segments, decoding performance after SAME-based augmentation is superior. Overall, TDCA exhibits significantly better classification performance than eTRCA-K. The proposed method significantly improves ITR over all single-branch pathways across all data lengths, confirming the effectiveness of multi-pathway fusion.
Table 5.
ITR (bits/min) of different single-branch pathways on the Benchmark dataset.
3.3. Robustness of the Proposed Method Across Datasets
For the BETA dataset, as shown in Table 6 and Table 7, the proposed method achieves the highest classification accuracy and ITR across data lengths from 0.5 s to 1 s. As illustrated in Figure 4 and Figure 5, the impact of data length on SSVEP decoding is minimal for the proposed method (with the most concentrated results, see Figure 4), and the shortest data length required to achieve the optimal ITR is 0.6 s (see Figure 5). On the BETA dataset, the proposed method achieves the highest ACC and ITR across all data lengths, with significant differences from all baseline methods.
Table 6.
Comparison of the ACC (%) achieved by different methods on the BETA dataset.
Table 7.
Comparison of the ITR (bits/min) achieved by different methods on the BETA dataset.
Figure 4.
Comparison of ACC (%) of 8 different methods on the BETA dataset.
Figure 5.
Comparison of ITR (bits/min) of 8 different methods on the BETA dataset.
As shown in Figure 6, for the proposed method, the difference in decoding accuracy (Benchmark-Beat/Benchmark) between the two datasets at data lengths of 0.5 s, 0.6 s, 0.7 s, 0.8 s, 0.9 s, and 1 s is 16.19%, 16.23%, 15.95%, 15.84%, 14.97%, and 14.06%, respectively. For OS-SSVEP, the corresponding differences are 16.45%, 16.85%, 16.08%, 16.20%, 15.41%, and 14.66%; for Ensemble DNN: 11.70%, 13.79%, 16.02%, 17.91%, 18.15%, and 17.37%; for Compact CNN: 13.87%, 18.59%, 16.01%, 18.59%, 22.32%, and 23.79%; for LST: 26.67%, 30.74%, 33.46%, 33.67%, 29.89%, and 30.21%; for TRCA: 14.57%, 16.53%, 21.33%, 19.25%, 19.45%, and 19.08%; for FBCCA: 11.81%, 4.44%, 4.76%, 6.76%, 9.86%, and 11.76%; and for msCCA: 16.29%, 17.82%, 20.06%, 22.25%, 22.70%, and 22.85%. From these results, it can be observed that for data lengths greater than 0.7 s, the accuracy difference between the two datasets is the second smallest for the proposed method (the smallest is FBCCA, which has extremely low decoding accuracy and limited practical applicability). The same approach was applied to ITR calculation, yielding results consistent with the accuracy comparison, thereby demonstrating the robustness of the proposed method for cross-dataset SSVEP decoding.
Figure 6.
Comparison of decoding accuracy across different data lengths on the Benchmark and BETA datasets.
3.4. Ablation Study
As shown in Table 8, on the Benchmark dataset, for longer data lengths (1 s), removing the Double-Area network module, eTRCA-K + SAME, eTRCA-K + PLTS, TDCA + SAME, and TDCA + PLTS resulted in a decrease in classification accuracy from 94.39% to 91.82% (−2.57%), 94.21% (−0.18%), 94.18% (−0.21%), 94.30% (−0.09%), and 94.27% (−0.12%), respectively. For shorter data lengths (e.g., 0.6 s), accuracy dropped from 84.22% to 77.41% (−6.81%), 83.84% (−0.39%), 83.92% (−0.30%), 84.07% (−0.15%), and 83.91% (−0.31%), indicating that the Double-Area network has the greatest impact on overall model performance, especially for shorter data lengths.
Table 8.
ACC (%) in the ablation study on the Benchmark dataset.
From the perspective of data augmentation, removing PLTS-related modules (eTRCA-K + PLTS and TDCA + PLTS) led to a decrease in accuracy from 94.39% to 94.10% (−0.29%) at 1 s and from 84.22% to 83.15% (−1.07%) at 0.6 s. Removing SAME-related modules (eTRCA-K + SAME and TDCA + SAME) resulted in performance reductions from 94.39% to 93.71% (−0.68%) at 1 s and from 84.22% to 82.93% (−1.29%) at 0.6 s, suggesting that, among individual augmentation modules, SAME has a greater overall impact on model performance than PLTS.
The ITR results presented in Table 9 are consistent with those in Table 7 and are not repeated here.
Table 9.
ITR (bits/min) in the ablation study on the Benchmark dataset.
4. Discussion
This study presents an elegant framework that integrates five SSVEP decoding pathways, each with distinct strengths and complementary advantages. The framework leverages the strengths of different pathways under varying data conditions (e.g., data length, signal-to-noise ratio) while mitigating their weaknesses, achieving high-performance, robust, and data-length-insensitive SSVEP decoding.
4.1. Advantages of the Proposed Method
The proposed method does not simply stack similar models; rather, it creatively fuses data-driven deep learning models (Double-Area network) with traditional discriminative models (eTRCA-K and TDCA). The Double-Area network comprises a Filter Path and a MRLS path. The Filter Path employs multiple parallel band-pass filters covering different frequency bands where SSVEP signals may occur, ensuring comprehensive extraction of frequency-domain features. The MRLS path directly uses sine–cosine templates as references for transformation, projecting the data toward ideal SSVEP response patterns and effectively highlighting phase-locked information related to the stimulus frequency. By fusing features from different layers of the two paths, the Double-Area network can simultaneously exploit low-level (detailed) and high-level (abstract) information, enriching the hierarchical representation of features. Consequently, the Double-Area network can automatically learn complex, high-level nonlinear features from raw data without relying on strong assumptions, offering greater adaptability. In contrast, TDCA and eTRCA-K are based on traditional signal correlation and spatial filtering techniques, with clear underlying principles, and are highly effective when the data meet their model assumptions. These two types of discriminative paradigms are fundamentally complementary: when one model underperforms due to specific data characteristics, the other can still provide stable performance, thereby ensuring the overall robustness of the system.
Of course, the high performance achieved in this study also relies on a dual-pronged data augmentation strategy. To address the critical issue of having very few calibration samples in the target domain, two augmentation methods are introduced. SAME augmentation effectively expands the training dataset by adding noise to the template signals that matches the channel signals, enhancing noise robustness and preventing the model from overfitting to the limited samples. PLTS augmentation cleverly exploits the periodic nature of SSVEP signals to generate multiple homologous training samples from a single trial. This not only increases the data volume but allows the model to learn the temporal dynamics and stability (sequential characteristics) of SSVEP signals, thereby enhancing temporal generalization capability.
After applying SAME or PLTS augmentation, TDCA’s inter-class optimization strategy converts the supplemented temporal dynamic features into class-discriminative information, enabling the synergistic use of dynamic information and class boundaries, and addressing issues of insufficient feature dimensionality and class overlap. The eTRCA-K replaces the traditional correlation coefficient with Kendall correlation analysis, which is more sensitive to nonlinear temporal correlations. Compared with standard eTRCA, this modification leads to improved SSVEP decoding performance (see Figure 7).
Figure 7.
Comparison of eTRCA and eTRCA-K. The shaded regions represent the standard deviation. The asterisks in the figure indicate significant differences determined by Paired T-tests (* denotes p < 0.05, ** denotes p < 0.01, *** denotes p < 0.001).
4.2. Impact of Data Length on SSVEP Decoding
One of the most prominent advantages of the proposed method is its insensitivity to data length, which mainly arises from the multi-level, multi-mechanism synergistic effect. First, the Double-Area network, trained end-to-end, can learn to extract highly discriminative micro-patterns or instantaneous phase features from limited data points, maintaining good performance even with short data segments. Within the Double-Area network, the MRLST transformation is based on least-squares fitting to the ideal template. Even with short data segments, as long as the basic response pattern is present, this transformation can effectively separate signal from noise and map it to the template space, providing fundamental stability.
Secondly, methods such as eTRCA-K and TDCA rely on inter-signal correlations, and longer data provide more cycles and more stable statistical properties, making the computed signal templates and correlation measures more reliable, thereby significantly improving performance (as shown in Figure 8).
Figure 8.
Classification ACC (%) and ITR (bits/min) achieved by single-branch pathways on the Benchmark dataset.
Thirdly, the SAME and PLTS augmentation strategies reduce interference from random neural activity. SAME first computes the average SSVEP template from the original calibration data and then estimates the mixing matrix based on this template. With medium- to long-length data, the template SNR is higher and the estimation error of the mixing matrix is smaller, allowing the generated artificial EEG signals to accurately match the frequency and phase characteristics of real SSVEP signals, avoiding feature distortion caused by cycle truncation. The main challenge with short- to medium-length data is the complex variation in temporal dynamics and lower robustness, coupled with insufficient individual specificity, which can lead to matching bias. PLTS extracts the time segments with the highest similarity via optimal temporal shifts, supplementing missing temporal dynamics in short- to medium-length data. The synergistic effect of SAME and PLTS thus reduces interference and enhances the feature representation used for classification.
4.3. Analysis of Model Robustness and Generalization
In terms of handling noise and data uncertainty, SAME augmentation introduces noise designed according to the channel variance, exposing the model to simulated realistic noise during training and thereby enhancing its robustness. From the perspective of multi-pathway fusion, even if one or several branches make incorrect predictions due to noise interference, the final fused output can still be correct as long as the other branches make accurate judgments. This greatly reduces the probability of failure caused by local issues.
Regarding information utilization, PLTS and SAME augmentation fully exploit the value of a very small amount of labeled data, mitigating overfitting in small-sample learning and ensuring that the model is adequately trained in the target domain. These factors collectively enable the proposed method to demonstrate stronger robustness when confronted with noise and uncertain data.
5. Limitations and Future Work
Although the proposed method was evaluated on the Benchmark and BETA datasets, both use closely related 40-target joint frequency-phase modulation paradigms. Therefore, the current results demonstrate robustness across participant cohorts and acquisition settings within similar paradigms, rather than broad generalization to substantially different BCI configurations. Moreover, phase coding was not treated as an independent experimental factor. Future studies will evaluate the framework using more diverse target numbers, controlled frequency-phase combinations, stimulation protocols, acquisition devices, and hybrid or multimodal paradigms.
The present evaluation was based on offline public datasets collected under controlled conditions and therefore did not capture motion artifacts, environmental noise, attention fluctuations, inter-session variability, or other challenges of online BCI operation. Under the one-shot Benchmark protocol, the experimental calibration procedure for each stimulus class requires approximately 6 s. Therefore, the total calibration time increases approximately linearly with the number of stimulus classes, reaching approximately 240 s (4 min) when all 40 classes are included. However, electrode preparation, subjective fatigue, visual discomfort, and user acceptance were not measured. Future work will conduct online and session-level usability evaluations incorporating these practical factors.
Finally, the five-path ensemble introduces additional computational cost, while the SAME and PLTS augmentation parameters are manually specified and the current group-balanced fusion uses fixed weights. Although this parameter-free fusion reduces the risk of overfitting under one-shot calibration, it cannot adapt to variations in pathway reliability across participants, trials, or data lengths. Future work will investigate lightweight inference, data-driven augmentation-parameter selection, and confidence-aware or data-length-dependent fusion using an independent validation protocol to avoid overfitting and test-data leakage.
6. Conclusions
In this study, we proposed an individual-adapted SSVEP decoding framework that integrates hybrid data augmentation, collaborative discrimination, and a Double-Area network to address the fundamental challenges of limited calibration data and substantial inter-subject variability in SSVEP-based BCI systems. Building upon a multi-path ensemble architecture, the proposed method incorporates SAME-based static feature augmentation and PLTS-based temporal-dynamic enhancement to enrich subject-specific calibration data, while eTRCA-K and TDCA are employed to strengthen feature extraction across individuals. In parallel, the Double-Area network introduces a complementary deep hierarchical representation that captures nonlinear spatiotemporal patterns and multi-harmonic structures crucial for SSVEP decoding. By jointly modeling subject-level adaptation, harmonic-sensitive feature learning, and pathway-level fusion, the proposed framework effectively mitigates the sensitivity of SSVEP decoding to data length and enhances both robustness and generalization across subjects and datasets. This design addresses the core challenges of data scarcity, individual variability, and temporal instability in SSVEP-based BCIs, enabling more accurate, reliable, and user-friendly target recognition for real-world BCI applications.
Author Contributions
J.N.: writing—review and editing, supervision, validation. S.Z. (Shuyao Zhai): writing—original draft, methodology, data curation. S.Z. (Siyuan Zhang): supervision, validation, conceptualization. Y.Z.: formal analysis, resources, software. All authors have read and agreed to the published version of the manuscript.
Funding
This work was supported by the Key Science and Technology Program of Henan Province under Grants 262102211079; the Postgraduate Education Reform and Quality Improvement Project of Henan Province under Grant YJS2023JD67; the Key Science Research Project of Colleges and Universities in Henan Province of China under Grant 25A520003; the Post-Doctoral Foundation of Henan Province under Grant HN2025151; the Zhengzhou Youth Science and Technology Talent Program under Grant 20248651.
Data Availability Statement
Public Benchmark and BETA SSVEP datasets were utilized in this work; their official open-access download links can be found in the main text. The source code implementing the proposed five-path ensemble decoding framework is publicly available at GitHub: https://github.com/HMCBI-LAB-ZZULI/Individual-Adapted-SSVEP-Decoding-HACD, accessed on 11 February 2026. The software environment consisted of Python 3.10.0, PyTorch 1.12.1, CUDA 11.3, and cuDNN 8.3.2. Further assistance can be obtained from the corresponding author for any reasonable inquiries.
Conflicts of Interest
The authors declare no conflict of interest.
References
- Lin, Z.; Zhang, C.; Wu, W.; Gao, X. Frequency recognition based on canonical correlation analysis for SSVEP-based BCIs. IEEE Trans. Biomed. Eng. 2006, 53, 2610–2614. [Google Scholar] [CrossRef] [Scilit]
- Chen, X.; Wang, Y.; Gao, S.; Jung, T.-P.; Gao, X. Filter bank canonical correlation analysis for implementing a high-speed SSVEP-based brain–computer interface. J. Neural Eng. 2015, 12, 046008. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Xia, M.; Chen, K.; Xu, P.; Yao, D. Progresses and prospects on frequency recognition methods for steady-state visual evoked potential. J. Biomed. Eng. 2022, 39, 192–197. [Google Scholar]
- Nakanishi, M.; Wang, Y.; Chen, X.; Wang, Y.-T.; Gao, X.; Jung, T.-P. Enhancing detection of SSVEPs for a high-speed brain speller using task-related component analysis. IEEE Trans. Biomed. Eng. 2017, 65, 104–112. [Google Scholar] [CrossRef] [Scilit]
- Wong, C.M.; Wan, F.; Wang, B.; Wang, Z.; Nan, W.; Lao, K.F.; Mak, P.U.; Vai, M.I.; Rosa, A. Learning across multi-stimulus enhances target recognition methods in SSVEP-based BCIs. J. Neural Eng. 2020, 17, 016026. [Google Scholar] [CrossRef] [Scilit]
- Xu, M.; Wu, Q.; Xiong, W.; Xiao, X.; Ming, D. Research on encoding and decoding algorithms for medium/high-frequency SSVEP-based brain-computer interface. J. Signal Process. 2022, 38, 1881–1891. [Google Scholar]
- Liu, B.; Chen, X.; Shi, N.; Wang, Y.; Gao, S.; Gao, X. Improving the performance of individually calibrated SSVEP-BCI by task-discriminant component analysis. IEEE Trans. Neural Syst. Rehabil. Eng. 2021, 29, 1998–2007. [Google Scholar] [CrossRef] [Scilit]
- Kwak, N.-S.; Müller, K.-R.; Lee, S.-W. A convolutional neural network for steady state visual evoked potential classification under ambulatory environment. PLoS ONE 2017, 12, e0172578. [Google Scholar] [CrossRef]
- Lawhern, V.J.; Solon, A.J.; Waytowich, N.R.; Gordon, S.M.; Hung, C.P.; Lance, B.J. EEGNet: A compact convolutional neural network for EEG-based brain–computer interfaces. J. Neural Eng. 2018, 15, 056013. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Waytowich, N.; Lawhern, V.J.; Garcia, J.O.; Cummings, J.; Faller, J.; Sajda, P.; Vettel, J.M. Compact convolutional neural networks for classification of asynchronous steady-state visual evoked potentials. J. Neural Eng. 2018, 15, 066031. [Google Scholar] [CrossRef] [Scilit]
- Guney, O.B.; Oblokulov, M.; Ozkan, H. A deep neural network for ssvep-based brain-computer interfaces. IEEE Trans. Biomed. Eng. 2021, 69, 932–944. [Google Scholar] [CrossRef] [Scilit]
- Pan, Y.; Chen, J.; Zhang, Y.; Zhang, Y. An efficient CNN-LSTM network with spectral normalization and label smoothing technologies for SSVEP frequency recognition. J. Neural Eng. 2022, 19, 056014. [Google Scholar] [CrossRef] [Scilit]
- Deng, Y.; Sun, Q.; Wang, C.; Wang, Y.; Zhou, S.K. TRCA-Net: Using TRCA filters to boost the SSVEP classification with convolutional neural network. J. Neural Eng. 2023, 20, 046005. [Google Scholar] [CrossRef] [Scilit]
- Li, D.; Huang, Y.; Luo, R.; Zhao, L.; Xiao, X.; Wang, K.; Yi, W.; Xu, M.; Ming, D. Enhancing detection of SSVEPs using discriminant compacted network. J. Neural Eng. 2025, 22, 016043. [Google Scholar] [CrossRef] [Scilit]
- Deng, Y.; Ji, Z.; Wang, Y.; Zhou, S.K. OS-SSVEP: One-shot SSVEP classification. Neural Netw. 2024, 180, 106734. [Google Scholar] [CrossRef] [Scilit]
- Luo, R.; Xu, M.; Zhou, X.; Xiao, X.; Jung, T.-P.; Ming, D. Data augmentation of SSVEPs using source aliasing matrix estimation for brain–computer interfaces. IEEE Trans. Biomed. Eng. 2022, 70, 1775–1785. [Google Scholar] [CrossRef] [Scilit]
- Mai, X.; Ai, J.; Wei, Y.; Zhu, X.; Meng, J. Phase-locked time-shift data augmentation method for SSVEP brain-computer interfaces. IEEE Trans. Neural Syst. Rehabil. Eng. 2023, 31, 4096–4105. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Chen, X.; Gao, X.; Gao, S. A benchmark dataset for SSVEP-based brain–computer interfaces. IEEE Trans. Neural Syst. Rehabil. Eng. 2016, 25, 1746–1752. [Google Scholar] [CrossRef] [Scilit]
- Liu, B.; Huang, X.; Wang, Y.; Chen, X.; Gao, X. BETA: A Large Benchmark Database Toward SSVEP-BCI Application. Front. Neurosci. 2020, 14, 627. [Google Scholar] [CrossRef] [Scilit]
- Bian, R.; Wu, H.; Liu, B.; Wu, D. Small data least-squares transformation (sd-LST) for fast calibration of SSVEP-based BCIs. IEEE Trans. Neural Syst. Rehabil. Eng. 2022, 31, 446–455. [Google Scholar] [CrossRef] [Scilit]
- Convolutions, C.-W. Channelnets: Compact and efficient convolutional neural networks via. IEEE Trans. Pattern Anal. Mach. Intell. 2021, 43, 2570–2581. [Google Scholar] [CrossRef] [Scilit]
- Guney, O.B.; Ozkan, H. Transfer learning of an ensemble of DNNs for SSVEP BCI spellers without user-specific training. J. Neural Eng. 2023, 20, 016013. [Google Scholar] [CrossRef] [Scilit]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.







