1. Introduction
Alzheimer’s disease (AD) and frontotemporal dementia (FTD) are major neurodegenerative disorders that are associated with heterogeneous and overlapping cognitive and behavioral manifestations, particularly in early disease stages, which complicates reliable clinical assessment [
1,
2,
3,
4], while structural and metabolic imaging modalities such as MRI and FDG-PET are well established [
5,
6], they are costly and limited in temporal sensitivity to functional neural changes. Electroencephalography (EEG), by contrast, provides a low-cost and temporally precise measure of large-scale neural dynamics and has been widely investigated as a complementary biomarker for dementia [
7,
8,
9].
Most prior EEG-based studies on dementia have focused on the eyes-closed (EC) condition, where relatively stable posterior alpha rhythms enable reliable spectral characterization [
10,
11]. Under EC, AD-related alterations are often characterized by reproducible changes in oscillatory power, coherence, or complexity, implicitly assuming the existence of disease-specific patterns that are sufficiently stable to be captured by averaging or covariance-based representations [
12]. In this setting, a wide range of methods—including power spectral density (PSD) features, entropy-based complexity measures, and Riemannian geometry-based covariance representations—have reported promising classification performance for AD and FTD [
10,
13,
14]. More recently, deep learning approaches such as convolutional neural networks (CNNs) and Transformers have further improved accuracy under EC conditions [
12,
15,
16,
17].
However, this assumption does not readily extend to eyes-open (EO) conditions. Under EO stimulation, posterior alpha rhythms are suppressed, and stimulus-driven responses become highly nonstationary. In AD, visual entrainment is often fragmented and unstable, reflecting impaired large-scale synchronization and deficient alpha reactivity [
18,
19]. FTD exhibits different but partially overlapping alterations, further complicating discrimination [
20]. As a result, conventional short-window spectral summaries that perform well under EC conditions degrade substantially under EO, and only a small number of studies have attempted EO-based dementia classification [
21]. The recently released OpenNeuro dataset ds006036, recorded under EO photic stimulation, highlights both the diagnostic potential and the methodological challenges of this setting.
A further challenge concerns evaluation methodology. It is now well recognized that EEG classifiers evaluated under non-subject-wise splits can severely overestimate performance. When re-evaluated under strict leave-one-subject-out (LOSO) validation, many models—including deep neural networks—exhibit pronounced performance degradation and even near-zero accuracy for certain subjects [
22,
23]. Recent large-scale reanalyses have shown that even LOSO accuracy can be misleading, as apparent generalization may collapse under repeated or nested subject-wise validation (e.g., N-LOSO), revealing poor robustness across unseen individuals [
22]. These findings underscore that accuracy alone is insufficient for assessing clinical reliability in heterogeneous populations such as AD and FTD.
A recent study has demonstrated that, under EC conditions, carefully designed feature representations—particularly those based on the Riemannian geometry of covariance matrices—can achieve reliable subject-wise performance even under strict LOSO-CV validation, despite limited dataset sizes [
14]. These findings indicate that, when disease-related EEG patterns are relatively stable, appropriate feature construction can partially compensate for data scarcity and support generalization across subjects. In such EC settings, reliability is largely driven by the degree to which an individual subject’s representation aligns with characteristic low-frequency connectivity patterns.
Under EO conditions, however, reliability assessment faces an additional challenge beyond classification accuracy. Because subject-level representations are less likely to conform to stable, disease-specific templates, misclassifications often arise from qualitatively unstable or fragmented neural responses rather than marginal deviations from a well-defined pattern. As a result, reliability in EO settings becomes inherently asymmetric: incorrect predictions may not correspond to low-confidence samples near a decision boundary but instead reflect subjects whose neural dynamics fundamentally diverge from the learned structure. This property motivates the need for reliability analyses that explicitly characterize miss-class behavior, rather than relying solely on pattern similarity or distance-based confidence measures.
Meanwhile, a complementary line of research has shifted attention away from the exclusive choice of individual features or classifiers toward the organization of the feature space itself. Rather than treating each EEG epoch as an independent sample to be directly classified, these approaches first cluster extracted representations to identify representative patterns and subsequently encode new samples by their similarity or projection onto these learned prototypes. Such prototype-based or dictionary-based representations have been explored in EEG analysis through Bag-of-Words models, codebook learning, and medoid-based clustering, demonstrating improved robustness to inter-subject variability and enhanced interpretability compared with direct feature classification [
24,
25,
26,
27].
Building on this perspective, we propose a white-box EEG framework that explicitly targets the challenges of EO dementia analysis under strict LOSO evaluation. Building on Dynamic Mode Decomposition (DMD), which captures spatio-temporal neural dynamics beyond stationary spectral content [
28], we introduce a clustering-based pattern projection approach. Rather than discarding unstable or fragmented responses, the proposed method clusters DMD-derived representations from training data to identify recurring dynamical motifs—including characteristic patterns of breakdown—and encodes each epoch by its similarity to medoid-based prototypes. In this way, heterogeneous and nonstationary EO responses are transformed into a structured representation that emphasizes geometric alignment with group-level dynamical patterns rather than raw spectral magnitude.
To address these challenges, we propose a unified framework for EEG-based AD detection that integrates (i) DMD-based long-term summarization of temporally segmented EEG, (ii) a Bag-of-Words-inspired clustering and pattern projection scheme using medoid prototypes to capture recurring disease-relevant dynamics, and (iii) reliability-aware subject-wise evaluation with margin-based analysis to assess generalization under strict LOSO validation. This integrated approach, termed the DMD–CPP (Dynamic Mode Decomposition–Clustered Pattern Projection) framework, is designed to provide an interpretable and generalizable representation of nonstationary EEG dynamics while mitigating the overestimation of performance commonly observed in deep learning-based models.
The main contributions of this work are threefold. First, we address the underexplored problem of EO-based AD discrimination by interpreting the characteristic breakdown of AD responses not as noise but as a recurrent dynamical structure. Using DMD and clustering, these breakdown patterns are consolidated into representative prototype groups and modeled through a projection-based representation. Second, we adopt strict LOSO validation and demonstrate that subject-level reliability cannot be inferred from accuracy alone, motivating the need for confidence-aware analysis. Third, we introduce a margin-based reliability assessment that reveals how the proposed DMD-based clustering framework improves confidence, separability, and controllability at the subject level, particularly for AD-related decisions.
2. Materials and Methods
Details of the dataset, feature construction, classification protocol, and experimental setup are provided in this section. The EO photostimulation EEG dataset is introduced, including subject demographics and the inclusion/exclusion criteria, by which ten stimulus-related epochs per subject are obtained. DMD is applied to 2 s slices, and the resulting mode-magnitude representations are temporally summarized within each epoch using simple statistical descriptors to construct fixed-length mode-based feature vectors. To assess reliability beyond nominal accuracy, a margin-based analysis is performed to compare absolute decision margins between correctly classified and misclassified epochs under LOSO validation. For fair comparison, a baseline pipeline based on principal component analysis (PCA) [
29] is included. The baseline is evaluated alongside the proposed
DMD-CPP method, in which mode-based representations are combined with class-specific clustering and projection to enable robust EEG classification.
2.1. Dataset
The publicly available EO photostimulation EEG dataset on OpenNeuro (dataset ID: ds006036, v1.0.4; DOI:10.18112/openneuro.ds006036.v1.0.4 (
https://openneuro.org/datasets/ds006036/versions/1.0.4, accessed on 5 February 2026)) was analyzed. The dataset includes 88 individuals (36 AD, 23 FTD, 29 Normal Control (CN)), recorded with 19 scalp electrodes (10–20 system; Fp1, Fp2, F3, F4, C3, C4, P3, P4, O1, O2, F7, F8, T3, T4, T5, T6, Fz, Cz, Pz) at 500 Hz (10 µV/mm). During clinical acquisition, intermittent photic stimulation was administered in nominal 5 Hz increments (5, 10, 15, 20, and—where tolerated—up to 30 Hz); however, the ordering, dwell times, and upper limits were allowed to vary across subjects and recording blocks under routine conditions. Detailed event annotations are provided in the dataset’s release files. Cognitive and neuropsychological functioning was assessed using the international Mini-Mental State Examination (MMSE), which yields scores from 0 to 30, with lower values reflecting greater cognitive impairment. The Alzheimer’s disease (AD) cohort comprised 36 participants (11 males, 25 females) with a mean age of 66.2 years (SD = 7.5) and a mean MMSE score of 17.4 (SD = 4.6). The frontotemporal dementia (FTD) group included 23 participants (12 males, 11 females), with an average age of 64.4 years (SD = 7.3) and a mean MMSE score of 22.6 (SD = 2.7). The cognitively normal (CN) group consisted of 29 individuals (17 males, 12 females), with a mean age of 68.3 years (SD = 4.9); all CN participants achieved the maximum MMSE score of 30. All experimental procedures were conducted in accordance with ethical standards and were approved by the Scientific and Ethical Committee of the Aristotle University of Thessaloniki and AHEPA University Hospital (protocol no. 142/12-04-2023).
Although the dataset provides both raw and preprocessed EEG recordings, many prior studies have relied on the preprocessed signals, which already incorporate noise filtering and artifact removal (see [
21]). Following this convention, we used the preprocessed EEG data in the present study. Accordingly, no additional preprocessing was applied during feature extraction, as the provided signals were already suitable for analysis. Moreover, because our objective was to examine brain responses elicited by visual stimulation, the analysis was restricted to EEG segments corresponding to visual stimulus events. For each subject, time intervals during which visual stimuli were presented were identified, and only the EEG data within these intervals were extracted for subsequent analysis.
As shown in
Table 1, usable durations vary across subjects. Each epoch uses 20 s (10 consecutive 2 s segments). To construct ten epochs per subject, we require at least 19 non-overlapping 2 s segments (i.e., ≥38 s of usable stimulus-related data). Recordings shorter than this threshold were excluded. Note that, unlike EC resting, our EO recordings include visually driven activity (photic entrainment) mixed with spontaneous fluctuations and small ocular/motion events. Segments from different EO stimulation frequencies are therefore analyzed together to assess whether diagnostic patterns generalize beyond any single stimulation condition, rather than reflecting EC-style resting-state dynamics.
Because photic-stimulation frequencies differ across subjects, 20 s epochs may contain mixed frequency segments. A category-based sensitivity analysis reported in [
28] found no association between epoch-level stimulus composition and diagnostic group. Accordingly, stimulation was treated as a nuisance factor, justifying a uniform slice-based DMD pipeline (see Appendices 6.1 and 6.2 in [
28]).
2.2. Feature Extraction
This subsection introduces the construction of fixed-size, mode-based descriptors of stimulus-related EEG that preserve sensor-space structure across time. Each recording was partitioned into non-overlapping 2 s segments; epochs were defined as 20 s windows comprising 10 consecutive segments, with start indices uniformly spaced across subjects, so that each subject contributed 10 uniformly distributed epochs. DMD was applied to every 2 s segment to obtain rectified mode matrices together with their associated eigenfrequencies. Within each epoch, mode-magnitude matrices were aggregated across segments and summarized using temporal statistics (element-wise mean), yielding fixed-length representations suitable for subsequent pattern learning and classification.
The overall pipeline proceeds as follows: Stage 1 constructs mode-based feature vectors from 20 to s epochs by aggregating DMD mode magnitudes across time; Stage 2 learns compact, class-specific pattern dictionaries via hierarchical clustering and medoid selection; finally, each sample is mapped to cosine-similarity features for downstream classification.
2.2.1. Dynamic Mode Decomposition
DMD is a data-driven method that represents multichannel signals as superpositions of coherent spatio-temporal modes with characteristic eigenvalues [
30]. In EEG data, where the number of channels is typically much smaller than the number of temporal samples, an extended (stacked) DMD formulation is adopted to enrich the state representation by concatenating consecutive time samples. Under this formulation, the signal admits the modal expansion
where
denotes the multichannel EEG signal at time
t,
are the DMD modes,
are the corresponding continuous-time eigenvalues, and
are the modal amplitudes determined by the initial condition. Details of the augmented DMD construction are provided in
Appendix A.1.
2.2.2. DMD Configuration and Post–Processing
From the stimulus-related EEG in
Table 1, each recording was partitioned into non-overlapping 2 s segments, and one epoch was defined as a sequence of ten consecutive segments (20 s). Let
denote the number of usable 2 s segments for a subject. To ensure uniform coverage of the stimulus interval while avoiding overlap at the segment level, ten epochs were placed at evenly spaced start indices over the range
:
This requires
(i.e., at least 38 s of usable data); recordings below this threshold were excluded (IDs: 15, 21, 64, 65, and 78; see
Table 1). Avoiding overlap is particularly important in stimulus-locked analyses, as overlapping windows can introduce spurious entrainment and related confounds [
31,
32].
As illustrated in
Figure 1, DMD was performed once for each non-overlapping 2 s segment within the stimulus onset–offset interval. Epochs served solely as an aggregation and indexing layer: each 20 s epoch collected the ten segment-level DMD feature maps falling within its span. Consequently, although epoch windows may visually overlap, no additional DMD computation was introduced. For each segment, DMD was applied using the stacked formulation with a fixed stacking parameter and truncation size, yielding a finite set of dynamic modes associated with eigenfrequencies
as defined in (
1). The resulting channel-level DMD modes are collected as
where each
represents the spatial pattern of the
j-th dynamic mode across the
EEG channels. The truncation size of the reduced-order approximation determines the total number of retained modes. To mitigate sensor-uniform contributions and volume-conduction-related inflation, the DMD modes
in (
2) were further refined by removing sensor-uniform components within each mode while preserving the inter-sensor phase relationships. This refinement suppresses stimulus-locked harmonic contamination without altering the relative phase structure across EEG channels. Unless stated otherwise,
henceforth denotes the rectified DMD mode matrix used in all subsequent analyses.
Figure 1d schematically summarizes this step within the overall epoch-level processing pipeline.
The rectification process, which recovers the original
M-channel representation from the augmented DMD modes, is described in
Appendix A.2. For completeness, we note that the formulation admits complex-valued DMD modes with phase information as a general representation to enable future phase-aware extensions. In this study, however, we focus on magnitude-based descriptors to emphasize the spatial participation patterns of the modes, and phase information is therefore not used in the experiments.
2.2.3. Stage 1—Epoch-Level Descriptor Construction
Each 20 s epoch is first partitioned into non-overlapping segments of length 2 s. For each segment, DMD is applied to the multichannel EEG, yielding a set of rectified DMD modes with associated eigenfrequencies. Only modes with frequencies in the interval are retained, thereby excluding very low- and very high-frequency components.
From the retained modes of each segment, a non-negative mode-image is constructed by stacking channel-wise magnitudes and rescaling the mode dimension to a fixed width
P via interpolation. Here,
P denotes the number of normalized mode indices used to represent the frequency-ordered DMD spectrum on a common grid. This results in a sequence of rectangular matrices, one per segment, that encode how strongly each EEG channel loads onto the ordered DMD modes. These segment-level representations are then aggregated across the
T segments within an epoch by computing their elementwise mean, thereby summarizing epoch-level mode activity. Finally, the aggregated representation is linearly rescaled to the unit interval, yielding the epoch-level descriptor
where
M denotes the number of EEG channels and
P denotes the dimensionality of the mode axis after interpolation (set to
in our experiments).
Let
denote the epoch-level descriptor of the
n-th epoch, as defined in (
3). Each descriptor is converted into a representation by column-wise stacking,
Let
denote an index set specifying a subset of epochs; the corresponding
design matrix is defined as
where each column
in (
4) represents the vectorized epoch-level descriptor. This matrix serves as the input to clustering and dictionary (basis) learning. The specific construction of the index set
—for example, according to a training split, a diagnostic group, or a validation protocol—is described separately in the subsequent section.
2.2.4. Stage 2—Basis Learning from the Design Matrix
In Stage 2, basic learning is performed on collections of epoch-level descriptors selected according to a given criterion. Let
denote an index set specifying a subset of epochs, and let
be the corresponding design matrix defined in (
5).
Hierarchical clustering is a classical approach for identifying structure in high-dimensional data through recursive partitioning based on pairwise dissimilarities [
25]. In this work, we employ hierarchical divisive clustering with complete linkage under cosine dissimilarity as a mechanism for selecting representative basis elements from the design matrix
. Specifically, each column of
corresponds to a vectorized epoch-level descriptor, denoted by
. Similarity between two such descriptors,
and
, is measured using cosine similarity,
and complete linkage defines the dissimilarity between two subsets as the maximum pairwise dissimilarity between their elements [
33]. Based on this criterion, a hierarchical tree is constructed and recursively partitioned from the top level.
Recursive partitioning proceeds by splitting a subset only when it satisfies minimum support and separation conditions, ensuring that the resulting groups correspond to coherent regions of the descriptor space rather than noise-driven artifacts. Specifically, a subset is further divided only if
- (a)
Its size exceeds a support threshold ;
- (b)
The corresponding dendrogram split height exceeds , and;
- (c)
Both resulting subsets contain at least samples.
The subsets of
that satisfy these criteria are retained as surviving groups,
which define candidate regions for representative basis selection.
From each surviving group
in (
6) obtained by hierarchical partitioning, a single representative descriptor is selected as the medoid. Under cosine dissimilarity, the medoid is defined as the sample that minimizes the total dissimilarity to all other members of the group,
where
. Collecting the selected medoids yields the dictionary
where each column of
corresponds to a representative descriptor selected from one surviving group.
The medoid-based selection is adopted instead of a centroid-based alternative for two reasons. First, the medoid corresponds to an actual observed epoch, ensuring that each basis element represents a physically realizable EEG pattern rather than a virtual average. Second, medoids are inherently robust to outliers and skewed cluster geometries, which commonly arise in high-dimensional mode-based descriptors [
34,
35].
Given a vectorized epoch descriptor
in (
4) and the dictionary
defined in (
7), the relationship between
and the
k-th dictionary atom (medoid) is quantified by their cosine similarity,
The resulting value
constitutes the
k-th component of the dictionary-based feature vector associated with
.
2.3. Classification
We employ support vector machines (SVMs) for supervised classification [
36]. Because the mode-based representation described in Stages 1–2 is high-dimensional, we adopt the linear SVM as our primary classifier (All classifiers are implemented using MATLAB (R2025b) with the fitcsvm function using a linear kernel.) Linear SVMs provide a robust margin-based baseline, require no kernel-scale tuning, and remain stable in the small-
N, large-
p regime—where
N denotes the number of training samples and
p the dimensionality of the feature space—typical of subject-level EEG studies.
Throughout this work, all learning steps are restricted to training subjects only; no information from the test subjects is used in dictionary learning, normalization, or classifier training.
Figure 2 summarizes the complete feature–assembly and classification workflow.
2.3.1. Classification Procedure
Fix a binary classification task
(AD vs. CN, FTD vs. CN or AD vs. FTD) with the class set
Stage 1 and Stage 2 together provide, for each epoch
n, a mode-based vectorized descriptor
as in (
4), and the split-wise design matrices, defined as instances of the general sub-design matrix in (
5),
where
denotes the index set of epochs assigned to split
(with
). For a given class
, let
denote the index set of training epochs belonging to class
c. The corresponding class-restricted design matrix is given by
which serves as the input to the hierarchical divisive clustering procedure. The resulting surviving groups yield medoids and the class-specific dictionary
as defined in (
7). This dictionary learning step is carried out once per class
and
, using training data only, and is shared by both the training and test evaluations.
Given a dictionary
and a vectorized descriptor
, its projection onto
is computed by cosine similarity according to (
8), yielding a
-dimensional feature vector for that sample. Collecting these projections over all samples in the split
produces the feature block
For task
, the Stage 2 representation is obtained by concatenating the two class-specific blocks:
Thus,
is the sole input to the downstream classifier for task
.
To ensure comparability across subjects and prevent scale bias, each feature dimension is standardized using statistics from the training split only. Let
and
denote the per-feature mean and standard deviation of
. We apply the affine normalization
A linear SVM is trained on
using the labels in
, and test predictions are obtained by applying the trained classifier to
. All reported metrics (accuracy, precision, recall, F1, and confusion matrices) follow the validation protocol described in the next subsection.
2.3.2. Validation Methodology, Classification Tasks, and Metrics
To rigorously assess generalization while preventing subject-specific leakage, we employ a LOSO validation strategy. In each fold, one subject s is held out for testing, and all epochs of s are excluded from training. The model is trained on the remaining subjects and evaluated once on the held-out subject. This process is repeated until every subject in the task has served as the test fold.
An important consideration in EEG analysis is that subjects often differ substantially in recording length. As a result, shorter recordings yield a smaller number of epochs with a higher degree of overlap, which undermines the statistical independence of epoch-level samples and introduces asymmetric bias across subjects. For this reason, epoch-level performance metrics may be unreliable and should not be interpreted as reflecting true generalization at the subject level.
Despite this limitation, epoch-level results are reported in parallel for indirect comparison with prior EEG studies that adopt non-subject-wise validation schemes, such as leave-N-segments-out or epoch-level cross-validation. In contrast, all primary analyses, statistical evaluations, and interpretations of the proposed method—including the discussion of reliability and generalization—are based exclusively on subject-level summaries, which constitute the clinically and methodologically meaningful unit of analysis.
For each task, predictions from all LOSO folds are pooled by summing the subject-specific confusion matrices. This aggregated epoch-level confusion matrix contains the total , , , and counts accumulated over the entire evaluation. Because each subject contributes the same number of test epochs, simple summation ensures a correct and unbiased aggregation.
All reported metrics—accuracy, precision, recall, and F1—are computed exclusively from the aggregated confusion matrix. For example,
with F1 defined analogously.
In addition to epoch-level performance, we evaluated diagnostic performance at the subject level, which represents the clinically relevant unit under LOSO validation. For each test subject, the trained classifier produces multiple epoch-wise predictions. These predictions are aggregated by majority voting to yield a single subject-level decision. A subject is considered correctly classified if more than half of its test epochs are assigned to the true diagnostic class.
Subject-level accuracy is defined as the proportion of subjects whose aggregated predictions match their ground-truth labels. Under LOSO validation, each subject contributes exactly one binary outcome (correct or incorrect), and the resulting accuracy estimate therefore follows a binomial sampling model. To quantify the statistical uncertainty associated with the finite number of subjects, we report 95% confidence intervals for subject-level accuracy using the Wilson score method.
We evaluate three binary discrimination tasks under LOSO: AD vs. CN, FTD vs. CN, and AD vs. FTD. In all tasks, the positive class is defined as AD for AD vs. CN, FTD for FTD vs. CN, and AD for AD vs. FTD.
2.4. Decision-Margin Analysis
To assess model reliability beyond accuracy, we performed a decision-margin analysis based on the decision scores produced by the linear SVM. For each test epoch under LOSO evaluation, the classifier outputs a signed decision margin
m, defined as the distance to the separating hyperplane. The magnitude
is commonly interpreted as a measure of prediction confidence [
35,
37].
2.4.1. Subject-Level Margin Aggregation
Because clinical decisions are made at the subject level, epoch-wise margins were first summarized within each subject. For a given subject s, let denote the signed margin of epoch e. Subject-level descriptors were obtained by aggregating these epoch-wise margins within each subject. Specifically, the confidence magnitude of subject s was defined as the median of the absolute epoch-wise margins, . All subsequent analyses are performed on these subject-level summaries rather than on individual epochs.
2.4.2. Outcome Grouping and Margin Descriptors
Subjects were grouped according to their classification outcome into four categories: true positive (TP), false negative (FN), true negative (TN), and false positive (FP). For each outcome group, subject-level margin behavior was characterized using three complementary descriptors, each computed by first aggregating epoch-wise margins within subjects and then comparing the resulting subject-level summaries across groups.
- 1.
Confidence magnitude (within-subject). For each subject, confidence magnitude was defined as the median of the absolute margins across epochs, . This quantity summarizes the typical distance of that subject’s epoch-wise representations from the decision boundary. Group-level comparisons were then performed on these subject-level confidence magnitudes.
- 2.
Within-subject dispersion. For each subject, variability of epoch-wise margins was quantified using the interquartile range . This descriptor captures how stable or fluctuating the subject’s margins are across epochs. Group-level differences were assessed by comparing these subject-level IQR values across outcome groups.
- 3.
Within-subject sign consistency. For each subject, we computed the proportion of epochs with positive margins, , and summarized directional stability using the index . Values close to 0 indicate frequent sign changes across epochs, whereas values approaching indicate that the subject’s epoch-wise margins consistently favor one decision direction. These subject-level indices were subsequently compared across outcome groups.
2.4.3. Statistical Analysis
The proposed framework constructs projection features anchored to disease-related break patterns learned during clustering. Under this design, subjects whose data genuinely exhibit such break patterns are expected to show systematically different margin behavior from those that do not. In particular, subjects correctly identified as patients (TP) and those incorrectly projected as patients (FP) are expected to differ in their subject-level margin characteristics, as the latter reflect spurious or unstable matches to the learned disease anchors. Accordingly, our primary interest is to assess whether meaningful differences exist between these outcome groups, with comparisons between TN and FN considered complementary.
To this end, for each subject-level margin descriptor, we tested the null hypothesis that the two outcome groups do not differ in their central tendency. Group differences were summarized using the median difference, , which captures the separation between the typical subject-level values of the two groups. Statistical significance was assessed using permutation tests, in which subjects were randomly reassigned between the two groups while preserving the original group sizes. This procedure evaluates whether the observed median difference is larger than would be expected by chance alone, under the assumption that group membership carries no systematic information about the descriptor.
All tests were two-sided and conducted at a significance level of
, reflecting the exploratory nature of subject-level reliability analysis under limited and imbalanced sample sizes. In addition to statistical significance, effect size was quantified using Cliff’s
, a nonparametric measure of stochastic dominance that estimates the probability that a randomly selected subject from one group has a larger descriptor value than a randomly selected subject from the other group, minus the reverse probability. Cliff’s
ranges from
to 1, with values near zero indicating substantial overlap between groups and larger absolute values indicating stronger separation [
38].
2.5. Comparing Algorithm: PCA-Based Mode Features
Our main pipeline combines mode-based Stage 1 summaries with class-specific dictionary learning via hierarchical clustering and medoid extraction, followed by classification using a linear SVM. To isolate and assess the contribution of this Stage 2 design, we construct a simpler baseline that employs the same Stage 1 mode-based representation but replaces the clustering-based dictionary learning with PCA. Importantly, both pipelines use the same linear SVM classifier so that any performance differences can be attributed specifically to the choice of feature projection—clustering-based basis learning versus PCA—rather than to differences in the classifier itself.
2.5.1. Shared Mode-Based Stage 1 Representation
The PCA-based baseline uses the same Stage 1 features as the main pipeline. For each 20 s epoch, DMD is applied to the 2 s segments, and the resulting vectorized epoch-level descriptor
is constructed as in (
3). For a given binary classification task
with class set
, we construct task-specific design matrices separately for the training and test splits. Let
denote the data split, and let
be the index set of all subjects assigned to split
under LOSO validation. From this set, we select only those subjects whose labels belong to
, yielding the task-restricted index set
.
The resulting task-specific design matrix is defined as
where each column corresponds to the vectorized descriptor of a subject belonging to one of the two classes in task
.
2.5.2. PCA-Based Dimensionality Reduction
For each binary task
, PCA is fitted using only the training design matrix
in (
14) to avoid information leakage. After mean-centering, principal components are retained to explain 95% of the total variance, yielding a low-dimensional linear subspace that captures the dominant variation in the mode-based representations. Both training and test samples are then projected onto this task-specific subspace using the same projection learned from the training data.
Following projection, each retained component is standardized using statistics estimated exclusively from the training split, and the same normalization parameters are applied to the test split. This procedure ensures a fair comparison with the proposed clustering-based pipeline by preserving identical data splits, classifiers, and normalization rules while differing only in the choice of feature projection method.
2.6. Experimental Setup
All postprocessing, feature extraction, and classification were implemented in MATLAB. This section summarizes the experimental design, including the construction of analysis epochs, the feature-extraction settings, and the evaluation protocol.
2.6.1. Epoch Construction
Each subject’s EEG was segmented into 20 s epochs, each consisting of ten consecutive non-overlapping 2 s segments extracted from stimulus intervals. Under the LOSO scheme, we constructed ten epochs per subject. This epoch length and count reflect a trade-off: longer or more numerous epochs increase discriminability but reduce the number of usable subjects, whereas shorter or fewer epochs yield less reliable representations. We therefore adopted an intermediate setting balancing performance and generalization.
2.6.2. Feature Settings
Feature-extraction hyperparameters were fixed a priori and tuned independently of the downstream classification task. DMD was applied to each segment using a stacked formulation with stack size
and truncation rank
, as in (
A1). These parameters were selected to balance spectral resolution, numerical stability, and computational tractability; a detailed rationale for these choices is provided in
Appendix A.1. The resulting DMD modes were summarized into fixed-size epoch-level descriptors with mode-axis resolution
in (
3), yielding matrices of size
, where
denotes the number of EEG channels.
For Stage 2 dictionary learning, clustering was performed on
-normalized mode-based vectors using cosine dissimilarity. A hierarchical divisive clustering procedure was employed with a cluster size threshold
, minimum cluster size
, and stopping height
, using complete linkage throughout. These values were determined empirically, reflecting a trade-off between marginal performance gains achievable with stricter thresholds and the substantially higher computational cost they incur. Accordingly, this configuration was adopted as a practical setting for the proposed mode-based representation. The notation and parameter settings used throughout the proposed framework are summarized in
Table 2.
2.6.3. Evaluation
Classification was performed with the positive class defined as AD in AD vs. CN tasks, FTD in FTD vs. CN tasks, and AD in AD vs. FTD tasks. All learning components (dictionary construction, PCA baselines, standardization, and classifier training) were restricted to training subjects in the LOSO split.
4. Discussion
This section summarizes the main findings of the proposed DMD-CPP framework and discusses their methodological implications under EO stimulation. We first highlight task-dependent performance characteristics observed across AD vs. CN, FTD vs. CN, and AD vs. FTD classifications. We then describe the properties of the class-specific medoid patterns learned through clustering and examine how these representations shape decision-margin behavior. Finally, we discuss the limitations of the current study and relate our findings to prior EEG-based dementia research.
4.1. Summary of the Analytical Approach
This study investigates whether stimulus-interval EO EEG can provide diagnostically meaningful information for neurodegenerative diseases when represented through data-driven dynamic patterns derived from DMD. Using a publicly available OpenNeuro dataset, the analysis was restricted to visually driven stimulus periods, from which a uniform set of ten non-overlapping 20 s epochs per subject was constructed to ensure balanced subject-level evaluation. Extended DMD with temporal stacking was applied to each 2 s segment to extract channel-level modes that capture transient spatiotemporal dynamics beyond stationary spectral descriptors. The extracted mode-based summaries were assembled into fixed-length representations and organized using cosine-based hierarchical divisive clustering. This procedure yielded class-specific medoid dictionaries that encode representative dynamic patterns learned directly from the training data. Each epoch was then projected onto these dictionaries via cosine similarity, forming the DMD-CPP feature representation. Linear SVMs were trained under a strict LOSO protocol to ensure subject-wise generalization without information leakage.
To contextualize the proposed approach, two simplified DMD-based pipelines were evaluated: one that retained clustering-based projection on direct mode summaries and another that replaced clustering with PCA-based dimensionality reduction. In addition, the results were compared with a previously reported baseline using classical machine-learning models on short-window PSD features. Together, these comparisons enable the examination of the role of pattern-level clustering, projection, and margin-based reliability independently of nominal accuracy.
4.2. Headline Findings
Across the three binary classification tasks, the proposed DMD-CPP framework exhibited consistent yet strongly task-dependent behavior under EO photostimulation, highlighting limitations of conventional interpretations based solely on signal preservation or spectral clarity.
Under EC conditions, prior studies have consistently reported that AD vs. CN classification achieves the highest performance, supported by relatively stable and disease-specific low-frequency connectivity patterns, while FTD vs. CN shows moderately reduced accuracy due to greater heterogeneity, and AD vs. FTD remains the most challenging task owing to substantial overlap between dementia subtypes [
11,
12,
14,
27]. This EC-based performance hierarchy has often served as an implicit reference for interpreting EEG-based dementia results.
Under EO conditions, however, this hierarchy is typically disrupted. Alzheimer-related oscillatory activity is more severely attenuated and unstable during visual stimulation, leading many EO EEG studies to report poorer performance for AD vs. CN than for FTD vs. CN [
21,
28]. In contrast to this prevailing view, the proposed method yielded a larger relative improvement for AD vs. CN than for FTD vs. CN, elevating AD vs. CN performance to a level comparable with FTD vs. CN. Importantly, this improvement was observed although Alzheimer’s disease is known to exhibit pronounced disruption of visually driven neural dynamics under EO conditions, where conventional physiological signatures are often considered difficult to capture. By comparison, the FTD vs. CN task showed only moderate improvement, with residual variability and weaker separability. This behavior is consistent with the known heterogeneity of FTD and the partial overlap between FTD and control EEG dynamics, which limits the formation of sharply separable representations under clustering-based learning.
Taken together, these findings indicate that the improved margin behavior observed in AD-related tasks cannot be straightforwardly explained by preserved or cleaner EEG dynamics. Rather, they point to the need for an alternative analytical perspective that goes beyond conventional spectral or signal-quality-based accounts of EO EEG, motivating a reconsideration of how disease-related information is expressed and captured under stimulus-driven conditions.
4.3. Interpreting Medoid Patterns
Before discussing the margin asymmetries induced by DMD-CPP, it is necessary to clarify how the learned medoid patterns should be understood.
Figure 3 visualizes class-specific medoids learned by the proposed framework, and the following description should be read as a characterization of observed pattern structure rather than as a definitive physiological interpretation. In this study, medoids are not regarded as biologically prototypical or canonical EEG signatures. Instead, they represent recurrent configurations of oscillatory degradation, reflecting how spatiotemporal organization is altered or weakened under disease-specific conditions during EO stimulation.
Across all three groups (CN, AD, and FTD), a shared characteristic is that higher-frequency components associated with visually driven activity remain relatively localized, whereas lower-frequency structure becomes attenuated, fragmented, or unstable. The primary distinction across groups lies not in the presence of a fixed spectral template but in the manner by which low-frequency information diminishes and interacts with higher-frequency activity.
For both AD and FTD, low-frequency modes are generally reduced in strength relative to CN, and many medoids exhibit limited low-frequency continuity. In the upper-ranked medoids, some patterns from AD and FTD appear visually similar to CN, reflecting subject-specific or idiosyncratic realizations that may contribute to classification ambiguity at the subject level. Such patterns are consistent with the observed difficulty of subtype discrimination and the presence of atypical samples in the AD vs. FTD task. As medoid rank decreases toward more general representatives, clearer group-dependent tendencies emerge. In AD, low-frequency structure often appears increasingly blurred or erased, suggesting a progressive collapse of organized low-frequency dynamics. In contrast, FTD medoids tend to show a more uniform weakening of low-frequency components without converging toward a single dominant degradation pattern. Although these degradation modes differ in appearance, their internal similarity allows clustering to identify representative medoids that summarize common modes of breakdown.
CN medoids exhibit a different behavior. Across ranks, patterns remain heterogeneous and subject-specific, yet low-frequency components retain comparatively richer structure and continuity. Rather than converging toward a failure mode, CN medoids reflect normal inter-individual variability and preserved low-frequency organization supporting transitions to higher-frequency activity.
Taken together, these observations indicate that DMD-CPP medoids do not encode fixed disease templates but instead summarize recurrent ways in which oscillatory structure weakens under EO stimulation. This descriptive characterization provides the basis for understanding subsequent task-dependent margin behavior, which is examined in the following subsection.
4.4. Methodological Interpretation: Clustering-Based Basis Learning Under EO Conditions
Compared with the EC state, the EO photic condition is generally considered more challenging for differentiating AD from healthy controls when using conventional spectral representations [
21,
39]. Under EO stimulation, posterior alpha rhythms are attenuated and less regular, and visually driven responses in AD often exhibit reduced or unstable entrainment, characterized by diminished harmonic power and weakened interhemispheric coherence relative to controls [
19,
40]. Such instability of steady-state visual responses has been interpreted as a breakdown of large-scale synchronization and impaired neural coupling in AD, thereby limiting the discriminative power of short-window spectral features.
As a consequence, features derived directly from short-window spectral power, covariance, or PSD estimates tend to show increased variability and weaker class separation under EO conditions than under EC resting states [
11]. This degradation reflects not only alpha suppression but also the fragmented and stimulus-dependent nature of EO spectral patterns, which vary substantially across epochs and subjects [
20]. Within this context, the clustering component of the proposed framework plays a central methodological role. Rather than assuming the existence of stable or prototypical disease-specific spectral templates, high-dimensional DMD-based representations are aggregated through hierarchical divisive clustering to form a compact dictionary of basis exemplars (medoids). These medoids summarize recurrent modes of spatiotemporal organization present in the training data, including characteristic patterns of oscillatory degradation. Each new epoch is subsequently projected onto these learned bases using cosine similarity, yielding normalized coordinates that quantify structural alignment with class-specific modes of organization or disorganization.
This clustering–projection mechanism is particularly suited to EO conditions, where AD-related responses tend to appear fragmented and inconsistent at the single-epoch level. By emphasizing shared geometric structure across epochs, clustering reduces the influence of transient amplitude fluctuations and inter-subject variability, while projection expresses each sample relative to group-level reference modes. Importantly, these reference modes are not interpreted as biologically preserved signatures but as reproducible structural configurations learned from data.
The larger gains observed for AD vs. CN—relative to the more modest improvements for FTD vs. CN—are therefore consistent with differences in how low-frequency degradation patterns manifest across diseases. In AD, EO-related disorganization appears to recur in a sufficiently structured manner across subjects and epochs to be consolidated through clustering. In contrast, greater heterogeneity and fewer recurring degradation modes in FTD limit the degree to which clustering can stabilize the representation. Overall, these results suggest that clustering-based basis learning can recover discriminative structure from EO EEG by organizing recurring modes of oscillatory disruption that are otherwise treated as noise in conventional analyses.
4.5. Interpretation of Asymmetric Margin Structure
An important observation of this study is the asymmetric margin structure induced by the proposed DMD-CPP framework under EO photostimulation. Specifically, margin separability differs across tasks and classes, with AD-related decisions exhibiting more pronounced and controllable margin behavior than those involving CN or FTD. This asymmetry is not fully explained by nominal classification accuracy alone but reflects how class-dependent structure is organized through clustering-based basis learning.
Clustering-based representations are not inherently robust when classes lack recurrent or consolidatable structure. When patterns are highly heterogeneous or sample sizes are limited, clustering may preserve idiosyncratic realizations rather than forming stable prototypes. From this perspective, the relatively weak margin structure observed for FTD vs. CN is consistent with the known clinical and neurophysiological heterogeneity of FTD, whose EEG alterations are less stereotyped and more regionally variable than those of AD [
41]. Combined with a smaller sample size, this heterogeneity constrains the formation of stable cluster anchors and limits margin-based confidence stratification.
By contrast, the behavior observed for AD suggests a different structural regime. Although AD does not exhibit a stable EEG pattern in the traditional sense, it is associated with relatively consistent breakdowns of large-scale oscillatory coupling, particularly in posterior and long-range networks [
42,
43]. Under EO conditions, this breakdown has been reported as fragmented visual entrainment, reduced phase coherence, and impaired interregional synchronization [
44,
45], while such responses appear irregular at the single-epoch level, their structural characteristics may recur across epochs and subjects. Both DMD + PCA and DMD-CPP capture this separability in the AD vs. FTD task, as reflected by significant margin differences under both representations. This suggests that low-frequency degradation patterns in AD and FTD, although both pathological, remain distinguishable at a structural level. The additional clustering stage in DMD-CPP appears to further organize these recurring AD-related configurations into more stable reference modes.
As a result, control epochs rarely align strongly with AD-related bases. When CN epochs are misclassified as AD, their similarity to AD anchors remains low, leading to systematically smaller decision margins. From a reliability standpoint, this implies that false-positive AD decisions tend to be associated with low confidence, making them amenable to margin-based screening or rejection. Importantly, this behavior should not be interpreted as evidence of preserved biological signatures. Rather, it indicates that the mode of oscillatory breakdown in AD may itself constitute a reproducible structural pattern that clustering-based representations can exploit under strict LOSO validation.
Conversely, AD epochs misclassified as CN exhibit more variable margins, consistent with the diffuse and heterogeneous structure of the CN feature space under EO conditions. Normal responses span a broad range of stimulus-dependent and subject-specific dynamics, limiting the formation of a strong central attractor and constraining margin separation near the decision boundary.
Overall, these results indicate that the primary contribution of DMD-CPP lies not in uniformly increasing accuracy but in reshaping decision geometry in a task- and class-dependent manner. Clustering-based basis learning is most effective when pathological processes give rise to recurrent structural configurations, as observed for AD under EO stimulation. When such recurrence is limited, as in FTD, the benefits of clustering diminish. This behavior underscores both the potential and the limitations of clustering-driven representations under strict LOSO evaluation.
4.6. Limitations
Several limitations of this study should be acknowledged.
First, EEG signals exhibit substantial inter- and intra-subject variability arising from vigilance fluctuations, medication effects, and recording-session factors. Under the LOSO protocol, a small subset of subjects displayed margin patterns that were inconsistent with their clinical labels. Such cases are likely attributable to individual neurophysiological idiosyncrasies rather than to systematic model failure, a well-known challenge in subject-level EEG classification [
46,
47,
48].
Second, although the DMD-CPP framework stabilizes heterogeneous EO responses through clustering-based basis learning, its effectiveness depends on the availability of sufficiently recurrent class-specific structures in the training data. In particular, the smaller sample size and higher physiological heterogeneity of the FTD cohort may have limited the formation of stable and representative prototypes, thereby constraining margin-based separability for FTD-related tasks.
Relatedly, inspection of the learned medoid patterns (
Figure 3) reveals clear class-dependent differences in the spectral breadth and density of DMD modes across frequency bands. AD, FTD, and CN exhibit distinct distributions in terms of how many modes are retained within low- and high-frequency ranges, reflecting different forms of spectral organization or degradation. However, the current framework does not explicitly encode or normalize these band-wise mode-count differences. As a result, subjects whose overall spectral geometry deviates from the dominant class-specific mode distribution—despite sharing similar clinical characteristics—may be misclassified. This limitation suggests that some errors arise not from a lack of discriminative structure but from an incomplete utilization of band-dependent mode complexity in the decision process.
Third, DMD is computationally demanding for long continuous recordings. To ensure tractability, our analysis relied on short (2-s) segments aggregated into 20 s epochs, which may underrepresent slower coupling dynamics or long-range coordination processes [
30]. Future work incorporating multi-scale or adaptive windowing strategies may help capture such effects more comprehensively.
Finally, the present study focused on DMD-derived representations combined with cosine-based clustering and prototype projection. We did not investigate whether the observed margin asymmetries are specific to DMD-based features or would also emerge when similar clustering–projection schemes are applied to alternative EEG representations, such as spectral, covariance-based, or data-driven latent features. Consequently, it remains an open question whether the reliability gains observed here reflect a unique advantage of the DMD-CPP formulation or a more general property of prototype-based feature learning in high-dimensional EEG spaces.
4.7. Positioning with Respect to Prior EEG-Based Dementia Studies
Most prior EEG-based dementia studies have focused on EC conditions, where relatively stable posterior alpha rhythms enable reliable characterization using spectral power, covariance structure, or deep learning-based representations. Under such settings, averaging-based or stationary descriptors are often effective, and extensive comparative evaluations have been reported. However, these assumptions do not readily extend to EO stimulation, where alpha suppression, stimulus-dependent nonstationarity, and fragmented oscillatory responses substantially reduce the robustness of conventional spectral features [
21,
39].
To date, EEG-based dementia classification studies conducted explicitly under EO photostimulation remain extremely limited. For the OpenNeuro ds006036 dataset analyzed in this work, prior EO-based classification studies are essentially restricted to two reports from the data-providing group: an earlier study based on short-window PSD features combined with conventional machine-learning classifiers [
21], and a more recent CNN-based approach [
28]. While these studies demonstrate the feasibility of EO-based discrimination, they primarily report nominal classification performance and do not provide analyses of subject-level reliability, decision confidence, or margin-based error structure under strict LOSO validation. As a result, constructing a direct quantitative comparison across EO-based approaches remains challenging.
In this context, one of the central contributions of the present study is the analysis of decision reliability through subject-level margin statistics under LOSO validation. To the best of our knowledge, margin-based reliability analyses—particularly those examining how decision confidence differs between correct and incorrect subject-level outcomes—have rarely been reported in EEG-based dementia studies, under either EO or EC conditions. Accordingly, there are currently no established benchmarks against which the confidence structure revealed by the proposed framework can be directly compared.
Rather than positioning the proposed DMD-CPP framework as an accuracy-driven alternative to existing methods, we emphasize its role as a reliability-oriented analytical approach for challenging EO EEG data. By organizing DMD-derived representations through clustering-based basis learning and projection, the method provides a structured feature space in which subject-level decision margins can be examined and interpreted. In particular, the framework enables systematic investigation of when margin-based confidence is informative and when it breaks down, as observed in the contrasting behaviors between AD vs. CN and FTD-related tasks. In this sense, the primary contribution of the present work lies not in claiming superiority over scarce EO-based baselines but in offering a principled framework for characterizing the reliability and limitations of EO-based dementia classification, especially for Alzheimer’s disease.
4.8. Future Directions
Building on the present findings, future work will shift the focus from detecting prototypical EEG patterns to explicitly characterizing pathological breakdown of neural dynamics. In EO conditions, Alzheimer’s disease is not primarily associated with the emergence of a stable alternative pattern but rather with the progressive collapse, fragmentation, or loss of low-frequency organization that is typically preserved in healthy subjects. This property fundamentally challenges conventional learning-based classifiers, which are designed to identify recurring and coherent templates.
Methodologically, we aim to develop models that treat such breakdown phenomena themselves as informative structures. Rather than forcing pathological EEG to match a representative prototype, future frameworks will quantify the degree, persistence, and variability of dynamical collapse across time and frequency. In this direction, we will investigate multi-scale and incremental variants of DMD to capture slow drifts, cross-epoch instability, and long-range temporal dependencies without incurring prohibitive computational cost.
Beyond DMD, the proposed reliability-oriented perspective will be extended to a broader class of nonlinear dynamical descriptors, including Riemannian covariance geometry, entropy-based measures, empirical mode decomposition (EMD), and related state-space representations. The goal is not only to improve nominal classification accuracy but to construct unified measures that explicitly distinguish between stable physiological organization and pathological disintegration. Such measures are expected to provide meaningful explanations for both correct and failed classifications, thereby supporting confidence-aware and interpretable decision-making.
5. Conclusions
This study proposed a DMD-CPP framework for analyzing photic-stimulation EEG under EO conditions. By combining mode-based representations with clustering-driven basis learning, the framework constructs a structured and interpretable feature space that is robust to the substantial inter- and intra-subject variability inherent in EO EEG.
Importantly, the primary contribution of this work does not lie solely in algorithmic performance gains but in the insight that Alzheimer’s disease—despite lacking a stable or stereotyped EEG signature—may exhibit a consistent structure of breakdown that is reproducible across subjects. Although individual EO EEG realizations from AD patients differ markedly, the manner in which low-frequency organization degrades and high-frequency components dominate appears to follow a coherent geometric pattern. The DMD-CPP framework captures this phenomenon by learning prototype representations that encode how neural dynamics fail, rather than what a canonical disease pattern looks like.
Unlike conventional spectral or covariance-based approaches, which often treat such fragmented EO responses as noise, the proposed clustering–projection mechanism aggregates these heterogeneous realizations into a set of representative medoid bases. Each epoch is then expressed relative to these bases, enabling reliable subject-level discrimination under strict LOSO evaluation—even when test subjects exhibit patterns not seen during training.
Empirically, this property manifests as strong margin-based reliability for AD-related discrimination. In particular, samples that are ambiguous or atypical (e.g., CN→AD errors) are systematically assigned low decision margins, indicating conservative and controllable behavior rather than overconfident misclassification. Crucially, this margin separation emerges even though LOSO evaluation ensures that test subjects’ EEG patterns differ from those used to construct the medoids. This suggests that the learned representations capture a generalizable failure structure of AD rather than subject-specific or idiosyncratic features.
While several limitations remain—including the computational cost of DMD and the absence of exhaustive comparisons with alternative EEG representations—the present findings demonstrate that prototype-based learning over DMD-derived dynamics offers a robust and interpretable pathway for EEG analysis under challenging EO conditions.