Abstract
Alzheimer’s disease (AD) is a serious neurodegenerative disorder that can severely affect behavior and thinking patterns, and is accompanied by frequent memory loss. The early diagnosis of AD is essential, as this can benefit the patient, but detecting AD is a complex process due to the nature of its associated clinical data. Electroencephalography (EEG) serves as a promising and cost-effective technique for analyzing AD-related brain activity patterns. In this work, a consolidated framework for detecting AD using EEG signals and hybrid models is proposed that uses a dataset that is available online. For the feature extraction module, five efficient techniques—Principal Component Analysis (PCA), Kernel Partial Least Squares (KPLS), Kriging Model, Isomap, and K-means clustering—are used. For feature selection, with the help of biomimetics-based concepts, three efficient algorithms are used: hybrid Cuckoo Search Optimization–Rat Swarm Optimization (CSO-RSO), Zebra Optimization (ZOA), and hybrid Gravitational Search Algorithm–Particle Swarm Optimization (GSA-PSO). Four interesting hybrid classifiers are utilized here to detect AD using EEG signals—hybrid Extreme Learning Machine–Adaboost (ELM–Adaboost), hybrid Classification and Regression Trees–Adaboost (CART–Adaboost), and hybrid weighted broad learning system-based Adaboost (HWBLSA), followed by a hybrid machine learning classification model with a soft voting technique—and, finally, these are compared with other standard machine learning classifiers. The highest classification accuracy of 98.71% is found when the Kriging Model feature extraction concept is combined with the hybrid GSA-PSO feature selection method and classified with the ELM–Adaboost classifier.
1. Introduction
AD, a prominent neurological disorder of the brain, gradually causes degeneration of neuronal cells [1]. Its origins remain unclear, but patients suffering from AD experience symptoms of cognitive decline, memory loss, and sometimes hallucinations. AD is often a progressive disease that develops slowly with few symptoms that do not affect the daily lives of patients, but later can lead to severe deterioration, affecting the patient and their family to a great extent [2]. In the initial stages of AD, its phenotype follows a pattern of Mild Cognitive Impairment (MCI) and is identified by memory loss. Early diagnosis is important so that the condition can be treated and managed efficiently and—to a certain extent—successfully; however, there is no cure for AD [3]. Distinguishing AD symptoms from signs of normal aging is quite difficult; brain tissue is usually examined to diagnose the condition. However, non-invasive techniques are still being investigated that may enable precise and definitive AS diagnoses, and the medical community is hopeful that successful results will be achieved soon [4]. With the aid of various brain imaging techniques alongside psychological tests, AD can be diagnosed, but this still depends on multiple factors including the skill of the neurologist [5]. Understanding and interpreting the correlation between biomarkers in diagnostic tests is still very difficult, and sometimes time-consuming and expensive. Thus, EEG serves as a wonderful tool as it is widely available, easy to use and maintain, and inexpensive; for these reasons, it is widely preferred in the research community [6].
Due to the physiological activity of neurons, electrical potential is generated, allowing this activity to be recorded using EEG [7]. Cell membranes can be depolarized easily, which helps to generate electric currents, creating waves that can be detected using scalp electrodes. The activity of a single neuron cannot be recorded, but when a group of neurons acts synchronously, it can be captured as an EEG signal [8]. A trained neurologist inspects the EEG signals visually, but this is a time-consuming task as it involves a lot of noise and artifacts. Multiple neural functions exhibit non-linear dynamics, and so more sophisticated techniques must be utilized to assess the behavior of such signals [9]. Computational analysis of EEG signals has been very successful, and various techniques have shown promising results in the analysis and classification of brain disorders. Various techniques such as time–frequency analysis, information theory analysis, and graph theory analysis are used to analyze and distinguish healthy and unhealthy subjects [10]. The EEG rhythm of people diagnosed with AD shows a slow pattern, and the complexity of EEG signals can vary greatly. The intention of this study is to apply a framework with efficient feature extraction and selection and hybrid classification models to successfully detect AD [11]. Before presenting the proposed models, the most prominent studies performed in this field for the classification of AD are discussed in the following.
The classification of AD patients with the help of EEG signal processing was reviewed by Fiscon et al. [12]. A comprehensive review on resting-state EEG for diagnosis and progressive assessment of AD was completed by Cassani et al. [13]. EEG signal processing techniques were combined with supervised techniques and discrete Fourier and wavelet transforms for Alzheimer’s patient classification by Fiscon et al., who reported a high accuracy of 92% [14]. EEG modulation spectral patch features were used to diagnose and detect the severity of AD by Cassani and Falk, who obtained an accuracy of 88% [15]. Bi and Wang employed EEG spectral images along with deep learning to analyze early AD diagnosis, where spectral topography maps were analyzed with a spike convolutional deep Boltzmann machine, reporting an accuracy of 95.04% [16]. With reference to physiological aging, Vecchio et al. used many innovative EEG biomarkers along with a machine learning model to classify AD, for which an accuracy of 95% was obtained [17]. Ieracitano et al. employed a novel multimodal machine learning concept using continuous wavelet transform (CWT) with bispectrum features and a Multi-Layer Perceptron (MLP) classifier, yielding a classification accuracy of 89.22% [18]. Conventional machine learning and recurrent neural network (RNN) models were used by Seo et al. for the full classification of AD patients, obtaining an average accuracy of 70.97% [19]. A resting-state EEG signal utilizing CWT tiled topographical images and AlexNet-based Convolutional Neural Networks (CNNs) were used by Huggins et al., who reported an accuracy of 98.9% [20]. For the automated diagnosis of brain disorders like AD and schizophrenia, Alves et al. used the concept of EEG functional connectivity and deep learning, with 100% accuracy [21]. Pirrone et al. described a supervised machine learning concept utilizing power spectrum density, short-time Fourier transform, and K-nearest neighbors (KNNs) for AD detection, reporting a classification accuracy of 86% [22]. The concept of EMD with Hjorth parameters, Kruskal–Wallis analysis, and SVM was analyzed by Puri et al., who reported an accuracy of 92.9% [23]. For the early diagnosis of AD, Rossini et al. developed integrated biomarkers along with machine learning techniques, implementing graph theory with SVM, and reported a classification accuracy of 95% [24].
Dogan et al. employed primate brain pattern-based automated detection of AD using EEG signals, with 100% accuracy reported [25]. A graph neural network (GNN) approach with functional connectivity nodes was employed by Klepl et al. for AD classification analysis, with an accuracy of 84.7% [26]. Some efficient computational techniques for analyzing EEG signals with respect to AD classification were discussed in detail by Vicchietti et al. [27]. Resting-state EEG was used by Zheng et al. to diagnose AD by integrating complexity, spectrum, and synchronization signal features, and a high classification accuracy of 95.86% was reported [28]. Detailed EEG-dependent classification of the stages of AD and MCI was achieved by Calub et al. [29]. With the aid of low-complexity orthogonal wavelet filter banks and SVM, the automatic detection of AD from EEG signals with a classification accuracy of 98.6% was reported by Puri et al. [30]. Sen et al. classified AD using EEG signals and deep learning, in which the concept of proper rotation components was extracted from the EEG signals and then implemented with 1D-CNN. An accuracy of 94% was obtained [31]. The resting-state EEG microstate features for AD classification were obtained with conventional machine learning classifiers, with a high accuracy of 99.22% reported by Yang et al. [32]. To quantify communication between electrode pairs for the efficient classification of AD and frontotemporal dementia, Ma et al. used Support Vector Machines (SVMs), reporting a high accuracy of 96.6% [33]. A comprehensive analysis of the discriminative features of EEG-based classification of AD and frontotemporal dementia was performed by Rostamikia et al. [34]. A Lattice 123 pattern for automated AD using EEG signals was employed by Dogan et al., who reported a classification accuracy of more than 98% [35]. Two channel EEG features were analyzed by Jang et al. for dementia classification with an Extreme Gradient Boosting model, reporting a balanced accuracy of 97.05% [36]. EEG signals and a few images of clock drawing tests were used with ensemble learning for AD classification, and this interesting strategy was proposed by Huh et al. [37]. The concept of dual attention and Optuna-optimized SVM with deep learning methods was presented by Arikan et al. for AD classification using EEG data [38]. The concept of synchrosqueezing transform and deep transfer learning was used for AD detection by Jain and Srivastava, who reported a high classification accuracy of 98.5% [39]. The concept of EEG phase synchronization was used by Cao et al. for AD analysis, in which a brain network analysis was constructed and a graph convolutional network was used with an average classification accuracy of 77.8% [40]. A multiscale temporal deep network was used for AD classifiers for EEG by Zini et al. [41], and in another study utilizing PSO, dimensionality reduction and conventional machine learning classifiers were used by Lopez and Varas, where a classification accuracy of over 95% was obtained [42]. Additionally, hippocampal microstructural signatures that unveil the stage-specific pathways in Alzheimer’s disease progression were discussed by Yu et al. in [43], and the efficacy of brain power mapping concepts with optimized deep learning for analyzing EEG data was discussed in detail by Chen et al. in [44]. In this work, the following workflow is proposed. Once the basic pre-processing step is completed using Independent Component Analysis (ICA), the work is conducted as follows.
- (i)
- For the feature extraction module, five efficient techniques like Principal Component Analysis (PCA), Kernel Partial Least Squares (KPLS), Kriging Model, Isomap, and K-means clustering techniques are used.
- (ii)
- For feature selection, three efficient algorithms are used: hybrid Cuckoo Search Optimization–Rat Swarm Optimization (CSO-RSO), Zebra Optimization (ZOA), and hybrid Gravitational Search Algorithm–Particle Swarm Optimization (GSA-PSO).
- (iii)
- Four interesting hybrid classifiers are utilized here to detect AD using EEG signals: hybrid Extreme Learning Machine–Adaboost (ELM–Adaboost), hybrid Classification and Regression Trees–Adaboost (CART–Adaboost), and hybrid weighted broad learning system-based Adaboost (HWBLSA), followed by a hybrid machine learning classification model with soft voting technique—and finally, these are compared with other standard machine learning classifiers.
Figure 1 is a simplified schematic representation of the entire workflow employed in this research.
Figure 1.
Simplified illustration of the overall workflow.
This manuscript’s structure is organized as follows: Section 2 discusses the feature extraction techniques and Section 3 discusses the feature selection techniques employed in this study. The classifiers used in this work are discussed in Section 4 and the results and a discussion of them are provided in Section 5. This paper is concluded in Section 6.
2. Feature Extraction Techniques Utilized in This Work
The feature extraction schemes PCA, KPLS, Kriging Model, Isomap, and K-means Clustering are explained in this section.
2.1. PCA
For achieving optimal performance, effective features must be chosen, and PCA aids greatly in this process [45]. In the field of multivariate statistics, it is a popular method and the most famously used unsupervised technique for choosing features. The dimensionality of the data is mitigated by PCA so that only the important attributes remain in the data. The total number of variables can be reduced with the utility of orthogonal pairings used with a wide range of dissimilarity. The important subset of the dataset can be chosen by the PCA so that it can be categorized clearly and easily. The main idea of PCA lies in the projection principle. The original data are present where is the number of columns and can be predicted into a specific subspace with fewer size elements, represented as , and the data integrity of it is maintained easily. The implementation of PCA is as follows:
From a set of dimensions to a set of dimensions, the feature dimensionality is mitigated using pre-processing and dimensionality reduction. The mean and variance of the data are standardized during pre-processing, and in the second phase, the construction of the covariance matrix and eigenvectors is framed. Based on the following equation, the mean and standard deviation are computed so that the input features are standardized:
where the number of cases is represented as and the datapoints are represented as . The parameter is substituted with . Unit variance is obtained so that each vector is transformed as follows:
Every is replaced with . The covariance matrix is now calculated as follows:
The eigenvalues and eigenvectors of are computed successfully. The eigenvalues are diminished by managing and setting the eigenvectors. The eigenvectors are chosen and the highest eigenvalues obtained are used to produce the components . Using and the following equation, the data are converted to the new subspace as follows:
where represents a vector representing a sample and represents the converted sample in the novel subspace. The number of attributes contributes a lot to the performance and computation of PCA, as it specifies the datapoints clearly.
2.2. KPLS
When the conventional PLS is improved, it then becomes KPLS [46]. In a high-dimensional space, the input variables present in the latent information can be extracted successfully by introducing the Kernel. It is presumed that the training dataset comprises both input and output variable matrices represented as and , where the sample dimension is indicated as and the sample size is indicated as . The mapping function is indicated as , where the training dataset is mapped to a high-dimensional space from the original space. The mapped vectors corresponding to and are indicated by and , respectively. The expression of Kernel function is as follows:
where the Gaussian Kernel function is indicated as and is expressed as follows:
where the Kernel parameter is specified by . The input variable matrix can be converted into a Gram matrix with the help of the above transformation and is expressed as follows:
The decomposition of the training data based on the non-linear PLS technique is expressed as follows:
For the input variable matrix, the loading matrix is specified as and for the output variable matrix, the loading matrix is specified as . Residual matrices are represented as and , respectively. A significant specification of and is expressed as follows:
where the total number of latent variables is represented as . and are scaled to a zero-mean level so that the decomposition becomes easier and is specified as follows:
where an identity matrix is specified as and it has -dimensionality. A vector with length and value 1 is specified by and a vector with length and value 1 is specified by . For the input variable matrix of both test data and training data, the respective Kernel latent matrix is specified as follows:
2.3. Kriging Model
The random process models can be constructed using an interpolation technique called the Kriging Model [47]. The values of the data correlations are fully analyzed and a spatial correlation is exhibited between the neighboring points so that the variable values can be projected well. A lot of unbiased estimates can be provided by the Kriging Model for the unknown information, as follows:
where the estimated values are represented as . The observed value is represented as and its specific weights are expressed as at location , where the number of sampling points is denoted by and it helps to assess and estimate the total number of measurement points in the technique. Mostly, a semi-variogram function is introduced by the Kriging technique so that the correlation can be characterized accurately. The weights can be computed so that the spatial characteristics of variables can be identified sufficiently by the Kriging method, as follows:
where the two points in space are identified as and . The distance between them is represented as , and the expectation operator is denoted as . At a location , the values of the random variable are specified by .
2.4. Isomap
The concept of classical scaling is quite successful in multiple applications but as it must always retain Euclidean distances, it can be a drawback in some applications. The classical scaling concept does not address the neighboring datapoints and so it finds it difficult to deal with datapoints present near the manifold whose size is much greater than the interpoint distance. To overcome this drawback, Isomap was proposed as it is a versatile technique enabling the curvilinear distance present in between the datapoints to be preserved [48]. Over the manifold, the measurement between two points is analyzed as the distance of curvilinear nature. Under the Isomap concept, the curvilinear distances for datapoint are calculated by means of assembling a neighboring graph so that each datapoint can be connected easily with its respective nearest neighbor in the dataset . Between these two points, the curvilinear distance can be estimated and is represented as the shortest distance between the two points in the entire graph; it is established with the help of Dijkstra’s algorithm. Between all the datapoints in , the curvilinear distance is calculated so that a pairwise curvilinear distance matrix is found. The concept of classical scaling is implemented and the low-dimensional specifications of the datapoints are computed on the obtained curvilinear distance matrix. Though Isomap can sometimes suffer from topological instability, it can still be managed well by adjusting the parameters of Dijkstra’s algorithm slightly.
2.5. K-Means Clustering
The samples with similar properties are collected separately and the samples with dissimilar properties are collected separately with the aid of clustering techniques. In the implementation of data analysis schemes, the K-means clustering algorithm is quite popular as it is easily implemented. K-means clustering is also highly scalable and flexible and so it is applied to multiple objective optimization schemes [49]. In different ways, the centroid of the K-means clustering can be defined easily, as follows:
so that the mean value can be assigned precisely to the different objects in the cluster. The Euclidean distance helps to assess the distance between the specifications of the cluster and the object and it is represented as follows:
The distance between all the datapoints and its respective cluster centers is considered and validated using the clustering quality parameter. The sum of the distance can be mathematically expressed as follows:
where the number of clusters is denoted by . In the cluster, the total number of datapoints is indicated by . The total number of clusters in the K-means clustering technique must be manually specified as it is quite sensitive to the cluster center process initially. The sensitivity measures of the K-means clustering process can be improved greatly but this involves a highly computational procedure. The K-means clustering technique has good global search capabilities and so it can also be implemented for many multi-objective optimization problems.
3. Feature Selection Techniques Utilized in This Work
Biomimetics is an interdisciplinary field that analyzes the time-tested patterns of nature and helps to develop sustainable strategies and innovative solutions to various engineering and mathematical problems [50]. Biological structures and processes are studied thoroughly so that sustainable and highly optimized designs can be developed and implemented successfully in this field. Nature, in a biomimetics context, can be considered a model, mentor, or efficient measure to judge the sustainability of innovations. Biomimetics applications can be implemented in other fields like the natural sciences, architecture, medicine, and robotics [50]. The main feature selection techniques utilized in this work are the hybrid CSO-RSO, ZOA, and hybrid GSA-PSO algorithm.
3.1. Hybrid CSO with RSO Algorithm
3.1.1. CSO
One of the famous optimization algorithms used is CSO and it is highly influenced by the attitude of brood parasitism present in the environment [51]. Cuckoo birds do not have the habit of constructing independent nests as they generally bury their eggs in other nests and nurture them carefully. To trace the nest position, levy flights are utilized instead of random walks. Once the egg is determined by the host as not its own egg, then it can discard the egg based on the situation. The expression of Levy flight is given as follows:
The step length is indicated by and the f Levy circulation is specified as follows:
The step length is indicated by , and the above Equation (23) signifies the mean and its corresponding variance. If the nest is good and has a large majority of support from the birds, it progresses towards the next step. A hybrid of local and global random walks is employed by this algorithm and it can be assessed by the switching parameters.
where the new progression is expressed as and the earlier progression is expressed as . The chosen solutions are expressed as and . The Heaviside side is specified by and the random number is expressed by . For each cuckoo, Levy flight is used to assess the random walk, as follows:
The specifies the random walk length and is often projected as a Levy circulation, which has a variance and mean for each , and is represented as follows:
where the probability of obtaining the Levy random number is specified by .
3.1.2. RSO
Inspired by the behavior of rats, this algorithm was developed and is widely used [52]. Rats display multiple movements such as jumping, chasing, and tumbling, and at times, rats can become aggressive in certain environments based on various triggering factors. To enhance the speed and local search ability of the algorithm, RSO is hybridized with CSO so that the optimal feature subset can be assessed accurately.
3.1.3. Implementation of the CSO-RSO Algorithm
The simplified implementation of the CSO-RSO algorithm is explained as follows.
- Step 1: Start process:
The population of the rats is initialized and then, the population of host nests of cuckoo search is also initialized, where and is represented as follows:
where the rat’s position is assigned as . is a parameter generated to assign the random numbers in the process.
- Step 2: Generation of random numbers:
Once the initialization process is completed, random features are generated with the help of this hybrid algorithm, as follows:
where .
where is a random number assigned in the range of [0, 1]. is responsible for controlling the total number of iterations.
- Step 3: Fitness function evaluation:
For each search agent, the fitness value is estimated. If is identified as a good search agent, then is satisfied. The exploration of a good search agent is completed here.
- Step 4: Update of position of search agent:
With the help of Equation (30), the search agent position of the rat is updated as follows:
where the updated rat location is expressed by .
- Step 5: Optimal feature subset determination:
The search agent fitness value is computed and updated, and finally, is also updated. Using Equations (26) and (30), the best feature subset is chosen.
- Step 6: End of process:
Once the conditions are met, the process is terminated; otherwise, it is repeated until the criteria are met. Therefore, the output of the hybrid CSO-RSO will assess the optimal feature set which is to be fed to the classifiers. Figure 2 shows the overall model of the hybrid CSO-RSO algorithm.
Figure 2.
Overall model of the hybrid CSO-RSO algorithm.
3.2. ZOA
This intelligent optimization algorithm was inspired by the movement exhibited by zebras in nature [53]. Zebras generally live in communities, enjoying the company of each other. However, they exhibit two important survival strategies: the foraging technique and the defense technique. Based on these two survival mechanisms, the ZOA was developed. The implementation of this algorithm can be explained as follows:
Depending on the population concept, ZOA was developed as it comprises many zebra individuals. For representing the solution space, the ideal candidate solutions are identified by these zebra individuals. The values of the related decision variables help to assess the search space position. Every zebra is assessed as a mathematical vector and with the help of a matrix, it can be visually expressed as follows. The population matrix of ZOA is expressed as follows:
where the specification of the zebra population is identified by . The zebra is specified as and the -dimensional decision mode of the zebra is indicated as . In the population, the members are represented as and the total number of decision nodes is represented as . To any problem, a potential solution is specified by every zebra. Once the decision variable value is accessed, its fitness value is computed. A vector is obtained if it is substituted into the fitness function and is expressed as follows:
where the fitness value vector is specified by . For the zebra, the fitness value is indicated by . For each member, the fitness value is evaluated and the candidate solutions are analyzed. Finally, the perfect candidate solutions are accurately identified.
- Phase 1: Foraging Mode
Grass is a staple food for zebras, some of which have peculiar foraging habits. The plain zebra is represented as the pioneer zebra in this work, and they exhibit a poor level of nutrition. The ecological niche is promoted by this behavior in zebras and so a novel foraging space is given to other species that can gnaw on the lower layer of grass, and this group exhibits a high level of nutrition. During the foraging stage, the zebra positions are modeled mathematically as follows:
Before updating the zebra, the value of the -dimension decision mode can be expressed as .
After updating the dimension for the zebra, the updated value is represented as . The random number is represented as and lies in the range of [0, 1]. The pioneer zebra is specified as and the value of its dimension is represented as . , where represents a random number in the range of [0, 1] and so . Depending on the first stage, the updated position of the zebra is specified as . Before updating the zebra, the corresponding fitness value is represented as . The novel fitness value with respect to the novel position of the zebra is represented as . If the obtained fitness value is high, then the updated one replaces the original position.
- Phase 2: Defense Mode
The threat of predators must be efficiently dealt with by the zebras as they face a high threat from lions, cheetahs, hyenas, tigers, wolves, and wild dogs. The zebras utilize the escape mode when facing target predators like lions and tigers and are mathematically expressed as follows:
When faced with the threat of small-sized predators like wild dogs, zebras tend to initiate a system so that they can scare them off easily, and this is represented as follows:
where the -dimension decision node of the zebra is represented as before being updated. The updated value of the -dimension decision node of the zebra is represented as . Constant is represented as and is set as 0.01 in our experiment and the random number is represented by and is in the range of [0, 1]. The current iteration is represented by and maximum iteration is specified as . specifies the randomly chosen individual zebra and expresses the value of its dimension. , when is a random number in the range of [0, 1] and so . The idea of choosing any of the strategies is detected by and is done in the interval of [0, 1], which is again randomly generated. Depending on the second stage, the zebra’s updated position is indicated as . Before updating the zebra, the corresponding fitness value is identified as . Based on the novel positions of the zebra, the respective fitness value is expressed by . The updated position can replace the original position only if the new fitness value is better; otherwise, it remains the same.
3.3. GSA-PSO Algorithm
The GSA and PSO are hybridized together, the process of which is explained in the following section.
3.3.1. GSA
Based on Newton’s gravity law and the laws of motion, GSA was developed [54]. With the aid of particle movements, the initial best solution is found by this algorithm in the entire group. Based on the gravity law, the progression of the mutual attraction within the particles is achieved. Depending on this rule, the particles can be effectively searched for in the process. With the movement of the particles, there is good improvement in the optimal solution and it can rise with the total number of iterations. We can assume there is a system which comprises particles in the entire search space, and it is represented as follows:
where the location of a particular particle is represented as . Among the particles and , the gravitation at a particular iteration is expressed as follows:
where indicates the inertial masses of the active force particle and indicates the inertial masses of the passive force particle . The Euclidean length is specified by and a small constant is specified by . At a particular iteration , the attractional constant is expressed by and is mathematically modeled as
The total amount of iterations is expressed by . The and are constant values and can have values ranging from [20 to 100], respectively. To enhance the stochastic properties of the algorithm, the total force equals the sum of forces of all the other particles in the GSA and is represented as follows:
where the stochastic gradient is expressed as and lies in the range of [0, 1]. For the particle at iterations, the acceleration is expressed by and represented as follows:
where, for the particle, the inertia mass is represented by . The velocity and phase of the particle is renewed after every iteration and it is represented as follows:
Depending on the fitness value, the computation of inertia and gravity mass is as follows:
where the fitness value of the particle is represented as during the iteration process. For a maximization problem, the following equation is utilized in GSA as follows:
3.3.2. PSO
To obtain the optimal solution in the PSO, the movement of the birds flying is utilized thoroughly in the search space [55]. The optimal solution is nothing but the best particle and it is obtained along its path. In a target space, a population of particles is present with dimension , where the particle indicates a vector with a specific dimension and is represented as follows:
where the position of the particle is indicated by . For the particle, the flight speed is indicated as follows:
where the particle speed is detailed by . The particle searches for the best position and is denoted as and represented as follows:
where the particle’s best position is represented as . The best position for all the particles is represented as and expressed as follows:
where the optimum value for all the particles is expressed as . Once the optimal values are traced, then the speed and particles are renewed as follows:
where and specify the study rate, and and specify the stochastic random process lying in the range of [0, 1].
3.3.3. Hybridizing the GSA-PSO Algorithm
The inherent advantages of both PSO and GSA are combined; then, this hybrid algorithm is built. A simplified illustration of this is shown in Figure 3. The new updated speed and positions are modeled as follows:
where the acceleration of particle is expressed as during an iteration . The overall procedure of GSA-PSO is as follows:
Figure 3.
Simplified illustration of the GSA-PSO algorithm.
- (1)
- All the particles are randomly initialized.
- (2)
- The fitness evaluation function is constructed.
- (3)
- The parameter value fed to every classifier is optimized by the feature selection optimization technique.
- (4)
- The construction of fitness function is as follows:where the amount of data is represented by and the entire error is denoted by . denotes the total number of subsets.
- (5)
- A fitness value is attained for every iteration.
- (6)
- Then, the fitness value of the present iteration is compared to the fitness value of the previous iteration and only the best fitness value is considered and updated.
- (7)
- Compute and renew the values of based on Equations (39)–(42), (46), and (51).
- (8)
- Compute the speed of the particle based on (54).
- (9)
- Compute the position of the particle based on (55).
- (10)
- Implement step 2 to step 9 until the stopping criteria is met.
- (11)
- Obtain the best solution and terminate the process.
4. Classifiers Utilized in This Work
In this study, we utilize four hybrid classifiers: the Collaborative Adaboost–ELM classification model, the CART-based Adaboost classification model, the hybrid weighted broad learning system-based Adaboost (HWBLSA) classifier, and a hybrid machine learning classification model with a soft voting technique.
4.1. Collaborative Adaboost–ELM Algorithm
This section explains the nuances of the ELM, Adaboost, and the Collaborative Adaboost–ELM classification model.
- (A)
- ELM:
A widely used machine learning technique is ELM [56]. There is always a random distribution of the input weights and bias of the hidden layer in this algorithm and so it differs greatly from the conventional machine learning techniques. Also, there is no need to adjust these parameters often, which is one of its greatest advantages. A high learning speed is obtained by ELM and good performance is achieved often when using this technique. The basic details of ELM are as follows:
Assuming a dataset with distinct samples , the output function can be represented for hidden nodes and an activation function , as follows:
The threshold of the hidden layer is indicated as and the output is denoted as . The output weights and input weights are specified as follows:
A zero error should be obtained between and parameters, and it is represented as follows:
- (B)
- Adaboost:
The Adaboost algorithm is a successful pattern recognition algorithm, in which multiple weak predictions are hybrid so that a strong predictor can be established effectively [57]. For the samples, the distribution weights can be high when the error is maximum when training the Adaboost. Vice versa, the distribution weights can be low when the error is minimum when training the Adaboost. Depending on the distribution of new weight, the samples are trained so that the predicted output can be largely improved. The computation and procedure are as follows:
- Step 1: Assume a sample set . For the samples, the weight distribution is specified as at the iteration.
- Step 2: At the initial iteration, the weight distribution is assigned as when .
- Step 3: Depending on the distribution of weights, the predicted output of is computed.
- Step 4: The forecasting error can be computed as
- Step 5: The proportional error is computed as
- Step 6: The connection weights are computed aswhere .
- Step 7: The weight distribution is now updated as
- Step 8: The final predictions after a certain number of iterations is attained by
- (C)
- Collaborative framework model of ELM–Adaboost:
The framework of this Collaborative Adaboost–ELM model is shown in Figure 4.
Figure 4.
Simplified illustration of the Collaborative Adaboost–ELM model.
- Step 1: The features are fed inside the collaborative framework model as test and training sets.
- Step 2: To deal with the mathematical manipulation of the dataset, it is entirely normalized with the following equation:where denotes the data present before normalization and denotes the data obtained after normalization process. and indicate the minimum and maximum values of the original feature set.
- Step 3: To establish multiple ELMs, the Adaboost algorithm is used, which helps to control the weight of various ELMs.
- Step 4: The predicted output of ELMs is computed so that the relevant errors are calculated effectively.
- Step 5: Once this is done, the weight distributions are updated, respectively.
- Step 6: The individual outputs of the ELMs are summarized successfully with the aid of connection weights.
- Step 7: Finally, a good ensemble output is obtained in terms of the desired performance metric.
4.2. CART-Based Adaboost Classifier
To analyze Alzheimer’s disease classification, the CART-based Adaboost classifier is used. The performance of the classifier is improved by integrating CART with the Adaboost algorithm.
- (A)
- CART procedure:
For regression and classification problems, one commonly used machine learning technique is the decision tree algorithm. A flowchart-like tree model of this type is initially generated so that the algorithm partition is generated successfully. A root node represents the start in the decision tree and then it is followed by child and leaf nodes so that it can have all the essential data required to train the model. The data are partitioned in a recursive manner and assessed by the tree structure so that the splitting criteria can be successful, as the goal is to obtain a leaf node. For the analysis of a decision tree, splitting criteria are highly necessary. CART is one such type of decision tree algorithm implemented to successfully prove the splitting criteria [58]. The Gini Index (GI) is utilized as the splitting criterion, which analyzes the impurity of a sample to a large extent. For a sample set in a specific node, the GI is computed as follows:
To achieve high-purity samples, the GI should be small. In every child node, the weighted average of the GI is reduced by the splitting quality attribute. The maximization of is expressed as follows:
where the number of categories is indicated as in the node and represents the attribute of the data split mode. The proportion of the category in the node is represented by . The difference in impurity before and after a split is indicated by . The sample set of the left child node is represented by and the proportion of the left child node is represented by . The sample set of the right child node is represented by , and the proportion of the right child node is represented by . In the CART technique, there is no necessity to choose independent variables initially, which represents one of the greatest advantages of using the CART technique. When every node undergoes splitting, the best variables are chosen and utilized. When the correlation between the input and output parameters is not known, CART can be utilized as it interprets the results quantitatively and qualitatively and so it is used for both regression and classification purposes.
- (B)
- Implementation of CART–Adaboost classifier:
A single CART model cannot perform exceedingly well as it has some constraints due to the limited structure of the generated trees. Good accuracy is not obtained if the tree structure is very simple and so ensemble techniques are utilized here in the form of the CART-based Adaboost classification model, where the base learner of the Adaboost algorithm is found by the CART procedure. A simplified illustration of this procedure is shown in Figure 5. To improve the accuracy of the model, the sample weights are adjusted continuously during training of the CART–Adaboost classification process. To obtain the ultimate prediction result, the weightage of every output in CART is given a lot of importance in this classification model.
Figure 5.
Simplified illustration of CART–Adaboost classifier.
4.3. Hybrid Weighted Broad Learning System-Based Adaboost (HWBLSA) Classifier
The concept of the weighted broad learning system is integrated with the Adaboost classifier and it forms an HWBLSA classifier.
- (A)
- Concept of Weighted Broad Learning System
There are three layers present in the broad learning system (BLS), and the enhancement nodes and feature nodes are contained in the hidden layer [59]. The training samples are represented as and its respective label is expressed as . From the feature nodes, the input samples are proceeded as follows:
where the mapping function is indicated as . The bias is specified as and the generated weight is represented in a random manner as . The total number of mapping windows is concatenated in the row direction and now, the feature nodes can be specified as follows:
Through efficiently managing feature nodes, enhanced nodes can be attained as follows:
where the activation function is represented as . The randomly generated bias is specified as and the randomly generated weight is specified as . The attainment of enhancement nodes can be expressed as follows:
The following relationship has been developed by BLS and it is expressed as follows:
where the true value of samples is represented as . The weight expresses the output weight and is directly connected to the output layer. The weight serves as a bridge between the enhancement and feature nodes and can be solved by ridge regression, as follows:
where the regularization parameter is represented as .
For the datasets whose samples are imbalanced, an accurate prediction is not possible using general models. So, cost-sensitive techniques are used to improvise them and, with the aid of cost matrices, the classification issue can be solved. The original cost-sensitive matrix is expressed as follows:
where expresses the risk of misclassification of a sample and . When the cost of the matrix increases, the error discrimination can be largely reduced. For the BLS, the optimization of the objective function is expressed as follows:
BLS incorporates the cost-sensitive matrix and so it is expressed as follows:
where the cost-sensitive matrix can be generated in a random manner. The classification error can be alleviated easily when using this method and so imbalanced datasets can be dealt with very well. The training error is accumulated iteratively and the majority of the class samples will be trained easily. The samples with the majority class are generally favored by the classification results. The samples which have more error space are generally considered as minority class samples. A huge weight is picked for the minority class samples and a small weight is picked for the majority class samples. Over time, a training bias can be managed well between the minority and majority samples. A weight is multiplied to every sample so that a good training bias is achieved and is expressed as follows:
where is the original norm condition. The total number of samples is expressed by . The average of the entire classes is expressed by the parameter. The extent of continuous equilibrium of the sample is measured by the term . To the samples of the majority class, the boundary is extended and its value is always small in between the minority and majority classes’ samples, so that good accuracy can be achieved. The objective function can be modified by the addition of weights so that it is expressed as follows:
where , , .
The ultimate output weight of the WBLS is expressed as follows:
- (B)
- HWBLSA classifier
Multiple individual classifiers are combined so that a robust classifier can be created; in machine learning terms, this is known as ensemble learning. For the various classifiers, the diversity can be leveraged so that the classification performance can be greatly improved when compared to that of a single classifier. Bagging, boosting, and AdaBoost are some commonly used ensemble algorithms. The weak classifiers are trained iteratively by this AdaBoost algorithm so that their predictions are hybrid, allowing them to obtain a strong classifier. Weight can be adjusted adaptively by AdaBoost so that higher weights can be assigned to misclassified samples. To the samples which are currently classified, lower weights can be assigned in the process. During every iteration, the samples are challenged so that the overall accuracy is enhanced. Through a process called weighted voting, the weak classifiers are combined by Adaboost so that their collective predictions are used and a versatile classifier is created. The weighted sum or contributions of the independent classifiers are considered so that the final decision can be made and the accurate ones can be given more priority. An exponential loss function is utilized on the sample set, which is expressed as follows:
where denotes the number of classification categories. To assess the base classifiers, the sample set type is analyzed and the probability magnitude is measured directly in this process. In the dataset, the imbalance characteristics are tried and balanced as much as possible so that the minority classes can be better discriminated. The distribution characteristics are analyzed well by suitable weights so that a higher accuracy is obtained. This is achieved by hybridizing the Adaboost concept with the WBLS. The base classifier considered is WBLS, and after several iterations using multiple base classifiers, it can ultimately be integrated as a robust HWBLSA classifier.
- (C)
- HWBLSA implementation
The weights are initialized by the HWBLSA classifier. The base classifier is iterated successfully a total of times, where the inputs are given as and , respectively, from . WBLS is considered the initial base classifier and an approximate probabilistic model is generated as follows:
With the help of a probabilistic model, the new classifier can be synthesized as follows:
The weights are updated as follows:
where the weights are normalized utilizing .
Until the required number of iterations are obtained, the above process is continued. For the classifier, the final output is represented as follows:
4.4. Hybrid Machine Learning Models with Soft Voting Technique
Machine learning plays a major role in the prompt detection and analysis of Alzheimer’s disease using EEG signals. Correlations and patterns can be discerned easily by machine learning models and so it is highly useful in analyzing any disease. The collective strength of the different machine learning models can be leveraged; thus, a good classification accuracy can be achieved with these ensemble learning techniques. These hybrid models can be fine-tuned by adjusting the hyperparameters depending on the necessity of the datasets and feature selection algorithms used. In this work, Logistic Regression (LR), SVM, Random Forest (RF), and KNN are used and hybridized together with a soft voting technique.
By means of employing the soft voting criteria, the final prediction can be assessed depending on the class with the largest probability [60]. When these models are hybridized together, the bias and variance can be reduced and so the overall robustness and accuracy can be improved well. A reliable and balanced decision can be made when employing the soft voting technique, as it ensures that there are very few or negligible misclassifications since multiple classifiers are engaged in determining the output. Therefore, an accurate and stable prediction analysis can be performed; thus, it is employed widely in the medical diagnosis of many disorders, outperforming the standalone classifiers or conventional models. Depending on the inherent characteristics of machine learning models, some can struggle with overfitting problems, an inability to deal with complex patterns, and interpretability issues, but once all the classifiers are hybrid and a soft voting strategy is implemented, the strength of these classifiers is then leveraged and a solid output can be bought out, implying a very good performance. Figure 6 shows a simplified illustration of the proposed hybrid machine learning model with the soft voting technique.
Figure 6.
Simplified illustration of the hybrid machine learning model with soft voting technique.
5. Results and Discussion
To analyze the proposed framework, a publicly available EEG signal dataset for detecting and classifying AD was utilized [61]. EEG signals were acquired from nine participants (eight healthy individuals and one with AD). The participants were made to sit comfortably in front of a monitor in a dimly lit room. A good distance was maintained between the participants when analyzing the experiments. A cathode ray tube (CRT) monitor was used to present a visual stimulus with the help of E-prime 2 software so that eye movements could be monitored easily. Each trial started with a fixation cross that was followed by a stimulus presentation accompanied by a warning tone. The participants completed a discrimination task, and the stimulus was presented for 300 ms. In experiment 1, the participants were asked to indicate whether the presented stimulus was a house, a face, or a scrambled image. In experiment 2, the stimuli comprised visual faces with fearful or neutral expressions, and in experiment 3, the stimuli were unfamiliar or very famous faces. Detecting amnesia or agnosia in the broader context of face recognition deficits in patients with AD using EEG signals was the main intention of the work, and so for convenience, experiment 1 employed 1225 healthy classes and 325 AD sample values and experiment 2 utilized 1200 healthy classes and 350 AD sample values. The overall dataset was quite small and imbalanced, so the proposed work was versatile, with no need to incorporate deep learning as the overall data size was very small. A repeated 10-fold cross-validation technique was employed throughout the work for the analysis. For performing the experiment, MATLAB R2022a software installed on a system with 32 GB of main memory, an i5 processor, and the Windows 11 Operating System was utilized.
As far as the KPLS technique is concerned, a Gaussian Kernel was utilized, and for the Kriging Model, the length scale was set to 0.5 and the overall variance was set to 0.2 for the entire process. The number of neighbors used in the Isomap algorithm was set as 10 and for the K-means clustering technique, the number of clusters was set to around 20 in the experiment. For the CSO-RSO algorithm, the parameters of CSO were set as follows: the population size was 100, the step size was 0.5, the Levy exponent was 1.5, and the probability of discovery was assigned in the range of [0, 1]. For RSO, the population size was again set as 100 and the total number of iterations was expressed as 200 and the number of dimensions was assigned depending on the distance and search space. The hunting mode of rats was set in the range of [0, 1]. For ZOA, the population size was set as 50, the maximum iteration was set as 100, and the random number range was set in the order of [0, 1], which helped model the movement of the zebras. For the hybrid GSA-PSO algorithm, the parameters of GSA were as follows: the population size was set as 100 and the gravitational constant was set as 0.5, the maximum iterations were set to 200, and the acceleration coefficients were set in the range of [0, 10] depending on the search space boundaries. For PSO, the population size was assigned as 100 and the inertia weight was initially set to 0.5. The cognitive and social coefficient values were assigned as 0.2 and 0.4, respectively, and the maximum number of iterations was set to 200 in the experiment. For the LR classifier, the multinomial class was selected and for SVM, the Kernel used in this work was polynomial. The value of C utilized in the SVM classifier was set to 2 so that fewer errors were obtained in the process and a good regularization could be achieved. For RF, the number of estimators used is 250 and the max_depth parameter set was 100, and for the KNN classifier, the total number of weights set was 5. The maximum iteration limit of 500 was set in the conventional machine learning classification process so that a good convergence was achieved.
According to the analyses completed on the data for experiment 1, a high classification accuracy of 92.22% is obtained when Kriging Model feature extraction is combined with CSO feature selection and classified with the CART–Adaboost classifier (Table 1); a high classification accuracy of 88.23% is obtained when KPLS Model feature extraction is combined with RSO feature selection and classified with the HWBLSA classifier (Table 2); a high classification accuracy of 97.89% is obtained when Kriging Model feature extraction is combined with hybrid CSO-RSO feature selection and classified with the ELM–Adaboost classifier (Table 3); a high classification accuracy of 95.98% is obtained when KPLS Model feature extraction is combined with ZOA feature selection and classified with the HWBLSA classifier (Table 4); a high classification accuracy of 93.89% is obtained when Kriging Model feature extraction is combined with GSA feature selection and classified with the ELM–Adaboost classifier (Table 5); a high classification accuracy of 92.56% is obtained when Kriging Model feature extraction is combined with PSO feature selection and classified with the HWBLSA classifier (Table 6); and lastly, a high classification accuracy of 96.56% is obtained when Kriging Model feature extraction is combined with hybrid GSA-PSO feature selection and classified with the HWBLSA classifier (Table 7).
Table 1.
Analysis of accuracy of the CSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 1.
Table 2.
Analysis of accuracy of the RSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 1.
Table 3.
Analysis of accuracy of the hybrid CSO-RSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 1.
Table 4.
Analysis of accuracy of the ZOA feature selection scheme with the feature extraction models and hybrid classifiers—experiment 1.
Table 5.
Analysis of accuracy of the GSA feature selection scheme with the feature extraction models and hybrid classifiers—experiment 1.
Table 6.
Analysis of accuracy of the PSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 1.
Table 7.
Analysis of accuracy of the Hybrid GSA-PSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 1.
According to the analysis of the data for experiment 2, a high classification accuracy of 94.89% is obtained when KPLS feature extraction is combined with CSO feature selection and classified with the HWBLSA classifier (Table 8). A high classification accuracy of 91.09% is obtained when RSO feature extraction is combined with RSO feature selection and classified with the HWBLSA classifier (Table 9). A high classification accuracy of 98.65% is obtained when KPLS feature extraction is combined with hybrid CSO-RSO feature selection and classified with the ELM–Adaboost classifier (Table 10). A high classification accuracy of 96.44% is obtained when KPLS Model feature extraction is combined with ZOA feature selection and classified with the ELM–Adaboost classifier (Table 11). Table 12 shows that a high classification accuracy of 91.78% is obtained when KPLS Model feature extraction is combined with GSA feature selection and classified with a hybrid classifier with the soft computing technique (Table 12). A high classification accuracy of 90.78% is obtained when Kriging Model feature extraction is combined with PSO feature selection and classified with the ELM–Adaboost classifier (Table 13). Lastly, a high classification accuracy of 98.71% is obtained when Kriging Model feature extraction is combined with hybrid GSA-PSO feature selection and classified with the ELM–Adaboost classifier (Table 14).
Table 8.
Analysis of accuracy of the CSO feature selection scheme with feature extraction models and hybrid classifiers—experiment 2.
Table 9.
Analysis of accuracy of the RSO feature selection scheme with feature extraction models and hybrid classifiers—experiment 2.
Table 10.
Analysis of accuracy of the hybrid CSO-RSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 2.
Table 11.
Analysis of accuracy of the ZOA feature selection scheme with the feature extraction models and hybrid classifiers—experiment 2.
Table 12.
Analysis of accuracy of the GSA feature selection scheme with the feature extraction models and hybrid classifiers—experiment 2.
Table 13.
Analysis of the accuracy of PSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 2.
Table 14.
Analysis of accuracy of the hybrid GSA-PSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 2.
On examining Figure 7, it is evident that a high classification accuracy of 97.89% is obtained when Kriging Model feature extraction is combined with hybrid CSO-RSO feature selection and classified with the ELM–Adaboost classifier when analyzing the data for experiment 1. On examining Figure 8, a high classification accuracy of 96.56% is obtained when Kriging Model feature extraction is combined with hybrid GSA-PSO feature selection and classified with the HWBLSA classifier when analyzing the data for experiment 1. In Figure 9, it can be seen that a high classification accuracy of 96.44% is obtained when KPLS Model feature extraction is combined with ZOA feature selection and classified with ELM–Adaboost classifier when analyzing the data for experiment 2. When Figure 10 is examined, a high classification accuracy of 98.67% is obtained when Kriging Model feature extraction is combined with hybrid GSA-PSO feature selection and classified with the ELM–Adaboost classifier when the analysis is completed on the data from experiment 2.
Figure 7.
Performance analysis of accuracy of the hybrid CSO-RSO feature selection scheme with the feature extraction models and hybrid classifiers—Experiment 1.
Figure 8.
Performance analysis of accuracy of the hybrid GSA-PSO feature selection scheme with feature extraction models and hybrid classifiers—experiment 1.
Figure 9.
Performance analysis of accuracy of the ZOA feature selection scheme with the feature extraction models and hybrid classifiers—experiment 2.
Figure 10.
Performance analysis of accuracy of the hybrid GSA-PSO feature selection scheme with the feature extraction models and hybrid classifiers—experiment 2.
As far as statistical analysis is concerned, when a two-sided Wilcoxon test was executed, a good confidence level was achieved for all the features. When the Friedman test was executed, the feature values developed a good variation amongst themselves, thereby proving fit for classification. Since only machine learning techniques were used in the experiment, the overall computational complexity was attained at only. The computational time for the proposed models was calculated as follows:
As is evident from Table 15, a low computational time of 5.003 s was obtained for the combination of KPLS + hybrid CSO-RSO feature selection + ELM–Adaboost classification model. The next best computational time of 5.609 s was obtained for the combination of the Kriging Model + hybrid GSA-PSO feature selection + ELM–Adaboost classification model. A comparatively high computational time of 10.914 s has been obtained for the combination of KPLS + RSO feature selection + HWBLSA classification model.
Table 15.
Computational time analysis for the best-performing models.
5.1. Comparison with Previous Works
Only one study has been reported that has used this specific dataset previously for AD detection and classification for the sake of performance comparison. However, a few important results for AD detection with other datasets are discussed for the readers’ detailed understanding. In reference [15], 20 healthy controls and 34 patients with AD were used and spectral feature extraction with the SVM technique was employed; the authors reported a classification accuracy of 88.1%. In [17], 120 healthy controls and 175 patients with AD were used and electromagnetic tomography with SVM was employed; a classification accuracy of 95% was reported. In [21], 24 healthy and 24 patients with AD were analyzed and the method employed was Pearson’s correlation with CNN; 100% accuracy was achieved. In [25], 11 healthy controls and 12 patients with AD were used and the method employed was a novel primate brain pattern with KNN classifier; the authors obtained a classification accuracy of 100%. In [35], eight healthy controls and one patient with AD were analyzed and the method employed the Lattice 123 concept with standard machine learning classifiers; the analysis showed a classification accuracy of 98.37% for experiment 1 and of 99.62 for experiment 2. We compared our results to those obtained by the authors of [35], as we used a similar dataset to that used by them and the results were computed and compared only for experiments 1 and 2. The proposed framework showed that the best results were as follows: The highest classification accuracy of 98.71% was obtained when Kriging Model feature extraction was combined with hybrid GSA-PSO feature selection and classified with the ELM–Adaboost classifier. This is according to the analysis of the dataset for experiment 2. The second highest classification accuracy of 98.65% is obtained when KPLS feature extraction is combined with hybrid CSO-RSO feature selection and classified with ELM–Adaboost classifier when the analysis is done for experiment 2 of the dataset. The third highest classification accuracy, of 97.89%, was obtained when Kriging Model feature extraction was combined with hybrid CSO-RSO feature selection and classified with ELM–Adaboost classifier when the data for experiment 1 were analyzed. The fourth highest classification accuracy of 96.56% was obtained when Kriging Model feature extraction was combined with hybrid GSA-PSO feature selection and classified with the HWBLSA classifier when the analysis is completed for experiment 1 of the dataset. The obtained results show good quality in terms of classification accuracy and seem to be robust and versatile, and so this framework has the potential to be applied to analyze other neurological disorders as well.
5.2. Study Limitations
For an ensemble model involving feature extraction, feature selection using standard methods, and biomimetics-based models followed by classification using conventional and hybrid machine learning models, one must pay careful attention to every module to avoid producing erroneous results. When dealing with feature extraction, issues may arise due to generalization, lack of interpretability, information loss, and computational intensity; thus, careful attention must be paid to it. Also, when dealing with feature selection, problems may occur such as a high cardinality bias, an inability to handle multicollinearity issues, overfitting issues, subjective bias, and unclear causality; therefore, appropriate measures must be taken when selecting efficient features. Any time the concept of biomimetics is used in research, various factors must be analyzed, such as sustainability issues, evolution constraints, knowledge gaps, complexity measures, and application bottlenecks. Thus, the application of biomimetics in certain fields needs careful analysis and experimentation. Finally, when machine learning is used, factors such as algorithm bias, computational cost and complexity, and time, as well as its ability to adapt dynamically, must be considered carefully when analyzing the experiment. A high level of analytical logic is always required when handling ensemble models, as hybridization and combinations are employed to ensure that the best results are consistently produced.
6. Conclusions and Future Work
AD is a neurological disease and patients suffering from it exhibit many symptoms, including memory loss, hallucinations, and the inability to perform basic daily tasks. Various genetic and environmental factors, as well as trauma and age, among others, contribute to the development of this disease. There is no specific or definite diagnostic test for AD; based on certain biomarkers and the clinician’s assessment, a treatment is decided upon. Artificial intelligence (AI)-based automated detection of AD assisted by EEG signals is widely used in the scientific research community. In this work, a consolidated framework is proposed that employs very interesting feature selection techniques and hybrid classifiers along with feature extraction schemes, and it is tested on an AD database. Following analysis, the proposed framework showed that the highest classification accuracy, at 98.71%, is achieved when Kriging Model feature extraction is combined with hybrid GSA-PSO feature selection and classified with the ELM–Adaboost classifier. The second highest classification accuracy, of 98.65%, is obtained when KPLS feature extraction is combined with hybrid CSO-RSO feature selection and classified with the ELM–Adaboost classifier. Future research should aim to modify the feature selection schemes and incorporate hybridization in the classification methods, so that a much higher accuracy can be achieved. Future studies should also aim to improve the overall model in terms of stability, reliability, and overall versatility, such that it can be applied to a variety of other biosignal processing and medical image processing datasets to analyze other neurological disorders and health issues.
Author Contributions
S.K.P.—conceptualization, methodology, software, validation, formal analysis, investigation, resources, data curation. D.-O.W.—validation, formal analysis, writing—review and editing, visualization, supervision, project administration, funding acquisition. All authors have read and agreed to the published version of the manuscript.
Funding
This research was supported by Hallym University Research Fund, 2026 (HRF-202603-009).
Data Availability Statement
A publicly available dataset was used to perform the experiment and the database can be found in “Mazzi C, Massironi G, Sanchez-Lopez J, De Togni L, Savazzi S (2020) Face recognition deficits in a patient with Alzheimer’s disease: amnesia or agnosia? The importance of electrophysiological markers for differential diagnosis. FrontAgingNeurosci12:580609” [61].
Conflicts of Interest
The authors declare no competing interests.
References
- Palacios-Navarro, G.; Buele, J.; Gimeno Jarque, S.; Bronchal García, A. Cognitive Decline Detection for Alzheimer’s Disease Patients Through an Activity of Daily Living (ADL). IEEE Trans. Neural Syst. Rehabil. Eng. 2022, 30, 2225–2232. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Marvi, F.; Chen, Y.-H.; Sawan, M. Alzheimer’s Disease Diagnosis in the Preclinical Stage: Normal Aging or Dementia. IEEE Rev. Biomed. Eng. 2025, 18, 74–92. [Google Scholar] [CrossRef] [Scilit]
- Eke, C.S.; Jammeh, E.; Li, X.; Carroll, C.; Pearson, S.; Ifeachor, E. Early Detection of Alzheimer’s Disease with Blood Plasma Proteins Using Support Vector Machines. IEEE J. Biomed. Health Inform. 2021, 25, 218–226. [Google Scholar] [CrossRef] [Scilit]
- Atitallah, S.B.; Driss, M.; Boulila, W.; Koubaa, A. Enhancing Early Alzheimer’s Disease Detection Through Big Data and Ensemble Few-Shot Learning. IEEE J. Biomed. Health Inform. 2025, 29, 6451–6462. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Escudero, J.; Ifeachor, E.; Zajicek, J.P.; Green, C.; Shearer, J.; Pearson, S. Machine Learning-Based Method for Personalized and Cost-Effective Detection of Alzheimer’s Disease. IEEE Trans. Biomed. Eng. 2013, 60, 164–168. [Google Scholar] [CrossRef] [Scilit]
- Kwak, M.G.; Mao, L.; Zheng, Z.; Su, Y.; Lure, F.; Li, J. A Cross-Modal Mutual Knowledge Distillation Framework for Alzheimer’s Disease Diagnosis: Addressing Incomplete Modalities. IEEE Trans. Autom. Sci. Eng. 2025, 22, 14218–14233. [Google Scholar] [CrossRef] [Scilit]
- Mitra, U.; Rehman, S.U. ML-Powered Handwriting Analysis for Early Detection of Alzheimer’s Disease. IEEE Access 2024, 12, 69031–69050. [Google Scholar] [CrossRef] [Scilit]
- Irfan, M.; Shahrestani, S.; Elkhodr, M. Early Detection of Alzheimer’s Disease Using Cognitive Features: A Voting-Based Ensemble Machine Learning Approach. IEEE Eng. Manag. Rev. 2023, 51, 16–25. [Google Scholar] [CrossRef] [Scilit]
- Hassan, N.; Miah, A.S.M.; Suzuki, T.; Shin, J. Gradual Variation-Based Dual-Stream Deep Learning for Spatial Feature Enhancement with Dimensionality Reduction in Early Alzheimer’s Disease Detection. IEEE Access 2025, 13, 31701–31717. [Google Scholar] [CrossRef] [Scilit]
- Gallego-Viñarás, L.; Mira-Tomás, J.M.; Gaeta, A.M.; Pinol-Ripoll, G.; Barbé, F.; Olmos, P.M.; Muñoz-Barrutia, A. Alzheimer’s Disease Detection in EEG Sleep Signals. IEEE J. Biomed. Health Inform. 2025, 29, 948–959. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Li, W.; Zhao, Y.; Chen, X.; Xiao, Y.; Qin, Y. Detecting Alzheimer’s Disease on Small Dataset: A Knowledge Transfer Perspective. IEEE J. Biomed. Health Inform. 2019, 23, 1234–1242. [Google Scholar] [CrossRef] [Scilit]
- Fiscon, G.; Weitschek, E.; Felici, G.; Bertolazzi, P.; De Salvo, S.; Bramanti, P.; De Cola, M.C. Alzheimer’s disease patients’ classification through EEG signals processing. In Proceedings of the 2014 IEEE Symposium on Computational Intelligence and Data Mining (CIDM), Orlando, FL, USA, 9–12 December 2014; pp. 105–112. [Google Scholar]
- Cassani, R.; Estarellas, M.; San-Martin, R.; Fraga, F.J.; Falk, T.H. Systematic review on resting-state EEG for Alzheimer’s disease diagnosis and progression assessment. Dis. Markers 2018, 2018, 5174815. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Fiscon, G.; Weitschek, E.; Cialini, A.; Felici, G.; Bertolazzi, P.; De Salvo, S.; Bramanti, A.; Bramanti, P.; De Cola, M.C. Combining EEG signal processing with supervised methods for Alzheimer’s patients’ classification. BMC Med. Inform. Decis. Mak. 2018, 18, 35. [Google Scholar] [CrossRef] [Scilit]
- Cassani, R.; Falk, T.H. Alzheimer’s disease diagnosis and severity level detection based on electroencephalography modulation spectral “patch” features. IEEE J. Biomed. Health Inform. 2019, 24, 1982–1993. [Google Scholar] [CrossRef] [Scilit]
- Bi, X.; Wang, H. Early Alzheimer’s disease diagnosis based on EEG spectral images using deep learning. Neural Netw. 2019, 114, 119–135. [Google Scholar] [CrossRef] [Scilit]
- Vecchio, F.; Miraglia, F.; Alù, F.; Menna, M.; Judica, E.; Cotelli, M.; Rossini, P.M. Classification of Alzheimer’s disease with respect to physiological aging with innovative EEG biomarkers in a machine learning implementation. J. Alzheimer’s Dis. 2020, 75, 1253–1261. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Ieracitano, C.; Mammone, N.; Hussain, A.; Morabito, F.C. A novel multi-modal machine learning based approach for automatic classification of EEG recordings in dementia. Neural Netw. 2020, 123, 176–190. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Seo, J.; Laine, T.H.; Oh, G.; Sohn, K.-A. EEG-Based Emotion Classification for Alzheimer’s Disease Patients Using Conventional Machine Learning and Recurrent Neural Network Models. Sensors 2020, 20, 7212. [Google Scholar] [CrossRef] [Scilit]
- Huggins, C.J.; Escudero, J.; Parra, M.A.; Scally, B.; Anghinah, R.; De Araújo, A.V.L.; Basile, L.F.; Abasolo, D. Deep learning of resting-state electroencephalogram signals for three-class classification of Alzheimer’s disease, mild cognitive impairment and healthy ageing. J. Neural Eng. 2021, 18, 046087. [Google Scholar] [CrossRef] [Scilit]
- Alves, C.L.; Pineda, A.M.; Roster, K.; Thielemann, C.; Rodrigues, F.A. EEG functional connectivity and deep learning for automatic diagnosis of brain disorders: Alzheimer’s disease and schizophrenia. J. Phys. Complex. 2022, 3, 025001. [Google Scholar] [CrossRef] [Scilit]
- Pirrone, D.; Weitschek, E.; Di Paolo, P.; De Salvo, S.; De Cola, M.C. EEG signal processing and supervised machine learning to early diagnose Alzheimer’s disease. Appl. Sci. 2022, 12, 5413. [Google Scholar] [CrossRef] [Scilit]
- Puri, D.; Nalbalwar, S.; Nandgaonkar, A.; Kachare, P.; Rajput, J.; Wagh, A. Alzheimer’s Disease detection using empirical mode decomposition and Hjorth parameters of EEG signal. In Proceedings of the 2022 International Conference on Decision Aid Sciences and Applications (DASA), Chiangrai, Thailand, 3–25 March 2022; IEEE: Piscataway, NJ, USA, 2022; pp. 23–28. [Google Scholar]
- Rossini, P.M.; Miraglia, F.; Vecchio, F. Early dementia diagnosis, MCI-to-dementia risk prediction, and the role of machine learning methods for feature extraction from integrated biomarkers, for EEG signal analysis. Alzheimer’s Dement. 2022, 18, 2699–2706. [Google Scholar] [CrossRef] [Scilit]
- Dogan, S.; Baygin, M.; Tasci, B.; Loh, H.W.; Barua, P.D.; Tuncer, T.; Tan, R.-S.; Acharya, U.R. Primate brain pattern-based automated Alzheimer’s disease detection model using EEG signals. Cogn. Neurodynamics 2022, 17, 647–659. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Klepl, D.; He, F.; Wu, M.; Blackburn, D.J.; Sarrigiannis, P. EEG-Based Graph Neural Network Classification of Alzheimer’s Disease: An Empirical Evaluation of Functional Connectivity Methods. IEEE Trans. Neural Syst. Rehabil. Eng. 2022, 30, 2651–2660. [Google Scholar] [CrossRef] [Scilit]
- Vicchietti, M.L.; Ramos, F.M.; Betting, L.E.; Campanharo, A.S.L.O. Computational methods of EEG signals analysis for Alzheimer’s disease classification. Sci. Rep. 2023, 13, 8184. [Google Scholar] [CrossRef] [Scilit]
- Zheng, X.; Wang, B.; Liu, H.; Wu, W.; Sun, J.; Fang, W.; Jiang, R.; Hu, Y.; Jin, C.; Wei, X.; et al. Diagnosis of Alzheimer’s disease via resting-state EEG: Integration of spectrum, complexity, and synchronization signal features. Front. Aging Neurosci. 2023, 15, 1288295. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Calub, G.I.A.; Elefante, E.N.; Galisanao, J.C.A.; Iguid, S.L.B.G.; Salise, J.C.; Prado, S.V. EEG-Based Classification of Stages of Alzheimer’s Disease (AD) and Mild Cognitive Impairment (MCI). In Proceedings of the 2023 5th International Conference on Bioengineering for Smart Technologies (BioSMART), Paris, France, 7–9 June 2023; pp. 1–6. [Google Scholar]
- Puri, D.V.; Nalbalwar, S.L.; Nandgaonkar, A.B.; Gawande, J.P.; Wagh, A. Automatic detection of Alzheimer’s disease from EEG signals using low-complexity orthogonal wavelet filter banks. Biomed. Signal Process Control 2023, 81, 104439. [Google Scholar] [CrossRef] [Scilit]
- Sen, S.Y.; Cura, O.K.; Yilmaz, G.C.; Akan, A. Classification of Alzheimer’s dementia EEG signals using deep learning. Trans. Inst. Meas. Control 2024, 47, 1353–1365. [Google Scholar] [CrossRef] [Scilit]
- Yang, X.; Fan, Z.; Li, Z.; Zhou, J. Resting-state EEG microstate features for Alzheimer’s disease classification. PLoS ONE 2024, 19, e0311958. [Google Scholar] [CrossRef] [Scilit]
- Ma, Y.; Bland, J.K.S.; Fujinami, T. Classification of Alzheimer’s Disease and Frontotemporal Dementia Using Electroencephalography to Quantify Communication between Electrode Pairs. Diagnostics 2024, 14, 2189. [Google Scholar] [CrossRef] [Scilit]
- Rostamikia, M.; Sarbaz, Y.; Makouei, S. EEG-based classification of Alzheimer’s disease and frontotemporal dementia: A comprehensive analysis of discriminative features. Cogn. Neurodynamics 2024, 18, 3447–3462. [Google Scholar] [CrossRef] [Scilit]
- Dogan, S.; Barua, P.D.; Baygin, M.; Tuncer, T.; Tan, R.S.; Ciaccio, E.J.; Fujita, H.; Devi, A.; Acharya, U.R. Lattice 123 pattern for automated Alzheimer’s detection using EEG signal. Cogn. Neurodynamics 2024, 18, 2503–2519. [Google Scholar] [CrossRef] [Scilit]
- Jang, K.I.; Kim, Y.I.; Ju, H.J.; An, S.J.; Park, P.W. Dementia classification using two-channel electroencephalography features. Sci. Rep. 2025, 15, 11513. [Google Scholar] [CrossRef] [Scilit]
- Huh, Y.J.; Park, J.H.; Kim, Y.J.; Kim, K.G. Ensemble Learning-Based Alzheimer’s Disease Classification Using Electroencephalogram Signals and Clock Drawing Test Images. Sensors 2025, 25, 2881. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Arikan, F.B.; Cetintas, D.; Aksoy, A.; Yildirim, M. A Deep Learning Approach to Alzheimer’s Diagnosis Using EEG Data: Dual-Attention and Optuna-Optimized SVM. Biomedicines 2025, 13, 2017. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Jain, S.; Srivastava, R. ‘Enhanced EEG-based Alzheimer’s disease detection using synchrosqueezing transform and deep transfer learning. Measurement 2025, 576, 105–117. [Google Scholar] [CrossRef] [Scilit]
- Cao, J.; Li, B.; Li, X. Identification of Alzheimer’s disease brain networks based on EEG phase synchronization. BioMed Eng. OnLine 2025, 24, 32. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zini, S.; Barbera, T.; Bianco, S.; Napoletano, P. Alzheimer’s disease classification from EEG using a multiscale temporal deep network. Biomed. Signal Process. Control. 2026, 114, 109321. [Google Scholar] [CrossRef] [Scilit]
- Lopez, M.E.; Varas, R.S. EEG-based classification framework for the detection of Alzheimer’s disease and wild cognitive impairment. Biomed. Signal Process. Control. 2026, 112, 108733. [Google Scholar] [CrossRef] [Scilit]
- Yu, P.; Shen, L.; Tang, L. Multimodal DTI-ALPS and hippocampal microstructural signatures unveil stage-specific pathways in Alzheimer’s disease progression. Front. Aging Neurosci. 2025, 17, 1609793. [Google Scholar] [CrossRef] [Scilit]
- Chen, L.; Wen, S.; Tang, Y. Brain Power mapping with optimized deep learning for EEG-based pilot fatigue detection. Biomed. Signal Process. Control. 2025, 114, 109284. [Google Scholar] [CrossRef] [Scilit]
- Trang, H.; Loc, T.H.; Nam, H.B.H. Proposed combination of PCA and MFCC feature extraction in speech recognition system. In Proceedings of the 2014 International Conference on Advanced Technologies for Communications (ATC 2014), Hanoi, Vietnam, 15–17 October 2014; pp. 697–702. [Google Scholar]
- Dhanjal, C.; Gunn, S.R.; Shawe-Taylor, J. Efficient Sparse Kernel Feature Extraction Based on Partial Least Squares. IEEE Trans. Pattern Anal. Mach. Intell. 2009, 31, 1347–1361. [Google Scholar] [CrossRef] [Scilit]
- Dong, M.; Cheng, Y.; Wan, L. A Novel Adaptive Bayesian Model Averaging-Based Multiple Kriging Method for Structural Reliability Analysis. IEEE Trans. Reliab. 2025, 74, 2185–2199. [Google Scholar] [CrossRef] [Scilit]
- Li, B.; He, Y.; Guo, F.; Zuo, L. A Novel Localization Algorithm Based on Isomap and Partial Least Squares for Wireless Sensor Networks. IEEE Trans. Instrum. Meas. 2013, 62, 304–314. [Google Scholar] [CrossRef] [Scilit]
- Gour, B.; Bandopadhyaya, T.K.; Sharma, S. ART neural network-based clustering method produces best quality clusters of fingerprints in comparison to Self-Organizing Map and K-Means Clustering Algorithms. In Proceedings of the 2008 International Conference on Innovations in Information Technology, Al Ain, United Arab Emirates, 16–18 December 2008; pp. 282–286. [Google Scholar]
- Yu, J.; Liu, L.; Wang, L.; Tan, M.; Xu, D. Turning Control of a Multilink Biomimetic Robotic Fish. IEEE Trans. Robot. 2008, 24, 201–206. [Google Scholar] [CrossRef] [Scilit]
- Liu, J.; Wang, D.; Luo, S. An Effective Constraint-Handling Improved Cuckoo Search Algorithm and Its Application in Aerodynamic Shape Optimization. IEEE Access 2020, 8, 139121–139142. [Google Scholar] [CrossRef] [Scilit]
- Lou, T.; Guan, G.; Yue, Z.; Wang, Y.; Tong, S. A Hybrid K-means Method based on Modified Rat Swarm Optimization Algorithm for Data Clustering. In Proceedings of the 2024 43rd Chinese Control Conference (CCC), Kunming, China, 28–31 July 2024; pp. 3689–3696. [Google Scholar]
- Trojovská, E.; Dehghani, M.; Trojovský, P. Zebra Optimization Algorithm: A New Bio-Inspired Optimization Algorithm for Solving Optimization Algorithm. IEEE Access 2022, 10, 49445–49473. [Google Scholar] [CrossRef] [Scilit]
- Bulut, N.E.; Dandil, E.; Yuzgec, U.; Duysak, A. CMACGSA: Improved Gravitational Search Algorithm Based on Cerebellar Model Articulation Controller for Optimization. IEEE Access 2025, 13, 20847–20870. [Google Scholar] [CrossRef] [Scilit]
- Zhao, Q.; Li, C. Two-Stage Multi-Swarm Particle Swarm Optimizer for Unconstrained and Constrained Global Optimization. IEEE Accesss 2020, 8, 124905–124927. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Q.; Tsang, E.C.C.; Hu, M.; He, Q.; Chen, D. Fuzzy Set-Based Kernel Extreme Learning Machine Autoencoder for Multi-Label Classification. In Proceedings of the 2021 International Conference on Machine Learning and Cybernetics (ICMLC), Adelaide, Australia, 4–5 December 2021; pp. 1–6. [Google Scholar]
- Wu, S.; Nagahashi, H. Parameterized AdaBoost: Introducing a Parameter to Speed Up the Training of Real AdaBoost. IEEE Signal Process. Lett. 2014, 21, 687–691. [Google Scholar] [CrossRef] [Scilit]
- Li, M. Application of CART decision tree combined with PCA algorithm in intrusion detection. In Proceedings of the 2017 8th IEEE International Conference on Software Engineering and Service Science (ICSESS), Beijing, China, 24–26 November 2017; pp. 38–41. [Google Scholar]
- Chen, C.L.P.; Liu, Z. Broad Learning System: An Effective and Efficient Incremental Learning System Without the Need for Deep Architecture. IEEE Trans. Neural Netw. Learn. Syst. 2018, 29, 10–24. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Y.; Wei, S.; Wang, Z.; Liu, H. Dual-Modal Gesture Recognition Using Adaptive Weight Hierarchical Soft Voting Mechanism. IEEE Trans. Cybern. 2025, 55, 1497–1508. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Mazzi, C.; Massironi, G.; Sanchez-Lopez, J.; De Togni, L.; Savazzi, S. Face recognition deficits in a patient with Alzheimer’s disease: Amnesia or agnosia? The importance of electrophysiological markers for differential diagnosis. Front. Aging Neurosci. 2020, 12, 580609. [Google Scholar] [CrossRef] [Scilit] [PubMed]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.









