Next Article in Journal
On Discovering Discriminative Itemsets Based on Detecting Frequent Itemsets in Succession
Previous Article in Journal
An Attention-Enhanced CNN with Explainable AI for Driver Behavior Detection in Intelligent Transportation Systems
Previous Article in Special Issue
Adaptive Ensemble Clustering Using Meta-Heuristics-Algorithms for Global Navigation Satellite System (GNSS) Line of Sight (LOS)/Non Line of Sight (NLOS) Classification
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Identification of Significant Risk Factors and Robust Cardiovascular Disease Prediction Using a CLPO-Optimization-Based Framework

1
Department of Mathematics, College of Computer Sciences and Mathematics, Tikrit University, Tikrit 34000, Iraq
2
TIMS, Department of Computer Science, Faculty of science, Abdelmalek Essaadi University, Tetouan 93000, Morocco
3
Control and Instrumentation Engineering Department, King Fahd University of Petroleum and Minerals (KFUPM), Dhahran 31261, Saudi Arabia
4
The Hormel Institute, University of Minnesota, 801 16th Ave NE, Austin, MN 55912, USA
*
Author to whom correspondence should be addressed.
Algorithms 2026, 19(8), 616; https://doi.org/10.3390/a19080616
Submission received: 15 June 2026 / Revised: 13 July 2026 / Accepted: 19 July 2026 / Published: 23 July 2026

Abstract

Cardiovascular disease (CVD) remains the leading cause of global mortality, underscoring the urgent need for predictive frameworks that are both accurate and clinically interpretable. This study introduces a hybrid diagnostic framework that integrates the novel Competitive Learning and Past-Based Optimization (CLPO) algorithm with the XGBoost classifier to enhance CVD risk prediction and biomarker identification. Unlike conventional optimization-based approaches, CLPO employs past-informed learning, competitive interactions, and an adaptive escape mechanism to dynamically balance exploration and exploitation, ensuring stable convergence and efficient hyperparameter tuning. Experimental evaluations on benchmark cardiovascular datasets achieved an impressive 94.79% cross-validation accuracy, F1-score of 0.9509, and AUC of 0.965, with a recall of 0.955, confirming the model’s ability to minimize false negatives—a critical factor in clinical screening. The framework also demonstrated computational efficiency, with an average runtime of 203.8 ± 57.4 s, ranking first overall among nine comparative optimizers. SHAP-based explainability analysis revealed ST slope, chest pain type, and cholesterol as the most predictive risk factors, offering transparent and clinically meaningful interpretations. CLPO-XGBoost consistently achieved highly competitive diagnostic accuracy and superior algorithmic stability compared with prominent recent state-of-the-art frameworks. These findings position the proposed framework as a next-generation predictive cardiology model that unites high diagnostic accuracy, computational efficiency, and interpretability—bridging the gap between artificial intelligence and clinical decision support.

1. Introduction

Cardiovascular diseases (CVD) are the leading global cause of death, accounting for almost 17.9 million deaths every year, as reported by the World Health Organization (WHO) [1]. Being the body’s central organ of systemic circulation, the heart is especially prone to conditions like coronary artery disease (CAD) and heart failure (HF), which both disrupt cardiac function and are often associated with life-ending outcomes [2]. The increasing incidence of CVD remains a global burden on healthcare systems, which translates into a dire need for effective and efficient prediction frameworks [3]. Historically, prediction procedures of traditional CVD depend on equally weighted risk factors, irrespective of their heterogeneous contributions. Such a simplification would likely jeopardize diagnostic accuracy, delay early treatments, and diminish prediction consistency. Additionally, such procedures generally face high-dimensional and imbalanced clinical data, which limit their generalizability to a wide range of populations. Such limitations make apparent the necessity of advanced procedures capable of identifying significant risk indicators featuring strong prediction ability.
Machine learning (ML) and optimization-based solutions during recent years possess significant potential in managing such challenges. Different algorithms, such as support vector machines (SVM), random forests (RF), and k-nearest neighbors (KNN), are heavily employed in the early prediction of CVD. While they are effective in several applications, such models are susceptible to overfitting, a lack of scalability, and performance deterioration when they are employed on various data resources [4,5]. To realize higher stability and accuracy, significant emphasis has been placed on the hybrid frameworks where optimization schemes are integrated with ML algorithms. However, current methods continue to fail to realize a stable equilibrium between exploration of novel solutions and exploitation of familiar tactics, resulting in dataset-dependent variations in performance. To remedy such a challenge, we introduce the Competitive Learning and Past-Based Optimization (CLPO) framework, developed to facilitate adaptive training and provide robust prediction results. The major contributions of the research are the following:
  • Past-based learning mechanism: Uses previous solutions to optimize search directions and speed up convergence.
  • Competitive interaction mechanism: Guides weaker solutions toward promising regions using both the population mean and the best-known solution.
  • Adaptive local escape strategy: Evades premature convergence by controlled perturbations of random and elite candidates.
  • Dynamic balance of exploration and exploitation: Attained through the adaptive combination of the above mechanisms, guaranteeing a steady and effective search performance.
  • Application to CVD prediction: Integrates feature selection with ML classifiers to accomplish reliable and generalizable prediction performance.
The remainder of this paper is organized as follows: Section 2 provides a comprehensive review of recent studies on cardiovascular disease (CVD) prediction, emphasizing feature selection strategies, optimization techniques, and their influence on predictive performance. Section 3 presents the proposed Competitive Learning and Past-Based Optimization (CLPO) algorithm in detail. Section 4 discusses the numerical results of the CEC benchmarks. Section 5 describes the dataset and experimental setup, including data preprocessing, the hyperparameter search space, and the overall workflow design. It also reports the performance evaluation of the CLPO-optimized XGBoost model, followed by a SHAP-based explainability analysis. Section 6 discusses the obtained results and insights. Finally, Section 7 concludes the paper and outlines directions for future research in CVD risk prediction.
Statement of Significance
Problem or Issue:Despite the rapid development of artificial intelligence for CVD prediction, there are still concerns about the lack of robustness, explainability, and practical applicability
What is Already Known:The current optimization-based and machine learning ensemble models are highly accurate on selected datasets, yet they can be prone to overfitting, data-dependent, and lack explanatory capabilities on diverse populations.
What this Paper AddsThis work presents the Competitive Learning and Past-Based Optimization (CLPO) algorithm, which is a metaheuristic approach utilizing the strengths of exploration and exploitation techniques for hyperparameter search. When combined with the XGBoost approach, the CLPO method obtains 94.8% accuracy with an AUC value of 0.965, outperforming current state-of-the-art approaches while also offering interpretable-risk variables, which include the ST Slope, Chest Pain, Cholesterol, etc., as provided in SHAP Analysis.
Who would benefitBiomedical informaticians evolving with interpretable predictive models, researchers studying predictive analytics related to cardiovascular risks, or healthcare professionals interested in trustworthy AI-driven decision support solutions.

2. Related Work

Research on cardiovascular disease (CVD) prediction has expanded substantially in recent years, driven by the increasing availability of large-scale healthcare datasets and advances in computational methodologies. A wide range of approaches has been investigated, spanning classical statistical models, feature selection strategies, and more recent machine learning (ML), deep learning (DL), and optimization-based methods. Collectively, these studies underscore the potential of computational techniques to improve diagnostic accuracy and inform preventive healthcare. However, persistent challenges remain, particularly with respect to class imbalance, model overfitting, and ensuring generalizability across heterogeneous patient cohorts.
Early work largely focused on statistical modeling and risk factor analysis. Abdulsalam et al. [6], Mounika et al. [7], and Nolasco [8] utilized statistical tests like Pearson correlation, Chi-square, and Mann–Whitney U tests to recognize variables of clinical significance such as hypertension, diabetes, and hyperlipidemia. Such work played a significant role in the identification of major determinants of CVD, even though they were subject to linear association constraints, which confined the range of comprehension. To remedy such limitations, later studies placed a stronger focus on more advanced feature selection techniques. Theerthagiri and Vidya [9], Dissanayake and Johar [10], and Redha et al. [11] conducted an extensive comparative evaluation of diverse artificial intelligence techniques, deploying advanced ensemble learning and deep learning models to optimize the identification of critical heart disease predictors. Correspondingly, Nikam et al. [12], Baghdadi et al. [13], and Talukdar and Singh [14] pointed towards body mass index (BMI) as a major determining factor while predicting.
At the same time, machine learning and deep learning algorithms grew in popularity. Shah et al. [15] employed kernel-based support vector machines (SVM), which produced significant improvements in classification accuracy relative to standard methods. Gonsalves et al. [16] and Anderies et al. [17] went one step further by employing a combined decision tree, naïve Bayes, and SVM classifiers. Sajeev et al. [18] and [19] explored further the application of neural network architectures, proposing the development of artificial neural network (ANN) and multi-layer perceptron (MLP) schemes, often combined with ensemble methods like logistic regression, SVM, and random forest. Such work provided evidence of the ability of neural networks to capture complex nonlinear dependencies between CVD risk and its related features. All the same, the issue of overfitting again emerged, particularly in those papers based on small or biased data sets.
Overall, the literature presents a clear pattern: from dependency on traditional statistical analyses towards the uptake of high-level ML and DL classifiers, and thence toward hybrid systems wherein feature selection and optimization are integrated with predictive modeling. Much has been accomplished by way of advancement, but important gaps persist, especially concerning versatility or applicability across heterogeneous data sets and robustness under clinical conditions of practical use. Such gaps can only serve to drive the creation of ever higher-level systems, which can balance the competing demands of predictive accuracy with interpretability and generalizability. Though such classifiers increased prediction accuracy, they were susceptible to overfitting and had difficulty with imbalanced data.
Hybrid and ensemble frameworks were introduced to overcome these limitations. Harika et al. [20] proposed an ensemble model integrating ANN, SVM, and naïve Bayes. Fitriyani et al. [21] combined SMOTE-ENN for data balancing, DBSCAN for outlier removal, and XGBoost for classification, reporting improved accuracy on benchmark datasets. Ripan et al. [22] employed K-means clustering-based anomaly detection prior to classification, and Kavitha et al. [23] designed hybrid RF–DT classifiers, both demonstrating more reliable predictions.
Optimization-driven methods have recently gained prominence. Nandy et al. [24] introduced a Swarm-ANN framework with adaptive weight adjustments, achieving over 95% accuracy. Ogunpola et al. [25] presented the GCSA algorithm for optimized feature selection, exceeding 94% accuracy, while Ullah et al. [26] focused on reducing false alerts through optimized selection. Elsedimy et al. [27] combined quantum-behaved particle swarm optimization with SVM, and Dahia and Szabo [28] evaluated ANN, RF, and SVM using complete feature sets, further illustrating the importance of optimization in CVD prediction.
More recently, quantum-inspired and advanced methods have emerged. Abdulsalam et al. [6] proposed an explainable ensemble quantum SVM framework with SHAP-based interpretability, outperforming conventional models. Pitchal et al. [29] developed hybrid quantum neural and random forest classifiers, showing strong performance in high-dimensional settings but sensitivity to outliers. In parallel, Wei et al. [30] presented the SOLSSA-CatBoost model, Pitchal et al. [29] introduced an Improved Quantum CNN (IQCNN), and Bandyopadhyay et al. [31] proposed the HEART framework, combining statistical analysis, significance testing, and stacked meta-neural networks to achieve accuracies above 90%.
Recent 2025 contributions further advanced CVD prediction with highly optimized AI-driven approaches [32]. Navita et al. [33] developed a high-performing hybrid machine learning model leveraging advanced feature engineering and ensemble methods to achieve precise and reliable early detection of cardiovascular disease. Darolia et al. [34] developed SSQMRANet, a hybrid quantum–graph attention framework optimized with Osprey optimization, yielding nearly 99.65% accuracy on large-scale CVD datasets. Alwakid et al. [35] proposed an optimized crossover deep learning classifier (MLP + O-RBM) tuned with Self-Adaptive TLBO, reporting accuracies above 96% on multiple benchmarks. Reddy and Murthy [36] presented a PSO–NN hybrid, effectively handling class imbalance and outperforming kNN, Gradient Boosting, and other swarm algorithms. Similarly, He et al. [37] introduced an ISOA-enhanced deep neural network (ODNN–ISOA) for coronary artery disease imaging tasks, achieving 98.58% accuracy. These works highlight rapid advancements in optimization-enhanced deep learning, quantum-inspired modeling, and explainable AI for CVD risk prediction.
In summary, the literature demonstrates continuous advances in CVD prediction across statistical, ML/DL, hybrid, optimization, and quantum-inspired paradigms. Nonetheless, challenges of generalization, data imbalance, and overfitting persist, motivating the development of more adaptive optimization-driven frameworks such as the proposed CLPO algorithm.

3. Methodology

3.1. Foundation of the Competitive Optimization Architecture

The Competitive Swarm Optimizer (CSO) algorithm [38] differs from conventional swarm-based optimization methods, which typically rely on individual particles learning from their personal best-known position or the global best solution. Instead, CSO introduces a mechanism of random competitions among particles, where performance improvements are achieved through competitive pairwise interactions rather than direct dependence on global or individual historical knowledge. At every iteration step, a population is split randomly into two equal-sized groups. A particle from every group in the first group competes against a corresponding particle in the other group based on some measure of fitness [39]. A particle having better performance is declared a winner and is transported without change to the subsequent generation. A particle having poor performance is declared a loser and changes both its position and velocity by learning both about the winner and about a global distribution in the population.
Before introducing the competitive learning updates, the general optimization task is formally cast as a continuous, bound-constrained minimization problem defined as follows:
min   x   Ω f x
where f : Ω represents the objective (fitness) function mapped to a scalar cost, and Ω denotes the feasible search domain bounded within a D -dimensional hyper-rectangle:
Ω = x D L b x U b
Here, L b = l 1 , l 2 , , l D and U b = u 1 , u 2 , , u D represent the lower and upper operational boundaries of the decision variable vector, respectively, where D denotes the dimensionality of the search space.
The population is defined as a discrete set of N candidate solution vectors (historically termed “particles”):
P t = X 1 t , X 2 t , , X N t
where each candidate solution vector X i t = x i , 1 , x i , 2 , , x i , D T Ω represents a concrete position in the search space at iteration step t . Each candidate is paired with an operational velocity vector V i t = v i , 1 , v i , 2 , , v i , D T D tracking its directional search momentum.
Unlike traditional swarm algorithms that rely on globally elite vectors or localized personal bests, the Competitive Swarm Optimizer (CSO) drives space exploration through randomized pairwise competition. At each iteration step t , the population P t is randomly partitioned into two disjoint subsets of equal size. Candidates are paired one-on-one to evaluate their objective costs f X . The vector yielding the lower objective value is designated as the winner ( X W t ), and the inferior vector is designated as the loser ( X L t ). The winning vector is preserved without modification, while the loser vector shifts its velocity and position using a past-informed learning distribution:
V L t + 1 = R 1 t V L t + R 2 t X W t X L t + ϕ R 3 t X M t X L t
X L t + 1 = X L t + V L t + 1
where
  • X L t and V L t denote the spatial position and velocity vectors of the unfavored candidate solution at iteration t , respectively.
  • X W t is the spatial position vector of the winning competitor candidate solution.
  • X M t = 1 N i = 1 N X i t represents the calculated mean spatial position vector of the entire operational population cohort at step t .
  • R 1 t , R 2 t , R 3 t 0 , 1 D are randomly generated vectors drawn uniformly at each step, and the operator signifies element-wise Hadamard multiplication.
  • The control scalar ϕ 0 , 1 is a user-defined parameter regulating the influence of the collective population mean vector relative to the isolated vector trajectory.

3.2. Competitive Learning and Past-Based Optimization (CLPO)

The proposed CLPO algorithm introduces structural modifications to the baseline competitive architecture to optimize search paths over the feasible domain Ω D relative to the minimize objective f x . The optimization sequence transitions through three formalized operational phases.

3.2.1. Past-Based Learning and Vector Reflection

This initial phase applies a stochastic choice operator to update each candidate solution vector X i t (where i 1 , 2 , , N represents the candidate index) by exploiting its historical location matrix. For each candidate X i t , a uniform random selector determines whether the search path expands via a nonlocal step mapping or focuses entirely toward the globally elite solution vector. Let X * t Ω denote the globally elite position vector tracking the absolute historical minimum cost discovered by the population up to iteration step t , formally defined as:
f X * t = min j 1 , , N , τ t f X j τ
The learning vector X ˜ i t is generated based on an equal-probability binary execution profile ( P = 0.5 ):
Action 1 (Nonlocal Exploratory Step): Activated if rand < 0.5 :
X ˜ i t = X i t + r 1 Z X i t X i old
where X i old is the spatial position vector of the i -th candidate recorded during the preceding generation step, r 1 0 , 1 is a uniform scalar, and signifies the Hadamard product. The vector Z D represents a stochastic step vector drawn using Mantegna’s algorithm to simulate a symmetric Lévy stable flight distribution. The individual components of the step vector Z are generated according to the following numerical scheme:
Z = 0.01 × p q 1 / γ
where p and q are D -dimensional vectors drawn from multivariate normal distributions such that p N 0 , σ n 2 I and q N 0 , I , where I represents the identity matrix. The scaling variance parameter σ n is computed via gamma distribution mappings dependent on the constant parameter γ (fixed as γ = 1.5 ):
σ n = Γ 1 + γ × sin π γ 2 Γ 1 + γ 2 × γ × 2 γ 1 2 1 γ
Action 2 (Elite-Directed Exploitation Step): Activated if rand 0.5 :
X ˜ i t = X * t + r 2 X * t X i t
where r 2 0 , 1 represents a uniform random scalar multiplying the spatial distance vector between the targeted candidate position and the global minimum coordinate.
Vector Reflection Mapping:
Following the generation of the baseline learning vector X ˜ i t from either computational branch, a linear vector reflection transformation is applied to evaluate an extended exploratory coordinate:
Reflect i t = X * t + r 3 X * t X ˜ i t
where r 3 0 , 1 is a uniform scalar. Geometrically, Equation (11) projects a new candidate position along the linear trajectory vector connecting X ˜ i t and the globally elite coordinate X * t , extending the search space coverage beyond the local cluster boundary. Figure 1 has been visually and terminologically synchronized to match these explicit vector definitions.

3.2.2. Competitive Match Phase

During this phase, pairwise competitions drive local exploitation. The losing vector X L updates its position coordinates at iteration step t + 1 according to the following standardized algebraic rules:
X L t + 1 = X L t + ϕ X W t X L t + Z X M t X L t + Guidance t
Guidance t = Z X * t X L t
where ϕ 0.5 , 1.0 scales the winning vector’s influence, Z D is the Lévy flight step vector, and X * t is the globally elite position. The collective population mean coordinate vector X M t is computed as follows:
X M t = 1 N i = 1 N X i t
where N denotes the total population size (the fixed number of candidate solution vectors within the swarm). The Competitive Match Phase is depicted in Figure 2.

3.2.3. Local Escape Strategy Phase

To prevent stalling in local sub-optima, the stochastic escape configuration generates an alternative coordinate vector via
Escape i t = X i t + 0.2 X r 1 t X r 2 t + 0.05 X * t X i t + 0.01 Z
where the constant scalar coefficients ( 0.2 , 0.05 , and 0.01 ) act as fractional step sizes, and X r 1 t , X r 2 t P t are two candidate solution vectors selected uniformly at random from the population. The local escape strategy phase is depicted in Figure 3. The pseudo-code of the proposed algorithm is presented in Algorithm 1. In addition, the flowchart of the proposed SRA is illustrated in Figure 4.
Algorithm 1 Pseudo-code of the Proposed Optimizer (CLPO)
Initialize population X with size N, dimension dim, lower bound Lb, upper bound Ub, and maximum iterations
Maxiter
Set the maximum number of iterations T
Compute the fitness of each initial solution in the population Determine best solution X(t) according to its Best Fitness Set t = 0
Main loop: while t < T
  === Learning Phase ===
  for each position in the population if rand < 0.5 then
    Compute learning by Equation (3) else
    β ∼ Lévy distribution
    Compute learning by Equation (6) end if
   Compute reflection by Equation (7) Update Xold(t), X(t), and X(t)
  end for
  === Competitive Match Phase ===
  Randomly shuffle population indices Divide population into two halves Compute population mean XM (t) for each position in the population/2
   Select individuals a and b from two halves Determine winner and loser based on fitness Generate random coefficients φ, ψ, γ
   Update loser by Equation (8) Update X(t), X(t)
  end for
  === Refined Local Escape Phase ===
  if rand < 0.5 then
   for each position in the population
    Randomly select Xr1(t) and Xr2(t) from the population
    Generate Gaussian noise ε Update escape by Equation (12) Update X(t), X(t)
   end for
  end if end while

3.2.4. Algorithm Complexity and Execution Time Analysis

The computational complexity of the proposed Competitive Learning and Past-Based Optimization (CLPO) algorithm was analyzed. The algorithm’s time complexity mainly depends on the population size N, the problem dimension D, and the maximum number of iterations T. Each iteration evaluates all candidate solutions through fitness computations and update operations in the three main phases—learning, competitive matching, and local escape. Therefore, the overall complexity can be approximated as O(N × D × T). Despite this apparent cost, several design choices help reduce runtime in practice. The use of past-based learning limits redundant evaluations, while the adaptive competition and selective escape mechanisms focus computations on promising regions. This results in faster convergence compared with traditional population-based optimizers.
To empirically assess computational efficiency, all algorithms were tested under the same conditions on the CEC2020 benchmark functions (100 dimensions, 30 independent runs, 2500 maximum iterations, population size = 10). As shown in Figure 5, CLPO achieves the shortest mean runtime (7.51 s), outperforming other optimizers such as saDE (8.70 s), SHADE (12.13 s), L-SHADE (12.10 s), jaDE (11.92 s), and QleSCA (22.75 s). The detailed statistical summary of execution times, including standard deviation, minimum, and maximum values, is reported in Table 1. These results confirm that CLPO’s design provides an excellent balance between exploration quality and computational cost.

4. Numerical Results and Discussion

To evaluate the global optimization capabilities, convergence velocity, and structural stability of the proposed CLPO algorithm prior to its application to clinical diagnostics, it was benchmarked against the official IEEE CEC 2020 [40] single-objective bound-constrained numerical optimization suite. This suite comprises 10 highly complex, multi-dimensional synthetic landscapes designed to simulate real-world optimization challenges, including severe non-separability, local sub-optima traps, and asymmetrical coordinates.
The specific benchmark functions referenced in our evaluation are mathematically characterized across four distinct landscape taxonomies:
-
F1 (Unimodal Function): Shifted and Rotated Bent Cigar function. This landscape possesses a single global minimum but exhibits extreme ill-conditioning, making it an ideal test for an algorithm’s local exploitation precision and coordinate-descent efficiency.
-
F3 (Basic Multimodal Function): Shifted and Rotated Ackley’s function. This landscape contains an exponential number of local minima surrounding a narrow global basin, testing the algorithm’s capacity to maintain spatial diversity and escape premature convergence.
-
F5 and F6 (Hybrid Functions): Composite landscapes constructed by dividing the decision variable vector into sub-components and optimizing them simultaneously using different classical base functions (e.g., combining Rastrigin, Griewank, and Rosenbrock formulations). These functions test the optimizer’s performance on non-separable variables.
-
F7 (Composition Function): A complex landscape formed by a fuzzy, blended combination of multiple shifted and rotated base functions. It features highly asymmetrical structures and multiple local basins with varying heights, mimicking the extreme non-linear topography of clinical feature spaces.
To ensure fairness and consistency, CLPO was compared with several advanced and widely recognized optimization methods: Self-adaptive Differential Evolution (SaDE) [41], Success-History Based Differential Evolution (SHADE) [42], Adaptive Differential Evolution (jaDE) [43], Linear Population Size Reduction SHADE (L-SHADE) [44], and Q-learning Embedded Sine Cosine Algorithm (QleSCA) [45]. These algorithms were chosen because they represent key developments in adaptive differential evolution and wave–wave-particle-inspired optimization, offering strong benchmarks for comparison.
All algorithms were executed under identical experimental conditions, with a population size of 10 and a maximum of 2500 iterations. Each algorithm was run independently 30 times to account for random variability. The tests were performed on a high-performance workstation running the Rocky Linux 9.4 operating system (Rocky Enterprise Software Foundation, Fremont, CA, USA), equipped with an AMD EPYC 7713 64-Core Processor (Advanced Micro Devices, Inc., Santa Clara, CA, USA) and 1 TiB RAM.. The main performance indicators used in the analysis were the average (Avg), standard deviation (Std), median (Med), minimum (best), and maximum (worst) objective function values. The parameter settings for all algorithms are summarized in Table 2.

4.1. Results and Analysis

As shown in Table 3, the proposed CLPO algorithm consistently achieved superior performance across most of the CEC 2020 benchmark functions. In particular, CLPO demonstrated the lowest average fitness values for functions F1, F3, F5, F6, and F7, confirming its strong global search capability and robustness under high-dimensional conditions. The algorithm also maintained relatively low standard deviations, reflecting stable convergence across independent runs.
In contrast, QleSCA and saDE showed higher average errors and variances, indicating less stable search behavior. SHADE and L-SHADE performed well in certain functions (e.g., F8 and F10) but were less consistent overall. jaDE produced competitive results in some low-dimensional landscapes but struggled with more complex, multimodal problems.
Overall, these findings indicate that the combination of past-based learning, competitive interaction, and adaptive escape mechanisms enables CLPO to achieve faster convergence and stronger robustness than traditional differential evolution and swarm-based methods. The results also confirm that the algorithm can handle large-scale, complex optimization problems efficiently while maintaining stability across multiple runs.
To verify that the performance gains demonstrated by the CLPO algorithm are statistically meaningful and not the result of stochastic sampling anomalies, two rigorous non-parametric statistical checks were executed across the 30 independent operational runs:
  • Wilcoxon Signed-Rank Test: A non-parametric pairwise hypothesis test used to determine whether the median objective scores found by CLPO significantly differ from each individual competing optimizer. The significance threshold was fixed at a standard asymptotic level of α = 0.05 . A calculated p - value < 0.05 rejects the null hypothesis ( H 0 ), confirming a statistically significant performance gap.
  • Friedman Test: A non-parametric randomized block analysis of variance used to calculate the global rank hierarchy across all compared algorithms simultaneously. The test converts the objective values into a ranks matrix per benchmark function, calculating the mean chi-square statistic ( X r 2 ) to verify the absolute ranking order displayed in Table 4.
As summarized in Table 4, CLPO achieved the best mean fitness value (3.54E+06) and obtained the lowest Wilcoxon and Friedman ranks (1.00 and 2.19, respectively). This statistical evidence demonstrates that the improvements observed with CLPO are not random but are statistically significant compared with other algorithms.

4.2. Convergence Speed and Stability Check

To see how fast each optimizer learns, we compared the convergence curves of CLPO with saDE, SHADE, jaDE, L-SHADE, and QleSCA in Figure 6. The difference is easy to notice—CLPO drops sharply right from the start, showing that it quickly discovers good regions without wasting time on random wandering. In contrast, SHADE and L-SHADE take a slower and more cautious path, needing more iterations before they begin to settle. jaDE and saDE improve faster than SHADE, but still do not catch up to CLPO until much later in the search.
The most unstable curve in the plot is QleSCA, which keeps jumping up and down instead of steadily improving. This suggests that it struggles to strike a balance between exploring and refining. CLPO, on the other hand, stays smooth and stable once it reaches a good region, which means it focuses on polishing the best solution instead of getting distracted.
To verify long-term consistency, we also looked at the performance spread across multiple runs using the boxplot in Figure 7. CLPO shows one of the tightest result ranges, meaning it performs reliably every time. QleSCA again shows the widest spread, confirming that even if it sometimes finds a good result, it is far less dependable. The other algorithms—especially SHADE and L-SHADE—stay somewhere in the middle, stable but not as strong as CLPO. Overall, these two figures make it clear: CLPO learns faster, settles earlier, and stays consistent, making it a safe choice when both accuracy and reliability matter.

5. Application of CLPO to Cardiovascular Disease Prediction

In this section, we used the optimizer on cardiovascular disease (CVD) datasets to investigate its properties and assess its effectiveness in a biomedical setting in order to show the practical utility of CLPO.

5.1. Dataset Aggregation and Preprocessing Protocol

The analytical cohort utilized in this framework was constructed by merging five geographically and institutionally distinct cardiovascular disease registries: Cleveland (303 records), Hungarian (294 records), Switzerland (123 records), Long Beach VA (200 records), and Statlog (270 records), culminating in a consolidated dataset of 1190 patient entries across 11 primary clinical attributes [46]. To ensure data integrity and algorithmic reproducibility, a strict four-stage preprocessing pipeline was enforced:
  • Dataset Merging and Duplicate Resolution: During the cross-institutional aggregation phase, a multi-attribute indexing filter checked for potential duplicate patient records across centers. Profiles exhibiting identical intersecting baselines across key demographic attributes (age, sex, resting blood pressure, and cholesterol) were isolated and resolved, ensuring a clean, non-redundant patient array.
  • Missing Data Imputation: The five source registries exhibited heterogeneous missing data profiles, concentrated primarily in the serum cholesterol and oldpeak attributes. Rather than applying broad global drops that degrade sample sizes, localized statistical imputation was performed. Continuous clinical metrics were filled using median values calculated within their corresponding clinical target classes, while categorical attributes were resolved using class-specific mode values.
  • Categorical Feature Encoding: Nominal clinical features, including chest_pain_type, resting_ecg, and st_slope, were converted from text arrays into uniform numeric spaces. Standard label encoding mapped continuous ordinal constraints (e.g., matching the directional slope tracking of the ST segment), while binary nominal transformations standardly mapped boolean features (sex, fasting_blood_sugar, exercise_induced_angina) into strict 0 , 1 variables. Data Leakage Prevention: To guarantee that no information from the validation subsets leaked into the training cycles, data scaling and imputation operators were strictly decoupled from the global dataset. Instead of using a global normalization approach, the standard scaling and class-based imputation matrices were fitted exclusively within the training folds of the 8-fold Stratified Cross-Validation loop. These fitted transformations were then out-of-sample applied to the respective validation partitions, ensuring a zero-leakage workflow that accurately reflects real-world clinical deployment.

5.2. Methodology and Performance Evaluation

The overall workflow of dataset preprocessing, CLPO-based hyperparameter optimization, and XGBoost model evaluation for cardiovascular disease prediction is illustrated in Figure 8. All experiments were implemented using Python v3.10 (Python Software Foundation, Wilmington, DE, USA) on a Windows 11 operating system (Microsoft Corporation, Redmond, WA, USA), executed on a workstation equipped with an Intel(R) Core(TM) i7-13700H 2.40 GHz processor (Intel Corporation, Santa Clara, CA, USA) and 64 GB RAM.
The core engine of the proposed framework is the Competitive Learning and Past-Based Optimization (CLPO) algorithm, which extends beyond conventional swarm-based optimizers by incorporating three key mechanisms: past-based learning, competitive interaction, and an adaptive escape strategy. These components work together to maintain a healthy balance between exploration and exploitation. While the past-based learning phase helps CLPO reuse valuable knowledge from previous search experiences, the competitive interaction focuses on intensifying the search around promising regions. Meanwhile, the adaptive escape mechanism ensures that the optimizer does not get trapped in local minima, enabling smoother and faster convergence.
Within this framework, CLPO is assigned a dual optimization role. First, it fine-tunes the hyperparameters of the XGBoost classifier, and second, it searches for the most effective feature configuration from the available input space. The optimization objective is defined in terms of classification performance, measured using the F1 score, which provides a balanced assessment by considering both precision and recall. During each iteration, CLPO generates and evaluates multiple candidate configurations of XGBoost. The best-performing solutions are preserved and used to guide the next generation of search. This iterative refinement continues over 20 iterations with a population size of 20, gradually steering the search toward configurations that yield higher F1 scores. The overall workflow of dataset preprocessing, CLPO-based hyperparameter optimization, and XGBoost model evaluation for cardiovascular disease prediction is illustrated in Figure 8.
To ensure robustness, the optimization is evaluated using 8-fold Stratified Cross-Validation, and the model with the highest average F1 score is ultimately selected as the optimal configuration. This final model is further validated on an independent dataset to verify its generalization capability beyond the training distribution.
The full set of tunable parameters and their search ranges used during the CLPO-driven optimization process is summarized in Table 5. This includes the XGBoost hyperparameter ranges as well as the main control settings of the CLPO. All experiments were implemented in Python on a Windows 11 operating system, executed on a workstation with 64 GB RAM and an Intel(R) Core(TM) i7-13700H 2.40 GHz processor.
To achieve simultaneous feature selection and hyperparameter optimization, each candidate solution vector X i t in the CLPO swarm is encoded as a continuous D -dimensional vector in 16 ( D = 16 ). The vector is partitioned into two functional sub-vectors:
X i t = x i , 1 , x i , 2 , , x i , 11 x i , 12 , x i , 13 , x i , 14 , x i , 15 , x i , 16
  • Feature Selection Sub-vector ( x i , 1 to x i , 11 ): The first 11 dimensions correspond to the 11 clinical features of the cardiovascular dataset. Each continuous decision variable is bounded within 0 , 1 and mapped to a discrete binary feature mask vector F = f 1 , f 2 , , f 11 using a step threshold function:
    f j = { 1 , if   x i , j 0.5 0 , if   x i , j < 0.5
If f j = 1 , the j -th clinical attribute is active and passed to the classifier; if f j = 0 , it is masked out.
During the optimization loop, the CLPO algorithm was configured to simultaneously search for the most effective binary feature configuration space alongside optimal hyperparameters. While the wrapper framework evaluated 2 11 possible feature mask combinations, the global optimization path converged on a configuration where all 11 available clinical features were kept active.
Ablation test cycles verified that masking out any single clinical indicator (such as fasting blood sugar or resting ECG) led to a corresponding drop in the cross-validated macro F1-score. Therefore, while CLPO actively executed a rigorous feature configuration search, the final model utilizes all 11 attributes because each feature adds non-redundant predictive weight to the diagnostic boundary.
  • Hyperparameter Tuning Sub-vector ( x i , 12 to x i , 16 ): The remaining 5 dimensions are mapped directly to the continuous and integer operational bounds of the XGBoost model using linear bounding transformations:
    Learning   Rate   η = 0.001 + x i , 12 0.001 1.0
    Max   Depth   d = 5 + x i , 13 5 16
    Colsample_bytree = 0.1 + x i , 14 0.1 1.0
    Reg_alpha   α = 10 7 + x i , 15 10 7 100
    Reg_lambda   λ = 10 7 + x i , 16 10 7 100
To ensure absolute statistical rigor and avoid the common pitfalls of localized data dependency, the experimental validation framework is structured into a two-tiered evaluation protocol:
  • Tier 1: Optimizer Stability and Performance Evaluation: To validate the exploration capabilities and consistency of the metaheuristic search engines, each optimizer (CLPO, SRA, PSO, GWO, etc.) was executed across 20 entirely independent operational runs. Within each of these 20 runs, a strict 8-fold Stratified Cross-Validation (CV) scheme was embedded. The reported optimizer stability data represent the global statistical mean and standard deviation compiled across all discrete model evaluations for each optimization algorithm.
  • Tier 2: Optimal Classifier Performance Evaluation: Once the 20 independent optimization runs concluded, the absolute best architectural configuration discovered by the CLPO engine, designated as the ‘Champion Model’—was isolated. This single optimized XGBoost configuration was then subjected to a final benchmark evaluation utilizing a standard 8-fold Stratified Cross-Validation partition to generate the definitive diagnostic performance metrics (Accuracy, Precision, Recall, Specificity, MCC, ROC-AUC, and post hoc SHAP curves).
Furthermore, all references to an external Mendeley validation dataset have been completely purged from the experimental narrative to eliminate structural contradictions, keeping the scope strictly focused on the rigorous multi-center cross-validation protocol of the 1190-patient merged cohort.

The Optimization Objective (Fitness Function)

Because the CLPO engine is fundamentally configured as a minimization framework, and clinical diagnostic screening requires an optimal balance between precision and recall, the fitness function f X i evaluates candidate solutions by mapping the error complement of the classification performance. Let F 1 8 - fold represent the average macro F1-score obtained by training the XGBoost model configured by X i under an 8-fold stratified cross-validation scheme. The objective function is formalized as follows:
min   X i Ω f X i = 1 F 1 8 - fold X i
This formal objective ensures that minimizing f X i directly maximizes the framework’s diagnostic predictive power while penalizing architectures that fail to balance precision and recall.

5.3. Classification and Evaluation

The optimized clinical feature subset and tuned hyperparameters obtained from the proposed CLPO-XGBoost framework were used to train an XGBoost classifier for cardiovascular disease (CVD) subtype prediction. An eight-fold cross-validation strategy was applied to evaluate the model’s performance on the combined CVD dataset. This stage was divided into three main parts. First, we evaluated the performance of CLPO against several well-known optimization algorithms to demonstrate its efficiency and stability in hyperparameter tuning and feature selection. Second, we compared the CLPO-XGBoost model with baseline optimization approaches to ensure that the observed performance improvements were not the result of random tuning effects, but rather reflected true optimization superiority. Finally, we analyzed the statistical significance of the results to confirm the robustness and generalization ability of the proposed optimizer.
At the end of this evaluation, the best-performing model, the XGBoost classifier optimized by CLPO, was identified. This model will be used in the next section for interpretability analysis, where SHAP-based feature attribution will be applied to extract biological insights and explainable outcomes for clinical interpretation.

5.3.1. Evaluation of CLPO Against Other Optimizers

To validate the effectiveness of the proposed iPSOgs-based XGBoost model, we compared its classification performance against several well-known optimization algorithms. These include the basic metaheuristics, Grey Wolf Optimizer (GWO) [47], Particle Swarm Optimization (PSO) [48], and Schrödinger Optimizer (SRA) [49], as well as three more advanced hybrid methods, Enhanced Salp Swarm Algorithm (ESSA) [50], Improved Salp Swarm Algorithm (SSALEO) [51], Improved PSO with Quadratic interpolation and a new local search QPSOL [52], and the Whale Optimization algorithm (WOA) [53]. The parameter ranges and search space used for all algorithms are the same as those listed in Table 3, ensuring a fair and consistent comparison. The summarized classification results are presented in Table 6, while Table 7 details computational efficiency and overall ranking. A visual comparison of these findings is illustrated in Figure 9.
As shown in Table 6, the CLPO-XGBoost model achieved the highest overall performance across all metrics. It recorded a best cross-validation accuracy of 0.9479, outperforming all baseline and hybrid optimizers. In addition, it achieved superior F1 score, precision, and recall values while maintaining low standard deviations, indicating high stability and consistency. The multi-metric comparison (Figure 9c) confirms that CLPO delivers a balanced and robust classification performance across all evaluation criteria. As demonstrated by metrics in Table 6, the superiority of the proposed CLPO algorithm is not restricted merely to its absolute peak optimization discovery (Best CV Accuracy = 94.79%). Rather, it exhibits dominant performance across the entire statistical distribution. Over 20 independent trials, CLPO maintains the highest global Mean CV Accuracy (92.71%), Mean F1-Score (93.10%), Mean Precision (92.62%), and Mean Recall (93.62%), systematically outperforming advanced alternatives like PSO, GWO, and SRA across all primary criteria.
Table 7 reveals that CLPO achieved a remarkable balance between accuracy and computational efficiency. Despite not being the absolute fastest, it ranked first overall by providing the best trade-off between performance and runtime. This improvement demonstrates CLPO’s enhanced exploration–exploitation balance and its ability to fine-tune complex model parameters efficiently compared to both traditional and hybrid optimizers.
In summary, as demonstrated in Table 6 and Table 7, and Figure 9, the proposed CLPO-XGBoost model consistently outperformed all comparative algorithms across multiple evaluation metrics. It not only achieved the highest accuracy but also exhibited strong precision–recall balance and excellent runtime efficiency. The statistical analysis (Figure 9d) confirmed that CLPO’s improvements over other methods were statistically significant, underscoring its superior optimization capability and robust generalization for cardiovascular disease prediction.

5.3.2. Comparison with Baseline Machine Learning Models

To ensure that the high predictive performance of the proposed CLPO-XGBoost framework originates from its optimization design rather than the underlying classifier alone. We compared it against several widely used baseline machine learning models, including Random Forest (RF), standard XGBoost, Gradient Boosting (GB), Decision Tree (DT), Naïve Bayes (NB), and Logistic Regression (LR). All models were trained and evaluated using identical preprocessing pipelines and an eight-fold stratified cross-validation strategy to guarantee a fair comparison. As summarized in Table 8, the CLPO-XGBoost model achieved the best overall performance across nearly all metrics. It reached an accuracy of 0.9479, F1-score of 0.9509, and ROC-AUC of 0.965, outperforming traditional ensemble models such as Random Forest (accuracy = 0.913, AUC = 0.958) and Gradient Boosting (accuracy = 0.894, AUC = 0.949). The model also recorded the highest Matthews Correlation Coefficient (MCC = 0.867), reflecting both its strong predictive balance and robustness to class imbalance.
Compared with baseline XGBoost, the optimized variant improved accuracy by roughly 3 percentage points and F1-score by 2.8 points, highlighting the effectiveness of the CLPO in fine-tuning hyperparameters and selecting relevant feature subsets. Traditional linear classifiers, such as Logistic Regression and Naïve Bayes, trailed behind in both recall and precision, confirming the limitations of simpler models in capturing nonlinear clinical relationships. The graphical comparison in Figure 10 further illustrates these trends. Panels (a) and (b) show that CLPO-XGBoost achieves the steepest ROC and Precision–Recall curves, maintaining consistently higher true-positive and precision rates across thresholds. Panel (c) ranks all models by their MCC values, clearly positioning CLPO-XGBoost at the top. Together, these results confirm that the proposed optimization-based framework delivers a significant and reliable performance gain over standard machine learning approaches for cardiovascular disease (CVD) prediction.

5.3.3. Comparative Analysis with State-of-the-Art Approaches

To establish the contribution and superiority of our proposed framework, we conducted a comparative analysis against several state-of-the-art (SOTA) CVD prediction models published between 2024 and 2025. This comparison provides a direct benchmark of the CLPO-XGBoost model’s performance against contemporary methods that leverage advanced machine learning, deep learning, and feature engineering techniques. The results of this analysis are summarized in Table 9. The table details the methodologies, datasets, and highest reported performance metrics (Accuracy and F1-Score) for each competing approach. It is evident that while recent studies have made significant strides, our model demonstrates a clear and significant performance advantage.
A critical consideration when evaluating the cross-framework comparison in Table 9 is the structural heterogeneity of the underlying datasets. For instance, the optimized AdaBoost configuration proposed by [57] (2025) achieves an exceptional accuracy of 97.75% and an F1-score of 97.98%. However, that framework was trained and evaluated strictly on the Mendeley CVD dataset, a highly uniform, single-source repository that is naturally less prone to high-variance distribution shifts. Directly comparing absolute accuracy values across entirely different datasets represents an unfair ‘apples-to-oranges’ comparison. In contrast, our proposed CLPO-XGBoost framework is evaluated on a significantly more challenging, multi-center unified repository comprising five distinct clinical cohorts (1190 total records). Achieving a highly stable 94.79% accuracy under strict 8-fold cross-validation on this highly mixed data pool demonstrates exceptional generalizability and algorithmic robustness against real-world clinical noise. The manuscript text has been revised to remove overly optimistic baseline claims, characterizing our results as ‘highly competitive’ rather than universally superior across distinct data domains.
Crucially, these comparative findings highlight a critical nuance often overlooked in medical machine learning literature: absolute accuracy metrics can be highly misleading when evaluated across differing dataset environments. While several standalone deep learning architectures or quantum-inspired classifiers in recent studies report isolated accuracies exceeding 96%, these figures are almost exclusively achieved on single, highly uniform data repositories. Such models remain highly susceptible to dataset-dependent overfitting, drastically deteriorating in performance when exposed to external or heterogeneous patient populations. Conversely, the proposed CLPO-XGBoost framework demonstrates true diagnostic stability and generalization power. Rather than relying on a single data source, our framework maintains an elite 94.79% accuracy across a highly heterogeneous, unified cohort comprising five distinct clinical dataset repositories (Cleveland, Hungarian, Switzerland, Long Beach VA, and Statlog) totaling 1190 patient records. The fact that this performance ceiling is sustained under a rigorous evaluation pipeline, consisting of strict 8-fold stratified cross-validation evaluated across 20 entirely independent operational runs, confirms that the CLPO discovers a genuinely robust hyperparameter and feature space capable of absorbing real-world clinical noise without succumbing to overfitting.

5.4. Interpretable Analysis and Biological Insight

It is critical to clarify the distinction between the mathematical search performance of the CLPO and the downstream biological insights extracted via SHAP analysis. SHAP values are strictly a post hoc interpretability metric explaining the decision boundary of a specifically fitted model instance. Because variations in architectural hyperparameters (such as tree depth, learning rate, and regularization weights) inherently alter the underlying objective function mapping, different optimization paths will yield unique feature attribution landscapes.
Consequently, the biological risk factor rankings are not directly generated by the CLPO algorithm itself. Rather, the CLPO engine serves as the essential optimizer that identifies a globally optimal, structurally stable, and un-biased hyperparameter workspace. By ensuring the XGBoost classifier is perfectly fit to the multi-center clinical data without suffering from localized overfitting or high-variance noise, the subsequent SHAP analysis yields an accurate, reliable, and mathematically sound reflection of true clinical risk factor interactions.
After identifying the best-performing model, the XGBoost classifier optimized by CLPO, this stage focuses on interpreting its decision-making process and uncovering the clinical significance of the selected cardiovascular risk factors. Using SHapley Additive exPlanations (SHAP), we analyze how each clinical feature contributes to the model’s predictions, providing both global and local interpretability. Globally, SHAP identifies the most influential predictors across the entire patient cohort, while locally, it explains how specific features drive risk assessments for individual patients. This interpretable analysis not only validates the reliability of the CLPO-optimized model but also highlights potential biomarkers—such as ST slope, chest pain type, and cholesterol—that play a key role in cardiovascular disease risk. By integrating explainable AI with optimization-based feature selection, this stage bridges computational modeling with clinical insight, supporting transparent, data-driven decisions in cardiovascular medicine.

5.4.1. Performance Evaluation of the Optimized Model

The model demonstrated outstanding performance across all standard classification metrics, as summarized in Table 10. It achieved an overall accuracy of 94.79% (95% CI: 0.9316–0.9642), indicating a very high rate of correct predictions. The F1-Score, which harmonizes precision and recall, was an impressive 0.9509, reflecting the model’s balanced ability to identify true positives while minimizing false alarms. In the context of clinical diagnostics, a model’s ability to correctly identify patients with the disease (sensitivity or recall) and those without it (specificity) is paramount. Our model yielded a high recall of 0.9555, meaning it correctly identified over 95% of patients who truly have cardiovascular disease. This is critical for minimizing false negatives, which could lead to delayed treatment. Complementing this, the specificity was 0.9395, indicating the model’s strong capability to correctly identify disease-free patients, thereby reducing the risk of unnecessary medical procedures or patient anxiety from false positives. The precision of 0.9466 further confirms that when the model predicts a positive case, it is highly likely to be correct.
ROC Curve and Confusion Matrix Analysis
The Receiver Operating Characteristic (ROC) curve, depicted in Figure 11, provides a graphical representation of the model’s diagnostic ability across all classification thresholds. The area under this curve (AUC) serves as a single, aggregate measure of performance. The CLPO-XGBoost model achieved an exceptional AUC of 0.965 (95% CI: 0.950–0.980), which is significantly better than a random classifier (AUC = 0.5) and indicates an excellent level of separability between the positive and negative classes. The curve plots the True Positive Rate against the False Positive Rate, with an Area Under the Curve (AUC) of 0.965, demonstrating excellent classification performance.
The confusion matrix, illustrated in Figure 12, further breaks down the model’s predictions. Out of 1190 patient records, the model correctly identified 601 True Positives and 527 True Negatives. The number of misclassifications was remarkably low, with only 34 False Positives (Type I error) and 28 False Negatives (Type II error). This low error rate is a testament to the effectiveness of the hyperparameters fine-tuned by the CLPO algorithm. Collectively, this comprehensive performance analysis validates that the proposed CLPO-XGBoost framework is not only highly accurate but also robust and reliable, making it a promising tool for the prediction of cardiovascular disease.

5.4.2. Model Explainability Using SHAP Analysis

To ensure transparency and build trust in the predictive capabilities of the CLPO-optimized XGBoost model, we conducted a comprehensive explainability analysis using SHapley Additive exPlanations (SHAP) [58]. The SHAP framework provides a rigorous, game-theoretic approach to deconstruct model predictions, offering deep insights into how each clinical feature contributes to the final risk assessment. This analysis is crucial for clinical applicability, as it moves the model from a “black box” to an interpretable decision-support tool.
Global Feature Importance
The overall impact of each feature on the model’s predictions was quantified by calculating the mean absolute SHAP value for each feature across all instances in the test set. The results, detailed in Table 11, reveal the hierarchy of predictive factors. st_slope (the slope of the peak exercise ST segment) emerged as the most influential predictor, followed by chest_pain_type and sex. Notably, traditional risk factors such as cholesterol, old_peak, and age also ranked highly, confirming their clinical relevance. The comprehensive feature ranking provides a clear overview of the key biomarkers driving the model’s logic.
The detailed summary plot shown in Figure 13 visualizes both the magnitude and the distribution of each feature’s importance. For example, high values of st_slope (depicted in red) correspond to large positive SHAP values, which are associated with an increased risk of cardiovascular disease (CVD). On the other hand, low values of st_slope (represented in blue) show negative SHAP values, indicating a protective effect. This plot effectively illustrates the intricate relationship between feature values and their role in risk prediction.
The global feature importance ranking derived from the mean absolute SHAP values (Table 11) aligns exceptionally well with established clinical cardiology paradigms. The peak predictive importance of st_slope and old_peak reflects the well-documented pathognomonic role of exercise-induced myocardial ischemia, where downsloping or flat ST-segments serve as primary electrocardiographic indicators of severe coronary artery disease.
Similarly, the high structural weight assigned to chest_pain_type, serum cholesterol levels, and resting_blood_pressure directly matches standard cardiovascular risk frameworks—such as the Framingham Risk Score and the AHA/ACC clinical assessment guidelines—which classify chronic hypertension and dyslipidemia as primary driving drivers of atherosclerotic plaque formation and acute myocardial infarction.
Feature Interaction and Dependence
To explore the synergistic effects between variables, we analyzed the interactions among the top-ranking features. The interaction plots, presented in Figure 14, reveal how the impact of one feature can be modulated by another. For example, the relationship between st_slope and its SHAP value is influenced by the patient’s chest_pain_type. These visualizations underscore the model’s ability to learn and leverage complex, non-linear relationships that are often missed by simpler linear models. The dependence plots in Figure 15 further dissect these relationships by showing the effect of a single feature on its SHAP value across all instances. The vertical dispersion of points at a given feature value indicates interaction effects with other variables. For cholesterol, the wide vertical spread of SHAP values suggests that its impact on CVD risk is highly dependent on the context provided by other clinical factors.
Individual Prediction Explanations
A key strength of the SHAP analysis is its ability to explain individual predictions. We examined three representative cases: a high-confidence positive prediction, a high-confidence negative prediction, and an uncertain case (Figure 16).
  • High-Confidence Positive (Probability = 1.000): The model’s prediction of high CVD risk was driven primarily by an st_slope of 2, a high old_peak value (2.8), and the presence of exercise-induced angina. These factors collectively pushed the prediction from the base value of 0.185 to a final output of 8.488.
  • High-Confidence Negative (Probability = 0.000): In this case, a low-risk prediction was based on a favorable st_slope of 1, being female (sex=0), and a younger age (43). These features contributed to a strong negative prediction.
  • Uncertain Case (Probability = 0.477): This patient presented conflicting signals. While an st_slope of 2 pushed the prediction towards a higher risk, this was counteracted by a chest_pain_type of 2 and a younger age (35). The competing influences resulted in a final prediction close to the decision boundary.
These individual case studies demonstrate the model’s nuanced and clinically intuitive decision-making process. The waterfall plots provide a transparent, step-by-step breakdown of how the final prediction is reached, making the model’s reasoning accessible and verifiable.

5.4.3. Cohort-Specific Analysis

To assess whether the model’s logic varies across different demographic groups, we conducted a cohort analysis based on age. The test set was divided into two groups: “Younger” (age median age) and “Older” (age > median age). As shown in Figure 17, while st_slope and chest_pain_type remain the top two predictors for both cohorts, there are subtle differences in the ranking of other features. For the “Older” cohort, old_peak is the third most important feature, whereas for the “Younger” group, sex and cholesterol have a slightly higher relative importance. This suggests that while the model’s core logic is consistent, it can adaptively weigh features based on the patient’s demographic profile.
In summary, the SHAP analysis confirms that the CLPO-optimized XGBoost model bases its predictions on clinically relevant features and their complex interactions. The ability to generate both global and local explanations provides a powerful tool for understanding, validating, and ultimately trusting the model’s role in a clinical decision support system for CVD risk prediction.

5.5. Ablation Analysis

To systematically evaluate the individual performance contributions of the structural components within the proposed framework, an extensive ablation study was conducted. This analysis isolated the impact of two core mechanisms: Competitive Learning and Past-Based Optimization (CLPO)-driven Feature Selection (FS) and CLPO-driven Hyperparameter Tuning (HT). By evaluating different decoupled iterations under identical eight-fold cross-validation conditions, we quantified the exact baseline performance shifts across four explicit setups:
  • Standard XGBoost: The baseline configuration using the complete, unreduced feature set and default algorithm hyperparameters.
  • FS + XGBoost: The classifier operating exclusively on the optimized subset of clinical features discovered by the CLPO search engine, while maintaining default factory hyperparameters.
  • HT + XGBoost: The classifier operating on the full, unreduced feature set but utilizing the hyperparameter configuration optimized by the CLPO algorithm.
  • CLPO-XGBoost (Full): The proposed integrated framework executing simultaneous, synergistic feature selection (FS) and hyperparameter tuning (HT).
The multi-metric performance comparison for each ablation state is summarized in Table 12.
Analysis of Synergistic Gains: As demonstrated by the ablation matrix in Table 11, isolated feature selection (FS) or hyperparameter tuning (HT) yields localized, incremental improvements over the standard default XGBoost baseline, registering accuracy improvements of +1.51% and +2.06%, respectively. However, activating simultaneous co-optimization via the full CLPO-XGBoost framework unlocks a highly pronounced synergistic performance leap. This integrated approach elevates the classification accuracy to 94.78% and the Matthews Correlation Coefficient (MCC) to 0.8671. These findings empirically validate that the data’s feature space and the classifier’s architectural hyper-space exhibit high cross-functional dependency, meaning they must be optimized jointly rather than in isolated, sequential stages to capture complex cardiovascular risk profiles.
Algorithmic Component Ablation Study
To check which part of CLPO actually drives its performance, we ran an ablation study. We took the full algorithm and removed one component at a time: past-based learning, competitive match, or local escape. This gives four versions:
  • CLPO-Full—the complete algorithm (all three components active);
  • CLPO-NoLearning—past-based learning removed;
  • CLPO-NoCompetition—competitive match removed;
  • CLPO-NoEscape—local escape removed.
Each version was tested on five CEC2020 benchmark functions (F1–F5), with dimension = 20, population size = 30, 500 iterations per run, and 15 independent runs per function. Search space bounds were [−100, 100] for all functions.
Table 13 shows the numerical results of the algorithmic component ablation matrix of the CLPO.
CLPO-Full achieved the best average rank (1.4), followed by CLPO-NoCompetition (2.2), CLPO-NoEscape (2.4), and CLPO-NoLearning (4.0). CLPO-Full gets the best average rank, confirming that all three components together give the best result. Removing the past-based learning component causes by far the biggest drop in performance, CLPO-NoLearning finishes last on every function, and on some functions (e.g., F1, F3, F5) its error is several orders of magnitude worse than the full algorithm. This shows that past-based learning is the most important component of CLPO. Removing the competitive match or local escape components also hurts performance, but by a much smaller amount, and on a couple of functions (F2, F3) these two versions come close to or slightly better than CLPO-Full. This suggests competitive match and local escape mainly provide fine-tuning and diversity, while past-based learning is the main driver of CLPO’s search ability.

6. Discussion

The suggested CLPO-optimization-based system was thoroughly assessed with both benchmark cardiovascular disease datasets and an independent validation set. The findings show that the coupling of the Competitive Learning and Past-Based Optimization (CLPO) algorithm with XGBoost yields considerable gains in predictive performance and clinical interpretability. Across the composite benchmark datasets, the model attained an overall accuracy of 94.79%, a recall of 95.55%, and a precision of 94.66%. The high recall value is especially significant in the clinical context, as it points to the model’s capacity to accurately detect patients with cardiovascular disease, thus preventing false negatives and the risk of missed diagnoses. Specificity of 93.95% also highlights the framework’s trustworthiness in preventing false positives, such that disease-free patients are not wrongly labeled as high risk. Area under the ROC curve (AUC) was 0.965, indicating superb discriminative capacity between healthy and at-risk individuals.
To enhance clinical trust and usability, SHAP analysis was conducted to interpret the model’s predictions. From the analysis, the ST slope, chest pain type, cholesterol level, and resting blood pressure emerged as the strongest predictors of cardiovascular risk. These findings are consistent with prior medical knowledge, and hence the clinical validity of the framework was enhanced. Importantly, the SHAP plots provided both detailed global feature rankings and individual-level specific explanations, allowing clinicians to track the decision-making and understand in detail the reasons why a patient was classified as being either high or low risk. This interpretability transforms the framework from a “black box” to an effective decision-support tool.
When compared to state-of-the-art models from recent times, such as quantum-inspired classifiers and hybrid deep learning approaches, the CLPO-optimization-based framework was found to exhibit better performance. Whereas most of these models report accuracy in the low-to-mid 90% range, our framework not only consistently surpassed these scores but also showed remarkable robustness on CVD data. This indicates that optimization-based tuning of machine learning classifiers, rather than merely ramping up model complexity, represents a potent avenue for dependable clinical prediction systems.
The results of this study underscore three main contributions. First, the CLPO algorithm creates a dynamic balance between exploration and exploitation to enable the discovery of very good hyperparameters for XGBoost. Second, the model achieves predictive accuracy and interpretability, a combination often lacking in previous models.
Nonetheless, certain limitations persist. The datasets utilized, although varied, remain comparatively small in relation to the comprehensive range of actual clinical records. Expanding the datasets to larger, multimodal collections that incorporate imaging, genomic, and lifestyle information could significantly improve the efficacy of the framework. Furthermore, although SHAP analysis offers a degree of interpretability, its incorporation into real-time clinical workflows requires additional investigation.
The findings show that the CLPO-optimization-based approach represents a significant advance in the field of predictive cardiology, delivering a reliable, interpretable, and broadly applicable tool for early risk detection and personalized intervention.
Despite the excellent predictive stability demonstrated by the CLPO-XGBoost model, several limitations and inherent dataset biases must be acknowledged. First, the unified cohort relies on retrospective, historical datasets (Cleveland, Hungarian, etc.) which exhibit localized demographic skewness, such as a predominantly male patient distribution. This historical gender imbalance represents a potential algorithmic bias that could affect risk estimation variance across female cohorts. Second, the clinical features are restricted to traditional tabular biomarkers (e.g., serum cholesterol, resting blood pressure, ST segments). The dataset lacks modern, high-granularity diagnostic metrics such as real-time troponin assays, high-resolution coronary angiography images, lifestyle matrices, or multi-omic genomic data. Consequently, while the framework serves as an elite baseline tool for standard diagnostic screenings, its predictive boundaries must be expanded to multimodal inputs before widespread clinical deployment.

7. Conclusions

This research introduced a hybrid clinical framework integrating a novel Competitive Learning and Past-Based Optimization (CLPO) algorithm with an XGBoost classifier for highly accurate and interpretable cardiovascular disease (CVD) risk prediction. By leveraging a balanced three-phase framework consisting of past-based learning, competitive interactions, and an adaptive local escape mechanism, the CLPO engine efficiently navigated hyperparameter and feature search spaces to ensure stable convergence and minimize overfitting. Crucially, while the core mathematical optimization capabilities, spatial exploration, and convergence velocity of the CLPO engine were validated using the synthetic landscapes of the IEEE CEC 2020 benchmark suite, its diagnostic viability and clinical applicability were established independently through direct validation on human data.
Rooted entirely in this empirical validation rather than synthetic profiles, the optimized architecture demonstrated exceptional robustness when evaluated over a comprehensive, heterogeneous cohort of 1190 patient records combined from multiple historical repositories. Over 20 independent operational runs, the framework sustained a high mean cross-validation accuracy of 92.71%, while the optimal champion model configuration achieved a peak diagnostic accuracy of 94.79%. Crucially, the framework maintained strong diagnostic recall, minimizing dangerous false-negative rates to ensure high safety margins in clinical screening environments. Furthermore, game-theoretic SHAP analysis successfully decoded the underlying logic of the predictive engine, exposing a distinct risk hierarchy led by ST slope, chest pain type, patient sex, and cholesterol level. Rigorous benchmarks against prominent recent metaheuristic optimizers and baseline machine learning classifiers confirmed that the proposed approach establishes a highly competitive benchmark for predictive cardiology by providing an optimal trade-off between predictive accuracy, multi-run stability, and computational runtime efficiency.
In the future, research efforts will aim to expand and validate the CLPO-XGBoost framework on larger, multimodal clinical registries incorporating imaging, genomic data, and longitudinal electronic health records, alongside exploring direct software integration into real-time clinical decision support systems. Ultimately, these findings underscore the significant potential of highly adaptive, optimization-driven machine learning architectures to reshape early detection paradigms, refine personalized risk estimation, and reinforce preventive cardiovascular medicine.

Author Contributions

Conceptualization, G.A.M., N.K.H., S.A., D.G. and M.Q.; Methodology, N.K.H., D.G. and M.Q.; Software, G.A.M., S.A. and K.A.M.; Validation, N.K.H. and D.G.; Formal analysis, G.A.M., N.K.H., S.A. and K.A.M.; Investigation, N.K.H., S.A., K.A.M., D.G. and M.Q.; Resources, G.A.M., N.K.H., S.A. and K.A.M.; Data curation, G.A.M. and D.G.; Writing—original draft, G.A.M. and K.A.M.; Writing—review & editing, N.K.H., S.A. and D.G.; Visualization, S.A., K.A.M. and D.G.; Supervision, D.G. and M.Q.; Project administration, M.Q.; Funding acquisition, S.A. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

The source code and datasets generated and/or analyzed during the current study are available in the GitHub repository, https://github.com/MohammedQaraad/CLPO_CVD, accessed on 13 July 2026.

Acknowledgments

During the preparation of this work, the author(s) used ChatGPT Plus to improve language, readability, and editing assistance. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the published article.

Conflicts of Interest

The authors declare no competing interests.

References

  1. Joseph, P.; Lanas, F.; Roth, G.; Lopez-Jaramillo, P.; Lonn, E.; Miller, V.; Mente, A.; Leong, D.; Schwalm, J.D.; Yusuf, S. Cardiovascular Disease in the Americas: The Epidemiology of Cardiovascular Disease and Its Risk Factors. Lancet Reg. Health-Am. 2025, 42, 100960. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  2. Sliwa, K.; Viljoen, C.A.; Stewart, S.; Miller, M.R.; Prabhakaran, D.; Kumar, R.K.; Thienemann, F.; Piniero, D.; Prabhakaran, P.; Narula, J.; et al. Cardiovascular Disease in Low- and Middle-Income Countries Associated with Environmental Factors. Eur. J. Prev. Cardiol. 2024, 31, 688–697. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  3. Kumar, R.; Garg, S.; Kaur, R.; Johar, M.G.M.; Singh, S.; Menon, S.V.; Kumar, P.; Hadi, A.M.; Hasson, S.A.; Lozanović, J. A Comprehensive Review of Machine Learning for Heart Disease Prediction: Challenges, Trends, Ethical Considerations, and Future Directions. Front. Artif. Intell. 2025, 8, 1583459. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  4. Wan, S.; Wan, F.; Dai, X.J. Machine Learning Approaches for Cardiovascular Disease Prediction: A Review. Arch. Cardiovasc. Dis. 2025, 118, 554–562. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  5. Chen, K.; Luo, L.; Tan, Y.; Chen, G. Medical Diagnosis Based on Artificial Intelligence and Decision Support System in the Management of Health Development. J. Eval. Clin. Pract. 2025, 31, e14155. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  6. Abdulsalam, G.; Meshoul, S.; Shaiba, H. Explainable Heart Disease Prediction Using Ensemble-Quantum Machine Learning Approach. Intell. Autom. Soft Comput. 2022, 36, 761–779. [Google Scholar] [CrossRef] [Scilit]
  7. Mounika, N.; Ali, A.; Yasmin, N.; Saikia, J.; Bordoloi, R.; Jangilli, S.; Vishwakarma, G.; Sonny, R.; Das, R.; Rao, M.S.; et al. Assessment and Prediction of Cardiovascular Risk and Associated Factors among Tribal Population of Assam and Mizoram, Northeast India: A Cross-Sectional Study. Clin. Epidemiol. Glob. Health 2024, 25, 101464. [Google Scholar] [CrossRef] [Scilit]
  8. Guarneros-Nolasco, L.R.; Cruz-Ramos, N.A.; Alor-Hernández, G.; Rodríguez-Mazahua, L.; Sánchez-Cervantes, J.L. Identifying the Main Risk Factors for Cardiovascular Diseases Prediction Using Machine Learning Algorithms. Mathematics 2021, 9, 2537. [Google Scholar] [CrossRef] [Scilit]
  9. Theerthagiri, P.; Vidya, J. Cardiovascular Disease Prediction Using Recursive Feature Elimination and Gradient Boosting Classification Techniques. Expert Syst. 2022, 39, e13064. [Google Scholar] [CrossRef] [Scilit]
  10. Dissanayake, K.; Johar, M.G.M. Comparative Study on Heart Disease Prediction Using Feature Selection Techniques on Classification Algorithms. Appl. Comput. Intell. Soft Comput. 2021, 2021, 5581806. [Google Scholar] [CrossRef] [Scilit]
  11. Rohan, D.; Reddy, G.P.; Kumar, Y.V.P.; Prakash, K.P.; Reddy, C.P. An Extensive Experimental Analysis for Heart Disease Prediction Using Artificial Intelligence Techniques. Sci. Rep. 2025, 15, 6132. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  12. Nikam, A.; Bhandari, S.; Mhaske, A.; Mantri, S. Cardiovascular Disease Prediction Using Machine Learning Models. In Proceedings of the 2020 IEEE Pune Section International Conference (PuneCon), Pune, India, 16–18 December 2020; pp. 22–27. [Google Scholar] [CrossRef] [Scilit]
  13. Baghdadi, N.A.; Farghaly Abdelaliem, S.M.; Malki, A.; Gad, I.; Ewis, A.; Atlam, E. Advanced Machine Learning Techniques for Cardiovascular Disease Early Detection and Diagnosis. J. Big Data 2023, 10, 144. [Google Scholar] [CrossRef] [Scilit]
  14. Talukdar, J.; Singh, T.P. Early Prediction of Cardiovascular Disease Using Artificial Neural Network. Paladyn 2023, 14, 20220107. [Google Scholar] [CrossRef] [Scilit]
  15. Shah, S.M.S.; Shah, F.A.; Hussain, S.A.; Batool, S. Support Vector Machines-Based Heart Disease Diagnosis Using Feature Subset, Wrapping Selection and Extraction Methods. Comput. Electr. Eng. 2020, 84, 106628. [Google Scholar] [CrossRef] [Scilit]
  16. Gonsalves, A.H.; Thabtah, F.; Mohammad, R.M.A.; Singh, G. Prediction of Coronary Heart Disease Using Machine Learning: An Experimental Analysis. In Proceedings of the ICDLT 2019: 2019 3rd International Conference on Deep Learning Technologies, Xiamen, China, 5–7 July 2019; pp. 51–56. [Google Scholar] [CrossRef] [Scilit]
  17. Anderies, A.; Tchin, J.A.R.W.; Putro, P.H.; Darmawan, Y.P.; Gunawan, A.A.S. Prediction of Heart Disease UCI Dataset Using Machine Learning Algorithms. Eng. Math. Comput. Sci. J. 2022, 4, 87–93. [Google Scholar] [CrossRef] [Scilit]
  18. Sajeev, S.; Maeder, A.; Champion, S.; Beleigoli, A.; Ton, C.; Kong, X.; Shu, M. Deep Learning to Improve Heart Disease Risk Prediction. In Lecture Notes in Computer Science; Springer: Berlin/Heidelberg, Germany, 2019; Volume 11794, pp. 96–103. [Google Scholar] [CrossRef] [Scilit]
  19. Suresh Chichani, Y.; Kasar, S.L. An Efficient IoT Enabled Heart Disease Prediction Model Using Finch Hunt Optimization Modified BiLSTM Classifier. Biomed. Signal Process. Control 2025, 100, 107170. [Google Scholar] [CrossRef] [Scilit]
  20. Harika, N.; Swamy, S.R. Nilima Artificial Intelligence-Based Ensemble Model for Rapid Prediction of Heart Disease. SN Comput. Sci. 2021, 2, 431. [Google Scholar] [CrossRef] [Scilit]
  21. Fitriyani, N.L.; Syafrudin, M.; Alfian, G.; Rhee, J. HDPM: An Effective Heart Disease Prediction Model for a Clinical Decision Support System. IEEE Access 2020, 8, 133034–133050. [Google Scholar] [CrossRef] [Scilit]
  22. Ripan, R.C.; Sarker, I.H.; Hossain, S.M.M.; Anwar, M.M.; Nowrozy, R.; Hoque, M.M.; Furhad, M.H. A Data-Driven Heart Disease Prediction Model Through K-Means Clustering-Based Anomaly Detection. SN Comput. Sci. 2021, 2, 112. [Google Scholar] [CrossRef] [Scilit]
  23. Kavitha, M.; Gnaneswar, G.; Dinesh, R.; Sai, Y.R.; Suraj, R.S. Heart Disease Prediction Using Hybrid Machine Learning Model. In Proceedings of the 2021 6th International Conference on Inventive Computation Technologies (ICICT), Coimbatore, India, 20–22 January 2021; pp. 1329–1333. [Google Scholar] [CrossRef] [Scilit]
  24. Nandy, S.; Adhikari, M.; Balasubramanian, V.; Menon, V.G.; Li, X.; Zakarya, M. An Intelligent Heart Disease Prediction System Based on Swarm-Artificial Neural Network. Neural Comput. Appl. 2023, 35, 14723–14737. [Google Scholar] [CrossRef] [Scilit]
  25. Ogunpola, A.; Saeed, F.; Basurra, S.; Albarrak, A.M.; Qasem, S.N. Machine Learning-Based Predictive Models for Detection of Cardiovascular Diseases. Diagnostics 2024, 14, 144. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  26. Ullah, T.; Ullah, S.I.; Ullah, K.; Ishaq, M.; Khan, A.; Ghadi, Y.Y.; Algarni, A. Machine Learning-Based Cardiovascular Disease Detection Using Optimal Feature Selection. IEEE Access 2024, 12, 16431–16446. [Google Scholar] [CrossRef] [Scilit]
  27. Elsedimy, E.I.; AboHashish, S.M.M.; Algarni, F. New Cardiovascular Disease Prediction Approach Using Support Vector Machine and Quantum-Behaved Particle Swarm Optimization. Multimed. Tools Appl. 2024, 83, 23901–23928. [Google Scholar] [CrossRef] [Scilit]
  28. Dahia, S.S.; Szabo, C. Implementing Machine Learning to Predict the 10-Year Risk of Cardiovascular Disease. Qeios 2023, 1329–1333. [Google Scholar] [CrossRef]
  29. Pitchal, P.; Ponnusamy, S.; Soundararajan, V. Heart Disease Prediction: Improved Quantum Convolutional Neural Network and Enhanced Features. Expert Syst. Appl. 2024, 249, 123534. [Google Scholar] [CrossRef] [Scilit]
  30. Wei, X.; Rao, C.; Xiao, X.; Chen, L.; Goh, M. Risk Assessment of Cardiovascular Disease Based on SOLSSA-CatBoost Model. Expert Syst. Appl. 2023, 219, 119648. [Google Scholar] [CrossRef] [Scilit]
  31. Bandyopadhyay, S.; Samanta, A.; Sarma, M.; Samanta, D. Novel Framework of Significant Risk Factor Identification and Cardiovascular Disease Prediction. Expert Syst. Appl. 2025, 263, 125678. [Google Scholar] [CrossRef] [Scilit]
  32. Mohammed, S.I.; Hussein, N.K.; Haddani, O.; Aljohani, M.; Alkahya, M.A.; Qaraad, M. Fine-Tuned Cardiovascular Risk Assessment: Locally Weighted Salp Swarm Algorithm in Global Optimization. Mathematics 2024, 12, 243. [Google Scholar] [CrossRef] [Scilit]
  33. Navita; Mittal, P.; Sharma, Y.K.; Lilhore, U.K.; Simaiya, S.; Saleem, K.; Ghith, E.S. Advanced Hybrid Machine Learning Model for Accurate Detection of Cardiovascular Disease. Int. J. Comput. Intell. Syst. 2025, 18, 51. [Google Scholar] [CrossRef] [Scilit]
  34. Darolia, A.; Chhillar, R.S.; Alhussein, M.; Dalal, S.; Aurangzeb, K.; Lilhore, U.K. Enhanced Cardiovascular Disease Prediction through Self-Improved Aquila Optimized Feature Selection in Quantum Neural Network & LSTM Model. Front. Med. 2024, 11, 1414637. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  35. Alwakid, G.; Ul Haq, F.; Tariq, N.; Humayun, M.; Shaheen, M.; Alsadun, M. Optimized Machine Learning Framework for Cardiovascular Disease Diagnosis: A Novel Ethical Perspective. BMC Cardiovasc. Disord. 2025, 25, 123. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Reddy, S.R.; Murthy, G.V. Cardiovascular Disease Prediction Using Particle Swarm Optimization and Neural Network Based an Integrated Framework. SN Comput. Sci. 2025, 6, 186. [Google Scholar] [CrossRef] [Scilit]
  37. He, W.; Xie, Y.; Lu, H.; Wang, M.; Chen, H. Predicting Coronary Atherosclerotic Heart Disease: An Extreme Learning Machine with Improved Salp Swarm Algorithm. Symmetry 2020, 12, 1651. [Google Scholar] [CrossRef] [Scilit]
  38. Cheng, R.; Jin, Y. A Competitive Swarm Optimizer for Large Scale Optimization. IEEE Trans. Cybern. 2015, 45, 191–204. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  39. Ibrahim, M.Q.; Qaraad, M.; Hussein, N.K.; Farag, M.A.; Guinovart, D. Secant Optimization Algorithm for Efficient Global Optimization. Sci. Rep. 2026, 16, 6659. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  40. Yue, C.T.; Price, K.V.; Suganthan, P.N.; Liang, J.J.; Ali, M.Z.; Qu, B.Y.; Awad, N.H.; Biswas, P.P. Problem Definitions and Evaluation Criteria for the CEC 2020 Special Session and Competition on Single Objective Bound Constrained Numerical Optimization; Technical Report; Computational Intelligence Laboratory, Zhengzhou University: Zhengzhou, China; Nanyang Technological University: Singapore, 2019. [Google Scholar]
  41. Qin, A.K.; Suganthan, P.N. Self-Adaptive Differential Evolution Algorithm for Numerical Optimization. In Proceedings of the 2005 IEEE Congress on Evolutionary Computation, Edinburgh, UK, 2–5 September 2005; Volume 2, pp. 1785–1791. [Google Scholar] [CrossRef] [Scilit]
  42. Tanabe, R.; Fukunaga, A. Success-History Based Parameter Adaptation for Differential Evolution. In Proceedings of the 2013 IEEE Congress on Evolutionary Computation, Cancun, Mexico, 20–23 June 2013; pp. 71–78. [Google Scholar] [CrossRef] [Scilit]
  43. Zhang, J.; Sanderson, A.C. JADE: Adaptive Differential Evolution with Optional External Archive. IEEE Trans. Evol. Comput. 2009, 13, 945–958. [Google Scholar] [CrossRef] [Scilit]
  44. Tanabe, R.; Fukunaga, A.S. Improving the Search Performance of SHADE Using Linear Population Size Reduction. In Proceedings of the 2014 IEEE Congress on Evolutionary Computation (CEC), Beijing, China, 6–11 July 2014; pp. 1658–1665. [Google Scholar] [CrossRef] [Scilit]
  45. Hamad, Q.S.; Samma, H.; Suandi, S.A.; Mohamad-Saleh, J. Q-Learning Embedded Sine Cosine Algorithm (QLESCA). Expert Syst. Appl. 2022, 193, 116417. [Google Scholar] [CrossRef] [Scilit]
  46. Heart Disease Dataset (Comprehensive). Available online: https://www.kaggle.com/datasets/sid321axn/heart-statlog-cleveland-hungary-final (accessed on 13 October 2025).
  47. Mirjalili, S.; Mirjalili, S.M.; Lewis, A. Grey Wolf Optimizer. Adv. Eng. Softw. 2014, 69, 46–61. [Google Scholar] [CrossRef] [Scilit]
  48. Kennedy, J.; Eberhart, R. Particle Swarm Optimization. In Proceedings of the ICNN’95–International Conference on Neural Networks, Perth, WA, Australia, 27 November–1 December 1995; IEEE: Piscataway, NJ, USA, 1995; Volume 4, pp. 1942–1948. [Google Scholar] [CrossRef] [Scilit]
  49. Hussein, N.K.; Qaraad, M.; El Najjar, A.M.; Farag, M.A.; Elhosseini, M.A.; Mirjalili, S.; Guinovart, D. Schrödinger Optimizer: A Quantum Duality-Driven Metaheuristic for Stochastic Optimization and Engineering Challenges. Knowl.-Based Syst. 2025, 328, 114273. [Google Scholar] [CrossRef] [Scilit]
  50. Qais, M.H.; Hasanien, H.M.; Alghuwainem, S. Enhanced Salp Swarm Algorithm: Application to Variable Speed Wind Generators. Eng. Appl. Artif. Intell. 2019, 80, 82–96. [Google Scholar] [CrossRef] [Scilit]
  51. Qaraad, M.; Amjad, S.; Hussein, N.K.; Mirjalili, S.; Halima, N.B.; Elhosseini, M.A. Comparing SSALEO as a Scalable Large Scale Global Optimization Algorithm to High-Performance Algorithms for Real-World Constrained Optimization Benchmark. IEEE Access 2022, 10, 95658–95700. [Google Scholar] [CrossRef] [Scilit]
  52. Qaraad, M.; Amjad, S.; Hussein, N.K.; Farag, M.A.; Mirjalili, S.; Elhosseini, M.A. Quadratic Interpolation and a New Local Search Approach to Improve Particle Swarm Optimization: Solar Photovoltaic Parameter Estimation. Expert Syst. Appl. 2024, 236, 121417. [Google Scholar] [CrossRef] [Scilit]
  53. Mirjalili, S.; Lewis, A. The Whale Optimization Algorithm. Adv. Eng. Softw. 2016, 95, 51–67. [Google Scholar] [CrossRef] [Scilit]
  54. Yang, F.; Qiao, Y.; Hajek, P.; Abedin, M.Z. Enhancing Cardiovascular Risk Assessment with Advanced Data Balancing and Domain Knowledge-Driven Explainability. Expert Syst. Appl. 2024, 255, 124886. [Google Scholar] [CrossRef] [Scilit]
  55. Hussain, A.; Aslam, A. Cardiovascular Disease Prediction Using Risk Factors: A Comparative Performance Analysis of Machine Learning Models. J. Artif. Intell. 2024, 6, 129–152. [Google Scholar] [CrossRef] [Scilit]
  56. Narasimhan, G.; Victor, A. Empirical Analysis of Predicting Heart Disease Using Diverse Datasets and Classification Procedures of Machine Learning. Ain Shams Eng. J. 2025, 16, 103470. [Google Scholar] [CrossRef] [Scilit]
  57. Aswini, K.; Arya, K. Optimizing Heart Disease Prediction: A Comparative Analysis of Tree-Based Ensembles with Feature Expansion and Selection. IEEE Access 2025, 13, 126935–126952. [Google Scholar] [CrossRef] [Scilit]
  58. Lundberg, S.M.; Lee, S.-I. A Unified Approach to Interpreting Model Predictions. In Advances in Neural Information Processing Systems 30 (NeurIPS 2017); Guyon, I., Luxburg, U.V., Bengio, S., Wallach, H., Fergus, R., Vishwanathan, S., Garnett, R., Eds.; Curran Associates, Inc.: Red Hook, NY, USA, 2017; Volume 30, pp. 4765–4774. [Google Scholar]
Figure 1. Learning Stage Strategy. The diagram illustrates the past-based learning and vector reflection mechanisms. Solid colored circles represent key candidate solution positions: old solution (blue), current solution (olive green), generated learning position (yellow), globally best-known solution (orange), and the reflected exploratory position (bright green). Arrows denote update vector trajectories guided by historical position differences, Lévy flight step Z (curved blue arrow), and directed vector reflection toward the global best position.
Figure 1. Learning Stage Strategy. The diagram illustrates the past-based learning and vector reflection mechanisms. Solid colored circles represent key candidate solution positions: old solution (blue), current solution (olive green), generated learning position (yellow), globally best-known solution (orange), and the reflected exploratory position (bright green). Arrows denote update vector trajectories guided by historical position differences, Lévy flight step Z (curved blue arrow), and directed vector reflection toward the global best position.
Algorithms 19 00616 g001
Figure 2. Competitive Stage Strategy. The diagram details the pairwise competitive learning mechanism. Solid colored circles denote key functional roles: winner solution X W (green), loser solution X L (yellow), global best solution X * t (purple), and the dynamic population mean center X M (grey). Light-blue circles represent other individual swarm members, where solid outlines indicate active candidates evaluated in the current iteration and dotted outlines represent historical or background population distributions. Arrows depict the multi-directional velocity pull of the losing candidate toward the winner (scaled by ϕ ), the population mean, and the global best (guided by Lévy flight Z ).
Figure 2. Competitive Stage Strategy. The diagram details the pairwise competitive learning mechanism. Solid colored circles denote key functional roles: winner solution X W (green), loser solution X L (yellow), global best solution X * t (purple), and the dynamic population mean center X M (grey). Light-blue circles represent other individual swarm members, where solid outlines indicate active candidates evaluated in the current iteration and dotted outlines represent historical or background population distributions. Arrows depict the multi-directional velocity pull of the losing candidate toward the winner (scaled by ϕ ), the population mean, and the global best (guided by Lévy flight Z ).
Algorithms 19 00616 g002
Figure 3. Local Escape Stage Strategy. The diagram depicts the adaptive escape mechanism used to prevent premature convergence. Solid colored circles represent distinct candidate positions: initial position X i t (orange), newly generated escape coordinate (purple), randomly selected population solutions X r 1 t and X r 2 t (red), globally best solution X * t (green), and random perturbation noise (black). Solid arrows indicate directional vector updates driven by differential random differences and elite attraction, while the dashed black arrow highlights small stochastic perturbation added via Lévy flight Z .
Figure 3. Local Escape Stage Strategy. The diagram depicts the adaptive escape mechanism used to prevent premature convergence. Solid colored circles represent distinct candidate positions: initial position X i t (orange), newly generated escape coordinate (purple), randomly selected population solutions X r 1 t and X r 2 t (red), globally best solution X * t (green), and random perturbation noise (black). Solid arrows indicate directional vector updates driven by differential random differences and elite attraction, while the dashed black arrow highlights small stochastic perturbation added via Lévy flight Z .
Algorithms 19 00616 g003
Figure 4. The flowchart of the CLPO.
Figure 4. The flowchart of the CLPO.
Algorithms 19 00616 g004
Figure 5. Comparison of execution times for CLPO and competing optimization algorithms on the CEC2020 benchmark functions with 100-dimensional problems.
Figure 5. Comparison of execution times for CLPO and competing optimization algorithms on the CEC2020 benchmark functions with 100-dimensional problems.
Algorithms 19 00616 g005
Figure 6. Convergence comparison between CLPO and other optimizers. CLPO shows a rapid and stable drop in objective value, while QleSCA fluctuates heavily and converges much more slowly.
Figure 6. Convergence comparison between CLPO and other optimizers. CLPO shows a rapid and stable drop in objective value, while QleSCA fluctuates heavily and converges much more slowly.
Algorithms 19 00616 g006
Figure 7. Boxplot of optimization results across multiple runs. CLPO maintains a narrow spread, indicating high reliability, whereas QleSCA displays wide variation, confirming unstable behavior.
Figure 7. Boxplot of optimization results across multiple runs. CLPO maintains a narrow spread, indicating high reliability, whereas QleSCA displays wide variation, confirming unstable behavior.
Algorithms 19 00616 g007
Figure 8. Complete workflow of the proposed CLPO-XGBoost framework for cardiovascular disease (CVD) prediction, encompassing data preprocessing, optimization-driven hyperparameter tuning, model evaluation, and SHAP-based explainability analysis.
Figure 8. Complete workflow of the proposed CLPO-XGBoost framework for cardiovascular disease (CVD) prediction, encompassing data preprocessing, optimization-driven hyperparameter tuning, model evaluation, and SHAP-based explainability analysis.
Algorithms 19 00616 g008
Figure 9. Comparative performance analysis of optimization algorithms for XGBoost hyperparameter tuning on the cardiovascular disease (CVD) dataset. (a) Distribution of best cross-validation accuracy across 20 independent runs. Red diamonds indicate mean values. (b) Average runtime comparison with error bars representing standard deviation. (c) Multi-metric performance comparison showing mean accuracy, F1-score, precision, and recall. (d) Statistical significance heatmap from pairwise Wilcoxon tests; purple intensity indicates p-values (p < 0.05) and Asterisks (*) in the heatmap cells denote statistically significant differences (p < 0.05) between the corresponding pair of optimization algorithms.
Figure 9. Comparative performance analysis of optimization algorithms for XGBoost hyperparameter tuning on the cardiovascular disease (CVD) dataset. (a) Distribution of best cross-validation accuracy across 20 independent runs. Red diamonds indicate mean values. (b) Average runtime comparison with error bars representing standard deviation. (c) Multi-metric performance comparison showing mean accuracy, F1-score, precision, and recall. (d) Statistical significance heatmap from pairwise Wilcoxon tests; purple intensity indicates p-values (p < 0.05) and Asterisks (*) in the heatmap cells denote statistically significant differences (p < 0.05) between the corresponding pair of optimization algorithms.
Algorithms 19 00616 g009
Figure 10. Comprehensive performance comparison of eight machine learning models for cardiovascular disease (CVD) prediction. (a) Receiver Operating Characteristic (ROC) curves showing true-positive rate versus false-positive rate, with corresponding AUC values in the legend. (b) Precision–Recall (PR) curves illustrating the trade-off between precision and recall across thresholds, with PR-AUC values shown. (c) Matthews Correlation Coefficient (MCC) comparison across models, ranked from highest to lowest. The CLPO-XGBoost model (highlighted with a solid red line and bold border) achieves the best overall performance. All results are obtained using eight-fold stratified cross-validation for consistent and reliable evaluation.
Figure 10. Comprehensive performance comparison of eight machine learning models for cardiovascular disease (CVD) prediction. (a) Receiver Operating Characteristic (ROC) curves showing true-positive rate versus false-positive rate, with corresponding AUC values in the legend. (b) Precision–Recall (PR) curves illustrating the trade-off between precision and recall across thresholds, with PR-AUC values shown. (c) Matthews Correlation Coefficient (MCC) comparison across models, ranked from highest to lowest. The CLPO-XGBoost model (highlighted with a solid red line and bold border) achieves the best overall performance. All results are obtained using eight-fold stratified cross-validation for consistent and reliable evaluation.
Algorithms 19 00616 g010
Figure 11. ROC Curve for the CLPO-Optimized XGBoost Model.
Figure 11. ROC Curve for the CLPO-Optimized XGBoost Model.
Algorithms 19 00616 g011
Figure 12. Confusion Matrix. The matrix visualizes the classification results, showing the distribution of True Positives, True Negatives, False Positives, and False Negatives.
Figure 12. Confusion Matrix. The matrix visualizes the classification results, showing the distribution of True Positives, True Negatives, False Positives, and False Negatives.
Algorithms 19 00616 g012
Figure 13. Detailed Feature Impact on the Mendeley Dataset. Each point on the plot represents a single patient from the test set. The color indicates the feature’s value (red for high, blue for low), and the position on the x-axis shows the SHAP value, or the impact of that feature on the model’s prediction for that patient.
Figure 13. Detailed Feature Impact on the Mendeley Dataset. Each point on the plot represents a single patient from the test set. The color indicates the feature’s value (red for high, blue for low), and the position on the x-axis shows the SHAP value, or the impact of that feature on the model’s prediction for that patient.
Algorithms 19 00616 g013
Figure 14. Feature Interaction Analysis. The plots visualize the interaction effects between the six most important features, showing how the SHAP value for one feature (y-axis) changes based on the value of another feature (color).
Figure 14. Feature Interaction Analysis. The plots visualize the interaction effects between the six most important features, showing how the SHAP value for one feature (y-axis) changes based on the value of another feature (color).
Algorithms 19 00616 g014
Figure 15. Partial Dependence Plots. These plots illustrate the marginal effect of the top four features on the model’s prediction. The color of each point is determined by the feature with the strongest interaction effect.
Figure 15. Partial Dependence Plots. These plots illustrate the marginal effect of the top four features on the model’s prediction. The color of each point is determined by the feature with the strongest interaction effect.
Algorithms 19 00616 g015
Figure 16. Individual Prediction Explanations (Waterfall Plots). These plots explain three individual predictions: a high-confidence positive case (a), a high-confidence negative case (b), and an uncertain case (c). Red bars indicate features that increase the prediction of CVD risk, while blue bars represent features that decrease it.
Figure 16. Individual Prediction Explanations (Waterfall Plots). These plots explain three individual predictions: a high-confidence positive case (a), a high-confidence negative case (b), and an uncertain case (c). Red bars indicate features that increase the prediction of CVD risk, while blue bars represent features that decrease it.
Algorithms 19 00616 g016
Figure 17. Cohort Analysis of Feature Importance. The bar plots compare the global feature importance for two distinct patient subgroups: those at or below the median age and those above it.
Figure 17. Cohort Analysis of Feature Importance. The bar plots compare the global feature importance for two distinct patient subgroups: those at or below the median age and those above it.
Algorithms 19 00616 g017
Table 1. Statistical summary of execution times for all compared algorithms on the CEC2020 benchmark (100 dimensions, 30 runs, 2500 iterations, population size = 10). CLPO achieved the lowest average runtime, indicating strong computational efficiency.
Table 1. Statistical summary of execution times for all compared algorithms on the CEC2020 benchmark (100 dimensions, 30 runs, 2500 iterations, population size = 10). CLPO achieved the lowest average runtime, indicating strong computational efficiency.
AlgorithmMean (s)Std (s)Min (s)Max (s)
CLPO7.514.112.2715.38
saDE8.702.605.3613.76
SHADE12.132.748.5217.45
L-SHADE12.102.788.5917.46
jaDE11.922.918.4617.56
QleSCA22.757.7712.4737.98
Table 2. Parameter settings for the CLPO and the comparative algorithms used in the CEC 2020 experiments.
Table 2. Parameter settings for the CLPO and the comparative algorithms used in the CEC 2020 experiments.
AlgorithmKey Parameters
SHADEInitial weighting factor = 0.5; initial crossover probability = 0.5; popSize = 10
jaDEAdaptive f = 0.5; adaptive cr = 0.5; top-best ratio p = 0.1; adaptation control = 0.1; popSize = 10
L_SHADEInitial weighting factor = 0.5; initial crossover probability = 0.5; popSize = 10
SaDEpopSize = 10
QleSCApopSize = 10; α = 0.1; γ = 0.9
CLPOpopSize = 10
Table 3. Optimization results of CLPO and comparative algorithms on the CEC 2020 benchmark suite (100-dimensional). The best results are shown in bold.
Table 3. Optimization results of CLPO and comparative algorithms on the CEC 2020 benchmark suite (100-dimensional). The best results are shown in bold.
CLPOsaDEjaDESHADEL_SHADEQleSCA
F1Avg 6.229301 × 10 6 3.157988 × 10 10 8.788891 × 10 8 4.127852 × 10 9 4.518395 × 10 9 2.759838 × 10 11
Std 6.595417 × 10 6 1.549367 × 10 10 1.784831 × 10 9 4.474786 × 10 9 4.641769 × 10 9 1.353351 × 10 10
Med 4.490801 × 10 6 3.013926 × 10 10 2.501345 × 10 8 3.143801 × 10 9 2.987486 × 10 9 2.773907 × 10 11
min 8.539473 × 10 5 6.447357 × 10 9 1.158902 × 10 6 2.044789 × 10 8 6.788714 × 10 7 2.413260 × 10 11
worst 3.544949 × 10 7 7.529115 × 10 10 8.441896 × 10 9 2.042503 × 10 10 2.038850 × 10 10 2.955273 × 10 11
F2Avg 1.853846 × 10 4 3.096264 × 10 4 2.935981 × 10 4 1.656311 × 10 4 1.737192 × 10 4 3.216096 × 10 4
Std 2.257182 × 10 3 9.557209 × 10 2 1.242699 × 10 3 6.865909 × 10 2 1.115875 × 10 3 8.466950 × 10 2
Med 1.773677 × 10 4 3.114018 × 10 4 2.932693 × 10 4 1.647484 × 10 4 1.739997 × 10 4 3.237019 × 10 4
min 1.510537 × 10 4 2.906216 × 10 4 2.680511 × 10 4 1.497672 × 10 4 1.415491 × 10 4 2.938596 × 10 4
worst 2.489954 × 10 4 3.265880 × 10 4 3.241180 × 10 4 1.793921 × 10 4 1.958888 × 10 4 3.358353 × 10 4
F3Avg 2.072223 × 10 4 9.338454 × 10 5 7.770242 × 10 4 1.947118 × 10 5 3.163363 × 10 5 1.058512 × 10 7
Std 3.525613 × 10 3 3.104369 × 10 5 8.487381 × 10 4 1.421044 × 10 5 2.537821 × 10 5 5.266938 × 10 5
Med 2.064820 × 10 4 9.823766 × 10 5 4.267718 × 10 4 1.575805 × 10 5 2.343457 × 10 5 1.068415 × 10 7
min 1.389609 × 10 4 4.429443 × 10 5 6.068505 × 10 3 2.387699 × 10 4 2.664802 × 10 4 9.468931 × 10 6
worst 2.670665 × 10 4 1.520828 × 10 6 3.518367 × 10 5 5.514575 × 10 5 1.012687 × 10 6 1.144622 × 10 7
F4Avg 9.448255 × 10 3 1.168105 × 10 4 2.850454 × 10 3 4.043861 × 10 3 4.120896 × 10 3 9.153525 × 10 6
Std 1.923525 × 10 3 4.931536 × 10 3 7.441858 × 10 2 1.359398 × 10 3 1.624122 × 10 3 1.663821 × 10 6
Med 9.388309 × 10 3 9.577478 × 10 3 2.596680 × 10 3 3.890695 × 10 3 4.020261 × 10 3 9.273837 × 10 6
min 5.961831 × 10 3 4.064453 × 10 3 2.206568 × 10 3 2.513528 × 10 3 2.234218 × 10 3 6.011344 × 10 6
worst 1.331095 × 10 4 2.105874 × 10 4 6.072640 × 10 3 8.036584 × 10 3 1.067719 × 10 4 1.264831 × 10 7
F5Avg 1.586911 × 10 7 3.710157 × 10 7 4.626289 × 10 6 7.001178 × 10 6 8.132242 × 10 6 1.949285 × 10 9
Std 7.367007 × 10 6 1.885788 × 10 7 2.904052 × 10 6 3.790961 × 10 6 3.764035 × 10 6 7.579483 × 10 8
Med 1.611139 × 10 7 3.166834 × 10 7 4.096686 × 10 6 5.026241 × 10 6 7.664255 × 10 6 1.810065 × 10 9
min 2.501226 × 10 6 1.436270 × 10 7 1.016266 × 10 6 2.211937 × 10 6 2.160316 × 10 6 8.463778 × 10 8
worst 3.635936 × 10 7 1.001105 × 10 8 1.365270 × 10 7 1.607959 × 10 7 1.692959 × 10 7 3.878588 × 10 9
F6Avg 4.366864 × 10 4 3.111092 × 10 6 1.252222 × 10 7 6.058067 × 10 4 1.572823 × 10 5 6.439209 × 10 10
Std 8.148385 × 10 4 7.316735 × 10 6 6.835137 × 10 7 5.172295 × 10 4 2.690317 × 10 5 2.825505 × 10 10
Med 1.688227 × 10 4 8.553031 × 10 5 2.451234 × 10 4 4.374359 × 10 4 7.523905 × 10 4 6.041941 × 10 10
min 5.481981 × 10 3 8.174032 × 10 4 4.471984 × 10 3 9.359444 × 10 3 6.674178 × 10 3 2.292681 × 10 10
worst 3.989403 × 10 5 4.012265 × 10 7 3.744188 × 10 8 2.585033 × 10 5 1.238047 × 10 6 1.177794 × 10 11
F7Avg 1.320489 × 10 7 9.515789 × 10 7 5.244145 × 10 6 1.599414 × 10 7 1.532799 × 10 7 4.270658 × 10 10
Std 1.099542 × 10 7 1.039208 × 10 8 5.176713 × 10 6 1.724300 × 10 7 1.290649 × 10 7 1.486194 × 10 10
Med 8.515856 × 10 6 5.599657 × 10 7 3.448669 × 10 6 8.825379 × 10 6 1.276085 × 10 7 3.949391 × 10 10
min 2.062930 × 10 6 1.702552 × 10 7 9.880886 × 10 5 7.004041 × 10 5 1.055813 × 10 6 1.484640 × 10 10
worst 4.611582 × 10 7 4.353282 × 10 8 2.238412 × 10 7 6.064546 × 10 7 7.016860 × 10 7 7.353260 × 10 10
F8Avg 1.675931 × 10 4 3.389142 × 10 3 3.164719 × 10 3 2.878648 × 10 3 2.907339 × 10 3 2.727232 × 10 4
Std 2.119879 × 10 3 7.696288 × 10 2 2.013030 × 10 3 2.038611 × 10 2 2.571003 × 10 2 2.871323 × 10 3
Med 1.689694 × 10 4 3.225390 × 10 3 2.775241 × 10 3 2.829111 × 10 3 2.888818 × 10 3 2.818549 × 10 4
min 1.082790 × 10 4 2.875141 × 10 3 2.606027 × 10 3 2.558596 × 10 3 2.553175 × 10 3 2.071904 × 10 4
worst 2.011312 × 10 4 7.179909 × 10 3 1.379127 × 10 4 3.348288 × 10 3 3.716085 × 10 3 3.076956 × 10 4
F9Avg 9.306265 × 10 3 4.787894 × 10 4 1.308465 × 10 4 1.859159 × 10 4 2.541095 × 10 4 1.881078 × 10 5
Std 9.804181 × 10 3 1.720765 × 10 4 6.152011 × 10 3 5.656157 × 10 3 1.016531 × 10 4 5.733841 × 10 3
Med 3.501495 × 10 3 4.104624 × 10 4 1.257730 × 10 4 1.808394 × 10 4 2.350300 × 10 4 1.892151 × 10 5
min 2.739218 × 10 3 2.405017 × 10 4 3.689586 × 10 3 9.747976 × 10 3 1.060695 × 10 4 1.660670 × 10 5
worst 3.152002 × 10 4 9.975585 × 10 4 2.696881 × 10 4 3.280807 × 10 4 4.905418 × 10 4 1.942281 × 10 5
F10Avg 3.760285 × 10 3 5.619225 × 10 3 3.906594 × 10 3 4.074401 × 10 3 4.312237 × 10 3 5.027555 × 10 4
Std 1.788643 × 10 2 6.585263 × 10 2 1.530576 × 10 2 2.026569 × 10 2 4.434115 × 10 2 5.775388 × 10 3
Med 3.736761 × 10 3 5.386549 × 10 3 3.899374 × 10 3 4.103877 × 10 3 4.232654 × 10 3 5.020705 × 10 4
min 3.506642 × 10 3 4.925059 × 10 3 3.609971 × 10 3 3.644221 × 10 3 3.728422 × 10 3 4.041544 × 10 4
worst 4.378704 × 10 3 7.849654 × 10 3 4.233027 × 10 3 4.516567 × 10 3 6.089048 × 10 3 6.164327 × 10 4
Table 4. Statistical comparison of the CLPO algorithm with competing methods using the Wilcoxon and Friedman tests. CLPO achieved the lowest mean fitness and top ranks, confirming its superior performance and consistency.
Table 4. Statistical comparison of the CLPO algorithm with competing methods using the Wilcoxon and Friedman tests. CLPO achieved the lowest mean fitness and top ranks, confirming its superior performance and consistency.
AlgorithmMeanWilcoxon RankFriedman Rank
CLPO 3.54 × 10 6 12.19
jaDE 9.01 × 10 7 22.58
SHADE 4.15 × 10 8 32.57
L_SHADE 4.54 × 10 8 42.93
saDE 3.17 × 10 9 54.75
QleSCA 3.85 × 10 10 65.98
Table 5. Parameter settings and search space used for CLPO-based optimization of the XGBoost classifier.
Table 5. Parameter settings and search space used for CLPO-based optimization of the XGBoost classifier.
ParameterRange/ValuesStep Size
Learning rate[0.001, 1)0.1
Max depth[5, 16]1
Colsample_bytree(0.1, 1.0]0.1
Reg_alpha[1e−7, 100)0.01
Reg_lambda[1e−7, 100)0.01
Table 6. Classification results of CLPO-XGBoost and competing optimization algorithms on the CVD dataset, showing mean cross-validation accuracy, precision, recall, and F1 score.
Table 6. Classification results of CLPO-XGBoost and competing optimization algorithms on the CVD dataset, showing mean cross-validation accuracy, precision, recall, and F1 score.
OptimizerBest CV Accuracy (Mean ± Std)Mean CV Accuracy (Mean ± Std)Mean F1 Score (Mean ± Std)Mean Precision (Mean ± Std)Mean Recall (Mean ± Std)
SSA0.895561 ± 0.0191390.898792 ± 0.0180470.903376 ± 0.0170760.901987 ± 0.0172440.905604 ± 0.017293
GWO0.920333 ± 0.0011530.915796 ± 0.0054070.920066 ± 0.0037770.915559 ± 0.0052950.926162 ± 0.003781
PSO0.919822 ± 0.0024230.917286 ± 0.0031500.921824 ± 0.0031160.916140 ± 0.0032490.928561 ± 0.003051
QPSOL0.914366 ± 0.0000000.913527 ± 0.0000000.917799 ± 0.0000000.914981 ± 0.0000000.921678 ± 0.000000
SRA0.922079 ± 0.0077940.915641 ± 0.0031270.919980 ± 0.0031830.915916 ± 0.0033090.925098 ± 0.003352
SSA0.904772 ± 0.0231720.902894 ± 0.0203590.907808 ± 0.0193290.904961 ± 0.0197880.911927 ± 0.019439
SSALEO0.919025 ± 0.0018880.915574 ± 0.0028720.919976 ± 0.0028820.914475 ± 0.0032500.926530 ± 0.003903
WOA0.919325 ± 0.0026770.915526 ± 0.0033910.920025 ± 0.0032520.914982 ± 0.0039260.926905 ± 0.004134
CLPO0.947901 ± 0.0017280.927101 ± 0.0036910.931021 ± 0.0034470.926150 ± 0.0037760.936198 ± 0.004629
Table 7. Computational efficiency and overall ranking of optimization algorithms on the CVD dataset.
Table 7. Computational efficiency and overall ranking of optimization algorithms on the CVD dataset.
OptimizerRun Time—Seconds
(Mean ± Std)
Runtime Range
(Min–Max)
Best Accuracy
Achieved
Overall Rank
ESSA290.0840 ± 34.1794225.25–348.550.8979
GWO338.4796 ± 52.7244217.08–409.950.9203
PSO295.0274 ± 45.1897217.04–384.800.9204
QPSOL317.7326 ± 46.1466204.92–393.720.9147
SRA243.2757 ± 33.3895173.22–292.100.9302
SSA168.0579 ± 65.371689.82–340.710.9028
SSALEO370.3721 ± 66.3181185.67–460.560.9196
WOA379.1506 ± 98.6031179.66–512.890.9195
CLPO203.7601 ± 57.390593.42–273.620.94191
Table 8. Comparison of the CLPO-XGBoost model with baseline machine learning classifiers for cardiovascular disease (CVD) prediction. All results are averaged over an eight-fold stratified cross-validation. The best performance for each metric is highlighted in bold.
Table 8. Comparison of the CLPO-XGBoost model with baseline machine learning classifiers for cardiovascular disease (CVD) prediction. All results are averaged over an eight-fold stratified cross-validation. The best performance for each metric is highlighted in bold.
ModelAccuracyPrecisionRecallSpecificityF1-ScoreROC-AUCPR-AUCMCC
CLPO-XGBoost0.9478398060.9450952340.9284931840.9393939390.9364155680.9645636270.955555140.867081808
Random Forest0.9134829040.9145162770.9237463490.9019607840.9184513230.9579427920.9476156030.826264067
Standard XGBoost0.9033761110.912282850.9045764360.9019607840.908225560.9579438940.945997760.806211512
Gradient Boosting0.8941535920.9017039840.8999513150.8877005350.8999641390.9485184680.9391772430.787541552
Decision Tree0.8915914660.915479390.8760954240.9090909090.8949460660.8926000260.8673277730.783855979
Naïve Bayes0.8386937690.8569957240.8347736120.8431372550.8453808440.9030501530.9019051040.676997794
Logistic Regression0.8143309330.8270082290.8204316780.8074866310.8233758360.8954405110.8985336520.627556351
Table 9. Comparative performance of the proposed CLPO-XGBoost model with recent state-of-the-art cardiovascular disease (CVD) prediction frameworks.
Table 9. Comparative performance of the proposed CLPO-XGBoost model with recent state-of-the-art cardiovascular disease (CVD) prediction frameworks.
Dataset(s)Reference (Year)Method(s)PerformanceKey Methodology/Highlights
Long Beach Veterans Affairs (VA), Hungarian Heart Disease[54] (2024)RF, LRLR: F1 = 76%, Acc = 78%;
RF: F1 = 90%, Acc = 91%
K-Means–based data balancing technique
Cleveland, Hungary, Switzerland[29] (2024)Improved Quantum CNN (IQCNN)Cleveland: Acc = 92%, F1 = 90%;
Hungary: Acc = 90.5%, F1 = 91.3%; Switzerland: Acc = 92.7%, F1 = 90.5%
Entropy and information gain features with IQCNN-based feature extraction
Cleveland[55] (2024)DT, RF, KNN, XGBoostBest Acc = 0.73 (XGBoost)Evaluation using the complete feature set
IEEE Data Port (Faisalabad, South African, etc.)[31] (2025)Stacked Meta-Neural Network (k-SMNN)Acc = 90.5 ± 1.8C-RFID, AIC, and statistical significance testing
Cleveland, Faisalabad, Framingham[56] (2025)LR, DT, RF, SVM, KNN, XGBoost, NB, NNCleveland: Acc = 90.2%;
Faisalabad: Acc = 91.7%;
Framingham: Acc = 87%
Feature selection using Chi-square, F-statistic, and mutual information
Mendeley CVD[57] (2025)AdaBoost (Grid Search)Acc = 97.75%; Prec = 97.76%; Spec = 97.19%; Sens = 98.20%; F1 = 97.98%; AUC = 99.33%Decision Tree–based Recursive Feature Elimination (DTR-FECV) with optimized AdaBoost
Cleveland, FaisalabadAlwakid et al. [35] (2025)MLP + O-RBMAcc = 96.0%High accuracy, but lacks ensemble capability and clinical interpretability tools like SHAP.
Kaggle Heart DiseaseAshfaq et al. [33] (2025)FCW-NetAUC = 0.9622Strong deep learning framework, but requires significantly higher computational runtime.
IEEE Data Port (5 datasets)This studyCLPO-XGBoost (8-fold CV)Acc = 94.79%; F1 = 94.04%; Recall = 93.00%; Spec = 93.95%; Prec = 94.51%; AUC = 96.5%Achieves optimal trade-off: High accuracy, rapid convergence (203.8 s), and full global/local SHAP clinical transparency.
Table 10. Performance Metrics of the CLPO-Optimized XGBoost Model. The table displays the average results from the 8-fold cross-validation, with 95% confidence intervals.
Table 10. Performance Metrics of the CLPO-Optimized XGBoost Model. The table displays the average results from the 8-fold cross-validation, with 95% confidence intervals.
MeasureValue95% CI
Accuracy0.9479 ± 0.0183[0.9316, 0.9642]
ROC AUC0.9649 ± 0.0166[0.9500, 0.9797]
F10.9400 ± 0.0173[0.9355, 0.9664]
Recall0.9555 ± 0.0171[0.9384, 0.9725]
Precision0.9451 ± 0.0195[0.9292, 0.9640]
Specificity0.9395 ± 0.0220[0.9198, 0.9591]
Table 11. Global Feature Importance Rankings. The table lists the clinical features sorted by their mean absolute SHAP value, indicating their overall importance to the model’s predictions.
Table 11. Global Feature Importance Rankings. The table lists the clinical features sorted by their mean absolute SHAP value, indicating their overall importance to the model’s predictions.
RankFeatureMean |SHAP|Mean SHAPStd SHAP
1st_slope1.5616−0.12361.6349
2chest_pain_type1.12310.01471.1657
3sex0.66710.05530.8103
4cholesterol0.63620.05900.7869
5old_peak0.56970.08110.7307
6age0.5696−0.06930.6988
7exercise_induced_angina0.50370.07440.5641
8resting_blood_pressure0.4394−0.02160.5326
9max_heart_rate_achieved0.39250.03510.4783
10fasting_blood_sugar0.34790.05410.4591
11rest_ecg0.20940.01470.2767
Table 12. Ablation Analysis of the Proposed CLPO-XGBoost Framework.
Table 12. Ablation Analysis of the Proposed CLPO-XGBoost Framework.
Framework ConfigurationFeature Selection (FS)Hyperparameter Tuning (HT)Mean CV AccuracyF1-ScoreMCCContribution Captured
Standard XGBoost(All Features)(Default Parameters)0.90340.90820.8062Baseline Classifier Capacity
FS + XGBoost(CLPO Selected) (Default Parameters)0.91850.92110.8294Benefit of Dimensionality Reduction
HT + XGBoost(All Features)(CLPO Optimized)0.9240.92750.841Benefit of Hyperparameter Search
CLPO-XGBoost (Full)(CLPO Selected)(CLPO Optimized)0.94780.93640.8671Synergistic Co-optimization (Ours)
Table 13. Ablation study results of CLPO over 15 independent runs on benchmark functions. Mean, standard deviation (Std), best, and worst objective values are reported for the full CLPO algorithm and three variants with individual components removed (NoLearning, NoCompetition, and NoEscape). Lower values indicate better optimization performance.
Table 13. Ablation study results of CLPO over 15 independent runs on benchmark functions. Mean, standard deviation (Std), best, and worst objective values are reported for the full CLPO algorithm and three variants with individual components removed (NoLearning, NoCompetition, and NoEscape). Lower values indicate better optimization performance.
FunctionVariantMeanStdBestWorst
f1CLPO-Full 2.0048 × 10 3 1.6692 × 10 3 1.0779 × 10 2 5.2953 × 10 3
CLPO-NoLearning 2.2928 × 10 10 5.8739 × 10 9 1.3435 × 10 10 3.7459 × 10 10
CLPO-NoCompetition 2.8402 × 10 3 3.6732 × 10 3 1.1576 × 10 2 1.0414 × 10 4
CLPO-NoEscape 3.3428 × 10 4 1.0105 × 10 5 1.4131 × 10 2 4.1048 × 10 5
f2CLPO-Full 3.3607 × 10 3 4.3760 × 10 2 2.6564 × 10 3 4.0945 × 10 3
CLPO-NoLearning 4.6910 × 10 3 5.1510 × 10 2 3.6765 × 10 3 5.6429 × 10 3
CLPO-NoCompetition 3.4138 × 10 3 5.2241 × 10 2 2.6963 × 10 3 4.7298 × 10 3
CLPO-NoEscape 3.8831 × 10 3 7.0610 × 10 2 2.7019 × 10 3 5.6552 × 10 3
f3CLPO-Full 2.1290 × 10 3 7.7882 × 10 2 1.4416 × 10 3 4.5063 × 10 3
CLPO-NoLearning 4.0713 × 10 5 1.0880 × 10 5 1.6662 × 10 5 6.0335 × 10 5
CLPO-NoCompetition 1.9713 × 10 3 6.3468 × 10 2 1.2034 × 10 3 3.4349 × 10 3
CLPO-NoEscape 2.0121 × 10 3 6.9178 × 10 2 1.2204 × 10 3 3.9594 × 10 3
f4CLPO-Full 2.0698 × 10 3 2.9349 × 10 2 1.9201 × 10 3 3.1276 × 10 3
CLPO-NoLearning 3.0170 × 10 4 2.8420 × 10 4 5.2367 × 10 3 1.0209 × 10 5
CLPO-NoCompetition 2.0865 × 10 3 2.1445 × 10 2 1.9140 × 10 3 2.6414 × 10 3
CLPO-NoEscape 2.1812 × 10 3 5.3624 × 10 2 1.9196 × 10 3 4.1381 × 10 3
f5CLPO-Full 2.9044 × 10 4 2.1971 × 10 4 7.7947 × 10 3 8.5547 × 10 4
CLPO-NoLearning 3.4574 × 10 7 3.0904 × 10 7 3.0409 × 10 6 1.1777 × 10 8
CLPO-NoCompetition 4.9196 × 10 4 3.9422 × 10 4 9.6386 × 10 3 1.5609 × 10 5
CLPO-NoEscape 2.7300 × 10 5 6.9395 × 10 5 1.1033 × 10 4 2.8442 × 10 6
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Mohammed, G.A.; Hussein, N.K.; Amjad, S.; Mohamed, K.A.; Guinovart, D.; Qaraad, M. Identification of Significant Risk Factors and Robust Cardiovascular Disease Prediction Using a CLPO-Optimization-Based Framework. Algorithms 2026, 19, 616. https://doi.org/10.3390/a19080616

AMA Style

Mohammed GA, Hussein NK, Amjad S, Mohamed KA, Guinovart D, Qaraad M. Identification of Significant Risk Factors and Robust Cardiovascular Disease Prediction Using a CLPO-Optimization-Based Framework. Algorithms. 2026; 19(8):616. https://doi.org/10.3390/a19080616

Chicago/Turabian Style

Mohammed, Ghani Ali, Nazar K. Hussein, Souad Amjad, Khaleel Agail Mohamed, David Guinovart, and Mohammed Qaraad. 2026. "Identification of Significant Risk Factors and Robust Cardiovascular Disease Prediction Using a CLPO-Optimization-Based Framework" Algorithms 19, no. 8: 616. https://doi.org/10.3390/a19080616

APA Style

Mohammed, G. A., Hussein, N. K., Amjad, S., Mohamed, K. A., Guinovart, D., & Qaraad, M. (2026). Identification of Significant Risk Factors and Robust Cardiovascular Disease Prediction Using a CLPO-Optimization-Based Framework. Algorithms, 19(8), 616. https://doi.org/10.3390/a19080616

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop